Image processing method, device and equipment

By pre-processing and AI processing of the original image in the front-end device, high-frequency information of image data is enhanced, and image quality is maintained during the encoding and compression process, the problem of image quality reduction after encoding of the front-end device is solved, and high-quality image transmission and storage are achieved.

CN119946286AActive Publication Date: 2025-05-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510435814.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

When the front-end device encodes the original image, the image quality is reduced, resulting in the target image quality decoded by the back-end device is lower than that of the original image.

Method used

In the front-end device, the original image is preprocessed, converted into image data suitable for the AI ​​model, the image data is processed through the AI ​​model, high-frequency information is enhanced, and the processed image data is stored in the storage medium. Then, the stored image data is encoded and compressed and sent to the backend device.

Benefits of technology

The AI ​​model enhances the high-frequency information in the image data, ensuring that even if the image quality is reduced during the encoding and compression process, the decoded target image quality can still be equivalent to the original image, improving the image quality effect of storage or display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946286A_ABST
    Figure CN119946286A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method, device and equipment, and the method comprises the steps: carrying out the preprocessing of an obtained original image, and obtaining first image data, and the preprocessing is used for converting the original image into the first image data matched with an AI model; performing AI processing on the first image data through the AI model to obtain second image data, wherein the AI processing is used for enhancing high-frequency information in the first image data; storing the second image data in a specified storage medium; and reading the stored image data from the specified storage medium, encoding and compressing the stored image data to obtain a target code stream, sending the target code stream to a back-end device, and obtaining a target image and storing or displaying the target image by the back-end device based on the target code stream. Through the technical scheme of the invention, the image quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, device and equipment. Background Art

[0002] Front-end devices (such as smartphones, IoT devices, cameras, etc.) are also called end-side devices. After acquiring the original image, the front-end device can encode the original image to obtain a bitstream and transmit the bitstream to the back-end device (such as a storage device, display device, etc.). After receiving the bitstream, the back-end device can decode the bitstream to obtain the target image, store the target image, or display the target image.

[0003] When the front-end device encodes the original image to obtain the bitstream, the encoding process is actually the compression process of the original image. Although it can reduce the amount of data transmission, it will also cause the image quality of the original image to deteriorate. Therefore, when the back-end device decodes the bitstream to obtain the target image, the image quality of the target image is lower than that of the original image, resulting in poor image quality of the stored or displayed target image. Summary of the invention

[0004] The present application provides an image processing method, which is applied to a front-end device, and the method comprises: Pre-processing the acquired original image to obtain first image data, wherein the pre-processing is used to convert the original image into first image data adapted to an artificial intelligence AI model; Performing AI processing on the first image data by the AI ​​model to obtain second image data, wherein the AI ​​processing is used to enhance high-frequency information in the first image data; storing the second image data in a designated storage medium; The stored image data is read from the designated storage medium, the stored image data is encoded and compressed to obtain a target code stream, the target code stream is sent to a back-end device, and the back-end device obtains a target image based on the target code stream, and stores or displays the target image.

[0005] The present application provides an image processing device, which is applied to a front-end device, and the device includes: A pre-processing module, used to pre-process the acquired original image to obtain first image data, wherein the pre-processing is used to convert the original image into first image data adapted to an artificial intelligence AI model; An AI processing module, configured to perform AI processing on the first image data through the AI ​​model to obtain second image data, wherein the AI ​​processing is used to enhance high-frequency information in the first image data; A storage module, used for storing the second image data in a designated storage medium; The sending module is used to read the stored image data from the specified storage medium, encode and compress the stored image data to obtain a target code stream, and send the target code stream to the back-end device, which obtains the target image based on the target code stream and stores or displays the target image.

[0006] The present application provides an electronic device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the image processing method of the above example of the present application.

[0007] The present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the image processing method of the above example of the present application is implemented.

[0008] The present application provides a machine-readable storage medium, which stores machine-executable instructions that can be executed by a processor; wherein the processor is used to execute the machine-executable instructions, and when the machine-executable instructions are executed, the image processing method of the above example of the present application is implemented.

[0009] It can be seen from the above technical solutions that in the embodiment of the present application, after the front-end device obtains the original image, it pre-processes the original image to obtain the first image data, performs AI processing on the first image data through the AI ​​model to obtain the second image data, and the AI ​​processing is used to enhance the high-frequency information in the first image data. On this basis, the front-end device encodes and compresses the image data to obtain the target code stream, sends the target code stream to the back-end device, and the back-end device decodes the target code stream to obtain the target image. Since the high-frequency information in the image data is enhanced by the AI ​​model, even if the encoding and compression process reduces the image quality of the image data (the high-frequency information in the image data will be lost, resulting in reduced image quality), when the back-end device decodes the target code stream to obtain the target image, the image quality of the target image can be equivalent to the image quality of the original image, thereby improving the image quality effect of the stored or displayed target image, and can improve the image quality. By processing the real-time data stream (i.e., the video stream) with the AI ​​model deployed on the front-end device, the image quality of the real-time data stream can be improved, and the video image encoding and decoding can be helped to obtain higher video image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a flowchart of an image processing method in one embodiment of the present application; Figure 2is a flowchart of an image processing method in one embodiment of the present application; Figure 3 It is a schematic diagram of the pre-compression process in one embodiment of the present application; Figure 4A is a schematic diagram of pixel filling for non-edge blocks in one implementation of the present application; Figure 4B is a schematic diagram of pixel filling of edge blocks in one implementation of the present application; Figure 5A is a schematic diagram of performing image transformation on first image data in one implementation of the present application; Figure 5B is a schematic diagram of performing image expansion on first image data in one implementation of the present application; Figure 5C is a schematic diagram of performing semantic transformation on first image data in one implementation of the present application; Figure 5D is a schematic diagram of performing AI processing on first image data in one implementation of the present application; Fig. 6A is a schematic diagram of image data storage in one embodiment of the present application; Figure 6B is a schematic diagram of a storage method of two tiles in one implementation of the present application; Figure 6C is a schematic diagram of image data display in one embodiment of the present application; Figure 7 is a schematic structural diagram of an image processing device in one embodiment of the present application; Figure 8 It is a hardware structure diagram of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION

[0011] In the embodiment of the present application, an image processing method is proposed, which can be applied to front-end devices, such as smart phones, Internet of Things devices, cameras, etc., and there is no restriction on the type of the front-end device. Figure 1 FIG. 1 is a flow chart of the image processing method, which may include: Step 101: Pre-process the acquired original image to obtain first image data, where the pre-processing is used to convert the original image into first image data that is compatible with an AI (artificial intelligence) model.

[0012] Step 102: Perform AI processing on the first image data through the AI ​​model to obtain second image data, where the AI ​​processing is used to enhance high-frequency information in the first image data.

[0013] Step 103: Store the second image data in a designated storage medium.

[0014] Step 104: read the stored image data from the designated storage medium, encode and compress the stored image data to obtain a target code stream, and send the target code stream to the back-end device, which obtains a target image based on the target code stream and stores or displays the target image.

[0015] Exemplarily, the pre-processing may include but is not limited to at least one of color space transformation, pixel value mapping, image block division, and image rearrangement. Pre-processing the acquired original image to obtain the first image data may include but is not limited to: if the pre-processing includes color space transformation, and the first color space of the original image is a color space not supported by the AI ​​model, the original image may be transformed in color space to obtain image data in a second color space, and the second color space is a color space supported by the AI ​​model.

[0016] If the pre-processing includes pixel value mapping, and the first pixel value distribution of the original image is a pixel value distribution not supported by the AI ​​model, the original image can be pixel value mapped to obtain image data of a second pixel value distribution, and the second pixel value distribution is a pixel value distribution supported by the AI ​​model.

[0017] If the pre-processing includes image block division, the original image can be divided into blocks to obtain multiple initial image blocks, and the multiple initial image blocks may include edge blocks and non-edge blocks. The edge blocks include the boundaries of the original image, and the non-edge blocks do not include the boundaries of the original image. For each initial image block, the initial image block can be filled with pixels to obtain a target image block. The multiple target image blocks serve as input data of the AI ​​model, and the multiple target image blocks are processed in parallel by the AI ​​model.

[0018] If the pre-processing includes image rearrangement, and the first arrangement of the original image is an arrangement not supported by the AI ​​model (i.e., data arrangement), the original image can be rearranged to obtain image data of a second arrangement, and the second arrangement is an arrangement supported by the AI ​​model.

[0019] Exemplarily, AI processing may include but is not limited to at least one of image transformation, image expansion and semantic transformation. If AI processing includes image transformation, then, performing AI processing on the first image data through the AI ​​model to obtain the second image data may include but is not limited to: performing image transformation on the first image data through the AI ​​model to obtain transformed image data, the resolution of the transformed image data may be equal to or less than the resolution of the first image data, and the number of channels of the transformed image data may be equal to or greater than the number of channels of the first image data; wherein the image transformation is used to separate high-frequency information and low-frequency information in the first image data. Wherein, the image transformation may include but is not limited to at least one of interpolation transformation (such as nearest neighbor interpolation transformation, mean transformation, bilinear interpolation transformation, adaptive interpolation transformation, bicubic interpolation transformation, etc.), wavelet transformation, Gaussian pyramid transformation, Laplace pyramid transformation and sobel filter transformation. The transformed image data is input into the image processing neural network in the AI ​​model, and the transformed image data is operated by the image processing neural network to obtain the second image data; wherein the resolution of the second image data may be equal to or less than the resolution of the first image data, and the image processing neural network is used to enhance the high-frequency information in the first image data.

[0020] Exemplarily, AI processing may include but is not limited to at least one of image transformation, image expansion and semantic transformation. If AI processing includes image expansion, then, performing AI processing on the first image data through the AI ​​model to obtain the second image data may include but is not limited to: inputting the first image data into the image processing neural network in the AI ​​model, operating the first image data through the image processing neural network to obtain the image data to be enhanced, the resolution of the image data to be enhanced may be equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced may be equal to or less than the number of channels of the first image data. Performing image expansion on the first image data through the AI ​​model to obtain the expanded image data corresponding to the image data to be enhanced, the resolution of the expanded image data is equal to or less than the resolution of the first image data, and the number of channels of the expanded image data is equal to or less than the number of channels of the first image data; wherein the expanded image data is used to enhance the high-frequency information in the image data to be enhanced; wherein the image expansion may include but is not limited to at least one of guided filtering, image edge enhancement, image contrast enhancement, and histogram equalization. The image data to be enhanced and the expanded image data are fused to obtain the second image data, and the resolution of the second image data is equal to or less than the resolution of the first image data.

[0021] Exemplarily, AI processing may include but is not limited to at least one of image transformation, image expansion and semantic transformation. If AI processing includes semantic transformation, then performing AI processing on the first image data through an AI model to obtain the second image data may include but is not limited to: inputting the first image data into the image processing neural network in the AI ​​model, operating the first image data through the image processing neural network to obtain the image data to be enhanced, the resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data. Inputting the first image data into the semantic processing neural network in the AI ​​model, operating the first image data through the semantic processing neural network to obtain structured semantic information, and the structured semantic information is used to describe the semantics of the image data to be enhanced. Generate the second image data based on the image data to be enhanced and the structured semantic information, that is, the second image data may include the image data to be enhanced and the structured semantic information.

[0022] Exemplarily, AI processing may include but is not limited to at least one of image transformation, image expansion and semantic transformation. If AI processing includes image transformation, image expansion and semantic transformation at the same time, then, performing AI processing on the first image data through the AI ​​model to obtain the second image data may include but is not limited to: performing image transformation on the first image data through the AI ​​model to obtain transformed image data, and the image transformation is used to separate high-frequency information and low-frequency information in the first image data; inputting the transformed image data into the image processing neural network in the AI ​​model, and operating the transformed image data through the image processing neural network to obtain the first image data to be enhanced, and the image processing neural network is used to enhance the high-frequency information in the first image data. Performing image expansion on the first image data through the AI ​​model to obtain expanded image data, and the expanded image data is used to enhance the high-frequency information in the first image data to be enhanced, and the first image data to be enhanced and the expanded image data are fused to obtain the second image data to be enhanced. Inputting the first image data into the semantic processing neural network in the AI ​​model, and operating the first image data through the semantic processing neural network to obtain structured semantic information, and the structured semantic information is used to describe the semantics of the second image data to be enhanced. The second image data is generated based on the second image data to be enhanced and the structured semantic information, that is, the second image data includes the second image data to be enhanced and the structured semantic information.

[0023] Exemplarily, if the first image data includes K target image blocks, the second image data includes K to-be-stored image blocks corresponding to the K target image blocks, K is a positive integer greater than 1; for each to-be-stored image block, the to-be-stored image block corresponds to a storage start address, and the interval between the storage start addresses of two adjacent to-be-stored image blocks is the width of the to-be-stored image block; the second image data corresponds to a storage span, and the storage span is greater than or equal to the sum of the widths of the K to-be-stored image blocks. On this basis, storing the second image data in a designated storage medium may include, but is not limited to: for each to-be-stored image block, when storing the i-th row of data of the to-be-stored image block in the designated storage medium, sequentially storing the i-th row of data starting from the target storage position; wherein the target storage position may be determined by the following formula: p=w+ stride*(i-1); wherein p may represent the target storage position, w may represent the storage start address corresponding to the to-be-stored image block, stride may represent the storage span, i may represent the number of rows of the to-be-stored image block, and i is a positive integer.

[0024] It can be seen from the above technical solutions that in the embodiment of the present application, after the front-end device obtains the original image, it pre-processes the original image to obtain the first image data, performs AI processing on the first image data through the AI ​​model to obtain the second image data, and the AI ​​processing is used to enhance the high-frequency information in the first image data. On this basis, the front-end device encodes and compresses the image data to obtain the target code stream, sends the target code stream to the back-end device, and the back-end device decodes the target code stream to obtain the target image. Since the high-frequency information in the image data is enhanced by the AI ​​model, even if the encoding and compression process reduces the image quality of the image data (the high-frequency information in the image data will be lost, resulting in reduced image quality), when the back-end device decodes the target code stream to obtain the target image, the image quality of the target image can be equivalent to the image quality of the original image, thereby improving the image quality effect of the stored or displayed target image, and can improve the image quality. By processing the real-time data stream (i.e., the video stream) with the AI ​​model deployed on the front-end device, the image quality of the real-time data stream can be improved, and the video image encoding and decoding can be helped to obtain higher video image quality.

[0025] The above technical solutions of the embodiments of the present application are described below in combination with specific application scenarios.

[0026] In an embodiment of the present application, a video image coding and decoding system is proposed, and the video image coding and decoding system may include a front-end device (the front-end device is used as a video encoding device of the video image coding and decoding system) and a back-end device (the back-end device is used as a video decoding device of the video image coding and decoding system). The front-end device may also be referred to as an end-side device, which is a device that is different from the cloud. The cloud-side device may use server resources to implement AI, while the end-side device uses local resources to implement AI, and does not use server resources to implement AI.

[0027] For example, the terminal side device can be a smart phone, an IoT device, a camera, a vehicle-mounted device, an intelligent driving device, an access control device, a conference tablet, an intelligent interactive all-in-one machine, etc. There is no restriction on the type of the terminal side device, and devices with image acquisition and image transmission functions can be called terminal side devices.

[0028] The backend device is a device that communicates with the frontend device, can receive images from the frontend device, and then process the images. The backend device can be a storage device (a device for storing images) or a display device (a device for displaying images). There is no restriction on the type of the backend device.

[0029] In the above application scenario, an image processing method is proposed in the embodiment of the present application, see Figure 2 FIG. 1 is a flow chart of the image processing method, which may include: Step 201: The front-end device obtains the original image.

[0030] For example, the front-end device can obtain a real-time data stream, which can be a video stream. Each frame of the image in the real-time data stream can be used as the original image, or part of the images in the real-time data stream can be used as the original image, and there is no restriction on the source of the original image.

[0031] Step 202: The front-end device pre-compresses the original image based on the configuration information to obtain image data.

[0032] Exemplarily, an AI model can be configured on a front-end device, and the AI ​​model can be used to pre-compress the original image based on the configuration information to obtain image data. By pre-compressing the original image, image data that is more conducive to subsequent encoding compression and compression recovery can be obtained. By pre-compressing the original image, the high-frequency information in the image data can be enhanced. Even if the encoding compression process reduces the image quality of the image data (the high-frequency information in the image data will be lost, resulting in reduced image quality), when the back-end device decodes the target code stream to obtain the target image, the image quality of the target image can be improved.

[0033] By configuring the AI ​​model on the front-end device instead of the back-end device, the AI ​​model can be executed on the front-end device to pre-compress the original image, which helps to process or make decisions on real-time data (original images) at the data source or on the front-end device that is "closer" to the data, thereby reducing data transmission between the front-end device and the back-end device, reducing data transmission delay, and reducing network bandwidth requirements.

[0034] When configuring an AI model on a front-end device, the AI ​​model can be a lightweight and efficient AI model, so that it can adapt to the needs of front-end devices with limited computing resources. Lightweight means that the amount of computing and memory required during the AI ​​model reasoning process is relatively low, and a certain model performance can still be maintained.

[0035] Step 203: The front-end device encodes and compresses the image data to obtain a target code stream.

[0036] Exemplarily, the front-end device can be used as a video encoding device of a video image encoding and decoding system, and a coding compression algorithm can be pre-configured. The front-end device can use the coding compression algorithm to encode and compress image data to obtain a target code stream. This embodiment does not limit this coding compression process.

[0037] Step 204: The front-end device sends the target code stream to the back-end device.

[0038] Step 205: The back-end device decodes the target code stream to obtain decoded data.

[0039] Exemplarily, the back-end device can be used as a video decoding device of a video image encoding and decoding system, and a decoding algorithm can be pre-configured. After receiving the target code stream, the back-end device can use the decoding algorithm to decode the target code stream to obtain decoded data. This decoding process is not limited in this embodiment.

[0040] Step 206: The back-end device compresses and restores the decoded data to obtain a target image.

[0041] Exemplarily, an AI model can be configured on the back-end device, and the decoded data can be compressed and restored by the AI ​​model to obtain a target image. The compression recovery can be the inverse process of pre-compression, that is, the AI ​​model of the back-end device is adapted to the AI ​​model of the front-end device. The AI ​​model of the front-end device is used to implement pre-compression, while the AI ​​model of the back-end device is used to implement compression recovery. By compressing and restoring the decoded data, the compressed and restored target image is adapted to the original image, that is, the target image is the same or approximately the same as the original image.

[0042] Step 207: The backend device stores or displays the target image. For example, if the backend device is a storage device, the target image can be stored; if the backend device is a display device, the target image can be displayed.

[0043] In a possible implementation, for step 202, see Figure 3 As shown in FIG. 1 , it is a schematic diagram of the pre-compression process. The following steps can be used to pre-compress the original image to obtain image data: Step 301: The front-end device pre-processes the original image to obtain first image data.

[0044] Exemplarily, pre-processing is used to convert the original image into first image data adapted to the AI ​​model based on the configuration information, so that the first image data can be the image data required by the AI ​​model.

[0045] Exemplarily, the pre-processing of the original image refers to processing the original image before encoding and transmitting it, such as downsampling processing, denoising processing, etc. The pre-processing is used to reduce the transmission bandwidth and improve the image quality. In this embodiment, pre-processing the original image refers to converting the original image into first image data that is compatible with the AI ​​model. Pre-processing may include, but is not limited to, at least one of color space transformation, pixel value mapping, image block division, and image rearrangement. The above are just a few examples of pre-processing, and there is no limitation to this. It is sufficient to be able to convert the original image into first image data that is compatible with the AI ​​model.

[0046] Exemplarily, the original image includes at least 2 channels of image data, or the original image includes at least 3 channels of image data. For example, the original image may include image data of a luminance (Luma) channel and image data of a chrominance (Chroma) channel. For example, the original image may include image data of a red (R) channel, image data of a green (G) channel, and image data of a blue (B) channel.

[0047] Case 1: If the first color space of the original image is a color space supported by the AI ​​model, the pre-processing does not include color space transformation. If the first color space of the original image is a color space not supported by the AI ​​model, the pre-processing includes color space transformation. Based on this, if the pre-processing includes color space transformation, the second color space supported by the AI ​​model is determined, and the original image is transformed into a color space to obtain image data of the second color space. For example, when performing color space transformation, the resolution of the image data in the second color space can be the same as the resolution of the original image in the first color space, and the number of channels of the image data in the second color space can be the same as the number of channels of the original image in the first color space.

[0048] Exemplarily, the configuration information may include a color space conversion formula, which is a conversion formula between the first color space and the second color space, and is used to convert an original image in the first color space into image data in the second color space. Based on this, the original image may be subjected to a color space conversion based on the color space conversion formula to obtain image data in the second color space.

[0049] For example, the first color space of the original image is the RGB color space, but the AI ​​model does not support the RGB color space, and the AI ​​model supports the YUV color space. Therefore, the original image can be transformed in color space to convert the original image in the RGB color space into image data in the YUV color space.

[0050] For example, the first color space of the original image is the YUV color space, but the AI ​​model does not support the YUV color space, and the AI ​​model supports the RGB color space. Therefore, the original image can be transformed in color space to convert the original image in the YUV color space into image data in the RGB color space.

[0051] Case 2: If the first pixel value distribution of the original image is a pixel value distribution supported by the AI ​​model, the pre-processing does not include pixel value mapping. If the first pixel value distribution of the original image is a pixel value distribution not supported by the AI ​​model, the pre-processing includes pixel value mapping. Based on this, if the pre-processing includes pixel value mapping, the second pixel value distribution supported by the AI ​​model is determined, and the original image of the first pixel value distribution (or the image data of the second color space) is pixel-mapped to obtain the image data of the second pixel value distribution.

[0052] For example, when performing pixel value mapping, the resolution of the image data of the second pixel value distribution can be the same as the resolution of the original image of the first pixel value distribution, and the number of channels of the image data of the second pixel value distribution can be the same as the number of channels of the original image of the first pixel value distribution.

[0053] Exemplarily, the configuration information may include a pixel value conversion formula, which is a conversion formula between the first pixel value distribution and the second pixel value distribution, and is used to convert an original image of the first pixel value distribution into image data of the second pixel value distribution. Based on this, the original image may be pixel-value mapped based on the pixel value conversion formula to obtain image data of the second pixel value distribution.

[0054] For example, the first pixel value distribution of the original image is 0~255, and the AI ​​model does not support the pixel value distribution of 0~255, and the AI ​​model supports the second pixel value distribution of 0~1. Therefore, the original image can be pixel value mapped (such as pixel value scaling operation) to scale the pixel values ​​of 0~255 to pixel values ​​of 0~1. The scaling method can use the pixel value conversion formula in the configuration information, and there is no restriction on this process.

[0055] For example, the first pixel value distribution of the original image is 0~255, the AI ​​model does not support the pixel value distribution of 0~255, and the AI ​​model supports the second pixel value distribution of -127~128. Therefore, the original image can be pixel value mapped (such as pixel value shift operation) to shift the pixel values ​​0~255 to pixel values ​​-127~128. The shifting method can use the pixel value conversion formula in the configuration information, and there is no restriction on this process.

[0056] Case 3: If the first arrangement of the original image is an arrangement supported by the AI ​​model, the pre-processing does not include image rearrangement. If the first arrangement of the original image is an arrangement not supported by the AI ​​model, the pre-processing includes image rearrangement. Based on this, if the pre-processing includes image rearrangement, the second arrangement supported by the AI ​​model is determined, and the original image of the first arrangement (or the image data of the second color space, or the image data of the second pixel value distribution) is rearranged to obtain the image data of the second arrangement.

[0057] For example, when rearranging images, the resolution of the image data in the second arrangement may be the same as or different from the resolution of the original image in the first arrangement, and the number of channels of the image data in the second arrangement may be the same as or different from the number of channels of the original image in the first arrangement.

[0058] Exemplarily, the configuration information may include an image rearrangement mode, which is a rearrangement mode between the first arrangement mode and the second arrangement mode, and the image rearrangement mode is used to rearrange the original image of the first arrangement mode into image data of the second arrangement mode. Based on this, the original image of the first arrangement mode may be rearranged based on the image rearrangement mode to obtain image data of the second arrangement mode.

[0059] For example, the first arrangement of the original image is interlaced RGB image data, but the AI ​​model does not support interlaced RGB image data, and the second arrangement supported by the AI ​​model is tiled RGB image data. Therefore, the interlaced RGB image data can be transformed into tiled RGB image data according to the image rearrangement method. For example, RGBRGB...RGB can be rearranged as RR..RGG..GBB..B.

[0060] Case 4: If the AI ​​model does not support parallel processing, the pre-processing does not include image block division. If the AI ​​model supports parallel processing, the pre-processing may include image block division, or the pre-processing may not include image block division. Based on this, if the pre-processing includes image block division, the original image (or image data of the second color space, or image data of the second pixel value distribution, or image data of the second arrangement method) can be divided into blocks to obtain multiple initial image blocks. The multiple initial image blocks may include edge blocks and non-edge blocks. The edge blocks may include the boundaries of the original image, and the non-edge blocks do not include the boundaries of the original image.

[0061] For each initial image block, the initial image block may be filled with pixels to obtain a target image block corresponding to the initial image block, so that multiple target image blocks may be obtained. Alternatively, for each initial image block, the initial image block may be used as a target image block corresponding to the initial image block.

[0062] For example, see Figure 4A As shown in FIG. 1 , it is a schematic diagram of pixel filling for non-edge blocks, where the non-edge blocks are initial image blocks that do not include the boundaries of the original image. For the target image block corresponding to the non-edge block, the target image block may include a tile area and an overlap area, where the tile area is the initial image block itself, and the overlap area is a pixel filling area, and the overlap area is obtained by filling the pixel values ​​of the original image in the overlap area, so that the initial image block and the overlap area can form the target image block.

[0063] For example, see Figure 4B As shown, it is a schematic diagram of pixel filling for edge blocks, where the edge block is the initial image block including the boundary of the original image. For the target image block corresponding to the edge block, the target image block may include a tile area, an overlap area and a padding area. The tile area is the initial image block itself, and the overlap area and the padding area are pixel filling areas. For the overlap area (i.e., the internal area of ​​the original image), the overlap area is obtained by filling based on the pixel values ​​of the original image in the overlap area. For the padding area (i.e., the external area of ​​the original image), the padding area may be filled based on the configured fixed pixel value, or the padding area may be filled based on the boundary pixel value of the original image, or the padding area may be filled in other ways, and there is no restriction on this filling method. In this way, the initial image block, the overlap area and the padding area may form the target image block.

[0064] See also Figure 4A and Figure 4BAs shown, the size of the initial image block is h*w, and the size of the target image block is (h+r)*(w+r). The configuration information includes a division size and a padding size. The division size is used to indicate the size of the initial image block, and the padding size is used to indicate the size of the pixel padding area. Based on this, based on the division size, the original image can be divided into initial image blocks of h*w size. Based on the padding size, the initial image block of h*w size can be padded with pixels to obtain a target image block of (h+r)*(w+r) size.

[0065] For example, for the initial image block or the target image block, the initial image block or the target image block can be a square image block, or a rectangular image block. For each target image block, the resolution of the target image block is smaller than the resolution of the original image.

[0066] Exemplarily, after the division to obtain multiple target image blocks, the multiple target image blocks are used as input data of the AI ​​model, so that the multiple target image blocks can be processed in parallel by the AI ​​model. Based on this set, the original image is divided into relatively small target image blocks through image block division, and each target image block can be processed independently, thereby allowing parallel processing and improving the processing efficiency of the AI ​​model.

[0067] At this point, step 301 is completed, and the first image data is obtained by pre-processing the original image.

[0068] Step 302: The front-end device performs AI processing on the first image data through the AI ​​model to obtain second image data, and the AI ​​processing is used to enhance the high-frequency information in the first image data based on the configuration information.

[0069] Exemplarily, an AI model is configured on a front-end device, and the AI ​​model performs AI processing on the first image data based on the configuration information to obtain the second image data. By performing AI processing on the first image data, the second image data that is more conducive to subsequent encoding compression and compression recovery is obtained. By performing AI processing on the first image data, the high-frequency information in the first image data can be enhanced to obtain the second image data.

[0070] For example, high-frequency information is a region of drastic changes in the first image data, and is a region of rapid local changes, as opposed to low-frequency information (smooth regions in the first image data), such as edges, textures, and details. High-frequency information is critical to image clarity, edge sharpness, and texture details. By enhancing high-frequency information, image clarity can be improved, and it is helpful to identify image edges and image texture features.

[0071] Based on this, in this embodiment, the high-frequency information in the first image data is enhanced through the AI ​​model. In this way, even if the encoding compression process will reduce the image quality of the image data (the high-frequency information in the image data will be lost, resulting in reduced image quality), the image quality effect of the target image can be improved.

[0072] Exemplarily, the AI ​​model includes at least one convolution layer and one activation layer. There is no limitation on the network structure of the AI ​​model, as long as the AI ​​model can enhance the high-frequency information in the first image data.

[0073] Exemplarily, when the AI ​​model performs AI processing on the first image data to obtain the second image data, the AI ​​processing can be real-time processing. Real-time processing means that the input data of the AI ​​model (i.e., the first image data) is streamed, and the AI ​​model immediately performs AI processing on the first image data after receiving the first image data, and outputs the inference result (i.e., the second image data) in a short time. Real-time processing requires that the AI ​​model inference has a low latency and can complete the processing of the input data within a limited time.

[0074] Exemplarily, the first image data may be image data of RGB channels, and the AI ​​model may be used to perform AI processing on at least one of the RGB channels, such as performing AI processing on the RGB channels simultaneously.

[0075] The first image data may be image data of a YUV channel, and the AI ​​model may be used to perform AI processing on at least one of the YUV channels, such as performing AI processing on the YUV channels simultaneously.

[0076] In this embodiment, performing AI processing on the first image data through the AI ​​model means enhancing the high-frequency information in the first image data to obtain the second image data. The AI ​​processing may include but is not limited to at least one of image transformation, image expansion, and semantic transformation. The above are just a few examples of AI processing, and there is no limitation to this, as long as the high-frequency information in the first image data can be enhanced.

[0077] Case 1: If the AI ​​processing includes image transformation, see Figure 5A As shown, it is a schematic diagram of the AI ​​model performing image transformation on the first image data to obtain the second image data. For example, the input data of the AI ​​model is the first image data and configuration information, and the output data of the AI ​​model is the second image data.

[0078] Exemplarily, the first image data is transformed through an AI model to obtain transformed image data, the resolution of the transformed image data may be equal to or less than the resolution of the first image data, and the number of channels of the transformed image data may be equal to or greater than the number of channels of the first image data.

[0079] For example, the first image data includes at least 2 channels of image data, such as a green channel image data a1, a red channel image data a2, and a blue channel image data a3. After the first image data is transformed by the AI ​​model, the transformed image data includes at least one green channel image data b1, that is, the number of channels of the green channel image data increases or remains unchanged, and the resolution of the image data b1 is equal to or less than the resolution of the image data a1. The transformed image data also includes at least one red channel image data b2, that is, the number of channels of the red channel image data increases or remains unchanged, and the resolution of the image data b2 is equal to or less than the resolution of the image data a2. The transformed image data also includes at least one blue channel image data b3, that is, the number of channels of the blue channel image data increases or remains unchanged, and the resolution of the image data b3 is equal to or less than the resolution of the image data a3.

[0080] When performing image transformation on the first image data through the AI ​​model, the configuration information may include an image transformation method, so that the first image data can be transformed based on the image transformation method. When performing image transformation on image data of multiple channels, the image transformation methods of different channels may be the same, or the image transformation methods of different channels may be different. For example, the same image transformation method may be used to perform image transformation on image data a1, image data a2, and image data a3. Alternatively, image transformation method 1 may be used to transform image data a1, image transformation method 2 may be used to transform image data a2, and image transformation method 3 may be used to transform image data a3.

[0081] When the first image data is transformed by the AI ​​model to obtain the transformed image data, the image transformation is used to separate the high-frequency information and the low-frequency information in the first image data. There is no restriction on this image transformation method, as long as the high-frequency information and the low-frequency information in the first image data can be separated.

[0082] For example, image transformation may include but is not limited to at least one of interpolation transformation (such as nearest neighbor interpolation transformation, mean transformation, bilinear interpolation transformation, adaptive interpolation transformation, bicubic interpolation transformation, etc.), wavelet transformation, Gaussian pyramid transformation, Laplace pyramid transformation and sobel filter transformation, and there is no restriction on this image transformation method. On this basis, the configuration information may include an image transformation method, which is used to indicate whether to use the nearest neighbor interpolation transformation to achieve image transformation, or to indicate to use the bilinear interpolation transformation to achieve image transformation, and so on, and there is no restriction on the content of this configuration information.

[0083] For example, if the image transformation method in the configuration information indicates the use of bilinear interpolation transformation, the AI ​​model is used to perform bilinear interpolation transformation on the first image data to obtain the transformed image data. If the image transformation method in the configuration information indicates the use of sobel filter transformation, the AI ​​model is used to perform sobel filter transformation on the first image data to obtain the transformed image data. If the image transformation method in the configuration information indicates the use of bilinear interpolation transformation and Gaussian pyramid transformation (that is, multiple transformations are used at the same time), the AI ​​model is used to perform bilinear interpolation transformation and Gaussian pyramid transformation on the first image data to obtain the transformed image data.

[0084] Exemplarily, after obtaining the transformed image data, the transformed image data can be input into the image processing neural network in the AI ​​model, and the transformed image data can be operated by the image processing neural network to obtain the second image data. For example, the AI ​​model includes an image processing neural network, and the image processing neural network includes at least one convolution layer and one activation layer, and there is no restriction on the network structure of the image processing neural network. When operating the transformed image data by the image processing neural network, convolution operations and activation operations can be performed on the transformed image data, and there is no restriction on the type of operation.

[0085] For example, the input data of the image processing neural network is the transformed image data, and the transformed image data includes at least 2 channels of image data. The output data of the image processing neural network is the second image data, and the second image data may include at least 2 channels of image data. The input data of the AI ​​model is the first image data, and the first image data may include at least 2 channels of image data. On this basis, the resolution of the second image data may be equal to or less than the resolution of the first image data.

[0086] For example, the first image data includes image data a1 of a green channel, image data a2 of a red channel, and image data a3 of a blue channel. The second image data includes image data c1 of a green channel and image data c2 of a red channel, then the resolution of image data c1 is equal to or less than the resolution of image data a1, and the resolution of image data c2 is equal to or less than the resolution of image data a2. Alternatively, the second image data includes image data c1 of a green channel and image data c3 of a blue channel, then the resolution of image data c1 is equal to or less than the resolution of image data a1, and the resolution of image data c3 is equal to or less than the resolution of image data a3. Alternatively, the second image data includes image data c2 of a red channel and image data c3 of a blue channel, then the resolution of image data c2 is equal to or less than the resolution of image data a2, and the resolution of image data c3 is equal to or less than the resolution of image data a3.

[0087] Exemplarily, when the transformed image data is operated by the image processing neural network, the image processing neural network is used to enhance the high-frequency information in the first image data. For example, the image change has separated the high-frequency information and the low-frequency information in the first image data, so that the transformed image data includes the high-frequency information in the first image data and the low-frequency information in the first image data, so that the high-frequency information in the first image data can be enhanced by the image processing neural network, that is, when the transformed image data is operated by the image processing neural network, the high-frequency information in the first image data can be enhanced.

[0088] Case 2: If AI processing includes image expansion, see Figure 5B As shown, it is a schematic diagram of the AI ​​model performing image expansion on the first image data to obtain the second image data. For example, the input data of the AI ​​model is the first image data and configuration information, and the output data of the AI ​​model is the second image data.

[0089] Exemplarily, the first image data may be input into an image processing neural network in an AI model, and the first image data may be operated by the image processing neural network to obtain the image data to be enhanced. For example, the AI ​​model may include an image processing neural network, and the image processing neural network may include at least one convolution layer and one activation layer. When the first image data is operated by the image processing neural network, a convolution operation and an activation operation may be performed on the first image data, and there is no restriction on the type of operation.

[0090] The image data to be enhanced includes at least 2 channels of image data, and the first image data includes at least 2 channels of image data. The resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data.

[0091] For example, the input data of the image processing neural network is the first image data, and the output data of the image processing neural network is the image data to be enhanced. The first image data includes image data a1 of the green channel, image data a2 of the red channel, and image data a3 of the blue channel. The image data to be enhanced includes image data d1 of the green channel and image data d2 of the red channel, the resolution of image data d1 is equal to or less than the resolution of image data a1, and the resolution of image data d2 is equal to or less than the resolution of image data a2. Alternatively, the image data to be enhanced includes image data d1 of the green channel and image data d3 of the blue channel, the resolution of image data d1 is equal to or less than the resolution of image data a1, and the resolution of image data d3 is equal to or less than the resolution of image data a3. Alternatively, the image data to be enhanced includes image data d2 of the red channel and image data d3 of the blue channel, the resolution of image data d2 is equal to or less than the resolution of image data a2, and the resolution of image data d3 is equal to or less than the resolution of image data a3.

[0092] Exemplarily, the first image data is expanded through an AI model to obtain expanded image data, the resolution of the expanded image data may be equal to or less than the resolution of the first image data, and the number of channels of the expanded image data may be equal to or less than the number of channels of the first image data.

[0093] For example, the first image data includes at least two channels of image data, such as green channel image data a1, red channel image data a2, and blue channel image data a3. The expanded image data includes at least one channel of image data, such as green channel image data e1 (the resolution of image data e1 is equal to or less than the resolution of image data a1), or red channel image data e2 (the resolution of image data e2 is equal to or less than the resolution of image data a2), or blue channel image data e3 (the resolution of image data e3 is equal to or less than the resolution of image data a3), or feature channel image data e4 (the resolution of image data e4 is equal to or less than the resolution of any image data).

[0094] When the first image data is expanded through the AI ​​model, the configuration information may include an image expansion method, so that the first image data can be expanded based on the image expansion method.

[0095] When the first image data is expanded by the AI ​​model to obtain the expanded image data, the image expansion is used to enhance the high-frequency information in the first image data. In this way, the expanded image data is used to enhance the high-frequency information in the enhanced image data. In other words, the expanded image data is enhanced data for the high-frequency information in the first image data. In this embodiment, there is no restriction on this image expansion method, as long as the high-frequency information in the first image data can be enhanced.

[0096] For example, image expansion may include but is not limited to at least one of guided filtering, image edge enhancement, image contrast enhancement, and histogram equalization, and there is no restriction on the image expansion method. On this basis, the configuration information may include an image expansion method, which is used to indicate that guided filtering is used to achieve image expansion, or that image edge enhancement is used to achieve image expansion, and so on.

[0097] For example, if the image expansion method in the configuration information is used to indicate the use of histogram equalization, the first image data can be subjected to histogram equalization through the AI ​​model to obtain the expanded image data. If the image expansion method in the configuration information is used to indicate the use of image edge enhancement, the first image data can be subjected to image edge enhancement through the AI ​​model to obtain the expanded image data, and so on.

[0098] Taking the example of obtaining extended image data by guided filtering of the first image data through the AI ​​model, guided filtering is performed based on the image data a1 of the green channel and the image data a2 of the red channel to obtain extended image data of the blue channel. This extended image data is used to enhance the high-frequency information in the image data to be enhanced in the blue channel, thereby introducing the information of the green channel and the red channel into the blue channel to enhance the high-frequency information of the blue channel. Guided filtering is performed based on the image data a1 of the green channel and the image data a3 of the blue channel to obtain extended image data of the red channel. This extended image data is used to enhance the high-frequency information in the image data to be enhanced in the red channel, and so on.

[0099] Exemplarily, after obtaining the image data to be enhanced and the image data after expansion, the image data to be enhanced and the image data after expansion can be fused to obtain the second image data. For example, the configuration information may include an image fusion method, so that the image data to be enhanced and the image data after expansion can be fused based on the image fusion method. For example, if there is image data to be enhanced of a blue channel and image data after expansion of a blue channel, the image data to be enhanced and the image data after expansion can be fused based on the image fusion method to obtain the second image data of the blue channel, and so on.

[0100] Obviously, since the expanded image data of the blue channel can enhance the high-frequency information in the image data to be enhanced of the blue channel, the high-frequency information can be enhanced by fusing the image data to be enhanced and the expanded image data, that is, the second image data is the enhanced image data.

[0101] For example, image fusion may include but is not limited to at least one of weighted fusion, splicing fusion, multiplication fusion, attention fusion, addition fusion, and convolution fusion, and there is no limitation on this. On this basis, the configuration information may include an image fusion mode, which is used to indicate that image fusion is implemented by weighted fusion, or that image fusion is implemented by attention fusion, and so on.

[0102] For example, if the image fusion method in the configuration information is used to indicate the use of weighted fusion, the enhanced image data and the expanded image data can be weighted fused to obtain the second image data. If the image fusion method in the configuration information is used to indicate the use of attention fusion, the enhanced image data and the expanded image data can be attention fused to obtain the second image data, and so on.

[0103] Exemplarily, after the second image data is obtained, the second image data includes at least two channels of image data, and the resolution of the second image data may be equal to or less than the resolution of the first image data.

[0104] Case 3: If AI processing includes semantic transformation, see Figure 5C As shown, it is a schematic diagram of the AI ​​model performing semantic transformation on the first image data to obtain the second image data. For example, the input data of the AI ​​model can be the first image data, and the output data of the AI ​​model can be the second image data.

[0105] Exemplarily, the first image data may be input into an image processing neural network in an AI model, and the first image data may be operated by the image processing neural network to obtain the image data to be enhanced. For example, the AI ​​model may include an image processing neural network, and the image processing neural network may include at least one convolution layer and one activation layer. When the first image data is operated by the image processing neural network, a convolution operation and an activation operation may be performed on the first image data, and there is no restriction on the type of operation.

[0106] The image data to be enhanced includes at least 2 channels of image data, and the first image data includes at least 2 channels of image data. The resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data.

[0107] Exemplarily, the first image data may be input into a semantic processing neural network in an AI model, and the first image data may be operated by the semantic processing neural network to obtain structured semantic information. For example, the AI ​​model may include a semantic processing neural network, and the semantic processing neural network includes at least one convolution layer and one activation layer, and there is no restriction on this network structure. When the first image data is operated by the semantic processing neural network, convolution operations and activation operations may be performed on the first image data.

[0108] For example, a semantic processing neural network can be pre-trained, the input data of the semantic processing neural network is the first image data, and the output data of the semantic processing neural network is structured semantic information. Therefore, after the first image data is input into the semantic processing neural network, the semantic processing neural network can process based on the first image data to obtain the structured semantic information corresponding to the first image data.

[0109] For example, the structured semantic information is used to describe the semantics of the image data to be enhanced (or the first image data), and the high-frequency information of the image data to be enhanced can be restored through the structured semantic information.

[0110] Structured semantic information is image semantic information, which is the content and meaning contained in the image data to be enhanced (or the first image data), such as the type of object, the action or state of the object, the theme or situation expressed by the image, and the subject in the image (such as people, animals, buildings or natural landscapes). Structured semantic information can be expressed through language, including natural language and symbolic language (mathematical language).

[0111] Exemplarily, after obtaining the image data to be enhanced and the structured semantic information, second image data can be generated based on the image data to be enhanced and the structured semantic information, that is, the second image data can include the image data to be enhanced and the structured semantic information, and the structured semantic information can be placed in front of the image data to be enhanced, and the structured semantic information can also be placed behind the image data to be enhanced.

[0112] Case 4: If AI processing includes image transformation, image expansion, and semantic transformation at the same time, see Figure 5D As shown, it is a schematic diagram of an AI model performing AI processing on first image data to obtain second image data.

[0113] The first image data is transformed by the AI ​​model to obtain transformed image data, and the image transformation is used to separate high-frequency information and low-frequency information in the first image data. The transformed image data is input into the image processing neural network in the AI ​​model, and the transformed image data is operated by the image processing neural network to obtain the first image data to be enhanced, and the image processing neural network is used to enhance the high-frequency information in the first image data. This process can be referred to in Case 1, and will not be repeated here.

[0114] The first image data is expanded by the AI ​​model to obtain expanded image data, and the expanded image data is used to enhance the high-frequency information in the first image data to be enhanced. The first image data to be enhanced (i.e., the image data output by the image processing neural network) and the expanded image data are fused to obtain the second image data to be enhanced. The process can be referred to in case 2 and will not be described here.

[0115] The first image data is input into the semantic processing neural network in the AI ​​model. The first image data is operated by the semantic processing neural network to obtain structured semantic information. The structured semantic information is used to describe the semantics of the second image data to be enhanced. The process can be referred to as Case 3 and will not be repeated here.

[0116] The second image data is generated based on the second image data to be enhanced and the structured semantic information, that is, the second image data may include the second image data to be enhanced and the structured semantic information.

[0117] At this point, step 302 is completed, and the second image data is obtained by performing AI processing on the first image data.

[0118] Step 303: The front-end device performs post-processing on the second image data.

[0119] Exemplarily, post-processing of image data refers to further processing of the image data after the image data is encoded, transmitted and decoded, such as decompression distortion, denoising, super-resolution, etc., to improve the subjective quality of the image. In this embodiment, post-processing the second image data refers to storing the second image data in a designated storage medium (such as a memory or a hard disk). For example, the second image data is a serial image data stream, and the second image data is rearranged according to a certain rule and then stored.

[0120] Exemplarily, during the storage process of the second image data, the second image data (i.e., the output data of the AI ​​model) is a serial image data stream. The serial image data stream can be rearranged according to certain rules and then stored in a designated storage medium (such as a memory or a hard disk, etc.) to facilitate access and reading in subsequent processes. In this way, the stored data is image data arranged according to certain rules.

[0121] Exemplarily, in the pre-processing process, if the image blocks are not divided, then the first image data is a complete image, and when the first image data is processed by AI to obtain the second image data, the second image data is also a complete image. Based on this, when storing the second image data in a designated storage medium, each row of data can be stored in sequence, for example, all the data of the first row is stored first, then all the data of the second row is stored, and so on, until all the data of the last row is stored.

[0122] Exemplarily, in the pre-processing process, if image block division is performed, then the first image data includes K target image blocks, K is a positive integer greater than 1, and when the first image data is subjected to AI processing to obtain the second image data, the second image data includes K to-be-stored image blocks corresponding to the K target image blocks. For example, AI processing is performed on the K target image blocks in parallel, and the processing process is shown in step 302. When AI processing is performed on each target image block, the image block to be stored corresponding to the target image block is obtained, so that all the image blocks to be stored constitute the second image data, that is, the second image data includes K to-be-stored image blocks.

[0123] When storing the second image data in a designated storage medium, for each image block to be stored, the image block to be stored corresponds to a storage start address, and the interval between the storage start addresses of two adjacent image blocks to be stored is the width of the image block to be stored. In addition, the second image data corresponds to a storage span, and the storage span is greater than or equal to the sum of the widths of the K image blocks to be stored. On this basis, for each image block to be stored, when storing the i-th row of data of the image block to be stored in the designated storage medium, the i-th row of data is stored sequentially starting from the target storage position; wherein the target storage position can be determined by the following formula: p=w+ stride*(i-1); wherein p represents the target storage position, w represents the storage start address corresponding to the image block to be stored, stride represents the storage span, i can represent the number of rows of the image block to be stored, and i is a positive integer.

[0124] See also Fig. 6AAs shown, it is a schematic diagram of storing multiple image blocks to be stored (recorded as image blocks A, B, ..., N to be stored) in a designated storage medium. First, the serial data of the image block A to be stored is stored, such as the first row of data (A1A1A1), the second row of data (A2A2A2), ..., the hth row of data (AhAhAh) of the image block A to be stored in sequence. Then, the serial data of the image block B to be stored is stored, such as the first row of data (B1B1B1), the second row of data (B2B2B2), ..., the hth row of data (BhBhBh) of the image block B to be stored in sequence. And so on, finally, the serial data of the image block N to be stored is stored, such as the first row of data (N1N1N1), the second row of data (N2N2N2), ..., the hth row of data (NhNhNh) of the image block N to be stored in sequence. In the encoding and compression process, the first row of data (A1A1A1) of the image block A to be stored is read, the first row of data (B1B1B1) of the image block B to be stored is read, ..., the first row of data (N1N1N1) of the image block N to be stored is read, so that the complete first row of data can be obtained. Then, the second row of data (A2A2A2) of the image block A to be stored is read, the second row of data (B2B2B2) of the image block B to be stored is read, ..., the second row of data (N2N2N2) of the image block N to be stored is read, so that the complete second row of data can be obtained. Similarly, the hth row of data (AhAhAh) of the image block A to be stored is read, the hth row of data (BhBhBh) of the image block B to be stored is read, ..., the hth row of data (NhNhNh) of the image block N to be stored is read, so that the complete hth row of data can be obtained.

[0125] Take the case where the N image blocks to be stored are 2 image blocks to be stored (referred to as tile A and tile B) as an example. Figure 6B As shown, it is a schematic diagram of the storage method of two tiles. For example, the configuration information includes the storage start address of each image block to be stored, such as the storage start address addA of tile A and the storage start address addB of tile B. The interval between the storage start addresses of two adjacent image blocks to be stored can be the width of the image block to be stored (that is, the previous image block to be stored), such as the interval between addB and addA can be the width wA of tile A, if there is tile C after tile B, then the interval between the storage start addresses addC and addB of tile C can be the width wB of tile B, and so on. The configuration information also includes the storage span stride, and the storage span stride is greater than or equal to the sum of the widths of all image blocks to be stored, such as the storage span stride is greater than or equal to the sum of the width wA of tile A and the width wB of tile B, that is, stride≥wA+wB.

[0126] See also Figure 4A and Figure 4B As shown, the target image block includes a tile area and a pixel filling area. Therefore, the image block to be stored also includes a tile area and a pixel filling area. When storing multiple image blocks to be stored, only the tile area of ​​the image block to be stored is stored, and there is no pixel filling area of ​​the image block to be stored.

[0127] The target image block A and the target image block B can be processed in parallel. The processing of the target image block A is recorded as AI model processing 1, and the serial data A1A1A1A2A2A2…AhAhAh of tile A is obtained, and the serial data of tile A is stored. For example, the target storage position of each row is determined, such as p=w+ stride*(i-1), p represents the target storage position, w represents the storage start address addA of tile A, and i represents the number of rows. Starting from the target storage position (addA) of the first row, the first row data A1A1A1 of the tile area (excluding the pixel filling area) of tile A is stored, starting from the target storage position (addA+ stride) of the second row, the second row data A2A2A2 of the tile area of ​​tile A is stored, starting from the target storage position (addA+stride*2) of the third row, the third row data A2A2A2 of the tile area of ​​tile A is stored, and so on.

[0128] In addition, the processing of the target image block B is recorded as AI model processing 2, and the serial data B1B1B1B2B2B2…BhBhBh of tile B can be obtained, and the serial data of tile B is stored. For example, the target storage position of each row is determined, such as p=w+ stride*(i-1), where w represents the storage start address addB of tile B. Starting from the target storage position (addB) of the first row, the first row data B1B1B1 of the tile area (excluding the pixel filling area) of tile B is stored, and starting from the target storage position (addB+ stride) of the second row, the second row data B2B2B2 of the tile area of ​​tile B is stored, and so on.

[0129] Through the above storage method, the two image blocks divided in the pre-processing can be re-stitched into the image before division (the size and shape of the image are restored to before division, but the image data has been processed by the AI ​​model). Figure 6B Only the storage method of two image blocks is illustrated in the figure, but by configuring the storage start address and storage span of each image block to be stored, it can be extended to the case of storing multiple image blocks to be stored, which will not be described in detail here.

[0130] In a possible implementation manner, when the second image data is post-processed, an image display may also be performed, see Figure 6C As shown, it is a schematic diagram of image display. In order to realize image display, the configuration information may include information for changing the scale (Scale) of image pixel values ​​and information for translating image pixel values ​​(Offset). In the image display process, the input data (i.e., the output data of the AI ​​model) includes at least one type of image data, and the output data is image data for encoding and compression (such as 8-bit images). The processing method is constrained by the configuration information, and different types of image data can use different processing methods. The configuration information includes information constraining image display processing. For example, the output data range of the convolutional neural network model is -0.5~0.5, and the configuration information contains two parameters, Scale=255, Offset=128.

[0131] In a possible implementation, for steps 203 to 205, the front-end device can read the stored image data from a specified storage medium, for example, see Fig. 6A As shown, the first row of data of the image block A to be stored is read (A1A1A1), the first row of data of the image block B to be stored is read (B1B1B1), …, the first row of data of the image block N to be stored is read (N1N1N1), and then, the second row of data of the image block A to be stored is read (A2A2A2), the second row of data of the image block B to be stored is read (B2B2B2), …, the second row of data of the image block N to be stored is read (N2N2N2), and so on.

[0132] The front-end device can encode and compress the stored image data to obtain a target code stream. The front-end device can send the target code stream to the back-end device. After receiving the target code stream, the back-end device can decode the target code stream to obtain decoded data, that is, decoded image data.

[0133] In a possible implementation, with respect to step 206, the backend device may compress and restore the decoded data to obtain a target image. The compression and restoration process may be understood as the inverse process of pre-compression.

[0134] For example, when the front-end device performs AI processing on the first image data through the AI ​​model to obtain the second image data, if the AI ​​processing includes image transformation (corresponding to case 1), then after receiving the decoded data, the back-end device can perform image inverse transformation on the decoded data to obtain the target image. For example, if the front-end device performs Gaussian pyramid transformation on the first image data through the AI ​​model, the back-end device performs Gaussian pyramid inverse transformation on the decoded data to the target image. If the front-end device performs interpolation transformation on the first image data through the AI ​​model, the back-end device performs inverse interpolation transformation on the decoded data to the target image.

[0135] For example, when the front-end device performs AI processing on the first image data through the AI ​​model to obtain the second image data, if the AI ​​processing includes image expansion (corresponding to case 2), then after receiving the decoded data, the back-end device can use the decoded data as the target image and no longer restore the decoded data.

[0136] For example, when the front-end device performs AI processing on the first image data through the AI ​​model to obtain the second image data, if the AI ​​processing includes semantic transformation (corresponding to case 2), then after receiving the decoded data, the back-end device can process the decoded data based on the AI ​​model to obtain the target image.

[0137] The AI ​​model may include an image processing neural network and a semantic restoration neural network, and the decoded data may include decoded image data and structured semantic information. Based on this, the decoded image data may be input into the image processing neural network in the AI ​​model, and the decoded image data may be operated on by the image processing neural network to obtain the image data to be enhanced. For example, the image processing neural network includes at least one convolution layer and one activation layer, and may perform convolution and activation operations on the decoded image data.

[0138] In addition, the structured semantic information can be input into the semantic restoration neural network in the AI ​​model, and the structured semantic information can be operated by the semantic restoration neural network to obtain semantically enhanced data. For example, the semantic restoration neural network can include at least one convolution layer and one activation layer, and there is no restriction on the network structure. In this way, convolution operations and activation operations can be performed on the structured semantic information.

[0139] The semantic restoration neural network can be pre-trained. The input data of the semantic restoration neural network is structured semantic information. The output data of the semantic restoration neural network is semantic enhancement data. The semantic enhancement data can be image data. The dimension of the semantic enhancement data is the same as the dimension of the image data to be enhanced. For example, if the image data to be enhanced is an M*N image, the semantic enhancement data can also be an M*N image. Based on this, after the structured semantic information is input into the semantic restoration neural network, the semantic enhancement data can be obtained.

[0140] After obtaining the image data to be enhanced and the semantic enhancement data, the image data to be enhanced and the semantic enhancement data can be fused to obtain the target image, such as by using at least one fusion method selected from weighted fusion, splicing fusion, multiplication fusion, attention fusion, addition fusion, and convolution fusion to fuse the image data to be enhanced and the semantic enhancement data. For example, by fusing the image data to be enhanced and the semantic enhancement data, the structured semantic information can be fused into the image data to be enhanced. Since the structured semantic information is used to describe the semantics of the image data to be enhanced, the high-frequency information of the image data to be enhanced can be restored through the structured semantic information, and the high-frequency information of the target image can be enhanced.

[0141] For example, when the front-end device performs AI processing on the first image data through the AI ​​model to obtain the second image data, if the AI ​​processing includes image transformation, image expansion and semantic transformation (corresponding to case 4) at the same time, then, after receiving the decoded data, the back-end device can perform image inverse transformation on the decoded data to obtain the first image data to be enhanced. The decoded image data can be input into the image processing neural network in the AI ​​model, and the decoded image data can be operated by the image processing neural network to obtain the second image data to be enhanced. The structured semantic information can be input into the semantic recovery neural network in the AI ​​model, and the structured semantic information can be operated by the semantic recovery neural network to obtain the semantic enhancement data. Then, the first image data to be enhanced and the second image data to be enhanced are fused to obtain the target image data to be enhanced, and the target image data to be enhanced and the semantic enhancement data are fused to obtain the target image.

[0142] In a possible implementation, with respect to step 207, after obtaining the target image, the backend device may directly store or display the target image. Alternatively, the backend device may perform post-processing (such as de-distortion, denoising, super-resolution, etc.) on the target image, and store or display the target image after post-processing.

[0143] It can be seen from the above technical solutions that in the embodiment of the present application, since the high-frequency information in the image data is enhanced by the AI ​​model, even if the encoding compression process will reduce the image quality of the image data (the high-frequency information in the image data will be lost, resulting in reduced image quality), when the back-end device decodes the target code stream to obtain the target image, the image quality of the target image can be equivalent to the image quality of the original image, thereby improving the image quality effect of the stored or displayed target image and improving the image quality. By processing the real-time data stream (i.e., video stream) with the AI ​​model deployed on the front-end device, the image quality of the real-time data stream can be improved, helping the video image encoding and decoding system to obtain higher video image quality.

[0144] Based on the same application concept as the above method, an image processing device is proposed in the embodiment of the present application and applied to a front-end device, see Figure 7 FIG. 1 is a schematic diagram of the structure of the device, wherein the device comprises: A pre-processing module 71 is used to perform pre-processing on the acquired original image to obtain first image data, wherein the pre-processing is used to convert the original image into first image data adapted to an artificial intelligence AI model; An AI processing module 72, configured to perform AI processing on the first image data through the AI ​​model to obtain second image data, wherein the AI ​​processing is used to enhance high-frequency information in the first image data; A storage module 73, used for storing the second image data in a designated storage medium; The sending module 74 is used to read the stored image data from the specified storage medium, encode and compress the stored image data to obtain a target code stream, and send the target code stream to the back-end device, which obtains the target image based on the target code stream and stores or displays the target image.

[0145] Exemplarily, the pre-processing includes at least one of color space transformation, pixel value mapping, image block division and image rearrangement; when the pre-processing module 71 performs pre-processing on the acquired original image to obtain the first image data, it is specifically used to: if the pre-processing includes color space transformation, and the first color space of the original image is a color space not supported by the AI ​​model, then the color space transformation is performed on the original image to obtain image data of the second color space, and the second color space is a color space supported by the AI ​​model; if the pre-processing includes pixel value mapping, and the first pixel value distribution of the original image is a pixel value distribution not supported by the AI ​​model, then the pixel value mapping is performed on the original image to obtain image data of the second pixel value distribution, and the second pixel value distribution is supported by the AI ​​model. Pixel value distribution; if the pre-processing includes image block division, the original image is divided into blocks to obtain multiple initial image blocks, and the multiple initial image blocks include edge blocks and non-edge blocks, the edge blocks include the boundary of the original image, and the non-edge blocks do not include the boundary of the original image; for each initial image block, the initial image block is pixel-filled to obtain a target image block; wherein, multiple target image blocks are used as input data of the AI ​​model, and the multiple target image blocks are processed in parallel by the AI ​​model; if the pre-processing includes image rearrangement, and the first arrangement of the original image is an arrangement not supported by the AI ​​model, the original image is rearranged to obtain image data of a second arrangement, and the second arrangement is an arrangement supported by the AI ​​model.

[0146] Exemplarily, the AI ​​processing includes at least one of image transformation, image expansion and semantic transformation, and the AI ​​processing module 72 performs AI processing on the first image data through the AI ​​model to obtain the second image data, and is specifically used for: if the AI ​​processing includes image transformation, the first image data is transformed through the AI ​​model to obtain transformed image data, the resolution of the transformed image data is equal to or less than the resolution of the first image data, and the number of channels of the transformed image data is equal to or greater than the number of channels of the first image data; the image transformation is used to separate high-frequency information and low-frequency information in the first image data; the image transformation includes at least one of interpolation transformation, wavelet transformation, Gaussian pyramid transformation, Laplace pyramid transformation and sobel filter transformation; the transformed image data is input into the image processing neural network in the AI ​​model, and the transformed image data is operated by the image processing neural network to obtain the second image data; the resolution of the second image data is equal to or less than the resolution of the first image data, and the image processing neural network is used to enhance the high-frequency information in the first image data.

[0147] Exemplarily, when the AI ​​processing module 72 performs AI processing on the first image data through the AI ​​model to obtain the second image data, it is specifically used to: if the AI ​​processing includes image expansion, the first image data is input into the image processing neural network in the AI ​​model, and the first image data is operated by the image processing neural network to obtain the image data to be enhanced, the resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data; the first image data is image expanded by the AI ​​model to obtain the expanded image data corresponding to the image data to be enhanced, the resolution of the expanded image data is equal to or less than the resolution of the first image data, and the number of channels of the expanded image data is equal to or less than the number of channels of the first image data; wherein the expanded image data is used to enhance the high-frequency information in the image data to be enhanced; wherein the image expansion includes at least one of guided filtering, image edge enhancement, image contrast enhancement, and histogram equalization; the image data to be enhanced and the expanded image data are merged to obtain the second image data, and the resolution of the second image data is equal to or less than the resolution of the first image data.

[0148] Exemplarily, when the AI ​​processing module 72 performs AI processing on the first image data through the AI ​​model to obtain the second image data, it is specifically used to: if the AI ​​processing includes semantic transformation, input the first image data to the image processing neural network in the AI ​​model, operate the first image data through the image processing neural network to obtain the image data to be enhanced, the resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data; input the first image data to the semantic processing neural network in the AI ​​model, operate the first image data through the semantic processing neural network to obtain structured semantic information, and the structured semantic information is used to describe the semantics of the image data to be enhanced; generate the second image data based on the image data to be enhanced and the structured semantic information.

[0149] Exemplarily, when the AI ​​processing module 72 performs AI processing on the first image data through the AI ​​model to obtain the second image data, it is specifically used for: if the AI ​​processing includes image transformation, image expansion and semantic transformation, performing image transformation on the first image data through the AI ​​model to obtain transformed image data, and the image transformation is used to separate high-frequency information and low-frequency information in the first image data; inputting the transformed image data into the image processing neural network in the AI ​​model, operating the transformed image data through the image processing neural network to obtain first image data to be enhanced, and the image processing neural network is used to enhance the high-frequency information in the first image data; performing image expansion on the first image data through the AI ​​model to obtain expanded image data, and the expanded image data is used to enhance the high-frequency information in the first image data to be enhanced, and fusing the first image data to be enhanced and the expanded image data to obtain the second image data to be enhanced; inputting the first image data into the semantic processing neural network in the AI ​​model, operating the first image data through the semantic processing neural network to obtain structured semantic information, and the structured semantic information is used to describe the semantics of the second image data to be enhanced; generating the second image data based on the second image data to be enhanced and the structured semantic information.

[0150] Exemplarily, if the first image data includes K target image blocks, the second image data includes K to-be-stored image blocks corresponding to the K target image blocks, where K is a positive integer greater than 1; for each to-be-stored image block, the to-be-stored image block corresponds to a storage start address, and the interval between the storage start addresses of two adjacent to-be-stored image blocks is the width of the to-be-stored image block; the second image data corresponds to a storage span, and the storage span is greater than or equal to the sum of the widths of the K to-be-stored image blocks; When storing the second image data in a designated storage medium, the storage module 73 is specifically used to: for each image block to be stored, when storing the i-th row of data of the image block to be stored in the designated storage medium, store the i-th row of data sequentially starting from the target storage position; wherein the target storage position is determined by the following formula: p=w+ stride*(i-1); p represents the target storage position, w represents the storage start address corresponding to the image block to be stored, stride represents the storage span, i represents the number of rows of the image block to be stored, and i is a positive integer.

[0151] Based on the same application concept as the above method, an electronic device (such as a front-end device) is proposed in the embodiment of the present application, see Figure 8 As shown, it includes: a processor 81 and a machine-readable storage medium 82, the machine-readable storage medium 82 stores machine-executable instructions that can be executed by the processor 81; the processor 81 is used to execute the machine-executable instructions to implement the image processing method disclosed in the above example of this application.

[0152] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the image processing method disclosed in the above example of the present application can be implemented.

[0153] The above-mentioned machine-readable storage medium may be any electronic, magnetic, optical or other physical storage device, which may contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof.

[0154] Based on the same application concept as the above method, an embodiment of the present application further provides a computer program product, which may include a computer program. When the computer program is executed by a processor, it implements the image processing method disclosed in the above example of the present application.

[0155] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of the claims of the present application.

Claims

1. An image processing method, characterized in that: Applied to a front-end device, the method comprises: Pre-processing the acquired original image to obtain first image data, wherein the pre-processing is used to convert the original image into first image data adapted to an artificial intelligence AI model; Performing AI processing on the first image data by the AI ​​model to obtain second image data, wherein the AI ​​processing is used to enhance high-frequency information in the first image data; storing the second image data in a designated storage medium; The stored image data is read from the designated storage medium, the stored image data is encoded and compressed to obtain a target code stream, the target code stream is sent to a back-end device, and the back-end device obtains a target image based on the target code stream, and stores or displays the target image.

2. The method according to claim 1, characterized in that The pre-processing includes at least one of color space conversion, pixel value mapping, image block division and image rearrangement; The step of pre-processing the acquired original image to obtain first image data includes: If the pre-processing includes color space transformation, and the first color space of the original image is a color space not supported by the AI ​​model, performing color space transformation on the original image to obtain image data in a second color space, where the second color space is a color space supported by the AI ​​model; If the pre-processing includes pixel value mapping, and the first pixel value distribution of the original image is a pixel value distribution not supported by the AI ​​model, pixel value mapping is performed on the original image to obtain image data of a second pixel value distribution, and the second pixel value distribution is a pixel value distribution supported by the AI ​​model; If the pre-processing includes image block division, the original image is divided into blocks to obtain a plurality of initial image blocks, wherein the plurality of initial image blocks include edge blocks and non-edge blocks, wherein the edge blocks include the boundary of the original image, and the non-edge blocks do not include the boundary of the original image; for each initial image block, the initial image block is pixel-filled to obtain a target image block; wherein the plurality of target image blocks are used as input data of the AI ​​model, and the plurality of target image blocks are processed in parallel by the AI ​​model; If the pre-processing includes image rearrangement, and the first arrangement of the original image is an arrangement not supported by the AI ​​model, then the original image is rearranged to obtain image data of a second arrangement, and the second arrangement is an arrangement supported by the AI ​​model.

3. The method according to claim 1, characterized in that The AI ​​processing includes at least one of image transformation, image expansion and semantic transformation. If the AI ​​processing includes image transformation, performing AI processing on the first image data by using the AI ​​model to obtain second image data includes: Performing an image transformation on the first image data through the AI ​​model to obtain transformed image data, wherein the resolution of the transformed image data is equal to or less than the resolution of the first image data, and the number of channels of the transformed image data is equal to or greater than the number of channels of the first image data; wherein the image transformation is used to separate high-frequency information and low-frequency information in the first image data; wherein the image transformation includes at least one of an interpolation transform, a wavelet transform, a Gaussian pyramid transform, a Laplace pyramid transform, and a sobel filter transform; The transformed image data is input into the image processing neural network in the AI ​​model, and the transformed image data is operated by the image processing neural network to obtain second image data; wherein the resolution of the second image data is equal to or less than the resolution of the first image data, and the image processing neural network is used to enhance the high-frequency information in the first image data.

4. The method according to claim 1, characterized in that: The AI ​​processing includes at least one of image transformation, image expansion and semantic transformation. If the AI ​​processing includes image expansion, performing AI processing on the first image data by using the AI ​​model to obtain second image data includes: Inputting the first image data to the image processing neural network in the AI ​​model, and operating the first image data through the image processing neural network to obtain image data to be enhanced, wherein the resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data; Performing image expansion on the first image data through the AI ​​model to obtain expanded image data corresponding to the image data to be enhanced, wherein the resolution of the expanded image data is equal to or less than the resolution of the first image data, and the number of channels of the expanded image data is equal to or less than the number of channels of the first image data; wherein the expanded image data is used to enhance high-frequency information in the image data to be enhanced; wherein the image expansion includes at least one of guided filtering, image edge enhancement, image contrast enhancement, and histogram equalization; The image data to be enhanced and the expanded image data are fused to obtain second image data, wherein a resolution of the second image data is equal to or less than a resolution of the first image data.

5. The method according to claim 1, characterized in that The AI ​​processing includes at least one of image transformation, image expansion and semantic transformation. If the AI ​​processing includes semantic transformation, performing AI processing on the first image data by using the AI ​​model to obtain the second image data includes: Inputting the first image data to the image processing neural network in the AI ​​model, and operating the first image data through the image processing neural network to obtain image data to be enhanced, wherein the resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data; Inputting the first image data into a semantic processing neural network in the AI ​​model, and operating the first image data through the semantic processing neural network to obtain structured semantic information, wherein the structured semantic information is used to describe the semantics of the image data to be enhanced; Second image data is generated based on the image data to be enhanced and the structured semantic information.

6. The method according to claim 1, characterized in that If the AI ​​processing includes image transformation, image expansion and semantic transformation, performing AI processing on the first image data by using the AI ​​model to obtain second image data includes: Performing image transformation on the first image data through the AI ​​model to obtain transformed image data, wherein the image transformation is used to separate high-frequency information and low-frequency information in the first image data; inputting the transformed image data into the image processing neural network in the AI ​​model, and operating the transformed image data through the image processing neural network to obtain first image data to be enhanced, wherein the image processing neural network is used to enhance the high-frequency information in the first image data; Performing image expansion on the first image data through the AI ​​model to obtain expanded image data, wherein the expanded image data is used to enhance high-frequency information in the first image data to be enhanced, and fusing the first image data to be enhanced and the expanded image data to obtain second image data to be enhanced; Inputting the first image data into a semantic processing neural network in the AI ​​model, and operating the first image data through the semantic processing neural network to obtain structured semantic information, and the structured semantic information is used to describe the semantics of the second image data to be enhanced; Second image data is generated based on the second image data to be enhanced and the structured semantic information.

7. The method according to claim 1, characterized in that If the first image data includes K target image blocks, the second image data includes K to-be-stored image blocks corresponding to the K target image blocks, where K is a positive integer greater than 1; for each to-be-stored image block, the to-be-stored image block corresponds to a storage start address, and the interval between the storage start addresses of two adjacent to-be-stored image blocks is the width of the to-be-stored image block; the second image data corresponds to a storage span, and the storage span is greater than or equal to the sum of the widths of the K to-be-stored image blocks; The storing the second image data in a designated storage medium comprises: For each image block to be stored, when storing the i-th row of data of the image block to be stored in the designated storage medium, the i-th row of data is stored sequentially starting from the target storage position; Among them, the target storage location is determined by the following formula: p=w+stride*(i-1); Among them, p represents the target storage position, w represents the storage start address corresponding to the image block to be stored, stride represents the storage span, i represents the number of rows of the image block to be stored, and i is a positive integer.

8. An image processing device, characterized in that: Applied to front-end equipment, the device comprises: A pre-processing module, used to perform pre-processing on the acquired original image to obtain first image data, wherein the pre-processing is used to convert the original image into first image data adapted to an artificial intelligence AI model; An AI processing module, configured to perform AI processing on the first image data through the AI ​​model to obtain second image data, wherein the AI ​​processing is used to enhance high-frequency information in the first image data; A storage module, used for storing the second image data in a designated storage medium; The sending module is used to read the stored image data from the specified storage medium, encode and compress the stored image data to obtain a target code stream, and send the target code stream to the back-end device, which obtains the target image based on the target code stream and stores or displays the target image.

9. The device according to claim 8, characterized in that in, The pre-processing includes at least one of color space transformation, pixel value mapping, image block division and image rearrangement; when the pre-processing module performs pre-processing on the acquired original image to obtain the first image data, it is specifically used to: if the pre-processing includes color space transformation, and the first color space of the original image is a color space not supported by the AI ​​model, then the color space transformation is performed on the original image to obtain image data of the second color space, and the second color space is a color space supported by the AI ​​model; if the pre-processing includes pixel value mapping, and the first pixel value distribution of the original image is a pixel value distribution not supported by the AI ​​model, then the pixel value mapping is performed on the original image to obtain image data of the second pixel value distribution, and the second pixel value distribution is a pixel value distribution supported by the AI ​​model. distribution; if the pre-processing includes image block division, the original image is divided into blocks to obtain a plurality of initial image blocks, the plurality of initial image blocks include edge blocks and non-edge blocks, the edge blocks include the boundary of the original image, and the non-edge blocks do not include the boundary of the original image; for each initial image block, the initial image block is pixel-filled to obtain a target image block; wherein the plurality of target image blocks are used as input data of the AI ​​model, and the plurality of target image blocks are processed in parallel by the AI ​​model; if the pre-processing includes image rearrangement, and the first arrangement mode of the original image is an arrangement mode not supported by the AI ​​model, the original image is rearranged to obtain image data of a second arrangement mode, and the second arrangement mode is an arrangement mode supported by the AI ​​model; Wherein, the AI ​​processing includes at least one of image transformation, image expansion and semantic transformation, and the AI ​​processing module performs AI processing on the first image data through the AI ​​model to obtain the second image data, which is specifically used for: if the AI ​​processing includes image transformation, the first image data is subjected to image transformation through the AI ​​model to obtain transformed image data, the resolution of the transformed image data is equal to or less than the resolution of the first image data, and the number of channels of the transformed image data is equal to or greater than the number of channels of the first image data; the image transformation is used to separate high-frequency information and low-frequency information in the first image data; the image transformation includes at least one of interpolation transformation, wavelet transformation, Gaussian pyramid transformation, Laplace pyramid transformation and sobel filter transformation; the transformed image data is input into the image processing neural network in the AI ​​model, and the transformed image data is operated by the image processing neural network to obtain the second image data; the resolution of the second image data is equal to or less than the resolution of the first image data, and the image processing neural network is used to enhance the high-frequency information in the first image data; When the AI ​​processing module performs AI processing on the first image data through the AI ​​model to obtain the second image data, it is specifically used to: if the AI ​​processing includes image expansion, input the first image data to the image processing neural network in the AI ​​model, operate the first image data through the image processing neural network to obtain the image data to be enhanced, the resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data; perform image expansion on the first image data through the AI ​​model to obtain the expanded image data corresponding to the image data to be enhanced, the resolution of the expanded image data is equal to or less than the resolution of the first image data, and the number of channels of the expanded image data is equal to or less than the number of channels of the first image data; wherein the expanded image data is used to enhance the high-frequency information in the image data to be enhanced; wherein the image expansion includes at least one of guided filtering, image edge enhancement, image contrast enhancement, and histogram equalization; fuse the image data to be enhanced and the expanded image data to obtain the second image data, the resolution of the second image data is equal to or less than the resolution of the first image data; When the AI ​​processing module performs AI processing on the first image data through the AI ​​model to obtain the second image data, it is specifically used to: if the AI ​​processing includes semantic transformation, input the first image data to the image processing neural network in the AI ​​model, operate the first image data through the image processing neural network to obtain the image data to be enhanced, the resolution of the image data to be enhanced is equal to or less than the resolution of the first image data, and the number of channels of the image data to be enhanced is equal to or less than the number of channels of the first image data; input the first image data to the semantic processing neural network in the AI ​​model, operate the first image data through the semantic processing neural network to obtain structured semantic information, and the structured semantic information is used to describe the semantics of the image data to be enhanced; generate the second image data based on the image data to be enhanced and the structured semantic information; When the AI ​​processing module performs AI processing on the first image data through the AI ​​model to obtain the second image data, it is specifically used to: if the AI ​​processing includes image transformation, image expansion and semantic transformation, perform image transformation on the first image data through the AI ​​model to obtain transformed image data, and the image transformation is used to separate high-frequency information and low-frequency information in the first image data; input the transformed image data to the image processing neural network in the AI ​​model, operate the transformed image data through the image processing neural network to obtain first image data to be enhanced, and the image processing neural network is used to enhance the high-frequency information in the first image data; perform image expansion on the first image data through the AI ​​model to obtain expanded image data, and the expanded image data is used to enhance the high-frequency information in the first image data to be enhanced, and fuse the first image data to be enhanced and the expanded image data to obtain the second image data to be enhanced; input the first image data to the semantic processing neural network in the AI ​​model, operate the first image data through the semantic processing neural network to obtain structured semantic information, and the structured semantic information is used to describe the semantics of the second image data to be enhanced; generate the second image data based on the second image data to be enhanced and the structured semantic information; Wherein, if the first image data includes K target image blocks, the second image data includes K to-be-stored image blocks corresponding to the K target image blocks, K is a positive integer greater than 1; for each to-be-stored image block, the to-be-stored image block corresponds to a storage start address, and the interval between the storage start addresses of two adjacent to-be-stored image blocks is the width of the to-be-stored image block; the second image data corresponds to a storage span, and the storage span is greater than or equal to the sum of the widths of the K to-be-stored image blocks; When storing the second image data in a designated storage medium, the storage module is specifically used to: for each image block to be stored, when storing the i-th row of data of the image block to be stored in the designated storage medium, sequentially store the i-th row of data starting from a target storage position; wherein the target storage position is determined by the following formula: p=w+ stride*(i-1); p represents the target storage position, w represents the storage start address corresponding to the image block to be stored, stride represents the storage span, i represents the number of rows of the image block to be stored, and i is a positive integer.

10. An electronic device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions executable by the processor; The processor is used to execute machine executable instructions to implement the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Visual saliency detection method based on semantic enhanced convolutional neural network

    CN110414513A

  • HEVC intra-frame coding compression performance optimization research combined with convolutional neural network

    CN111711817A

  • Image processing method and device and storage medium

    CN115460343A

  • Image coding method and related device

    CN116527922A

  • Video image enhancement method and device based on neural network, and electronic equipment

    CN118429202A

Cited By

  • Image quality enhancement and compression method, device and equipment based on deep learning

    CN120547350A