Method for generating an image and device therefor
By extracting features from the first image and combining them with high-quality features from the second image to generate the target image, the problem of low image quality after encoding is solved, and image quality is improved and artifacts are removed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2024-07-29
- Publication Date
- 2026-07-21
AI Technical Summary
In existing image or video encoding processes, the quality of the decoded image is low, with problems such as blurring, noise, blockiness, ringing, and color banding.
By extracting features from the first image and obtaining high-quality second image features using offline or online dictionary priors, the target image is generated by combining the first and second image features, compensating for information loss and distortion, and improving the quality of the decoded image.
It effectively removes distortion and artifacts during the encoding process, improves the quality and detail of the decoded image, and enhances the flexibility and adaptability of the image.
Smart Images

Figure CN119031084B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of multimedia technology, and specifically relates to an image generation method and apparatus. Background Technology
[0002] Encoding images or videos refers to removing intra-frame and inter-frame redundancy from the encoded frames of images or videos through encoding techniques, thereby achieving efficient data compression.
[0003] However, image encoding can produce distortion. Distortion refers to the loss or error of information introduced by data compression technology during the image or video encoding process. As a result, if the encoded bitstream information is decoded, the decoded image or video may have problems such as blurriness, noise, blockiness, ringing, and color banding, resulting in a decrease in image or video quality. Summary of the Invention
[0004] The purpose of this application is to provide an image generation method and apparatus that can solve the technical problem of low image or video quality after existing decoding.
[0005] In a first aspect, embodiments of this application provide a method for generating an image, the method comprising:
[0006] Obtain bitstream information, which is the information obtained when encoding the original image;
[0007] Extract first image features from the first image, where the first image is an image obtained by decoding the bitstream information;
[0008] Extract the second image features corresponding to the first image features from the second image, wherein the second image is the original image or a third image other than the original image, the third image has more texture information and edge information than the first image, and the third image has less noise and artifacts than the first image;
[0009] Generate a target image based on the first image features and the second image features.
[0010] Secondly, embodiments of this application provide an image generation apparatus, the apparatus comprising:
[0011] The acquisition module is used to acquire bitstream information, which is the information obtained when encoding the original image;
[0012] The first extraction module is used to extract first image features from the first image, wherein the first image is an image obtained by decoding the bitstream information;
[0013] The second extraction module is used to extract the second image features corresponding to the first image features from the second image, wherein the second image is the original image or a third image other than the original image, the third image has more texture information and edge information than the first image, and the third image has less noise and artifacts than the first image;
[0014] The generation module is used to generate a target image based on the first image features and the second image features.
[0015] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method provided in the first aspect.
[0016] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method provided in the first aspect.
[0017] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface, the communication interface and the processor being coupled together, the processor being used to run programs or instructions to implement the method provided in the first aspect.
[0018] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method as provided in the first aspect.
[0019] In the image generation method and apparatus of this application, bitstream information obtained by encoding the original image can be acquired, and then the decoding information can be decoded to obtain a first image. First image features are extracted from the first image, and then second image features corresponding to the first image features are acquired. Finally, a target image is generated based on the first and second image features. Since the second image features are high-quality image features, they can repair and enhance the first image. Therefore, this method can remove the distortion and artifacts caused by compression during the encoding process of the original image, improving the quality of the decoded image. Attached Figure Description
[0020] Figure 1 This is one of the schematic flowcharts of an image generation method provided in an embodiment of this application;
[0021] Figure 2 This is a second schematic flowchart of an image generation method provided in one embodiment of this application;
[0022] Figure 3This is the third schematic flowchart of an image generation method provided in one embodiment of this application;
[0023] Figure 4 This is the fourth flowchart illustrating an image generation method provided in one embodiment of this application;
[0024] Figure 5 This is a schematic diagram of the structure of an image generation apparatus provided in another embodiment of this application;
[0025] Figure 6 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application;
[0026] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0028] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0029] To address the aforementioned technical problems, this application provides an image generation method. The image generation method provided by this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0030] like Figure 1 As shown, Figure 1 This is a flowchart illustrating an image generation method according to an embodiment of this application. This application provides an image generation method, which may include:
[0031] S101, Obtain bitstream information, wherein the bitstream information is the information obtained when encoding the original image; in this embodiment, the bitstream information refers to the information obtained after encoding the original image, and the bitstream information may include the image information after compression of the original image and the encoding parameters.
[0032] In this embodiment, encoding refers to the process of converting image or video data into a specific format (usually a compressed format). The original image can be a single image or a matrix of pixels, such as a video frame. Specifically, the original image can be an uncompressed image or video frame obtained directly from the shooting device or image generation tool, or it can be an image or video frame restored from a compressed format. The original image can be a two-dimensional planar image or video frame, or it can be a three-dimensional image or video frame.
[0033] During the encoding process, the original image can be encoded using common encoding methods such as H.266 / VVC, H.265 / HEVC, H.264 / AVC, and JPEG, or it can be encoded based on a neural network. The encoded bitstream information can be transmitted over a network to the cloud for storage, or it can be stored locally on an electronic device.
[0034] S102, extract first image features from the first image, wherein the first image is an image obtained by decoding the bitstream information;
[0035] In this embodiment, after acquiring the bitstream information, the bitstream information can be decoded to obtain a first image. Compared to the original image, the decoded first image has distortion and quality degradation due to compression. After obtaining the first image, a feature extraction network can be used to extract first image features from the first image. The first image features are used to characterize the content and structure of the first image; they can also characterize the type and degree of distortion of the first image compared to the original image. The first image features may include features such as edges, textures, and frequency components of the first image.
[0036] S103, extract the second image features corresponding to the first image features from the second image, wherein the second image is the original image or a third image other than the original image, the third image has more texture information and edge information than the first image, and the third image has less noise and artifacts than the first image;
[0037] In this embodiment, since the first image is a low-quality image with distortion, the image details in the first image features are not rich enough, and it contains a lot of noise and artifacts. In order to improve the image quality of the decoded image, a second image feature that is similar in content to the first image feature can be obtained, and a new target image can be generated by using the second image feature and the first image feature together.
[0038] In this embodiment, the second image feature corresponding to the first image feature can be obtained through offline dictionary prior knowledge or online dictionary prior knowledge.
[0039] For example, if second image features are extracted using an offline dictionary prior, the offline dictionary can be constructed in advance. Specifically, multiple high-resolution third images can be acquired in advance, and multiple fourth image features can be extracted from these third images and stored in the offline dictionary. Then, whenever a low-quality first image feature is extracted from the decoded first image, a feature matching algorithm can be used in the offline dictionary to match the extracted first image feature with the multiple fourth image features in the offline dictionary, and find the second image feature that is most similar to the first image feature from among the multiple fourth image features.
[0040] As an alternative embodiment, such as Figure 2 As shown, the electronic device includes an encoding end and a decoding end. The specific process is as follows:
[0041] Encoding end
[0042] S201: Take the original image as input.
[0043] S202: The input raw image is compressed by the encoder to generate encoded bitstream information.
[0044] S203: Transmit the encoded bitstream information over a network or store it in the cloud for use at the decoding end.
[0045] Decoding end
[0046] S204: At the decoding end, the received bitstream information is decoded to generate a decoded frame (105). The decoded frame is the first image mentioned above. The decoded frame has distortion caused by compression.
[0047] S205: The first image after decoding, which has distortion and quality degradation due to compression.
[0048] S206: Extract the degraded image features, i.e., the first image features, from the decoded frame (105).
[0049] S207: The extracted first image features are used to find the corresponding high-resolution second image features in the offline dictionary prior (108).
[0050] S208: Pre-stored high-definition feature information is saved in the form of key-value pairs, which include key-value pairs of the first image feature and the second image feature.
[0051] S209: Expand and enhance the found second image features to generate high-quality third image features.
[0052] S210: The first image features and the extended second image features are fused to obtain the final high-quality target image.
[0053] S211: Output the processed target image and play it.
[0054] For example, if the second image features are extracted using an online dictionary prior, then the second image features need to be extracted from the original image corresponding to the first image. In this case, the second image features are the image features of the original image. Compared to Figure 2 The offline dictionary method shown exists, as does the online dictionary method. Figure 3 The difference shown can be made at the encoding end before encoding the original image:
[0055] S320 Online Dictionary Feature Compression: Extracts second image features from the input original image, compresses the second image features, and then stores the compressed second image features in the bitstream information. The compression method for the second image features can be neural network compression, principal component analysis, entropy coding, etc.
[0056] Then, on the playback device, you can:
[0057] The S321 online dictionary prior and S322 online dictionary feature processing are used to decode the compressed second image features in the bitstream information and obtain the initial second image features through principal component analysis inverse process, neural network expansion, upsampling, feature extraction network, etc.
[0058] In this way, after the electronic device acquires the bitstream information, it can decode the image information in the bitstream information to obtain the first image, and decompress the compressed second image features, that is, the second image features can be obtained through online dictionary feature processing.
[0059] Since the second image has richer details such as texture and edge information than the first image, the content of the features of the second image is similar to that of the first image, and the second image features have more image detail information than the first image features, and the second image features contain less noise and artifacts than the first image features, therefore, by combining the second image features and the first image features, a target image that is the same or similar to the first image in terms of content, but has richer details and less noise and artifacts can be generated.
[0060] Obtaining high-definition second image features through the aforementioned offline or online dictionaries can improve the quality of the decoded image, reduce noise and artifacts, and enhance the system's flexibility and adaptability during image and video decoding. Offline dictionaries do not require changes to the image encoding method; simply adding an offline dictionary at the decoding end can improve the clarity of the decoded image. Online dictionaries, on the other hand, can further enable the target image to restore the original image as closely as possible, while simultaneously resolving various inter-frame continuity and single-frame quality issues arising from encoding.
[0061] In some embodiments, obtaining the second image feature corresponding to the first image feature includes:
[0062] When the second image consists of multiple third images, at least one fourth image feature is extracted from each of the third images to obtain multiple fourth image features;
[0063] The fourth image feature that has the highest matching degree with the first image feature is selected from the plurality of fourth image features and used as the second image feature.
[0064] In this embodiment, if the second image features are extracted using an offline dictionary prior method, an offline dictionary including multiple fourth image features can be constructed in advance. The offline dictionary includes multiple fourth image features, and multiple high-definition third images can be obtained in advance. Then, at least one fourth image feature can be extracted from each third image to obtain multiple fourth image features.
[0065] Then, whenever a low-quality first image feature is extracted from the decoded first image, a feature matching algorithm can be used in the offline dictionary to match the extracted first image feature with multiple fourth image features in the offline dictionary, and find the second image feature that is most similar to the first image feature from among the multiple fourth image features. This feature matching can be performed using a neural network. The neural network can include a transformer network, a convolutional neural network, a recurrent neural network, or a combination thereof.
[0066] For example, in an offline dictionary, a unique index can be assigned to each fourth image feature. Then, by training a neural network model, the index of the second image feature with the highest matching degree of the first image feature can be predicted, thereby obtaining the second image feature based on the index.
[0067] In this embodiment, the high-resolution features corresponding to the first image features can be quickly obtained through the mapping relationship in the offline dictionary.
[0068] S104, Generate a target image based on the first image features and the second image features.
[0069] In this embodiment, after obtaining the first image features and the second image features, the extracted first image features and the corresponding second image features can be used to generate a higher-quality target image through a specific fusion or repair algorithm. This utilizes the high-quality second image features to compensate for information loss and distortion in the first image, thereby significantly improving image quality.
[0070] In this application, the bitstream information obtained by encoding the original image can be acquired, and then the decoding information can be decoded to obtain a first image. First image features are extracted from the first image, and then second image features corresponding to the first image features are obtained. Finally, a target image is generated based on the first and second image features. Since the second image features are high-quality image features, they can repair and enhance the first image. Therefore, this method can remove the distortion and artifacts caused by compression during the encoding process of the original image, improving the quality of the decoded image.
[0071] In some embodiments, S104 includes:
[0072] The second image feature is input into the feature expansion network to obtain the third image feature output by the feature expansion network, wherein the dimension of the third image feature is greater than the dimension of the second image feature;
[0073] Generate target image features based on the third image features and the first image features;
[0074] The target image is generated based on the features of the target image.
[0075] In this embodiment, the second image feature is an image feature stored in a pre-built dictionary. For ease of storage, the second image feature may be a compressed low-dimensional feature. Therefore, after obtaining the second image feature, it can be expanded to obtain a higher-dimensional third image feature. Image expansion can reconstruct or enhance the details and quality of the feature. The feature expansion network can be a neural network structure used to expand compressed or low-dimensional features into high-dimensional features. It gradually recovers high-frequency details and structural information in the image through a series of convolutions, deconvolutions, upsampling, and other operations. The dimension of the image feature refers to the number of elements or data points in the image. The larger the dimension of the image feature, the more information it contains. For example, compared to the second image feature, the third image feature can contain more image information in dimensions such as edges, textures, shapes, and colors.
[0076] For example, image features can be extracted from an image using a feature extraction network. This network extracts features through multiple convolutional operations, while a feature expansion network expands low-dimensional features into high-dimensional features through multiple deconvolutional operations, recovering high-frequency details and structural information from the image. Specifically, feature expansion can be performed using a first decoder, while extracting first image features from the first image can be accomplished using a first encoder. The forms of the first decoder and the first encoder can be symmetrical.
[0077] After obtaining the third image features through feature expansion, the third image features and the first image features can be fused into the target image features. Then, a decoder or generator network is used to convert the target image feature representation into the target image.
[0078] In this embodiment, since the second image feature is originally a high-definition image feature, the process of expanding it can further supplement and enhance the texture and edge information of the image. Therefore, the target image corresponding to the target image feature obtained by combining the third image feature and the first image feature can effectively remove artifacts and noise in the image compared to the first image, and obtain a high-quality target image.
[0079] In some embodiments, generating target image features based on the third image features and the first image features includes:
[0080] Obtain the fusion parameters of the first image features and the second image features;
[0081] Obtain the weight parameters of the third image feature;
[0082] The target image features are generated based on the fusion parameters, the third image features, and the weight parameters.
[0083] In this embodiment, after obtaining the first image feature and the third image feature, the third image feature and the first image feature can be fused into the target image feature so that the target image feature includes both the content of the first image feature and the content of the third image feature.
[0084] Specifically, the fusion process can be represented by the following formula:
[0085] F hd =F d +(a*F d +b)*w (1)
[0086] Where a and b are the fusion parameters of the first and second image features, which can be extracted through a neural network, w is the weight parameter of the third image feature, with a value between 0 and 1, and F d For the third image feature, F hd Features of the target image.
[0087] In the above formula, parameter 'a' controls the contribution of the second image feature in the fusion process, while parameter 'b' further adjusts the fusion process to ensure that the fused features better match the distribution of the target image. The weight parameter 'w' adjusts the proportion of the second and first image features in the final fusion result. The value ranges from 0 to 1, with values closer to 1 indicating a greater preference for the second image feature and closer to 0 indicating a greater preference for the first image feature.
[0088] Therefore, users can find a balance between fidelity and image quality by flexibly adjusting the weight parameter w, so that the output target image meets the user's needs.
[0089] In some embodiments, the bitstream information includes quantization parameters of the original image;
[0090] Before generating the target image based on the first image features and the second image features, the method further includes:
[0091] If the quantization parameter does not belong to the quantization threshold range, the parameter value of the weight parameter of the third image feature is adjusted to the first parameter value, wherein the quantization parameter corresponding to the first parameter value belongs to the quantization threshold range.
[0092] The step of generating a target image based on the first image features and the second image features includes:
[0093] The target image is generated based on the first parameter value, the first image feature, and the second image feature.
[0094] In this embodiment, during the encoding (compression) of the original image, quantization is a process of simplifying the transformed coefficients. The quantization parameter (QP) determines the size of the quantization step. A larger quantization parameter results in more information loss during image encoding, leading to a higher compression ratio, but also greater distortion. Therefore, during the encoding of the original image, a higher quantization parameter leads to a higher compression ratio but introduces greater distortion, resulting in decreased image quality and loss of detail. A lower quantization parameter retains more detail and has less distortion, but a lower compression ratio.
[0095] After the original image is encoded with preset quantization parameters, these quantization parameters can be stored in the bitstream information. If decoding of the bitstream information is required, the quantization parameters in the bitstream information are first obtained. Then, it is determined whether the quantization parameters belong to a preset quantization threshold range. If the quantization parameters belong to the preset quantization threshold range, the bitstream information can be decoded based on these quantization parameters. If the quantization parameters do not belong to the quantization threshold range, the quantization parameters need to be adjusted so that the adjusted quantization parameters fall within the quantization threshold range, and the bitstream information is then decoded based on the adjusted quantization parameters. A single quantization parameter can correspond to an entire image or a portion of an image.
[0096] Specifically, the quantization parameter can be adjusted by changing the weight parameter w of the third image feature, since the weight parameter is used to adjust the proportion of the second and first image features in the final fusion result. Increasing the weight parameter will decrease the quantization parameter accordingly, and decreasing the weight parameter will increase the quantization parameter accordingly.
[0097] In addition, the quantization parameters of the original image can be normalized to obtain normalized values, and the magnitude of the normalized values can be used to determine whether the quantization parameters belong to the quantization threshold range.
[0098] Therefore, if the quantization parameter does not belong to the quantization threshold range, the parameter value of the weight parameter of the third image feature can be adjusted from the initial parameter value to the first parameter value so that the quantization parameter corresponding to the first parameter value belongs to the quantization threshold range.
[0099] In this way, by dynamically adjusting the weight parameters to adjust the quantization parameters during the decoding process, the optimal balance between compression ratio and quality of image or video can be found.
[0100] In some embodiments, after generating the target image based on the first image features and the second image features, the method further includes:
[0101] Display the target image;
[0102] Upon receiving a user's input to modify the parameters of the target image, the system responds to the input by adjusting the value of the weight parameter of the third image feature to the value of the second parameter.
[0103] The target image is updated based on the second parameter value, the first image feature, and the second image feature.
[0104] In this embodiment, similarly, compared to Figure 2 As shown, this embodiment has the following characteristics: Figure 4 The difference shown is that after the original image is encoded with preset quantization parameters, these quantization parameters can be stored in the bitstream information. If decoding of the bitstream information is required, the quantization parameter QP(S215) in the bitstream information can be obtained first, and the target image can be generated and displayed based on this quantization parameter.
[0105] During the decoding process of bitstream information, higher quantization parameters will introduce higher distortion and noise, but can ensure that the target image and the original image have high similarity; lower quantization parameters will cause the target image and the original image to have lower similarity, but the distortion of the target image is smaller.
[0106] After displaying the target image, the user can observe the target image on the display interface and receive parameter modification input for user intensity adjustment (S216). In response to the parameter modification input, the quantization parameter QP is adjusted (S217). If the user believes that the realism of the target image is too low or the sharpness is too high, the user can input parameter modification to the electronic device, increasing the value of the weight parameter to decrease the quantization parameter, thereby increasing the proportion of the first image feature in the target image, thus improving the similarity between the target image and the original image or reducing the sharpness. Conversely, if the user believes that the sharpness of the target image is too low, the user can also input parameter modification to the electronic device, decreasing the value of the weight parameter to increase the quantization parameter, thereby increasing the proportion of the second image feature in the target image, thus improving the sharpness of the target image.
[0107] Specifically, users can modify the input to adjust the weight parameter value of the third image feature from the initial parameter value to the second parameter value, so that the target image corresponding to the second parameter value matches the user's needs.
[0108] Using the above method, users can dynamically adjust the quantization parameters during the decoding process by manually adjusting the weight parameters, so that the decoded image or video better meets the user's needs.
[0109] Figure 5 This is a schematic diagram of the structure of an image generation apparatus provided in another embodiment of this application, such as... Figure 5 As shown, the image generating apparatus may include:
[0110] The acquisition module 501 is used to acquire bitstream information, which is the information obtained when encoding the original image;
[0111] The first extraction module 502 is used to extract first image features from the first image, wherein the first image is an image obtained by decoding the bitstream information;
[0112] The second extraction module 503 is used to extract the second image features corresponding to the first image features from the second image, wherein the second image is the original image or a third image other than the original image, the third image has more texture information and edge information than the first image, and the third image has less noise and artifacts than the first image;
[0113] The generation module 504 is used to generate a target image based on the first image features and the second image features.
[0114] In this application, the bitstream information obtained by encoding the original image can be acquired, and then the decoding information can be decoded to obtain a first image. First image features are extracted from the first image, and then second image features corresponding to the first image features are obtained. Finally, a target image is generated based on the first and second image features. Since the second image features are high-quality image features, they can repair and enhance the first image. Therefore, this method can remove the distortion and artifacts caused by compression during the encoding process of the original image, improving the quality of the decoded image.
[0115] In another alternative example, the generation module 504 includes:
[0116] An extension unit is used to input the second image feature into a feature extension network to obtain a third image feature output by the feature extension network, wherein the dimension of the third image feature is greater than the dimension of the second image feature;
[0117] The first generation unit is configured to generate target image features based on the third image features and the first image features;
[0118] The second generation unit is used to generate the target image based on the target image features.
[0119] In another alternative example, the bitstream information includes quantization parameters of the original image;
[0120] The image generation apparatus includes:
[0121] The first adjustment module is used to adjust the parameter value of the weight parameter of the third image feature to a first parameter value when the quantization parameter does not belong to the quantization threshold range, wherein the quantization parameter corresponding to the first parameter value belongs to the quantization threshold range.
[0122] The generation module 504 further includes:
[0123] The third generation unit is used to generate the target image based on the first parameter value, the first image feature, and the second image feature.
[0124] In another alternative example, the image generating apparatus includes:
[0125] Display module, used to display the target image;
[0126] The second adjustment module is used to adjust the parameter value of the weight parameter of the third image feature to the second parameter value in response to the parameter modification input received from the user for the target image.
[0127] An update module is used to update the target image based on the second parameter value, the first image feature, and the second image feature.
[0128] In another alternative example, the second extraction module 503 includes:
[0129] The extraction unit is configured to extract at least one fourth image feature from each of the third images when the second image consists of multiple third images, thereby obtaining multiple fourth image features;
[0130] The acquisition unit is used to acquire the fourth image feature that has the highest matching degree with the first image feature from the plurality of fourth image features as the second image feature.
[0131] The image generation device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device to these types of devices.
[0132] The image generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0133] The image generation apparatus provided in this application embodiment can achieve Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0134] Optionally, such as Figure 6 As shown, this application embodiment also provides an electronic device 100, including a processor 110, a memory 119, and a program or instructions stored in the memory 119 and executable on the processor 110. When the program or instructions are executed by the processor 110, they implement the various processes of the above-described image generation method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0135] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0136] Please refer to the following: Figure 7, Figure 7 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. The electronic device 100 includes, but is not limited to, components such as: a radio frequency unit 121, a network module 122, an audio output unit 123, an input unit 124, a sensor 125, a display unit 126, a user input unit 127, an interface unit 128, a memory 129, and a processor 120.
[0137] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 120 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0138] Among them, the network module 122 is used to acquire bitstream information, which is the information obtained when encoding the original image;
[0139] Processor 120 is configured to extract first image features from a first image, wherein the first image is an image obtained by decoding the bitstream information;
[0140] Processor 120 is configured to extract second image features corresponding to the features of the first image from the second image, wherein the second image is the original image or a third image other than the original image, the third image has more texture information and edge information than the first image, and the third image has less noise and artifacts than the first image;
[0141] Processor 120 is configured to generate a target image based on the first image features and the second image features.
[0142] In this application, the bitstream information obtained by encoding the original image can be acquired, and then the decoding information can be decoded to obtain a first image. First image features are extracted from the first image, and then second image features corresponding to the first image features are obtained. Finally, a target image is generated based on the first and second image features. Since the second image features are high-quality image features, they can repair and enhance the first image. Therefore, this method can remove the distortion and artifacts caused by compression during the encoding process of the original image, improving the quality of the decoded image.
[0143] In another alternative example, the processor 120 is also used for:
[0144] The second image feature is input into the feature expansion network to obtain the third image feature output by the feature expansion network, wherein the dimension of the third image feature is greater than the dimension of the second image feature;
[0145] Generate target image features based on the third image features and the first image features;
[0146] The target image is generated based on the features of the target image.
[0147] In another alternative example, the bitstream information includes quantization parameters of the original image;
[0148] The processor 120 is also used for:
[0149] If the quantization parameter does not belong to the quantization threshold range, the parameter value of the weight parameter of the third image feature is adjusted to the first parameter value, wherein the quantization parameter corresponding to the first parameter value belongs to the quantization threshold range.
[0150] The target image is generated based on the first parameter value, the first image feature, and the second image feature.
[0151] In another alternative example, the display unit 126 is used to display the target image;
[0152] User input unit 127 is configured to, upon receiving a user's parameter modification input for the target image, respond to the parameter modification input and adjust the parameter value of the weight parameter of the third image feature to the second parameter value.
[0153] Processor 120 is configured to update the target image based on the second parameter value, the first image feature, and the second image feature.
[0154] In another alternative example, the processor 120 is configured to extract at least one fourth image feature from each of the third images when the second image is a plurality of third images, thereby obtaining a plurality of fourth image features;
[0155] Processor 120 is configured to obtain, from the plurality of fourth image features, the fourth image feature that has the highest matching degree with the first image feature as the second image feature.
[0156] It should be understood that, in this embodiment, the input unit 124 may include a graphics processing unit (GPU) 1241 and a microphone 1242. The GPU 1241 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 126 may include a display panel 1261, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 127 includes at least one of a touch panel 1271 and other input devices 1272. The touch panel 1271 is also called a touch screen. The touch panel 1271 may include a touch detection device and a touch controller. Other input devices 1272 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0157] The memory 129 can be used to store software programs and various data. The memory 129 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 129 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 129 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0158] Processor 120 may include one or more processing units; optionally, processor 120 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 120.
[0159] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image generation method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0160] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0161] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described image generation method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0162] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0163] This application provides a computer program product stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-described image generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0164] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0165] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0166] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for generating an image, characterized in that, include: Obtain bitstream information, which is the information obtained when encoding the original image; Extract first image features from the first image, where the first image is an image obtained by decoding the bitstream information; Extract the second image features corresponding to the first image features from the second image, wherein the second image is the original image or a third image other than the original image, the third image has more texture information and edge information than the first image, and the third image has less noise and artifacts than the first image; Generate a target image based on the first image features and the second image features; The step of generating a target image based on the first image features and the second image features includes: Obtain the fusion parameters of the first image features and the second image features; Obtain the weight parameters of the third image feature, which is obtained by inputting the second image feature into the feature expansion network and outputting the feature expansion network. The third image feature has more dimensions than the second image feature. Target image features are generated based on the fusion parameters, the third image features, and the weight parameters; Generate a target image based on the features of the target image.
2. The method according to claim 1, characterized in that, The bitstream information includes the quantization parameters of the original image; Before generating the target image based on the first image features and the second image features, the method further includes: If the quantization parameter does not belong to the quantization threshold range, the parameter value of the weight parameter of the third image feature is adjusted to the first parameter value, wherein the quantization parameter corresponding to the first parameter value belongs to the quantization threshold range. The step of generating a target image based on the first image features and the second image features includes: The target image is generated based on the first parameter value, the first image feature, and the second image feature.
3. The method according to claim 1, characterized in that, After generating the target image based on the first image features and the second image features, the method further includes: Display the target image; Upon receiving a user's input to modify the parameters of the target image, the system responds to the input by adjusting the value of the weight parameter of the third image feature to the value of the second parameter. The target image is updated based on the second parameter value, the first image feature, and the second image feature.
4. The method according to claim 1, characterized in that, Extracting the second image features corresponding to the first image features from the second image includes: When the second image consists of multiple third images, at least one fourth image feature is extracted from each of the third images to obtain multiple fourth image features; The fourth image feature that has the highest matching degree with the first image feature is selected from the plurality of fourth image features and used as the second image feature.
5. An image generation apparatus, characterized in that, include: The acquisition module is used to acquire bitstream information, which is the information obtained when encoding the original image; The first extraction module is used to extract first image features from the first image, wherein the first image is an image obtained by decoding the bitstream information; The second extraction module is used to extract the second image features corresponding to the first image features from the second image, wherein the second image is the original image or a third image other than the original image, the third image has more texture information and edge information than the first image, and the third image has less noise and artifacts than the first image; The generation module is used to generate a target image based on the first image features and the second image features; The generation module is specifically used to obtain the fusion parameters of the first image feature and the second image feature; obtain the weight parameters of the third image feature, wherein the third image feature is obtained by inputting the second image feature into a feature expansion network and outputting the feature expansion network, and the dimension of the third image feature is greater than that of the second image feature; generate a target image feature according to the fusion parameters, the third image feature and the weight parameters; and generate a target image according to the target image feature.
6. The apparatus according to claim 5, characterized in that, The bitstream information includes the quantization parameters of the original image; The image generation apparatus further includes: The first adjustment module is used to adjust the parameter value of the weight parameter of the third image feature to a first parameter value when the quantization parameter does not belong to the quantization threshold range, wherein the quantization parameter corresponding to the first parameter value belongs to the quantization threshold range. The generation module further includes: The third generation unit is used to generate the target image based on the first parameter value, the first image feature, and the second image feature.
7. The apparatus according to claim 5, characterized in that, The image generation apparatus further includes: Display module, used to display the target image; The second adjustment module is used to adjust the parameter value of the weight parameter of the third image feature to the second parameter value in response to the parameter modification input received from the user for the target image. An update module is used to update the target image based on the second parameter value, the first image feature, and the second image feature.
8. The apparatus according to claim 5, characterized in that, The second extraction module includes: The extraction unit is configured to extract at least one fourth image feature from each of the third images when the second image consists of multiple third images, thereby obtaining multiple fourth image features; The acquisition unit is used to acquire the fourth image feature that has the highest matching degree with the first image feature from the plurality of fourth image features as the second image feature.