Image processing method and related equipment

By splitting the image to be processed into multiple image blocks and processing in parallel, the existing diffusion model has solved the problem of slow processing speed and high hardware resource consumption in image processing, and more efficient image processing is achieved.

CN120088136APending Publication Date: 2025-06-03BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311640470.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing diffusion model has problems such as slow processing speed and high hardware resource consumption in image processing.

Method used

Parallel processing is achieved to improve processing speed by splitting the images to be processed into at least two first image blocks and inputting these image blocks into an image processing model to generate a target image.

Benefits of technology

It significantly improves the speed of image processing, reduces the consumption of hardware resources, and solves the problems of slow processing speed and high resource consumption in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088136A_ABST
    Figure CN120088136A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and related equipment. The method comprises the following steps: acquiring a to-be-processed image; based on the to-be-processed image, splitting to obtain at least two first image blocks; inputting the at least two first image blocks into an image processing model, and outputting to obtain at least two second image blocks; and generating a target image according to the at least two second image blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to an image processing method and related devices. Background Art

[0002] In the field of image processing, one research direction is how to generate high-resolution images from low-resolution images. For example, for an image, how to still display a clear image when magnified many times.

[0003] In related technologies, a diffusion model (DM for short) can be used to generate high-resolution images.

[0004] However, the inventors of the present disclosure have found that the diffusion model used in related technologies has problems of slow processing speed and high consumption of hardware resources. Summary of the Invention

[0005] The present disclosure provides an image processing method and related devices to solve or partially solve the above problems.

[0006] In a first aspect of the present disclosure, an image processing method is provided, including:

[0007] Obtain an image to be processed;

[0008] Based on the image to be processed, split it into at least two first image blocks;

[0009] Input the at least two first image blocks into an image processing model, and output at least two second image blocks;

[0010] Generate a target image according to the at least two second image blocks.

[0011] In a second aspect of the present disclosure, an image processing device is provided, including:

[0012] An obtaining module, configured to: obtain an image to be processed;

[0013] A splitting module, configured to: based on the image to be processed, split it into at least two first image blocks;

[0014] A processing module, configured to: input the at least two first image blocks into an image processing model, and output at least two second image blocks;

[0015] A generating module, configured to: generate a target image according to the at least two second image blocks.

[0016] In a third aspect of the present disclosure, a computer device is provided, including one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the programs include instructions for executing the method according to the first aspect.

[0017] In a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium containing a computer program is provided. When the computer program is executed by one or more processors, the processors are caused to execute the method according to the first aspect.

[0018] The image processing method and related devices provided in the embodiments of the present disclosure process by splitting the image to be processed into at least two first image blocks, so as to improve the processing speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are only the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 The schematic diagram of an exemplary system provided by the embodiments of the present disclosure is shown.

[0021] Figure 2A The flowchart of an exemplary method provided by the embodiments of the present disclosure is shown.

[0022] Figure 2B The flowchart of an exemplary method for processing the first image block according to the embodiments of the present disclosure is shown.

[0023] Figure 2C The flowchart of an exemplary method for generating a second image block according to a third image feature according to the embodiments of the present disclosure is shown.

[0024] Figure 2D The flowchart of an exemplary method for generating a second image block according to a second group of image features according to the embodiments of the present disclosure is shown.

[0025] Figure 3A The schematic diagram of the model architecture provided by the embodiments of the present disclosure is shown.

[0026] Figure 3B The schematic diagram of splitting an image according to the embodiments of the present disclosure is shown.

[0027] Figure 3CShows a schematic diagram of an exemplary image processing model provided by an embodiment of the present disclosure.

[0028] Figure 4A Shows a schematic diagram of an exemplary process for processing a first image block by an image processing model according to an embodiment of the present disclosure.

[0029] Figure 4B Shows a schematic diagram of an exemplary process for processing image features according to an embodiment of the present disclosure.

[0030] Figure 4C Shows a schematic diagram of another exemplary process for processing image features according to an embodiment of the present disclosure.

[0031] Figure 4D Shows a schematic diagram of stitching images according to an embodiment of the present disclosure.

[0032] Figure 4E Shows another schematic diagram of stitching images according to an embodiment of the present disclosure.

[0033] Figure 5 Shows a schematic diagram of the hardware structure of an exemplary computer device provided by an embodiment of the present disclosure.

[0034] Figure 6 Shows a schematic diagram of an exemplary device provided by an embodiment of the present disclosure. Detailed implementation manners

[0035] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0036] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0037] It is understandable that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0038] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operations of the technical solutions of the present disclosure according to the prompt message.

[0039] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0040] It is understandable that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0041] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 provided by an embodiment of the present disclosure.

[0042] As Figure 1 shown, the system 100 may include a terminal device 102, a terminal device 104, a server 106, and a database server 108. A medium (such as a network) for providing a communication link may be included between the terminal device 102, the terminal device 104, the server 106, and the database server 108. The network may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0043] Various application programs (APPs) may be installed on the terminal device 104. For example, image processing application programs, video conferencing application programs, reading application programs, video application programs, social application programs, payment application programs, web browsers, and instant messaging tools, etc. These application programs can all be used to display the information to be delivered.

[0044] The terminal devices 102 and 104 here can be either hardware or software. When the terminal devices 102 and 104 are hardware, they can be various electronic devices with a display screen, including but not limited to smartphones, tablets, e-book readers, MP3 players, laptop computers, and desktop computers (PCs), etc. When the terminal devices 102 and 104 are software, they can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (e.g., for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0045] The server 106 can be a server that provides various services, such as a background server that supports various applications displayed on the terminal devices 102 and 104. The database server 108 can also be a database server that provides various services. It can be understood that when the relevant functions of the database server 108 can be implemented in the server 106, the database server 108 may not be set in the system 100.

[0046] The server 106 and the database server 108 here can also be either hardware or software. When they are hardware, they can be implemented as a distributed server cluster composed of multiple servers, or they can be implemented as a single server. When they are software, they can be implemented as multiple software or software modules (e.g., for providing distributed services), or they can be implemented as a single software or software module. No specific limitation is made here.

[0047] It should be noted that the image processing method provided by the embodiments of the present disclosure can be executed by the terminal device 102 and / or the terminal device 104. It should be understood that Figure 1 the numbers of the terminal devices, users, servers, and database servers in

[0048] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, users, servers, and database servers.

[0049] In one embodiment, an image processing application or software can be installed in the terminal device 104, and the user 112 can use the application or software to process images. For example, the user can magnify an image and display the magnified image.

[0050] As Figure 1 shown, in some embodiments, the server 106 may further include multiple servers, for example, servers 106A to 106D. Each server can be used to deploy a machine learning model, and the machine learning models deployed in each server can be independent of each other. In other words, the machine learning models in multiple servers can be used to process data in parallel, thereby improving the processing efficiency. It can be understood that since the machine learning model may be large, one model may not be able to be deployed only in one server, but multiple distributed servers are needed to deploy one model. Anyway, in this embodiment, multiple servers in the server 106 can deploy multiple machine learning models respectively to process data in parallel.

[0051] Optionally, the pre-trained machine learning model can be deployed in the server 106. At this time, when the terminal device 104 processes an image, it can upload the image to the server 106 for processing, and the server 106 returns the processed image to the terminal device 104 for display. In some cases, if the pre-trained machine learning model can be lightweight, it can be deployed in the terminal device 104. When processing an image, the terminal device 104 can call the machine learning model stored locally to perform image processing without processing through the server 106.

[0052] It can be understood that when the machine learning model is deployed in the terminal device 104, the terminal device 104 may include multiple image processors (GPUs). Each image processor or every few image processors can be used to deploy a machine learning model, so that multiple machine learning models can be deployed in the terminal device 104 to process data in parallel.

[0053] As mentioned above, when the model in the related art completes a generative task (for example, generating a high-resolution image based on a low-resolution image), there are problems of slow processing speed and high consumption of hardware resources.

[0054] In view of this, the embodiments of the present disclosure provide an image processing method to solve or partially solve the above problems.

[0055] Figure 2A shows a schematic flowchart of an exemplary method 200 provided by the embodiments of the present disclosure. The method 200 can be used to process images. Optionally, the method 200 can be separately implemented by Figure 1 the terminal devices 102 and 104 respectively, or can be implemented by Figure 1 the system 100. The following will illustrate with the system 100 implementing the method 200.

[0056] As Figure 2AAs shown, the method 200 may further include the following steps.

[0057] In step 202, an image to be processed is obtained.

[0058] In some embodiments, the image to be processed may be an image region intercepted from an image with a general resolution (e.g., 1080×768). Optionally, the image region may be selected by the user. As an alternative embodiment, the image with the general resolution may be displayed on the display screen of the terminal device 104, and the user 112 may select, by means of touch or the like, the image region to be enlarged (which may also be regarded as a low-quality image LR or a low-resolution image) in the image as the image to be processed.

[0059] It can be understood that the acquisition method of the image is not limited to taking pictures with a camera, downloading through the network, retrieving from local, etc.

[0060] In step 204, based on the image to be processed, at least two first image blocks are split.

[0061] Figure 3A A schematic diagram of the model architecture 300 provided by an embodiment of the present disclosure is shown.

[0062] As Figure 3A shown, the image to be processed 302 may be split into multiple first image blocks. For example, the image blocks 302A to 302D.

[0063] In this step, the number of the first image blocks obtained by splitting may be arbitrary, and the specific splitting number is determined according to actual needs. For example, when it is necessary to improve the operation speed and the hardware device meets the requirements (for example, the number of the hardware (server, GPU) for deploying the image processing model is sufficient), the image to be processed may be split into more blocks. Conversely, the splitting number may be reduced. The following embodiments will be described by taking the image to be processed as being split into 4 first image blocks.

[0064] To ensure that the features of the region adjacent to the first image block in the image to be processed 302 can be learned when processing the split first image blocks, in some embodiments, when splitting the image to be processed 302, each first image block may have an overlapping region with its adjacent first image block. As Figure 3B shown, the shaded part of the image to be processed 302 shows the overlapping region between adjacent first image blocks, and the shaded part of the split first image blocks 302A to 302D shows the overlapping parts of the first image block with its adjacent first image block.

[0065] In step 206, the at least two first image blocks are input into an image processing model, and at least two second image blocks are output.

[0066] In some embodiments, the at least two first image blocks may be respectively input into at least two image processing models, and the at least two second image blocks are output.

[0067] As Figure 3A shown, after splitting the image 302 to be processed into 4 first image blocks 302A - 302D, these image blocks can be respectively sent to different image processing models 304A - 304D for processing, so as to respectively obtain 4 second image blocks 306A - 306D with improved resolution. In this way, by using different image processing models 304A - 304D to process the 4 first image blocks 302A - 302D respectively, the processing process of each first image block can be carried out in parallel, improving the inference speed and the operation speed.

[0068] In some embodiments, in order to obtain a super - resolution image (SR for short), the image processing models 304A - 304D may adopt a super - resolution model (i.e., a super - resolution model) to magnify the first image block, so as to obtain a target image with a higher resolution. It can be understood that the super - resolution model can be any model that can achieve super - resolution processing. For example, a convolutional neural network (CNN), a generative adversarial network (GAN), a variational auto - encoder (VAE), a normalization flow model (NF), a diffusion model (DM), and so on.

[0069] In particular, when the image processing models 304A - 304D adopt a diffusion model, if the image processing model magnifies the image by 4 times, the input and output resolutions of the diffusion model are 16 times those of generation and editing models (such as CNN, GAN). Therefore, the inference performance pressure is greater. By adopting the method of the embodiments of the present disclosure, the inference speed can be better improved and the video memory consumption can be reduced.

[0070] It can be understood that the image processing models 304A - 304D can be pre - trained models and can be deployed in different hardware devices to complete parallel processing tasks. For the specific model training process, various commonly used model training methods can be adopted, which will not be elaborated here.

[0071] It should also be understood that the above embodiments of parallel processing are merely exemplary. In some scenarios, considering that the model requires a large amount of resources, only one image processing model can be used to serially process the first image block, and the second image block can still be obtained.

[0072] In some embodiments, in order to further improve the processing speed, as Figure 2B shown, the step 206 of inputting the at least two first image blocks into at least two image processing models respectively and outputting at least two second image blocks may further include the following steps.

[0073] For each of the first image blocks, the following steps are performed:

[0074] In step 2062, based on the first image block, at least two first sub-image blocks are split.

[0075] Figure 3C The figure shows a schematic diagram of the exemplary image processing model 304A provided by the embodiments of the present disclosure.

[0076] Taking the image processing model 304A processing the first image block 302A as an example, as Figure 3C shown, the image processing model 304A can first further split the first image block 302A into at least two first sub-image blocks.

[0077] In some embodiments, similar to splitting the image 302 to be processed, there may also be an overlapping area between adjacent first sub-image blocks, so as to ensure that the features of the area adjacent to the first sub-image block in the first image block can be learned when processing the split first sub-image blocks. For the specific splitting method, refer to the appendix Figure 3B , which will not be elaborated here.

[0078] After splitting into at least two first sub-image blocks, in step 2064, the at least two first sub-image blocks can be respectively encoded to obtain at least two first image features corresponding to the at least two first sub-image blocks.

[0079] In this step, by further splitting the first image block 302A into at least two first sub-image blocks for separate encoding processing, the amount of data processed in a single encoding can be reduced, and the efficiency of the encoding processing can be improved.

[0080] As an optional embodiment, as Figure 3C shown, the image processing model 304A may further include an encoder 3042. Inputting the first sub-image block into the encoder 3042 can output the first image feature corresponding to the first sub-image block.

[0081] In some embodiments, the encoder 3042 can be the encoder of a variational autoencoder (VAE) and can jointly form a variational autoencoder with the decoder 3046, so that the image processing model 304A can achieve a better image magnification effect.

[0082] As an alternative embodiment, considering that the time required to serially (sequentially) process the first sub-image blocks may be relatively long, the at least two first sub-image blocks can be divided into at least two first sub-image block groups according to a preset quantity (e.g., 4, 8, 16, etc.) (e.g., Figure 3C the first sub-image block groups 3022A, 3024A), and then they are sequentially sent to the encoder 3042 for encoding processing in groups, thereby reducing the number of encoding processes and improving the processing speed.

[0083] Considering the network part of the VAE, there are a large number of self-attention layers, and the video memory complexity during inference has a relationship of O(n 2 ) with the resolution. In this embodiment, by further dividing the first image block 302A and sending it to the encoder 3042, a significant amount of video memory can be saved. With this block division strategy, in the case of an image with an input resolution of 512×512, the video memory of the image processing model 304A is reduced from 50.5G to 14.4G.

[0084] In step 2066, according to the at least two first image features, a second image block corresponding to the first image block is generated.

[0085] In some embodiments, the step 2066 of generating the second image block corresponding to the first image block according to the at least two first image features may further include:

[0086] Concatenating the at least two first image features into a second image feature;

[0087] Performing resampling on the second image feature to obtain a third image feature;

[0088] Generating the second image block corresponding to the first image block according to the third image feature.

[0089] In this embodiment, considering that the encoder part of the VAE first obtains the mean and variance of the image features through the model, and then performs resampling (formula transformation) according to the mean and variance output by the model to obtain the final image features. If resampling is directly performed on each first image feature separately, the obtained mean and variance will be too local and unable to reflect the global feature information, affecting the final image effect.

[0090] In view of this, as Figure 4AAs shown, in this embodiment, first, the output results (e.g., the first image features 402A - 402D) are stitched back to the second image feature 404 of the original size, and then resampling is performed on the overall information of the second image feature 404 to obtain the third image feature 406, so as to retain the overall sampling mean and variance results of the image and obtain better image features.

[0091] In some embodiments, as Figure 2C shown, the step of generating the second image block corresponding to the first image block according to the third image feature may further include the following steps.

[0092] In step 20662, the third image feature is split into at least two fourth image features.

[0093] Figure 4B shows a schematic diagram of an exemplary process for processing image features according to an embodiment of the present disclosure.

[0094] As Figure 4B shown, in this step, for the convenience of subsequent processing to reduce the processing complexity, after resampling the second image feature 404 to obtain the third image feature 406, a splitting process can be performed again. For example, the third image feature 406 is split into at least two fourth image features (e.g., the fourth image features 408A - 408D).

[0095] In some embodiments, similar to splitting the image 302 to be processed, there may also be an overlapping area between adjacent fourth image features, so as to ensure that the features of the area adjacent to the fourth image feature in the third image feature can be learned when processing the split fourth image features. For the specific splitting method, refer to Appendix Figure 3B , which will not be elaborated here.

[0096] In step 20664, the at least two fourth image features are divided into at least two first image feature groups, and each of the first image feature groups includes a preset number (e.g., 4, 8, 16, etc.) of the fourth image features.

[0097] As Figure 4B shown, in this step, for the convenience of subsequent processing and to improve the processing speed, the fourth image features 408A - 408D can be divided into the first image feature groups 410A and 410B.

[0098] It can be understood that Figure 4A and Figure 4B the number of image blocks, image features, and image feature groups in are only set for the convenience of illustration. In the actual processing process, the number of image blocks, image features, and image feature groups can be more or less.

[0099] In step 20666, the at least two first image feature groups are processed in sequence to obtain at least two second image feature groups that correspond one-to-one with the at least two first image feature groups (for example, Figure 4B second image feature groups 412A and 412B), and each of the second image feature groups includes a preset number (for example, 4, 8, 16, etc.) of fifth image features. Optionally, this step 20666 can be implemented using Figure 3C diffusion model 3044. Specifically, it can be implemented using a U-net model.

[0100] In some embodiments, when grouping, after chunking is completed, the fourth image features can be arranged in an unfolded manner (sequentially arranged), and then grouped according to the arrangement order and the preset number. For example, 4 fifth image features are sent to the network for processing each time according to the arrangement order. This processing method can avoid the situation of resource waste easily occurring at the end of each row after chunking.

[0101] Taking the third image feature with an input resolution of 512×512 as an example, in the case where the fourth image feature has an overlapping area, 25 fourth image features can be split, and correspondingly, the number of times the model runs is ceil(25 / 4) = 7 times (the last time is filled to 4 by replication). It is found through testing that compared with before optimization, the overall speed of the model is increased by 30%.

[0102] In step 20668, the fifth image features in the at least two second image feature groups are spliced into a sixth image feature.

[0103] After the model processing is completed, since two adjacent fourth image features have a partially overlapping area, two adjacent fifth image features have a partially overlapping area. If the fifth image features split into blocks are continued to be processed, it may cause problems with the accuracy of the model. Therefore, in this step, the fifth image features can be spliced back into the sixth image feature 416 of the original size. When splicing, for the feature values corresponding to the points in the splicing part formed by the overlapping area, the feature values of this point in the adjacent fifth image features can be weighted and calculated to obtain a weighted feature value, and the weighted feature value is used as the feature value of this point, thus completing the splicing.

[0104] In step 20670, according to the sixth image feature, a second image block corresponding to the first image block is generated.

[0105] In some embodiments, as Figure 2D shown, the step of generating a second image block corresponding to the first image block according to the sixth image feature can further include the following steps.

[0106] In step 20672, split the sixth image feature into at least two seventh image features.

[0107] Figure 4C FIG. shows a schematic diagram of another exemplary process for processing image features according to an embodiment of the present disclosure.

[0108] As Figure 4C shown, in this step, in order to facilitate subsequent processing to reduce the processing complexity, the sixth image feature 416 can be split again. For example, the sixth image feature 416 is split into at least two seventh image features (e.g., seventh image features 418A-418D).

[0109] In some embodiments, similar to splitting the image 302 to be processed, there may also be an overlapping area between adjacent seventh image features, so as to ensure that when processing the split seventh image features, the features of the area adjacent to the seventh image feature in the sixth image feature can be learned. For the specific splitting method, refer to the appendix Figure 3B , which will not be elaborated here.

[0110] In step 20674, divide the at least two seventh image features into at least two third image feature groups (e.g., Figure 4C third image feature groups 420A, 420B), and each of the third image feature groups includes a preset number of the seventh image features.

[0111] In step 20676, decode the at least two third image feature groups respectively (i.e., send each third image feature group into the decoder 3046 in sequence), and obtain at least two second sub-image blocks corresponding one-to-one to the seventh image features in the at least two third image feature groups.

[0112] As Figure 4C shown, after each third image feature group is input into the decoder 3046, a corresponding second sub-image block group can be obtained. For example, second sub-image block groups 3062A, 3062B, and each second sub-image block group may further include a preset number of second sub-image blocks.

[0113] Considering the network part of the VAE, there are a large number of self-attention layers, and the memory complexity during its inference is in an O(n 2 ) relationship with the resolution. In this embodiment, by further dividing the sixth image feature and sending it to the decoder 3046, the memory can be greatly saved.

[0114] In step 20678, splice the at least two second sub-image blocks into a second image block corresponding to the first image block.

[0115] In some embodiments, step 20678 of stitching the at least two second sub-image blocks into a second image block corresponding to the first image block may further include the following steps:

[0116] Overlap the overlapping regions of two adjacent second sub-image blocks with each other to complete the stitching, and the overlapping region forms a stitching part;

[0117] For the pixel points in the stitching part, perform weighted processing on the pixel values of the pixel points in the two adjacent second sub-image blocks to obtain a weighted pixel value, and use the weighted pixel value as the eigenvalue of the pixel point.

[0118] Figure 4D Shows a schematic diagram of a stitched image according to an embodiment of the present disclosure.

[0119] As Figure 4D shown, the shaded parts of the second sub-image blocks show the overlapping regions of adjacent second sub-image blocks. During stitching, the overlapping regions can be overlapped with each other to complete the stitching. Correspondingly, the overlapping region forms a stitching part, and the shaded part in the second image block 306A shows this stitching part.

[0120] For the pixel points in the stitching part, according to which second sub-image blocks' overlapping regions the pixel point is located in, perform weighted processing to obtain a weighted pixel value as the final pixel value of the pixel point.

[0121] For example, as Figure 4D shown, for a stitching part where only two second sub-image blocks overlap, the weights of the pixel values of the pixel point in each second sub-image block can be set to 0.5, and then weighted summation is performed to obtain the weighted pixel value. For a stitching part where four second sub-image blocks overlap, the weights of the pixel values of the pixel point in each second sub-image block can be set to 0.25, and then weighted summation is performed to obtain the weighted pixel value.

[0122] In step 208, as Figure 3A shown, the target image 306 can be generated according to the at least two second image blocks (for example, second image blocks 306A to 306D).

[0123] In some embodiments, as Figure 3A shown, generating the target image according to the at least two second image blocks includes: stitching the at least two second image blocks (for example, second image blocks 306A to 306D) into the target image 306.

[0124] Figure 4E Shows a schematic diagram of a stitched image according to an embodiment of the present disclosure.

[0125] AsFigure 4E As shown, compared with Figure 4D Similarly, the shaded parts of each of the second image blocks 306A - 306D show the overlapping regions of adjacent second image blocks. When splicing, the overlapping regions can be overlapped with each other to complete the splicing. Correspondingly, the overlapping regions form the splicing parts, and the shaded part in the target image 306 shows this splicing part.

[0126] For the pixel points in the splicing part, according to which second image blocks' overlapping regions the pixel point is located in, weighted processing is performed to obtain a weighted pixel value as the final pixel value of this pixel point.

[0127] For example, as Figure 4E shown, for a splicing part where only two second image blocks overlap, the weights of the pixel value of this pixel point in each second image block can be set to 0.5, and then weighted summation is performed to obtain the weighted pixel value. For a splicing part where four second image blocks overlap, the weights of the pixel value of this pixel point in each second image block can be set to 0.25, and then weighted summation is performed to obtain the weighted pixel value.

[0128] Thus, the generation process of the target image is completed.

[0129] It can be seen from the above embodiments that the image processing method provided by the embodiments of the present disclosure significantly improves the processing speed of the algorithm through the block optimization scheme.

[0130] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of these multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0131] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] The embodiments of the present disclosure also provide a computer device for implementing the above - mentioned method 200. Figure 5 The hardware structure diagram of an exemplary computer device 500 provided by the embodiments of the present disclosure is shown. The computer device 500 can be used to implementFigure 1 The server 106, servers 106A and 106B can also be used to implement Figure 1 the terminal devices 102 and 104. In some scenarios, the computer device 500 can also be used to implement Figure 1 the database server 108.

[0133] Such as Figure 5 shown, the computer device 500 may include: a processor 502, a memory 504, a network module 506, a peripheral interface 508, and a bus 510. Among them, the processor 502, the memory 504, the network module 506, and the peripheral interface 508 are communicatively connected to each other inside the computer device 500 through the bus 510.

[0134] The processor 502 can be a central processing unit (CPU), an image processor, a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. The processor 502 can be used to execute functions related to the technologies described in this disclosure. In some embodiments, the processor 502 may further include multiple processors integrated as a single logic component. For example, as Figure 5 shown, the processor 502 may include multiple processors 502a, 502b, and 502c.

[0135] The memory 504 can be configured to store data (such as instructions, computer code, etc.). As Figure 5 shown, the data stored in the memory 504 may include program instructions (such as program instructions for implementing the method 200 of the embodiments of this disclosure) and data to be processed (such as configuration files of other modules that the memory can store). The processor 502 can also access the program instructions and data stored in the memory 504 and execute the program instructions to operate on the data to be processed. The memory 504 may include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 504 may include a random access memory (RAM), a read-only memory (ROM), an optical disc, a magnetic disk, a hard disk, a solid state drive (SSD), a flash memory, a memory stick, etc.

[0136] The network interface 506 can be configured to provide communication with other external devices to the computer device 500 via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (such as Bluetooth, WiFi, Near Field Communication (NFC), etc.), a cellular network, the Internet, or a combination of the above. It can be understood that the type of the network is not limited to the above specific examples.

[0137] The peripheral interface 508 can be configured to connect the computer device 500 to one or more peripheral devices to enable information input and output. For example, the peripheral devices can include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, various sensors, etc., and output devices such as a display, a speaker, a vibrator, an indicator light, etc.

[0138] The bus 510 can be configured to transmit information between various components of the computer device 500 (such as the processor 502, the memory 504, the network interface 506, and the peripheral interface 508), such as an internal bus (such as a processor - memory bus), an external bus (USB port, PCI - E bus), etc.

[0139] It should be noted that although the architecture of the above - mentioned computer device 500 only shows the processor 502, the memory 504, the network interface 506, the peripheral interface 508, and the bus 510, in the specific implementation process, the architecture of the computer device 500 may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the architecture of the above - mentioned computer device 500 may also only include the components necessary to implement the solution of the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0140] The embodiments of the present disclosure also provide an image processing device. Figure 6 The schematic diagram of an exemplary device 600 provided by the embodiments of the present disclosure is shown. As Figure 6 shown, the device 600 can be used to implement the method 200 and can further include the following modules.

[0141] An acquisition module 602, configured to: acquire an image to be processed;

[0142] A splitting module 604, configured to: based on the image to be processed, split to obtain at least two first image blocks;

[0143] A processing module 606, configured to: input the at least two first image blocks into an image processing model and output at least two second image blocks;

[0144] A generation module 608, configured to: generate a target image according to the at least two second image blocks.

[0145] In some embodiments, the processing module 606 is configured to:

[0146] For the first image block, perform the following steps:

[0147] Based on the first image block, split it into at least two first sub-image blocks;

[0148] Encode the at least two first sub-image blocks respectively to obtain at least two first image features corresponding one-to-one to the at least two first sub-image blocks;

[0149] Generate a second image block corresponding to the first image block according to the at least two first image features.

[0150] In some embodiments, the processing module 606 is configured to:

[0151] Concatenate the at least two first image features into a second image feature;

[0152] Resample the second image feature to obtain a third image feature;

[0153] Generate a second image block corresponding to the first image block according to the third image feature.

[0154] In some embodiments, the processing module 606 is configured to:

[0155] Split the third image feature into at least two fourth image features;

[0156] Divide the at least two fourth image features into at least two first image feature groups, and each of the first image feature groups includes a preset number of the fourth image features;

[0157] Process the at least two first image feature groups in sequence to obtain at least two second image feature groups corresponding one-to-one to the at least two first image feature groups, and each of the second image feature groups includes a preset number of fifth image features;

[0158] Concatenate the fifth image features in the at least two second image feature groups into a sixth image feature;

[0159] Generate a second image block corresponding to the first image block according to the sixth image feature.

[0160] In some embodiments, the processing module 606 is configured to:

[0161] Split the sixth image feature into at least two seventh image features;

[0162] Divide the at least two seventh image features into at least two third image feature groups, where each of the third image feature groups includes a preset number of the seventh image features;

[0163] Decode the at least two third image feature groups respectively to obtain at least two second sub-image blocks that correspond one-to-one to the seventh image features in the at least two third image feature groups;

[0164] Stitch the at least two second sub-image blocks together to form a second image block corresponding to the first image block.

[0165] In some embodiments, the processing module 606 is configured to:

[0166] Overlap the overlapping regions of two adjacent second sub-image blocks with each other to complete the stitching, and the overlapping regions form stitching parts;

[0167] For the pixel points in the stitching parts, perform weighted processing on the pixel values of the pixel points in the two adjacent second sub-image blocks to obtain weighted pixel values, and use the weighted pixel values as the pixel values of the pixel points.

[0168] In some embodiments, the processing module 606 is configured to:

[0169] Use a diffusion model to process the at least two first image feature groups in sequence to obtain at least two second image feature groups that correspond one-to-one to the at least two first image feature groups.

[0170] In some embodiments, the generating module 608 is configured to: stitch the at least two second image blocks together to form the target image.

[0171] For the sake of convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0172] The device in the above embodiments is used to implement the corresponding method 200 in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0173] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method 200 described in any of the foregoing embodiments.

[0174] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0175] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the method 200 described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0176] Based on the same inventive concept, corresponding to the method 200 in any of the above embodiments, the present disclosure also provides a computer program product, which includes a computer program. In some embodiments, the computer program is executable by one or more processors to cause the processors to execute the method 200. Corresponding to the execution subject of each step in each embodiment of the method 200, the processor that executes the corresponding step can belong to the corresponding execution subject.

[0177] The computer program product of the above embodiment is used to cause the processor to execute the method 200 described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0178] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.

[0179] In addition, for simplicity of explanation and discussion, and so as not to make the embodiments of the present disclosure difficult to understand, well-known power / ground connections of integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be considered illustrative rather than restrictive.

[0180] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0181] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An image processing method, comprising: obtaining an image to be processed; based on the image to be processed, splitting to obtain at least two first image blocks; inputting the at least two first image blocks into an image processing model, and outputting at least two second image blocks; generating a target image according to the at least two second image blocks.

2. The method according to claim 1, wherein inputting the at least two first image blocks into at least two image processing models respectively, and outputting at least two second image blocks, comprising: for the first image block, performing the following steps: based on the first image block, splitting to obtain at least two first sub-image blocks; encoding the at least two first sub-image blocks respectively to obtain at least two first image features corresponding one-to-one to the at least two first sub-image blocks; generating a second image block corresponding to the first image block according to the at least two first image features.

3. The method according to claim 2, wherein generating a second image block corresponding to the first image block according to the at least two first image features, comprising: concatenating the at least two first image features into a second image feature; performing resampling on the second image feature to obtain a third image feature; generating a second image block corresponding to the first image block according to the third image feature.

4. The method according to claim 3, wherein generating a second image block corresponding to the first image block according to the third image feature, comprising: splitting the third image feature into at least two fourth image features; dividing the at least two fourth image features into at least two first image feature groups, each of the first image feature groups including a preset number of the fourth image features; processing the at least two first image feature groups in sequence to obtain at least two second image feature groups corresponding one-to-one to the at least two first image feature groups, each of the second image feature groups including a preset number of fifth image features; concatenating the fifth image features in the at least two second image feature groups into a sixth image feature; generating a second image block corresponding to the first image block according to the sixth image feature.

5. The method according to claim 4, wherein generating a second image block corresponding to the first image block according to the sixth image feature, comprising: splitting the sixth image feature into at least two seventh image features; dividing the at least two seventh image features into at least two third image feature groups, each of the third image feature groups including a preset number of the seventh image features; decoding the at least two third image feature groups respectively to obtain at least two second sub-image blocks corresponding one-to-one to the seventh image features in the at least two third image feature groups; concatenating the at least two second sub-image blocks into a second image block corresponding to the first image block.

6. The method according to claim 4, wherein concatenating the at least two second sub-image blocks into a second image block corresponding to the first image block, comprising: Overlap the overlapping regions of two adjacent second sub-image blocks with each other to complete splicing, and the overlapping regions form splicing parts; For the pixel points in the splicing parts, perform weighted processing on the pixel values of the pixel points in the two adjacent second sub-image blocks to obtain weighted pixel values, and use the weighted pixel values as the pixel values of the pixel points.

7. The method according to claim 4, wherein, Processing the at least two first image feature groups in sequence to obtain at least two second image feature groups corresponding one-to-one to the at least two first image feature groups, including: Using a diffusion model, processing the at least two first image feature groups in sequence to obtain at least two second image feature groups corresponding one-to-one to the at least two first image feature groups.

8. The method according to claim 1, wherein, Generating a target image according to the at least two second image blocks, including: Splicing the at least two second image blocks into the target image.

9. An image processing apparatus, comprising: An acquisition module configured to acquire an image to be processed; A splitting module configured to split, based on the image to be processed, into at least two first image blocks; A processing module configured to input the at least two first image blocks into an image processing model and output at least two second image blocks; A generating module configured to generate a target image according to the at least two second image blocks.

10. A computer device comprising one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the programs include instructions for executing the method according to any one of claims 1-8.

11. A non-volatile computer-readable storage medium containing a computer program, which when executed by one or more processors causes the processors to execute the method according to any one of claims 1-8.