Image processing method, device, storage medium and computer program product

By segmenting the image and setting optimization conditions, the problem of difficulty in extracting details when processing high-resolution images is solved, and image processing with higher quality and efficiency is achieved.

CN119559198BActive Publication Date: 2025-05-13INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510109129.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Existing artificial intelligence models are limited in resolution when processing complex visual tasks, making it difficult for the model to effectively extract image details, which in turn affects the image processing quality.

Method used

By segmenting the image according to the resolution of the image to be processed and the resolution of the visual encoder, optimization conditions are set to reduce the aspect ratio deviation of the sub-graph and constrain the sub-graph area difference value, thereby determining the appropriate segmentation method.

Benefits of technology

It effectively avoids the deformation of the image when adapting to the resolution of the visual encoder, reduces the calculation amount and information loss, and improves the quality and efficiency of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559198B_ABST
    Figure CN119559198B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and specifically discloses an image processing method, device, storage medium and computer program product. When segmenting an image to be processed according to the resolution of the image to be processed and the resolution of a visual encoder, a first optimization condition for reducing the deviation between the aspect ratio of a first sub-image and the aspect ratio of a second sub-image and a second optimization condition for constraining the area of ​​each second sub-image and the difference with the area of ​​the image to be processed are set, thereby ensuring that the aspect ratio of the second sub-image input to the visual encoder after segmentation is as small as possible compared with the aspect ratio of the image to be processed to avoid large deformation when adapting to the resolution of the visual encoder, and avoiding excessive calculation amount and loss of local integrity of the image due to too large a number of segmentations, thereby ensuring the processing quality of the image processing task of artificial intelligence from the perspective of quality and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image processing method, device, storage medium and computer program product. Background Art

[0002] Image processing technology is an important branch of artificial intelligence technology. It realizes visual processing of artificial intelligence technology by encoding the image features of the input image through a visual encoder and then inputting the image into the model for calculation, or combining multimodal features for calculation.

[0003] At present, artificial intelligence models are limited by resolution when processing complex visual tasks. That is, the visual encoder has an upper limit on its resolution, so the model is often designed to process input images of fixed size, which limits the model's ability to extract details from the image and leads to the loss of important information. In order to ensure the processing of details of the original image, the commonly used method is to split the original image into multiple sub-images that adapt to the resolution of the visual encoder for processing.

[0004] When segmenting the original image, how to determine the segmentation method affects the increase in model calculations and the quality of image processing.

[0005] How to determine the appropriate segmentation method for the original image in the image processing task to ensure the processing performance of the artificial intelligence image processing task from the perspective of quality and efficiency is a technical problem that technical personnel in this field need to solve. Summary of the invention

[0006] The object of the present invention is to provide an image processing method, device, storage medium and computer program product for ensuring the processing performance of artificial intelligence image processing tasks from the perspective of quality and efficiency.

[0007] In order to solve the above technical problems, the present invention provides an image processing method, comprising:

[0008] Get the image to be processed;

[0009] According to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition, the image to be processed is segmented to obtain a plurality of first sub-images corresponding to the image to be processed;

[0010] Processing the resolution of the first sub-image to the resolution of the visual encoder to obtain a second sub-image;

[0011] Inputting each of the second sub-images into the visual encoder for encoding and then into the target model to obtain an image processing result;

[0012] Among them, the first optimization condition is an optimization condition for reducing the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image, and the second optimization condition is an optimization condition for constraining the difference between the area of ​​each second sub-image and the area of ​​the image to be processed.

[0013] On the one hand, according to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition, the image to be processed is segmented to obtain a plurality of first sub-images corresponding to the image to be processed, including:

[0014] Determine, according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition, and the second optimization condition, a first number of sub-graphs for segmenting the image to be processed in a long dimension and a second number of sub-graphs for segmenting the image to be processed in a wide dimension;

[0015] The image to be processed is evenly divided without overlap according to the first number and the second number to obtain a plurality of the first sub-images corresponding to the image to be processed.

[0016] On the other hand, according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, determining to obtain a first number of sub-graphs for segmenting the image to be processed in the long dimension and a second number of sub-graphs for segmenting the image to be processed in the wide dimension, including:

[0017] Determining the number of sub-images according to the resolution of the image to be processed and the resolution of the visual encoder;

[0018] Determine a combination of the first number and the second number according to the number of sub-graphs, so that the product of the first number and the second number in the same group is the number of sub-graphs;

[0019] By using the first optimization condition and the second optimization condition, selecting the actual first quantity and the actual second quantity from the combination of the first quantity and the second quantity;

[0020] The image to be processed is subjected to uniform non-overlapping segmentation processing according to the first number and the second number to obtain a plurality of first sub-images corresponding to the image to be processed, including:

[0021] The image to be processed is subjected to non-overlapping uniform segmentation processing according to the actual first number and the actual second number, so as to obtain a plurality of the first sub-images corresponding to the image to be processed.

[0022] On the other hand, the first optimization condition is represented by an optimization objective function; the optimization goal of the optimization objective function is to minimize the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image.

[0023] On the other hand, the optimization objective function is to maximize the function value of the segmentation scoring function;

[0024] The segmentation scoring function is the reciprocal of the absolute value of the difference between the first logarithm and the second logarithm; the numerator of the real number of the first logarithm is the product of the width of the image to be processed and the second number, and the denominator of the real number of the first logarithm is the product of the length of the image to be processed and the first number; the real number of the second logarithm is the ratio of the width of the second sub-image to the length of the second sub-image.

[0025] On the other hand, the second optimization condition is that: the ratio of the difference between the area of ​​each of the second sub-images minus the area of ​​the image to be processed and the area of ​​the image to be processed is less than a first threshold;

[0026] According to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, calculating a first number of sub-graphs for segmenting the image to be processed in the long dimension and a second number of sub-graphs for segmenting the image to be processed in the wide dimension, including:

[0027] An optimization calculation is performed according to the optimization objective function and the second optimization condition to obtain the first quantity and the second quantity.

[0028] On the other hand, each of the second sub-images is input into the visual encoder for encoding and then input into the target model to obtain an image processing result, including:

[0029] Processing the resolution of the image to be processed to the resolution of the visual encoder to obtain a first image;

[0030] Inputting the first image into the visual encoder to obtain a first visual feature code;

[0031] Inputting the second sub-image into the visual encoder to obtain a second visual feature code;

[0032] After performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained;

[0033] The aggregated visual features are input into the target model to obtain the image processing result.

[0034] On the other hand, according to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition, the image to be processed is segmented to obtain a plurality of first sub-images corresponding to the image to be processed, including:

[0035] According to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, performing multiple rounds of segmentation processing on the image to be processed to obtain multiple groups of the first sub-images of different scales;

[0036] Inputting each of the second sub-images into the visual encoder for encoding and then inputting into the target model to obtain an image processing result, including:

[0037] Inputting the second sub-images corresponding to the first sub-images of each scale into the visual encoder respectively to obtain the second visual feature codes corresponding to the second sub-images of each scale;

[0038] After performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained, including:

[0039] The visual feature codes corresponding to the images of adjacent scales in the first image and the second sub-image of each scale are subjected to feature aggregation processing, and the obtained feature aggregation results are spliced ​​to obtain the aggregated visual features.

[0040] On the other hand, the number of rounds of multiple segmentation processes performed on the image to be processed is determined according to the resolution of the image to be processed.

[0041] On the other hand, after performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained, including:

[0042] Compressing the first visual feature code and the second visual feature code respectively to obtain the compressed first visual feature code and the compressed second visual feature code;

[0043] The compressed first visual feature code and the compressed second visual feature code are aggregated to obtain the aggregated visual feature.

[0044] On the other hand, after performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained, including:

[0045] The first visual feature encoding and the second visual feature encoding are cross-attention calculated to obtain the aggregated visual feature.

[0046] On the other hand, inputting the aggregated visual features into the target model to obtain the image processing result includes:

[0047] Compressing the aggregated visual features to obtain compressed aggregated visual features;

[0048] The compressed aggregated visual features are input into the target model to obtain the image processing result.

[0049] On the other hand, the image to be processed is a document image to be processed;

[0050] The aggregated visual features are compressed to obtain the compressed aggregated visual features, including:

[0051] Using a 1×a convolution kernel to compress the aggregated visual features to obtain compressed aggregated visual features;

[0052] Wherein, a is a positive integer greater than 1.

[0053] On the other hand, each of the second sub-images is input into the visual encoder for encoding and then input into the target model to obtain an image processing result, including:

[0054] Processing the resolution of the image to be processed to the resolution of the visual encoder to obtain a first image;

[0055] Inputting the first image into the visual encoder to obtain a first visual feature code;

[0056] Inputting the second sub-image into the visual encoder to obtain a second visual feature code;

[0057] After performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained;

[0058] Inputting the input text corresponding to the image to be processed into the text encoder of the target model to obtain text feature encoding;

[0059] The aggregated visual features and the text features are encoded and input into the target model to obtain the image processing result.

[0060] In order to solve the above technical problems, the present invention further provides an image processing device, comprising:

[0061] Memory for storing computer programs;

[0062] A processor is used to execute the computer program, and when the computer program is executed by the processor, the steps of any one of the image processing methods described above are implemented.

[0063] In order to solve the above technical problem, the present invention further provides a non-volatile storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the image processing method as described in any one of the above items are implemented.

[0064] In order to solve the above technical problem, the present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above image processing methods are implemented.

[0065] The image processing method provided by the present invention has the beneficial effect that when segmenting the image to be processed according to the resolution of the image to be processed and the resolution of the visual encoder, a first optimization condition for reducing the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image and a second optimization condition for constraining the area of ​​each second sub-image and the difference with the area of ​​the image to be processed are set, thereby ensuring that the aspect ratio of the second sub-image input to the visual encoder after segmentation is as small as possible compared with the aspect ratio of the image to be processed to avoid large deformation when adapting to the resolution of the visual encoder, and avoiding excessive number of segmentations resulting in excessive calculation and loss of local integrity of the image, thereby ensuring the processing quality of artificial intelligence image processing tasks from the perspective of quality and efficiency.

[0066] The image processing device, non-volatile storage medium and computer program product provided by the present invention have the above-mentioned beneficial effects, which will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0068] Figure 1 A flowchart of an image processing method provided by an embodiment of the present invention;

[0069] Figure 2 A schematic diagram of visual encoding without image segmentation provided by an embodiment of the present invention;

[0070] Figure 3 A schematic diagram of a first visual encoding method for image segmentation provided by an embodiment of the present invention;

[0071] Figure 4 A schematic diagram of a second visual encoding method for image segmentation provided by an embodiment of the present invention;

[0072] Figure 5A schematic diagram of multimodal image processing provided by an embodiment of the present invention;

[0073] Figure 6 A schematic diagram of the structure of an image processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0074] The core of the present invention is to provide an image processing method, device, storage medium and computer program product for ensuring the processing performance of artificial intelligence image processing tasks from the perspective of quality and efficiency.

[0075] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0076] To facilitate understanding of the technical solution provided by the embodiment of the present invention, some key terms used in the embodiment of the present invention are first explained here.

[0077] The image to be processed targeted by the embodiment of the present invention may also be referred to as an original image, which may be an image captured by a visual sensor such as a camera, or a video frame captured from a captured video.

[0078] Resolution is used to describe the clarity of an image, video or display device, and it indicates the number of pixels that can be displayed within a unit length. In various embodiments of the present invention, the resolution can be expressed in the form of the product of the number of pixels in the horizontal direction (also known as the image width W) and the number of pixels in the vertical direction (also known as the image height H). For example, 1920×1080 indicates that there are 1920 pixels in the horizontal direction and 1080 pixels in the vertical direction.

[0079] The visual encoder is a component used to process visual information. It converts image data into a format that can be input into an image processing model for calculation. Since the visual encoder is usually optimized based on images of a specific resolution during training, the visual encoder can only process input images of a fixed resolution. When the resolution of the image to be processed does not match the resolution of the visual encoder, the image to be processed needs to be processed to the resolution of the visual encoder before being input into the visual encoder. In this case, for images to be processed with a resolution higher than the resolution of the visual encoder, image detail information may be lost due to image compression.

[0080] To this end, the relevant field proposes to divide the image to be processed into smaller sub-images to be input into the visual encoder for encoding, so as to minimize the loss of image detail information when processing high-resolution images. In the segmentation process, the image to be processed is evenly and non-overlappingly divided into m rows and n columns of sub-images. If the resolution of the sub-image is smaller than the resolution of the visual encoder, the sub-image is upsampled to reach the resolution of the visual encoder and then input into the visual encoder for processing; if the resolution of the sub-image is larger than the resolution of the visual encoder, the sub-image is compressed to reach the resolution of the visual encoder and then input into the visual encoder for processing.

[0081] In the above segmentation process, a to-be-processed image is processed into multiple sub-images and then encoded, which obviously increases the computational complexity of the visual encoder and the image processing model. The higher the resolution of the to-be-processed image, the higher the computing resource consumption. If the to-be-processed image is only segmented into a few sub-images, resulting in the size of the segmented sub-images still exceeding the resolution limit of the visual encoder, then the sub-images need to be compressed before they can be input into the visual encoder, which may still result in loss of image detail information. In addition, if the aspect ratio of the sub-image is inconsistent with the resolution of the visual encoder, the aspect ratio of the sub-image will be deformed after the resolution of the sub-image is processed to the resolution of the visual encoder.

[0082] Therefore, the way to divide the image to be processed into sub-images affects the quality and efficiency of the image processing task.

[0083] In order to determine a suitable segmentation method for segmenting an image to be processed into sub-images so as to ensure the processing performance of an artificial intelligence image processing task from the perspectives of quality and efficiency, an image processing scheme provided by an embodiment of the present invention sets a first optimization condition for reducing the deviation between the aspect ratio of a first sub-image and the aspect ratio of a second sub-image and a second optimization condition for constraining the area of ​​each second sub-image and the difference with the area of ​​the image to be processed when segmenting the image to be processed according to the resolution of the image to be processed and the resolution of the visual encoder, thereby ensuring that the aspect ratio of the second sub-image input to the visual encoder after segmentation is as small as possible compared with the aspect ratio of the image to be processed to avoid large deformation when adapting to the resolution of the visual encoder, and avoiding excessive number of segmentations resulting in excessive computational complexity and loss of local image integrity, thereby ensuring the processing quality of the artificial intelligence image processing task from the perspectives of quality and efficiency.

[0084] The image processing method provided by the embodiment of the present invention is described below with reference to the accompanying drawings.

[0085] Figure 2 A schematic diagram of visual encoding without image segmentation provided by an embodiment of the present invention; Figure 3 A schematic diagram of the first visual encoding method for image segmentation provided in an embodiment of the present invention.

[0086] Figure 1 The present invention provides a flowchart of an image processing method.

[0087] like Figure 1 As shown, the image processing method provided by the embodiment of the present invention includes:

[0088] S101: Acquire an image to be processed;

[0089] S102: performing segmentation processing on the image to be processed according to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition to obtain a plurality of first sub-images corresponding to the image to be processed;

[0090] S103: Processing the resolution of the first sub-image into the resolution of the visual encoder to obtain a second sub-image;

[0091] S104: input each second sub-image into a visual encoder for encoding and then input into a target model to obtain an image processing result;

[0092] The first optimization condition is an optimization condition for reducing the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image, and the second optimization condition is an optimization condition for constraining the difference between the area of ​​each second sub-image and the area of ​​the image to be processed.

[0093] In the embodiment of the present invention, the image to be processed may be an image captured by a visual sensor such as a camera, or a video frame captured from a captured video. The target model may be an image processing model, or a multimodal model including an image processing model.

[0094] The image processing method provided by the embodiment of the present invention can be applied to the model training task of the image processing model, and can also be applied to the reasoning task of the image processing model. The image processing method provided by the embodiment of the present invention can be implemented based on one or more computing devices. When encoding and model calculating the multiple second sub-images corresponding to the image to be processed, multi-threading or multi-device parallel execution can be used to improve efficiency, and a centralized controller is set to summarize the image processing results corresponding to a single image to be processed.

[0095] For S101, the image to be processed is obtained, which may be obtained from a device that captures the original image, or the image to be processed provided by a user.

[0096] For S102, a segmentation method for the image to be processed is determined according to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition, and the image to be processed is segmented according to the segmentation method to obtain multiple first sub-images corresponding to the image to be processed.

[0097] The division method of the processed image can be represented by the number of first sub-images, or by the first number of sub-images for dividing the processed image in the long dimension and the second number of sub-images for dividing the processed image in the wide dimension.

[0098] When determining the slicing method of the sub-images to be segmented according to the resolution of the image to be processed and the resolution of the visual encoder of the target model, the embodiment of the present invention sets a first optimization condition to reduce the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image, and sets a second optimization condition to constrain the area of ​​each second sub-image and the difference with the area of ​​the image to be processed.

[0099] In an embodiment of the present invention, for an image to be processed with unknown image information distribution, a uniform non-overlapping segmentation method can be used. That is, in S102, the image to be processed is segmented according to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition, and the second optimization condition to obtain a plurality of first sub-images corresponding to the image to be processed, which may include: determining the first number of sub-images segmented in the long dimension and the second number of sub-images segmented in the wide dimension of the image to be processed according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition, and the second optimization condition; performing non-overlapping uniform segmentation on the image to be processed according to the first number and the second number to obtain a plurality of first sub-images corresponding to the image to be processed.

[0100] It can be understood that if the length and width of the image to be processed are integer multiples of the length and width of the resolution of the visual encoder, that is, they can be "divided", then the number of sub-images in the vertical direction can be determined by dividing the length value of the image to be processed by the length value of the resolution of the visual encoder, and the number of sub-images in the horizontal direction can be determined by dividing the width value of the image to be processed by the width value of the resolution of the visual encoder. At this time, the resolution of the first sub-image obtained is exactly the resolution of the visual encoder, which is an ideal segmentation scheme.

[0101] The image processing method provided by the embodiment of the present invention may also include: if the image to be processed can be divided evenly and without overlap into multiple first sub-images with the same resolution as the visual encoder, then each first sub-image is directly input into the visual encoder for encoding and then input into the target model to obtain the image processing result.

[0102] However, in actual situations, the above-mentioned "division" condition is often not met. Therefore, in S103, the first sub-image needs to be processed into a second sub-image with the same resolution as the visual encoder. For the first sub-image with a resolution lower than that of the visual encoder, its resolution can be aligned with the visual encoder by upsampling to obtain the second sub-image. For the first sub-image with a resolution higher than that of the visual encoder, its resolution can be aligned with the visual encoder by compression to obtain the second sub-image.

[0103] At this time, if the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image is large, the second sub-image will be deformed compared with the first sub-image, that is, the image formed by splicing the second sub-images will be deformed compared with the image to be processed, which will also cause the loss of the original image information of the image to be processed. Therefore, when performing the segmentation process in S102, the first optimization condition is used to reduce the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image.

[0104] However, if we only consider reducing the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image, that is, pursuing "division", it may lead to the result that more sub-images are screened out in the optimization calculation. At this time, although the aspect ratio of the smaller-sized first sub-image obtained after segmentation is closer to the resolution of the visual encoder, it will bring a significant increase in the amount of calculation and damage the local integrity of the image.

[0105] Therefore, when determining the segmentation method in S102, the second optimization condition is also used to constrain the difference between the area of ​​each second sub-image and the area of ​​the image to be processed, that is, the difference between the area of ​​each second sub-image and the area of ​​the image to be processed cannot be too large, so as to avoid the above-mentioned problem of restoring the aspect ratio but bringing additional calculation and destroying the local integrity of the image.

[0106] In some optional implementations of the embodiments of the present invention, determining the first number of sub-graphs for segmenting the image to be processed in the long dimension and the second number of sub-graphs for segmenting the image to be processed in the wide dimension according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition may include: performing optimization calculations according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition to obtain the first number and the second number. That is, the parameters of the optimization calculation may be set according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, and the first number and the second number may be solved by the optimization calculation to determine the segmentation method.

[0107] For S104, the second sub-image after segmentation and resolution adjustment is input into the visual encoder for encoding, and then input into the target model for processing to obtain an image processing result.

[0108] The image processing method provided by an embodiment of the present invention sets a first optimization condition for reducing the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image and a second optimization condition for constraining the area of ​​each second sub-image and the difference with the area of ​​the image to be processed when segmenting the image to be processed according to the resolution of the image to be processed and the resolution of the visual encoder, thereby ensuring that the aspect ratio of the second sub-image input to the visual encoder after segmentation is as small as possible compared with the aspect ratio of the image to be processed to avoid large deformation when adapting to the resolution of the visual encoder, and avoiding excessive number of segmentations resulting in excessive calculation and loss of local image integrity, thereby ensuring the processing quality of artificial intelligence image processing tasks from the perspectives of quality and efficiency.

[0109] In the above embodiment, the embodiment of the present invention determines the segmentation method of the image to be processed by setting the first optimization condition and the second optimization condition based on the resolution of the image to be processed and the resolution of the visual encoder of the target model, and can set the parameters of the optimization calculation according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, and solves the first number and the second number through the optimization calculation to determine the segmentation method.

[0110] In some other optional implementations of the embodiments of the present invention, the number of sub-images can be determined according to the resolution of the image to be processed and the resolution of the visual encoder, and then the actual segmentation method can be screened out according to the first optimization condition and the second optimization condition, thereby speeding up the efficiency of the optimization solution.

[0111] Then, according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, determining the first number of sub-images for dividing the image to be processed in the long dimension and the second number of sub-images for dividing the image to be processed in the wide dimension may include: determining the number of sub-images according to the resolution of the image to be processed and the resolution of the visual encoder; determining the combination of the first number and the second number according to the number of sub-images, so that the product of the first number and the second number in the same group is the number of sub-images; through the first optimization condition and the second optimization condition, screening out the actual first number and the actual second number from the combination of the first number and the second number. Performing a non-overlapping uniform segmentation process on the image to be processed according to the first number and the second number to obtain a plurality of first sub-images corresponding to the image to be processed may include: performing a non-overlapping uniform segmentation process on the image to be processed according to the actual first number and the actual second number to obtain a plurality of first sub-images corresponding to the image to be processed.

[0112] The number of sub-images is determined according to the resolution of the image to be processed and the resolution of the visual encoder, which may include: determining the total number of pixels of the image to be processed according to the resolution of the image to be processed, determining the total number of pixels of the visual encoder according to the resolution of the visual encoder, and obtaining the number of sub-images by dividing the total number of pixels of the image to be processed by the total number of pixels of the visual encoder and rounding up. The resolution of the image to be processed is , the resolution of the visual encoder is , then the above process can be expressed by the following formula:

[0113] ;

[0114] in, is the number of subgraphs, is the number of pixels in the horizontal direction of the resolution of the image to be processed, is the number of pixels in the vertical direction of the resolution of the image to be processed, is the number of pixels in the horizontal direction of the resolution of the visual encoder, is the number of pixels in the vertical direction of the resolution of the visual encoder, Indicates rounding up.

[0115] Determine the combination of the first number and the second number according to the number of sub-graphs, so that the product of the first number and the second number in the same group is the number of sub-graphs, which can be expressed by the following formula:

[0116] ;

[0117] in, Indicates a combination, Represents a set of segmentation methods, is the first quantity, is the second quantity, is the number of sub-graphs.

[0118] In the embodiment of the present invention, the first optimization condition may be represented by an optimization objective function; the optimization goal of the optimization objective function may be to minimize the deviation between the aspect ratio of the image to be processed and the aspect ratio of the second sub-image.

[0119] In some optional implementations of the embodiments of the present invention, the optimization objective function may be a function value of maximizing a segmentation scoring function. The segmentation scoring function is the inverse of the absolute value of the difference between the first logarithm and the second logarithm; the numerator of the real number of the first logarithm is the product of the width of the image to be processed and the second number, and the denominator of the real number of the first logarithm is the product of the length of the image to be processed and the first number; the real number of the second logarithm is the ratio of the width of the second sub-image to the length of the second sub-image.

[0120] The segmentation scoring function can be expressed as:

[0121] ;

[0122] The optimization objective function can be expressed as:

[0123] ;

[0124] in, Indicates the split method The score, is the number of pixels in the horizontal direction of the resolution of the image to be processed, is the number of pixels in the vertical direction of the resolution of the image to be processed, is the number of pixels in the horizontal direction of the resolution of the visual encoder, is the number of pixels in the vertical direction of the resolution of the visual encoder, is the first quantity, is the second quantity, represents the base 10 logarithm, Indicates the split method that determines the maximum score in the combination, represents the actual first quantity, Indicates the actual second quantity.

[0125] It can be understood that in addition to taking minimizing the deviation between the aspect ratio of the image to be processed and the aspect ratio of the second sub-image as the optimization goal, other forms of optimization objective functions can also be set. For example, the segmentation scoring function can be set to be positively correlated with the deviation between the aspect ratio of the image to be processed and the aspect ratio of the second sub-image, and the optimization objective function can be set to be a function with the goal of minimizing the segmentation scoring function.

[0126] The first optimization condition is used to constrain the difference between the area of ​​each second sub-image and the area of ​​the image to be processed.

[0127] In some optional implementations of the embodiments of the present invention, the second optimization condition may be: the ratio of the difference between the area of ​​each second sub-image and the area of ​​the image to be processed and the area of ​​the image to be processed is less than a first threshold. According to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, calculating the first number of sub-images for segmenting the image to be processed in the long dimension and the second number of sub-images for segmenting the image to be processed in the wide dimension may include: performing optimization calculations according to the optimization objective function and the second optimization condition to obtain the first number and the second number.

[0128] The second optimization condition can be expressed as follows:

[0129] ;

[0130] in, is the number of pixels in the horizontal direction of the resolution of the image to be processed, is the number of pixels in the vertical direction of the resolution of the image to be processed, is the number of pixels in the horizontal direction of the resolution of the visual encoder, is the number of pixels in the vertical direction of the resolution of the visual encoder, is the first quantity, is the second quantity, Indicates the first threshold.

[0131] First threshold is a hyperparameter and can be obtained through testing. The step of determining the first threshold may include: using image samples of the same resolution and type as the image to be processed to perform segmentation at different scales and then inputting them into the visual encoder and the target model for calculation, determining the target segmentation scale range according to the image processing results corresponding to each segmentation scale, and determining the first threshold according to the target segmentation scale range.

[0132] In some other optional implementations of the embodiments of the present invention, the second constraint condition may also be: the difference between the area of ​​each second sub-image and the area of ​​the image to be processed is less than a second threshold.

[0133] Then, through the first optimization condition and the second optimization condition introduced above, the actual first quantity and the actual second quantity are screened out from the combination of the first quantity and the second quantity through the first optimization condition and the second optimization condition, which can include: using the first optimization condition to score the combination and determine that the first quantity and the second quantity in the combination that meets the second optimization condition and has the highest score are the actual first quantity and the actual second quantity.

[0134] In addition to the above-mentioned method of performing joint optimization calculation according to the first optimization condition and the second optimization condition, in some other optional implementations of the embodiments of the present invention, the actual first quantity and the actual second quantity are screened out from the combination of the first quantity and the second quantity through the first optimization condition and the second optimization condition, which may also include: determining the scores corresponding to the combinations of the first quantity and the second quantity according to the first optimization condition, taking the first three numbers of combinations as candidate combinations in the order of the deviation of the aspect ratio of the first sub-image and the aspect ratio of the second sub-image corresponding to the scores from small to large, and selecting the final combination from the candidate combinations according to the second constraint condition to obtain the actual first quantity and the actual second quantity. Among them, the setting method of the first optimization condition can refer to the above description. The second optimization condition is to select the combination with the smallest difference between the area of ​​each second sub-image and the area of ​​the image to be processed from the candidate combination as the final combination.

[0135] Through the various implementations introduced in the embodiments of the present invention, the number of sub-images that divide the image to be processed from the width dimension and the length dimension respectively can be calculated, and the number of sub-images can make the converted second sub-image have a smaller deviation in aspect ratio and the overall image composed of each second sub-image compared to the image to be processed.

[0136] The above embodiment of the present invention introduces that a segmentation method of an image to be processed is determined by a first optimization condition for reducing the deviation between the aspect ratio of a first sub-image and the aspect ratio of a second sub-image and a second optimization condition for constraining the area of ​​each second sub-image and the difference between the area of ​​the image to be processed. In order to further optimize the image processing effect, in the image processing method provided in the embodiment of the present invention, each second sub-image is input into a visual encoder for encoding and then input into a target model to obtain an image processing result. It may also include: processing the resolution of the image to be processed into the resolution of the visual encoder to obtain a first image; inputting the first image into the visual encoder to obtain a first visual feature code; inputting the second sub-image into the visual encoder to obtain a second visual feature code; performing feature aggregation processing on the first visual feature code and the second visual feature code to obtain an aggregated visual feature; inputting the aggregated visual feature into the target model to obtain an image processing result.

[0137] In an embodiment of the present invention, a multi-view visual feature aggregation module may be designed to fuse multi-scale features to reduce model calculation overhead.

[0138] However, the segmentation method of the image to be processed determined by the first optimization condition and the second optimization condition may still lead to an unreasonable segmentation method. In order to reduce the probability of an unreasonable segmentation method, in the image processing method provided in the embodiment of the present invention, in S102, the image to be processed is segmented according to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition, to obtain multiple first sub-images corresponding to the image to be processed, which may include: according to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, the image to be processed is segmented for multiple rounds to obtain multiple groups of first sub-images of different scales. In S104, each second sub-image is input into the visual encoder for encoding and then input into the target model to obtain the image processing result, which may include: respectively inputting the second sub-image corresponding to the first sub-image of each scale into the visual encoder to obtain the second visual feature code corresponding to the second sub-image of each scale. At this time, after the first visual feature code and the second visual feature code are subjected to feature aggregation processing, the aggregated visual feature is obtained, which may include: the visual feature code corresponding to the image of the adjacent scale in the first image and the second sub-image of each scale is subjected to feature aggregation processing, and the obtained feature aggregation result is subjected to splicing processing to obtain the aggregated visual feature. That is, an embodiment of the present invention provides a multi-scale image segmentation solution, which segments the image to be processed at different scales, that is, segments the image again on the basis of the first segmentation, and aggregates the visual feature codes of each segmentation scale with the visual feature codes of the image to be processed to obtain aggregated visual features, so that the target model can observe the image to be processed from more scales (that is, at least three scales).

[0139] Figure 4 A schematic diagram of a second visual encoding method for image segmentation provided in an embodiment of the present invention.

[0140] like Figure 4 As shown in , if the image to be processed is segmented at two scales, the first sub-images corresponding to the two rounds of segmentation are obtained, and the first sub-image obtained by the second round of segmentation is a sub-image of the first sub-image obtained by the first round of segmentation. At this time, the second round of segmentation can be to subdivide the first sub-image segmented in the first round into half along each spatial dimension, thereby generating The first sub-image is processed into a second sub-image with the same resolution as the visual encoder as described in the above embodiment of the present invention, and used as the second sub-image corresponding to the second round of segmentation. Figure 4 In the segmentation method of , additional detail levels are considered through a second round of segmentation to capture detailed visual features through enlarged local resolution. Figure 4 The obtained two rounds of segmentation results and the image to be processed allow the target model to observe the input image from three different levels: global, and These images are then input into the visual encoder for feature extraction.

[0141] In the embodiment of the present invention, the number of rounds of multiple segmentation processing of the image to be processed can be determined according to the resolution of the image to be processed. The number of segmentation rounds can be positively correlated with the total number of pixels corresponding to the resolution of the image to be processed, that is, the higher the resolution, the more segmentation rounds.

[0142] After multiple rounds of segmentation, the visual feature encodings of adjacent scales are aggregated, and the aggregation results are concatenated to obtain the aggregated visual features.

[0143] In order to reduce the amount of calculation, the first visual feature code and the second visual feature code are subjected to feature aggregation processing to obtain the aggregated visual feature, which may include: respectively compressing the first visual feature code and the second visual feature code to obtain the compressed first visual feature code and the compressed second visual feature code; and aggregating the compressed first visual feature code and the compressed second visual feature code to obtain the aggregated visual feature. That is, before the visual features are aggregated, the visual features of each scale are compressed, thereby reducing the amount of calculation of the aggregation calculation.

[0144] In order to reduce the amount of calculation, the aggregated visual features are input into the target model to obtain the image processing result, which may include: compressing the aggregated visual features to obtain compressed aggregated visual features; and inputting the compressed aggregated visual features into the target model to obtain the image processing result. That is, after obtaining the aggregated visual features, the compressed aggregated visual features are input into the target model for calculation, thereby reducing the amount of calculation of the model calculation.

[0145] In some optional implementations of the embodiments of the present invention, the first visual feature code and the second visual feature code are subjected to feature aggregation processing to obtain an aggregated visual feature, which may include: the first visual feature code and the second visual feature code are subjected to feature splicing processing using a feature splicing method to obtain an aggregated visual feature.

[0146] Since the aggregation of feature splicing will lead to too high dimension and make the calculation too complicated, the first visual feature coding and the second visual feature coding can also be used to perform feature calculation to obtain aggregated features. That is, after the first visual feature coding and the second visual feature coding are subjected to feature aggregation processing, the aggregated visual features are obtained. It can also include: performing feature calculation on the features at the same position of the first visual feature coding and the second visual feature coding to obtain aggregated visual features. Among them, feature calculation can include but is not limited to feature addition calculation, feature mean calculation, channel weighted calculation, pixel weighted calculation, attention calculation, etc.

[0147] To improve the feature aggregation effect, in an embodiment of the present invention, the first visual feature code and the second visual feature code are subjected to feature aggregation processing to obtain an aggregated visual feature, and may also include: performing cross-attention calculation on the first visual feature code and the second visual feature code to obtain an aggregated visual feature.

[0148] In an embodiment of the present invention, after the image to be processed is segmented at multiple scales, each second sub-image is input into a visual encoder for encoding in S104 and then input into a target model to obtain an image processing result, which may include: starting from the second sub-image with the largest number of segmentation rounds among the first image and the second sub-images of each scale, visual coding features corresponding to the scale with a larger number of segmentation rounds are compressed and then fused with the visual coding features of the previous scale until they are fused into the first visual feature code of the first image to obtain an aggregated visual feature; the aggregated visual feature is compressed to obtain a compressed aggregated visual feature; the compressed aggregated visual feature is input into the target model to obtain an image processing result.

[0149] In the specific implementation, for two adjacent scales ( Scale and scale), for the Cross-attention pooling is applied to the visual features of the first scale to compress and retain detailed information, and the pooled features are combined with the The low-resolution feature aggregation of scale. is extracted from the resampler The visual characteristics of the scale, It is The number of sub-images of the scale, represents the number of queries for the resampler in the visual encoder, is the number of channels. In order to reduce the amount of calculation, the multi-view visual feature aggregation module provided by the embodiment of the present invention can be applied to implement Max pooling to obtain pooling features .

[0150] In order to reduce the information loss caused by the maximum pooling operation, the multi-view visual feature aggregation module provided by the embodiment of the present invention enhances the relevance of information at different scales through a lightweight cross-attention mechanism:

[0151] ;

[0152] ;

[0153] in, , , ,and is a learnable parameter, represents the first The visual characteristics of scale, represents the normalization function, Represents the first The visual characteristics of scale, represents the embedding dimension, Indicates transposed calculation. Pooling feature Corresponding to the query, the original features as key and value. is the aggregated visual feature. To simplify the expression, the layer normalization is omitted. Input to the target model for image processing calculation. With this approach, although multiple scales are considered in the image segmentation and feature extraction stages, the complexity of the model calculation remains unchanged because the multi-view visual feature aggregation module aggregates features of different scales into the visual features of the initial local scale.

[0154] In an embodiment of the present invention, a visual compression module may be designed to process the compression of visual features, such as the compression of aggregated visual features. Different types of images to be processed may be designed with different compression methods. The visual compression module may include a convolutional and fully connected layer structure for efficiently processing high-resolution images while ensuring that spatial information can be effectively utilized in the understanding and generation process.

[0155] The image processing method provided by the embodiment of the present invention can be applied to various types of images to be processed. The following is an introduction using a document image as an example.

[0156] When the image to be processed is a document image, compressing the aggregated visual features to obtain the compressed aggregated visual features may include: using a 1×a convolution kernel to compress the aggregated visual features to obtain the compressed aggregated visual features; wherein a is a positive integer greater than 1. That is, considering that the arrangement of characters in the document image is usually arranged in rows, a 1×a convolution kernel is set to compress the visual features in sequence according to the row order, so as to retain the spatial sensitivity of the visual features while compressing the high-resolution document image.

[0157] In other visual feature compression steps introduced in the above embodiments, compression can also be performed using the above method.

[0158] In practical applications, a can be 4, that is, a 1×4 convolution kernel is used to reduce the sequence length while retaining spatial information. The fully connected layer is used to map the processed visual features to the language embedding space, so that the text in the document is usually arranged along the line, which can effectively capture the sequential relationship of the text.

[0159] The convolution operation uses a convolution kernel with a kernel size of 1×4 and a step size of 4. The specific operation is:

[0160] ;

[0161] , ;

[0162] ;

[0163] in, Indicates An input sequence; Indicates The first element in the input sequence; Indicates The second element in the input sequence; Indicates The first elements; It means that after the convolution operation, The first new element; Represents a convolution operation. Since the convolution kernel size is 1×4 and the step size is 4, the convolution kernel covers four consecutive elements, that is, the input is 4 elements. , , , , and get an output element ; Represents the original length of the input sequence; A new sequence representing the output The first element of ; A new sequence representing the output The second element of A new sequence representing the output The last element of elements.

[0164] Such convolution operations can maintain semantic coherence and spatial relationships, thereby better utilizing the text information in the image during understanding and generation.

[0165] The image processing method provided by the embodiment of the present invention can be applied to a single-modal visual feature processing task or a multi-modal visual feature processing task. The multi-modal visual feature processing task can be an artificial intelligence task combining visual modality and text modality, such as a visual question answering task.

[0166] Then, in S104, each second sub-image is input into the visual encoder for encoding and then input into the target model to obtain an image processing result, which may include: processing the resolution of the image to be processed into the resolution of the visual encoder to obtain a first image; inputting the first image into the visual encoder to obtain a first visual feature code; inputting the second sub-image into the visual encoder to obtain a second visual feature code; performing feature aggregation processing on the first visual feature code and the second visual feature code to obtain an aggregated visual feature; inputting the input text corresponding to the image to be processed into the text encoder of the target model to obtain a text feature code; inputting the aggregated visual feature and the text feature code into the target model to obtain the image processing result.

[0167] Figure 5 A schematic diagram of multimodal image processing provided by an embodiment of the present invention.

[0168] like Figure 5 As shown, the multi-view image segmentation module provided in the embodiment of the present invention is used to perform multi-scale image segmentation on the image to be processed to obtain a multi-scale second sub-image, the image to be processed and the second sub-image of each round of segmentation are encoded using a visual encoder, and the visual features of each scale are aggregated using the multi-scale visual feature aggregation module provided in the embodiment of the present invention. Then, the aggregated visual features are compressed using the visual compression module provided in the embodiment of the present invention to obtain the final visual feature encoding (visual tokens). The input text is encoded using a text encoder to obtain a text feature encoding (text tokens). The text feature encoding and the visual feature encoding are input into a multimodal model for image processing to obtain an image processing result. The multimodal model can be a large language model. The image processing result can be an output text for a visual question answering task.

[0169] It should be noted that in the embodiments of the image processing methods of the present invention, some of the steps or features may be ignored or not executed. The hardware or software functional modules divided for the convenience of description are not the only implementation forms of the image processing methods provided in the embodiments of the present invention.

[0170] The above describes in detail various embodiments corresponding to the image processing method. On this basis, the present invention also discloses an image processing apparatus, device, non-volatile storage medium and computer program product corresponding to the above method.

[0171] The image processing device provided by the embodiment of the present invention may include:

[0172] A receiving unit, used for acquiring an image to be processed;

[0173] An image processing unit is used to segment the image to be processed according to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition to obtain a plurality of first sub-images corresponding to the image to be processed; and process the resolution of the first sub-image to the resolution of the visual encoder to obtain a second sub-image;

[0174] A model calculation unit, used for inputting each second sub-image into a visual encoder for encoding and then inputting the second sub-image into a target model to obtain an image processing result;

[0175] The first optimization condition is an optimization condition for reducing the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image, and the second optimization condition is an optimization condition for constraining the difference between the area of ​​each second sub-image and the area of ​​the image to be processed.

[0176] It should be noted that in each implementation of the image processing device provided by the embodiment of the present invention, the division of units is only a logical functional division, and other division methods can be used. The connection method between different units can be electrical, mechanical or other connection methods. The separated units can be located in the same physical location or distributed on multiple network nodes. Each unit can be implemented in the form of hardware or in the form of a software functional unit. That is, part or all of the units provided by the embodiment of the present invention can be selected according to actual needs and the corresponding connection method or integration method can be adopted to achieve the purpose of the embodiment of the present invention.

[0177] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, which will not be repeated here.

[0178] Figure 6 A schematic diagram of the structure of an image processing device provided by an embodiment of the present invention.

[0179] like Figure 6 As shown, the image processing device provided by the embodiment of the present invention includes: a memory 610, used to store a computer program 611; a processor 620, used to execute the computer program 611, and when the computer program 611 is executed by the processor 620, the steps of the image processing method provided by any of the above embodiments are implemented.

[0180] Among them, the processor 620 may include one or more processing cores, such as a 3-core processor, an 8-core processor, etc. The processor 620 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array. The processor 620 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 620 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 620 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0181] The memory 610 may include one or more non-volatile storage media, which may be non-transitory. The memory 610 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 610 is at least used to store the following computer program 611, wherein the computer program 611, after being loaded and executed by the processor 620, can implement the relevant steps in the image processing method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 610 may also include an operating system 612 and data 613, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 612 may be Windows or other types of operating systems. The data 613 may include, but is not limited to, the data involved in the above method.

[0182] In some embodiments, the image processing device may further include a display screen 630 , a power supply 640 , a communication interface 650 , an input / output interface 660 , a sensor 670 , and a communication bus 680 .

[0183] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation on the image processing device, and may include more or fewer components than those shown in the figure.

[0184] The image processing device provided by the embodiment of the present invention includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the steps of the image processing method provided by the above embodiment, and the effect is the same as above.

[0185] An embodiment of the present invention provides a non-volatile storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the image processing method provided in any one of the above embodiments can be implemented.

[0186] The non-volatile storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0187] For the introduction of the non-volatile storage medium provided in the embodiment of the present invention, please refer to the above method embodiment, and the effect thereof is the same as the image processing method provided in the embodiment of the present invention, and the present invention will not elaborate on it here.

[0188] An embodiment of the present invention provides a computer program product, including a computer program, which implements the steps of the image processing method provided in any one of the above embodiments when executed by a processor.

[0189] For an introduction to the computer program product provided by the embodiment of the present invention, please refer to the above method embodiment, and the effect thereof is the same as the image processing method provided by the embodiment of the present invention, and the present invention will not elaborate on it here.

[0190] The above is a detailed introduction to an image processing method, device, storage medium and computer program product provided by the present invention. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referenced to each other. For the devices, equipment, non-volatile storage media and computer program products disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the present invention.

[0191] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

Claims

1. An image processing method, characterized in that: include: Get the image to be processed; According to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition, the image to be processed is segmented to obtain a plurality of first sub-images corresponding to the image to be processed, including: determining the number of sub-images according to the resolution of the image to be processed and the resolution of the visual encoder; determining a combination of a first number of sub-images for segmenting the image to be processed in the long dimension and a second number of sub-images for segmenting the image to be processed in the wide dimension according to the number of sub-images, so that the product of the first number and the second number in the same group is the number of sub-images; and selecting the actual first number and the actual second number from the combination of the first number and the second number through the first optimization condition and the second optimization condition; Processing the resolution of the first sub-image to the resolution of the visual encoder to obtain a second sub-image; Inputting each of the second sub-images into the visual encoder for encoding and then into the target model to obtain an image processing result; Among them, the first optimization condition is an optimization condition for reducing the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image, and the second optimization condition is an optimization condition for constraining the difference between the area of ​​each second sub-image and the area of ​​the image to be processed.

2. The image processing method according to claim 1, characterized in that: Segmenting the image to be processed to obtain a plurality of first sub-images corresponding to the image to be processed includes: The image to be processed is subjected to non-overlapping uniform segmentation processing according to the actual first number and the actual second number to obtain a plurality of first sub-images corresponding to the image to be processed.

3. The image processing method according to claim 2, characterized in that: The first optimization condition is represented by an optimization objective function; the optimization goal of the optimization objective function is to minimize the deviation between the aspect ratio of the first sub-image and the aspect ratio of the second sub-image.

4. The image processing method according to claim 3, characterized in that: The optimization objective function is to maximize the function value of the segmentation scoring function; The segmentation scoring function is the reciprocal of the absolute value of the difference between the first logarithm and the second logarithm; the numerator of the real number of the first logarithm is the product of the width of the image to be processed and the second number, and the denominator of the real number of the first logarithm is the product of the length of the image to be processed and the first number; the real number of the second logarithm is the ratio of the width of the second sub-image to the length of the second sub-image.

5. The image processing method according to claim 3, characterized in that: The second optimization condition is that a ratio of a difference between an area of ​​each of the second sub-images minus an area of ​​the image to be processed and an area of ​​the image to be processed is less than a first threshold.

6. The image processing method according to claim 1, characterized in that: Inputting each of the second sub-images into the visual encoder for encoding and then inputting into the target model to obtain an image processing result, including: Processing the resolution of the image to be processed to the resolution of the visual encoder to obtain a first image; Inputting the first image into the visual encoder to obtain a first visual feature code; Inputting the second sub-image into the visual encoder to obtain a second visual feature code; After performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained; The aggregated visual features are input into the target model to obtain the image processing result.

7. The image processing method according to claim 6, characterized in that: According to the resolution of the image to be processed, the resolution of the visual encoder of the target model, the first optimization condition and the second optimization condition, the image to be processed is segmented to obtain a plurality of first sub-images corresponding to the image to be processed, including: According to the resolution of the image to be processed, the resolution of the visual encoder, the first optimization condition and the second optimization condition, performing multiple rounds of segmentation processing on the image to be processed to obtain multiple groups of the first sub-images of different scales; Inputting each of the second sub-images into the visual encoder for encoding and then inputting into the target model to obtain an image processing result, including: Inputting the second sub-images corresponding to the first sub-images of each scale into the visual encoder respectively to obtain the second visual feature codes corresponding to the second sub-images of each scale; After performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained, including: The visual feature codes corresponding to the images of adjacent scales in the first image and the second sub-image of each scale are subjected to feature aggregation processing, and the obtained feature aggregation results are spliced ​​to obtain the aggregated visual features.

8. The image processing method according to claim 7, characterized in that: The number of rounds of multiple segmentation processing performed on the image to be processed is determined according to the resolution of the image to be processed.

9. The image processing method according to claim 6, characterized in that: After performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained, including: Compressing the first visual feature code and the second visual feature code respectively to obtain the compressed first visual feature code and the compressed second visual feature code; The compressed first visual feature code and the compressed second visual feature code are aggregated to obtain the aggregated visual feature.

10. The image processing method according to claim 6, characterized in that: After performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained, including: The first visual feature encoding and the second visual feature encoding are cross-attention calculated to obtain the aggregated visual feature.

11. The image processing method according to claim 6, characterized in that: Inputting the aggregated visual features into the target model to obtain the image processing result includes: Compressing the aggregated visual features to obtain compressed aggregated visual features; The compressed aggregated visual features are input into the target model to obtain the image processing result.

12. The image processing method according to claim 11, characterized in that: The image to be processed is a document image to be processed; The aggregated visual features are compressed to obtain the compressed aggregated visual features, including: Using a 1×a convolution kernel to compress the aggregated visual features to obtain compressed aggregated visual features; Wherein, a is a positive integer greater than 1.

13. The image processing method according to claim 1, characterized in that: Inputting each of the second sub-images into the visual encoder for encoding and then inputting into the target model to obtain an image processing result, including: Processing the resolution of the image to be processed to the resolution of the visual encoder to obtain a first image; Inputting the first image into the visual encoder to obtain a first visual feature code; Inputting the second sub-image into the visual encoder to obtain a second visual feature code; After performing feature aggregation processing on the first visual feature code and the second visual feature code, an aggregated visual feature is obtained; Inputting the input text corresponding to the image to be processed into the text encoder of the target model to obtain text feature encoding; The aggregated visual features and the text features are encoded and input into the target model to obtain the image processing result.

14. An image processing device, characterized in that: include: Memory for storing computer programs; A processor is used to execute the computer program, and when the computer program is executed by the processor, the steps of the image processing method according to any one of claims 1 to 13 are implemented.

15. A non-volatile storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 13 are implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 13 are implemented.

Citation Information

Patent Citations

  • Image semantic segmentation method and device and terminal equipment

    CN110349164A

  • Map processing method and device, electronic equipment and storage medium

    CN118051574A