Image processing method and device, computing equipment and storage medium

By adaptively calculating the image magnification factor and performing magnification processing, the problems of resource waste and low efficiency in image processing in existing technologies are solved, and the recognition accuracy and efficiency of image processing models are improved.

CN121921180APending Publication Date: 2026-04-24BEIJING YUANLI WEILAI SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YUANLI WEILAI SCI & TECH CO LTD
Filing Date
2026-01-15
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies, when processing image data, especially when the image contains text information that is extremely small or has low contrast, cause key text information to become blurred or lost, affecting the recognition accuracy of intelligent models. Furthermore, existing magnification strategies lead to a waste of computing resources and low inference efficiency.

Method used

By determining the image size and text information, an adaptive image magnification factor is calculated, and the image is magnified under preset conditions. The image processing model is then used for recognition and processing, avoiding unnecessary computation and resource waste.

Benefits of technology

It improves the recognition accuracy and efficiency of image processing models, ensures the accuracy of image processing results, and reduces computational complexity and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921180A_ABST
    Figure CN121921180A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, computing equipment and a storage medium, and the image processing method comprises the steps: determining a to-be-processed image, and determining image size information and text information corresponding to the to-be-processed image, the text information being information of an image text contained in the to-be-processed image; calculating an image magnification factor corresponding to the to-be-processed image according to an image segmentation parameter corresponding to an image processing model, the image size information and the text information; when it is determined that the image magnification factor meets a preset magnification condition, magnifying the to-be-processed image according to the image magnification factor to obtain a magnified to-be-processed image; and inputting the amplified to-be-processed image into the image processing model to obtain an image processing result corresponding to the to-be-processed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the fields of artificial intelligence technology and image processing technology, and in particular to image processing methods and apparatus, computing devices and storage media. Background Technology

[0002] With the development of artificial intelligence technology, intelligent models can be used to process data such as images, text, audio, and video. These intelligent models can include machine learning models, neural network models, and large language models. However, when using intelligent models to process image data, if the image data contains text information that is extremely small or has low contrast, the crucial text information can become blurred, lost, or diluted at the pixel level. This makes it difficult for the intelligent model to recognize the text information within the image data, leading to a decrease in the accuracy of the image processing results. Therefore, an effective technical solution is urgently needed to address these problems. Summary of the Invention

[0003] In view of this, embodiments of this specification provide an image processing method. This specification also relates to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the aforementioned problems existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, an image processing method is provided, comprising: The image to be processed is determined, and the image size information and text information corresponding to the image to be processed are determined, wherein the text information is the image text information contained in the image to be processed; Based on the image segmentation parameters corresponding to the image processing model, the image size information, and the text information, calculate the image magnification factor corresponding to the image to be processed; If the image magnification factor meets the preset magnification conditions, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The magnified image to be processed is input into the image processing model to obtain the image processing result corresponding to the image to be processed.

[0005] According to a second aspect of the embodiments of this specification, an image processing apparatus is provided, comprising: An image detection module is configured to receive an image to be processed and determine the image size information and text information corresponding to the image to be processed, wherein the text information is the image text information contained in the image to be processed; The magnification decision module is configured to calculate the image magnification factor corresponding to the image to be processed based on the image segmentation parameters corresponding to the image processing model, the image size information, and the text information. The image magnification module is configured to, when it is determined that the image magnification factor meets the preset magnification conditions, perform a magnification operation on the image to be processed according to the image magnification factor to obtain the magnified image to be processed. The image processing module is configured to input the magnified image to be processed into the image processing model to obtain the image processing result corresponding to the image to be processed.

[0006] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above-described image processing method.

[0007] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the image processing method described above.

[0008] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the image processing method described above.

[0009] This specification provides an image processing method, comprising: determining an image to be processed, and determining image size information and text information corresponding to the image to be processed, wherein the text information is image text information contained in the image to be processed; calculating an image magnification factor corresponding to the image to be processed based on image segmentation parameters corresponding to an image processing model, the image size information, and the text information; when it is determined that the image magnification factor meets a preset magnification condition, performing a magnification operation on the image to be processed according to the image magnification factor to obtain a magnified image to be processed; and inputting the magnified image to be processed into the image processing model to obtain an image processing result corresponding to the image to be processed.

[0010] In the above method, after determining the image to be processed, the image size information and text information corresponding to the image to be processed can be determined. Based on the image segmentation parameters, image size information and text information corresponding to the image processing model used to process the image to be processed, the image magnification factor corresponding to the image to be processed is calculated. If the image magnification factor meets the preset magnification conditions, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The magnified image to be processed is then input into the image processing model so that the image processing model can recognize and process the magnified image to be processed, ensuring the processing accuracy of the image processing model and further ensuring the accuracy of the image processing results output by the image processing model. Attached Figure Description

[0011] Figure 1 This is a flowchart of an image processing method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of the enlarged title image in an image processing method provided in one embodiment of this specification; Figure 3 This is a flowchart illustrating an image processing method provided in one embodiment of this specification. Figure 4 This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification; Figure 5 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0012] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0013] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0014] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0015] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0016] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0017] Multimodal Large Models (MLLMs) are capable of simultaneously understanding and processing information from multiple different data types (or "modalities"). The most common modal combination is text (language) and images (vision), but their capabilities can be extended to other modalities such as audio, video, and 3D data. These models possess billions to hundreds of billions of parameters and are pre-trained on massive amounts of text and image data from the internet, enabling them to learn complex patterns and general feature representations across modalities. Based on the Transformer neural network architecture, MLLMs can effectively fuse and process information from different modalities. For example, upon seeing a picture of a cat, they can understand the textual concept of "cat," thus achieving multimodal understanding and perceiving the world comprehensively, much like humans. This powerful multimodal understanding capability has led to the widespread application of MLLMs in multiple fields. They can perform visual question answering, generating textual answers based on images and questions; and understand and respond to multimodal chatbots that include user input in multiple forms such as text, images, and speech.

[0018] Transformer is a deep learning model architecture based on a self-attention mechanism.

[0019] OCR: Optical Character Recognition, refers to the technology that automatically converts text content in an image into editable and searchable text.

[0020] In practical applications, multimodal large models can receive images and text as input and perform complex reasoning and content generation. When the text information in the image input to a multimodal large model is too small or has low contrast, key text information can be blurred, lost, or diluted at the pixel level. This prevents the subsequent visual encoder from extracting effective text features, thus affecting the inference performance of the multimodal large model and reducing the accuracy of the inference results. Typically, the image can be enlarged to reconstruct a high-resolution image from the original low-resolution image to restore lost details.

[0021] However, current image magnification methods typically employ a fixed magnification size. To ensure all subtle details are captured, a fixed high magnification ratio is used, or the original image is scaled to a fixed resolution before being used as model input. Alternatively, multiple attempts can be made to gradually increase the resolution from low to high until a satisfactory image is obtained. This fixed-size magnification strategy often involves unnecessary computation in most clear areas or areas without textual information, resulting in significant waste of computational resources and time. It also significantly increases the sequence length processed by multimodal models, thereby increasing inference latency and reducing inference efficiency. Furthermore, current inference processes for large multimodal models are mostly static and linear, lacking the ability to dynamically adjust preprocessing parameters based on input content. This fails to demonstrate a balance between accuracy requirements and efficiency constraints, thus limiting the versatility and performance ceiling of large multimodal models when handling diverse and heterogeneous inputs. Multiple attempts require multiple inference iterations, resulting in a long total inference time and reduced inference efficiency. Therefore, an effective technical solution is urgently needed to address these issues.

[0022] This specification provides an image processing method, and also relates to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.

[0023] Figure 1 A flowchart of an image processing method according to an embodiment of this specification is shown, which specifically includes the following steps: Step 102: Determine the image to be processed, and determine the image size information and text information corresponding to the image to be processed, wherein the text information is the image text information contained in the image to be processed.

[0024] In this context, the image to be processed can be understood as the original image that needs to be processed. The image to be processed may contain image text. The image size information may include the length and width of the image to be processed. The image text contained in the image to be processed can be understood as semantically meaningful text, symbols, numbers, or characters embedded or appearing at the pixel level of the image to be processed. The text information can be understood as the occupancy information of the image text within the image to be processed. The text information may include, for example, the number of characters in the image text, the area occupied by the image text in the image to be processed, or the height and size information of the image text, etc. This specification does not limit the specific examples in this embodiment.

[0025] Based on this, the image to be processed containing image text can be identified, and the length and width of the image to be processed, as well as the text information of the image text in the image to be processed, can be determined.

[0026] Optionally, in one embodiment of this specification, the text information corresponding to the image to be processed may be the text quantity information of the image text.

[0027] In practical applications, the image processing method provided in the embodiments of this specification can be applied to a client. Accordingly, the client can determine the image to be processed selected by the user and determine the image size information and text information corresponding to the image to be processed. The image processing method provided in the embodiments of this specification can also be applied to a server. Accordingly, the server can receive the image to be processed sent by the user through the client and determine the image size information and text information corresponding to the image to be processed.

[0028] Furthermore, the image processing method provided in the embodiments of this specification can be applied to the field of problem-solving. In this case, the image to be processed can be the image of the problem to be solved, and the image text included in the problem image can be the problem text. The problem image can be an image obtained by a user taking a picture of a problem in a textbook, workbook, or test paper, or it can be an image obtained by a user taking a screenshot of a problem in an electronic teaching platform. In addition, the image processing method provided in the embodiments of this specification can also be applied to other fields, such as image recognition, image retrieval, and image description, etc., which are not limited in this embodiment.

[0029] In specific implementation, determining the image size information and text information corresponding to the image to be processed includes: The image to be processed is subjected to size detection to obtain the image size information corresponding to the image to be processed; Text detection is performed on the image to be processed to obtain the text information of the image text contained in the image to be processed.

[0030] Specifically, when performing text detection on an image to be processed, the text region of the image text contained in the image to be processed can be extracted, the image text can be obtained from the text region, and the text quantity information of the image text can be counted as the text information of the image text.

[0031] In practical applications, OCR recognition tools can be used to perform size detection and text detection on the image to be processed, obtaining the image size information and text quantity information corresponding to the image to be processed. Alternatively, a pre-trained image detection model can be used to obtain the image size information and text quantity information of the image to be processed. This pre-trained image detection model can be, for example, a pre-trained machine learning model, a neural network model, a large language model, etc., or the size detection and text detection of the image to be processed can also be implemented through written program code. This specification does not limit the embodiments in this way.

[0032] In summary, by obtaining the image size and text information of the image to be processed, a data foundation is provided for subsequent calculation of the image magnification factor, which in turn provides a basis for subsequent magnification operations on the image to be processed.

[0033] Step 104: Calculate the image magnification factor corresponding to the image to be processed based on the image segmentation parameters corresponding to the image processing model, the image size information, and the text information.

[0034] Here, the image processing model can be understood as the model used to process the image to be processed. The image segmentation parameters corresponding to the image processing model can be understood as the segmentation parameters when the image processing model segments the image to be processed. The image magnification factor can be understood as the magnification factor of the image to be processed. For example, if the image magnification factor is 2, it means that the image to be processed will be magnified by 2 times.

[0035] Specifically, after determining the image size and text information corresponding to the image to be processed, the image magnification factor corresponding to the image to be processed can be calculated based on the image size information, text information, and image segmentation parameters corresponding to the image processing model.

[0036] In practical applications, image processing models can be multimodal large models. These models include a visual network layer that processes the input image. This visual network layer can segment the input image into multiple image blocks. The size of each image block is defined by the image segmentation parameters of the visual network layer. These parameters can be understood as the size of the segmented image blocks. For example, setting the segmented image block size to 10×10 means that for a 100×100 image, the visual network layer will segment the image into 100 image blocks.

[0037] In specific implementation, the step of calculating the image magnification factor corresponding to the image to be processed based on the image segmentation parameters corresponding to the image processing model, the image size information, and the text information includes: The image processing threshold corresponding to the image processing model is determined based on the image segmentation parameters corresponding to the image processing model. The image magnification factor corresponding to the image to be processed is calculated based on the image processing threshold, the image size information, and the text information.

[0038] The image processing threshold can be understood as the number of image blocks obtained when the image processing model segments the image to be processed.

[0039] Specifically, based on the image segmentation parameters corresponding to the image processing model and the image size information of the image to be processed, the image processing threshold corresponding to the image processing model can be calculated, and based on the image processing threshold, image size information and text quantity information, the image magnification factor corresponding to the image to be processed can be calculated.

[0040] In practical applications, when calculating the image magnification factor corresponding to the image to be processed, the image magnification factor can be calculated based on the image processing threshold, the length and width of the image to be processed, and the text quantity information. The image magnification factor $M$ = text quantity information ÷ (length × width ÷ image processing threshold).

[0041] Continuing with the previous example, when calculating the image processing threshold, if the image segmentation parameter is set to 10×10 and the image size of the image to be processed is 100×100, the visual network layer will segment the image into 100 image blocks. Therefore, the image processing threshold can be set to 100. Furthermore, for dynamic adjustments to different images, taking the image segmentation parameter of 10×10 as an example, if the image size of the image to be processed is 101×101, it is not possible to divide the image segmentation parameter by an integer. In this case, the image size can be rounded down, for example, 101×101 can be rounded down to 100×100, and the calculated image processing threshold will be 100. Alternatively, if the image size of the image to be processed is 109×109, the image size can be rounded down to 110×110.

[0042] In addition, the image processing threshold can also be determined based on other parameters of the image processing model, or it can be set based on historical image processing experience. This specification does not limit this.

[0043] In summary, by adaptively calculating the magnification factor of the image to be processed, it is easier to determine whether the image to be processed needs to be magnified based on the magnification factor. Furthermore, when it is determined that the image to be processed needs to be magnified, it can be magnified according to the magnification factor, thus achieving adaptive magnification of the image to be processed.

[0044] Step 106: If the image magnification factor meets the preset magnification conditions, magnify the image to be processed according to the image magnification factor to obtain the magnified image to be processed.

[0045] Specifically, after calculating the image magnification factor, and provided that the image magnification factor meets the preset magnification conditions, the image to be processed can be magnified according to the image magnification factor to obtain the magnified image to be processed.

[0046] In practical applications, the step of enlarging the image to be processed according to the image magnification factor, when the magnification factor meets the preset magnification conditions, to obtain the magnified image to be processed, includes: If the image magnification factor is determined to be greater than a preset magnification threshold, then the image magnification factor is determined to satisfy the preset magnification condition. Based on the image magnification factor, the image to be processed is magnified to obtain the magnified image to be processed.

[0047] The preset magnification threshold can be understood as a pre-set magnification factor threshold. If the magnification factor of the image is determined to be greater than the preset magnification threshold, it means that the image to be processed needs to be magnified.

[0048] Specifically, if the magnification factor of an image is greater than the preset magnification threshold, it means that the magnification factor of the image meets the preset magnification condition. At this time, the image to be processed can be magnified according to the magnification factor. The length of the image to be processed is multiplied by the magnification factor, and the width of the image to be processed is multiplied by the magnification factor to obtain the magnified image to be processed.

[0049] In practical applications, the preset magnification threshold can be 1. When the image magnification factor is greater than 1, the image to be processed is magnified. When the image magnification factor is less than or equal to 1, the image to be processed is not magnified and is directly input into the image processing model.

[0050] In addition, the image magnification factor can satisfy the preset magnification conditions, or it can be that the image magnification factor is within the preset magnification factor range. The preset magnification conditions can be determined according to actual needs, and the embodiments in this specification do not limit this.

[0051] In summary, by determining whether the image magnification factor meets the preset magnification conditions, and only when the magnification factor meets these conditions, the image to be processed is magnified. Computational resources are allocated as needed, significantly reducing the computational complexity and resource consumption of the inference process. Furthermore, based on the calculated image magnification factor, only one magnification operation is required, eliminating the need for repeated trial and error and further reducing resource consumption.

[0052] Further, the step of magnifying the image to be processed according to the image magnification factor to obtain the magnified image to be processed includes: Based on the image magnification factor, perform an interpolation magnification operation on the image to be processed to obtain a magnified image; or The image to be processed is input into an image magnification model. Based on the image magnification model, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The image magnification model is used to generate high-frequency information from low-frequency information in the image to be processed, and to magnify the image to be processed based on the high-frequency information.

[0053] Low-frequency and high-frequency information can be understood as data on different frequency components in an image. Low-frequency information can be used to describe smooth changes in an image, i.e., regions where color or brightness changes slowly. High-frequency information can be used to represent fast-changing parts of an image, such as edges, texture details, and noise.

[0054] Specifically, based on the bicubic interpolation algorithm, the image to be processed is interpolated and enlarged according to the image magnification factor to obtain the magnified image. Specifically, a mapping relationship can be established between the pixel coordinates of the image to be processed and the pixel coordinates of the magnified image based on the image magnification factor. According to this mapping relationship, the position of each pixel in the magnified image is traversed, and the pixel coordinates of each pixel in the image to be processed are calculated. Neighboring pixels centered on each pixel (e.g., 4×4 neighboring pixels) are extracted, and the weights of the neighboring pixels are calculated using a bicubic kernel function. The weights of the neighboring pixels are then weighted and summed to obtain the target pixel value of the center pixel. This calculation process is repeated for each pixel until the target pixel values ​​of all pixels are obtained, resulting in the magnified image to be processed.

[0055] In practical applications, an image magnification model can be understood as a model that performs image magnification based on a super-resolution algorithm. A super-resolution algorithm can recover or generate high-frequency information by utilizing low-frequency information in an image, thereby achieving image magnification. This image magnification model can be trained on a training dataset containing image sample pairs of low-resolution and high-resolution images. The image magnification model can be a deep learning model, a neural network model, etc., which is not limited in the embodiments of this specification. During the training process of the image magnification model, a low-resolution image can be input into the image magnification model to obtain the predicted high-resolution image output by the image magnification model. The model loss value is calculated based on the predicted high-resolution image and the corresponding high-resolution image of the low-resolution image. The model parameters of the image magnification model are adjusted according to the model loss value to achieve the training of the image magnification model. The image magnification model can automatically learn the complex mapping relationship between low-resolution and high-resolution images, including the recovery of high-frequency information such as texture details and edges. After the image magnification model is trained, the image to be processed and the image magnification factor can be input into the image magnification model. The image magnification model can perform a magnification operation on the image to be processed based on the image magnification factor to obtain the magnified image to be processed.

[0056] It is understandable that high-resolution images have a higher resolution than low-resolution images. In practical applications, any of the above magnification methods can be selected to magnify the image to be processed. Other algorithms for image magnification can also be used to magnify the image to be processed. This specification does not limit this aspect.

[0057] In summary, using the bicubic interpolation algorithm for magnification allows for rapid image upscaling, while employing a lightweight super-resolution algorithm results in higher accuracy of the magnified image. Furthermore, the resolution of the magnified image is the minimum necessary resolution, thus avoiding unnecessary pixel redundancy.

[0058] Step 108: Input the magnified image to be processed into the image processing model to obtain the image processing result corresponding to the image to be processed.

[0059] Specifically, after obtaining the magnified image to be processed, the magnified image to be processed can be input into the image processing model. Based on the image processing model, the image to be processed is processed to obtain the image processing result corresponding to the image to be processed, which is output by the image processing model.

[0060] In specific implementation, the step of inputting the magnified image to be processed into the image processing model to obtain the image processing result corresponding to the image to be processed includes: The magnified image to be processed is input into the image processing model, and the magnified image to be processed is encoded using a visual encoder to obtain the text features and visual features corresponding to the image to be processed. Based on the image processing model, feature processing is performed on the text features and the visual features to obtain the image processing result corresponding to the image to be processed.

[0061] The image processing model may include a visual encoder.

[0062] Specifically, after inputting the magnified image to be processed into the image processing model, the magnified image to be processed can be encoded based on the visual encoder to obtain the text features and visual features corresponding to the image to be processed. The text features and visual features are then fused to obtain the target fusion features. Based on the target fusion features, the image processing result corresponding to the image to be processed is determined.

[0063] In summary, because the image size information of the magnified image to be processed has been optimized, the visual encoder can efficiently extract clear text features, thereby improving the accuracy and precision of the image processing results.

[0064] Furthermore, when the image processing model is a multimodal large model, when the magnified image to be processed is input into the image processing model, the corresponding prompt text of the magnified image to be processed can also be input. The prompt text can be used to instruct the image processing model to process the image to be processed.

[0065] This explanation uses the title image as an example to illustrate the process; see [link / reference]. Figure 2 , Figure 2 The diagram illustrates an enlarged question image in an image processing method according to an embodiment of this specification. The prompt text corresponding to the question image may be "Please analyze the content of the image." Based on this, the enlarged question image and the prompt text can be input into an image processing model, and the resulting image processing result can be an answer to the question in the question image.

[0066] For example, the image size of the question image is 240×75, and the number of characters in the image text is 69, meaning there are 69 characters in the question image. The image processing threshold can be set to 28×28 = 784 based on the image processing model architecture. Therefore, 240×75 ÷ 784 = 22, and 69 ÷ 22 = 3, meaning the image magnification factor is 3. The question image can then be magnified 3 times based on this magnification factor, and the magnified question image and prompt text can be input together into the image processing model.

[0067] In summary, the above method, after determining the image to be processed, can determine the image size information and text information corresponding to the image to be processed. Based on the image segmentation parameters, image size information, and text information corresponding to the image processing model used to process the image to be processed, the image magnification factor corresponding to the image to be processed is calculated. If the image magnification factor meets the preset magnification conditions, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The magnified image to be processed is then input into the image processing model so that the image processing model can recognize and process the magnified image to be processed, ensuring the processing accuracy of the image processing model and further ensuring the accuracy of the image processing results output by the image processing model.

[0068] The following is in conjunction with the appendix Figure 3 Taking the image processing method provided in this specification as an example in the application of image reasoning, the image processing method will be further explained. Figure 3 This specification illustrates a flowchart of an image processing method according to an embodiment, which specifically includes the following steps: Step 302: Receive the image to be processed and the corresponding prompt text, and determine the image size information and text information corresponding to the image to be processed.

[0069] Specifically, it can receive the original image (i.e., the image to be processed) and optional input text (i.e., prompt text) from the user, and send the original image to the lightweight OCR analysis module. This module uses optimized (e.g., model quantization, pruning) high-speed algorithms to ensure that the analysis is completed with extremely low time overhead and does not become a bottleneck for inference latency. The lightweight OCR module performs text detection on the original image. Its core task is to count and output the number of texts recognized in the image (i.e., text information), and the length and width of the original image (i.e., image size information).

[0070] Step 304: Calculate the image magnification factor corresponding to the image to be processed based on the image processing threshold, image size information, and text information.

[0071] Specifically, the number of texts, the length and width of the original image obtained in step 302 are input into the adaptive magnification decision unit, which calculates the image magnification factor based on the preset image processing threshold, the number of texts, and the length and width of the original image.

[0072] Step 306: If the magnification factor of the image is greater than the preset magnification threshold, magnify the image to be processed according to the magnification factor to obtain the magnified image to be processed.

[0073] Specifically, the original image can be sent to the dynamic image scaling module, and a scaling operation is performed based on the image scaling factor. If the image scaling factor is greater than 1, the original image is scaled up; otherwise, no operation is performed. The scaling operation can use a fast bicubic interpolation algorithm, or a lightweight super-resolution algorithm when higher accuracy is required. The resulting image has an optimized resolution (both length and width are multiplied by the image scaling factor). This image resolution is the minimum necessary resolution, avoiding unnecessary pixel redundancy.

[0074] Step 308: Input the magnified image to be processed and the prompt text into the image processing model to obtain the image processing result corresponding to the image to be processed.

[0075] Specifically, the enlarged image and prompt text can be used as input to a multimodal large model (i.e., an image processing model). The visual encoder in the image processing model encodes the enlarged image. Because the size of the enlarged image is optimized, the visual encoder can efficiently extract clear text features with low computational cost. After fusing the visual features and text features of the enlarged image, the final image processing result is generated.

[0076] In summary, the above method, after determining the image to be processed, can determine the image size information and text information corresponding to the image to be processed. Based on the image segmentation parameters, image size information, and text information corresponding to the image processing model used to process the image to be processed, the image magnification factor corresponding to the image to be processed is calculated. If the image magnification factor meets the preset magnification conditions, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The magnified image to be processed is then input into the image processing model so that the image processing model can recognize and process the magnified image to be processed, ensuring the processing accuracy of the image processing model and further ensuring the accuracy of the image processing results output by the image processing model.

[0077] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 4 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. For example... Figure 4 As shown, the device includes: The image detection module 402 is configured to receive an image to be processed and determine the image size information and text information corresponding to the image to be processed, wherein the text information is the image text information contained in the image to be processed; The magnification decision module 404 is configured to calculate the image magnification factor corresponding to the image to be processed based on the image segmentation parameters corresponding to the image processing model, the image size information, and the text information. The image magnification module 406 is configured to, when it is determined that the image magnification factor meets the preset magnification conditions, perform a magnification operation on the image to be processed according to the image magnification factor to obtain the magnified image to be processed. The image processing module 408 is configured to input the magnified image to be processed into the image processing model to obtain the image processing result corresponding to the image to be processed.

[0078] In an optional embodiment, the amplification decision module 404 is further configured to: The image processing threshold corresponding to the image processing model is determined based on the image segmentation parameters corresponding to the image processing model. The image magnification factor corresponding to the image to be processed is calculated based on the image processing threshold, the image size information, and the text information.

[0079] In an optional embodiment, the image magnification module 406 is further configured to: If the image magnification factor is determined to be greater than a preset magnification threshold, then the image magnification factor is determined to satisfy the preset magnification condition. Based on the image magnification factor, the image to be processed is magnified to obtain the magnified image to be processed.

[0080] In an optional embodiment, the image magnification module 406 is further configured to: Based on the image magnification factor, perform an interpolation magnification operation on the image to be processed to obtain a magnified image; or The image to be processed is input into an image magnification model. Based on the image magnification model, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The image magnification model is used to generate high-frequency information from low-frequency information in the image to be processed, and to magnify the image to be processed based on the high-frequency information.

[0081] In an optional embodiment, the image detection module 402 is further configured to: The image to be processed is subjected to size detection to obtain the image size information corresponding to the image to be processed; Text detection is performed on the image to be processed to obtain the text information of the image text contained in the image to be processed.

[0082] In an optional embodiment, the image processing module 408 is further configured to: The magnified image to be processed is input into the image processing model, and the magnified image to be processed is encoded using a visual encoder to obtain the text features and visual features corresponding to the image to be processed. Based on the image processing model, feature processing is performed on the text features and the visual features to obtain the image processing result corresponding to the image to be processed.

[0083] In practical applications, the image detection module can be a lightweight OCR analysis module, used to quickly identify text regions in an image and quantify the amount of text. The input is the original image, and the output is the amount of text and the length and width of the original image. The magnification decision module can calculate an appropriate magnification factor based on the amount of text, the length and width of the original image, and a preset image processing threshold. The input is the amount of text and the length and width of the original image, and the output is the image magnification factor. The image magnification module can be a dynamic image scaling module, used to magnify the original image. The input is the original image and the image magnification factor, and the output is the magnified image. The multimodal large model can be used for inference on the magnified image. The input is the magnified image, and the output is the image processing result, which can be in text form. Before the multimodal large model inference, a computationally efficient lightweight module quickly obtains the amount of text in the input image as the core indicator for magnifying the image, realizing a lightweight adaptive decision mechanism for the amount of text. The adaptive preprocessing unit (including the aforementioned image detection module, magnification decision module, and image magnification module) is a pluggable, independent module, located at the beginning of the inference phase in time. Without changing the visual encoder architecture and input format of existing multimodal large models (which is still a complete image), it optimizes the quality and efficiency of the input image from the source, thereby improving the model's versatility and engineering applicability.

[0084] In summary, in the above-described device, after determining the image to be processed, the image size information and text information corresponding to the image to be processed can be determined. Based on the image segmentation parameters, image size information, and text information corresponding to the image processing model used to process the image to be processed, the image magnification factor corresponding to the image to be processed is calculated. If the image magnification factor meets the preset magnification conditions, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The magnified image to be processed is then input into the image processing model so that the image processing model can recognize and process the magnified image to be processed, ensuring the processing accuracy of the image processing model and further ensuring the accuracy of the image processing results output by the image processing model.

[0085] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0086] Figure 5 A structural block diagram of a computing device 500 according to an embodiment of this specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.

[0087] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0088] In one embodiment of this specification, the above-described components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0089] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.

[0090] The processor 520 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described image processing method.

[0091] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the image processing method described above.

[0092] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the image processing method described above.

[0093] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the image processing method described above.

[0094] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described image processing method.

[0095] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the image processing method described above.

[0096] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0097] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0098] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.

[0099] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0100] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An image processing method, characterized in that, include: The image to be processed is determined, and the image size information and text information corresponding to the image to be processed are determined, wherein the text information is the image text information contained in the image to be processed; Based on the image segmentation parameters corresponding to the image processing model, the image size information, and the text information, calculate the image magnification factor corresponding to the image to be processed; If the image magnification factor meets the preset magnification conditions, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The magnified image to be processed is input into the image processing model to obtain the image processing result corresponding to the image to be processed.

2. The method according to claim 1, characterized in that, The step of calculating the image magnification factor corresponding to the image to be processed based on the image segmentation parameters corresponding to the image processing model, the image size information, and the text information includes: The image processing threshold corresponding to the image processing model is determined based on the image segmentation parameters corresponding to the image processing model. The image magnification factor corresponding to the image to be processed is calculated based on the image processing threshold, the image size information, and the text information.

3. The method according to claim 1, characterized in that, The step of enlarging the image to be processed according to the image magnification factor, when the image magnification factor meets the preset magnification conditions, to obtain the magnified image to be processed, includes: If the image magnification factor is determined to be greater than a preset magnification threshold, then the image magnification factor is determined to satisfy the preset magnification condition. Based on the image magnification factor, the image to be processed is magnified to obtain the magnified image to be processed.

4. The method according to any one of claims 1-3, characterized in that, The step of magnifying the image to be processed according to the image magnification factor to obtain the magnified image to be processed includes: Based on the image magnification factor, perform an interpolation magnification operation on the image to be processed to obtain a magnified image; or The image to be processed is input into an image magnification model. Based on the image magnification model, the image to be processed is magnified according to the image magnification factor to obtain the magnified image to be processed. The image magnification model is used to generate high-frequency information from low-frequency information in the image to be processed, and to magnify the image to be processed based on the high-frequency information.

5. The method according to any one of claims 1-3, characterized in that, The process of determining the image size information and text information corresponding to the image to be processed includes: The image to be processed is subjected to size detection to obtain the image size information corresponding to the image to be processed; Text detection is performed on the image to be processed to obtain the text information of the image text contained in the image to be processed.

6. The method according to any one of claims 1-3, characterized in that, The step of inputting the magnified image to be processed into the image processing model to obtain the image processing result corresponding to the image to be processed includes: The magnified image to be processed is input into the image processing model, and the magnified image to be processed is encoded using a visual encoder to obtain the text features and visual features corresponding to the image to be processed. Based on the image processing model, feature processing is performed on the text features and the visual features to obtain the image processing result corresponding to the image to be processed.

7. An image processing apparatus, characterized in that, include: The image detection module is configured to receive an image to be processed and determine the image size information and text information corresponding to the image to be processed, wherein the text information is the image text information contained in the image to be processed; The magnification decision module is configured to calculate the image magnification factor corresponding to the image to be processed based on the image segmentation parameters corresponding to the image processing model, the image size information, and the text information. The image magnification module is configured to, when it is determined that the image magnification factor meets the preset magnification conditions, perform a magnification operation on the image to be processed according to the image magnification factor to obtain the magnified image to be processed. The image processing module is configured to input the magnified image to be processed into the image processing model to obtain the image processing result corresponding to the image to be processed.

8. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.