Image processing method, computing device, storage medium, and computer program product
By acquiring images, prompt text, and visual location information, and using a multimodal processing model to extract features for image processing, the problem of inaccurate recognition of complex graphic content in existing technologies is solved, achieving higher recognition accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2026-03-19
AI Technical Summary
Existing deep learning models struggle to accurately recognize complex text and image content, such as e-commerce product detail images and text, natural scene images and text, complex invoices in the financial and healthcare industries, and complex form data in the transportation and distribution industries, resulting in poor image recognition accuracy.
By acquiring the image to be processed, the associated prompt text, and visual location information, the target text features, image features, and location features are extracted. Then, a multimodal processing model is used for image processing to achieve accurate recognition of complex graphic and text content.
It achieves accurate recognition of complex graphic content and improves the accuracy of image processing results.
Smart Images

Figure CN2025108208_19032026_PF_FP_ABST
Abstract
Description
Image processing method, computing device, storage medium and computer program product
[0001] Cross-reference to related applications
[0002] The present disclosure claims priority to the Chinese patent application No. 2024112816955, filed on September 12, 2024, and entitled “Image processing method, computing device, storage medium and computer program product”, the entire content of which is incorporated herein by reference. TECHNICAL FIELD
[0003] Embodiments of the present specification relate to the technical field of computer technology, in particular to an image processing method, a computing device, a storage medium and a computer program product. BACKGROUND
[0004] At present, there are a large number of image documents in industrial production and commercial activities. These image documents, as the main carrier and medium of information, record a wealth of text and visual content. In order to understand the information recorded in the image documents, a deep learning model can be used to identify the image documents. However, for complex image-text content contained in the image documents, such as e-commerce detail image-text, natural scene image-text, complex bills in the financial and medical industries, and complex form data in the transportation and circulation industry, the current deep learning model is difficult to accurately identify, and the accuracy of the image recognition result is poor. Therefore, there is an urgent need for an effective technical solution to solve the above problems. SUMMARY
[0005] In view of this, the embodiments of the present specification provide two image processing methods. One or more embodiments of the present specification simultaneously relate to two image processing devices, a model training platform, a computing device, a computer readable storage medium and a computer program product to solve the technical defects existing in the prior art.
[0006] According to a first aspect of the embodiments of the present specification, an image processing method is provided, comprising:
[0007] obtaining a to-be-processed image, a prompt text associated with the to-be-processed image, and determining visual position information for the to-be-processed image;
[0008] obtaining a target text feature corresponding to the prompt text, a target image feature corresponding to the to-be-processed image, and a target position feature corresponding to the visual position information;
[0009] According to the target text feature, the target image feature and the target position feature, obtaining an image processing result corresponding to the to-be-processed image, wherein the image processing result is a result of processing the to-be-processed image according to the text content of the prompt text.
[0010] According to a second aspect of the embodiments of the present specification, an image processing apparatus is provided, comprising:
[0011] A first obtaining component configured to obtain a to-be-processed image, prompt text associated with the to-be-processed image, and determine visual position information of the to-be-processed image;
[0012] A second obtaining component configured to obtain target text features corresponding to the prompt text, target image features corresponding to the to-be-processed image, and target position features corresponding to the visual position information;
[0013] A third obtaining component configured to obtain an image processing result corresponding to the to-be-processed image according to the target text features, the target image features, and the target position features, wherein the image processing result is a result of processing the to-be-processed image according to text content of the prompt text.
[0014] According to a third aspect of the embodiments of the present specification, an image processing method is provided, comprising:
[0015] Receiving a to-be-processed image and prompt text associated with the to-be-processed image sent by a client, and receiving a point selection instruction sent by the client for the to-be-processed image, and determining visual position information of the to-be-processed image;
[0016] Obtaining target text features corresponding to the prompt text, target image features corresponding to the to-be-processed image, and target position features corresponding to the visual position information;
[0017] According to the target text features, the target image features, and the target position features, obtaining an image processing result corresponding to the to-be-processed image, wherein the image processing result is a result of processing the to-be-processed image according to text content of the prompt text;
[0018] Sending the image processing result to the client for display on a display interface of the client.
[0019] According to a fourth aspect of the embodiments of the present specification, an image processing apparatus is provided, comprising:
[0020] A receiving component configured to receive a to-be-processed image and prompt text associated with the to-be-processed image sent by a client, and receive a point selection instruction sent by the client for the to-be-processed image, and determine visual position information of the to-be-processed image;
[0021] A fourth obtaining component configured to obtain target text features corresponding to the prompt text, target image features corresponding to the to-be-processed image, and target position features corresponding to the visual position information;
[0022] The fifth obtaining component is configured to obtain an image processing result corresponding to the to-be-processed image according to the target text feature, the target image feature, and the target position feature, where the image processing result is a result of processing the to-be-processed image according to the text content of the prompt text;
[0023] The sending component is configured to send the image processing result to the client for display through a display interface of the client.
[0024] According to a fifth aspect of an embodiment of the present specification, a model training platform is provided, including a request interface component, a model training component, and a response component;
[0025] The request interface component is configured to receive a model training request, where the model training request includes model information of a to-be-trained machine learning model;
[0026] The model training component is configured to determine the to-be-trained machine learning model according to the model information, and perform model training on the to-be-trained machine learning model to obtain a trained multi-modal processing model, where the multi-modal processing model is used to process target text features of a prompt text, target image features of a to-be-processed image, and target position features corresponding to visual position information, to obtain an image processing result of the to-be-processed image, the prompt text is a prompt text associated with the to-be-processed image, and the visual position information is visual position information of the to-be-processed image;
[0027] The response component is configured to output the multi-modal processing model.
[0028] According to a sixth aspect of an embodiment of the present specification, a computing device is provided, including:
[0029] A memory and a processor;
[0030] The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which implement the steps of the above method when executed by the processor.
[0031] According to a seventh aspect of an embodiment of the present specification, a computer readable storage medium is provided, which stores computer programs / instructions, which implement the steps of the above method when executed by the processor.
[0032] According to an eighth aspect of an embodiment of the present specification, a computer program product is provided, including computer programs / instructions, which implement the steps of the above method when executed by the processor.
[0033] According to a ninth aspect of an embodiment of the present specification, a computer program is provided, which implements the steps of the above method when executed by the processor.
[0034] One embodiment of the specification realizes the determination of the visual position in the to-be-processed image, the target text feature corresponding to the prompt text, the target image feature corresponding to the to-be-processed image, and the target position feature corresponding to the visual position information by determining the to-be-processed image, the prompt text associated with the to-be-processed image, and the visual position information corresponding to the to-be-processed image. According to the multi-modal features of the target text feature, the target image feature, and the target position feature, the to-be-processed image is processed according to the text content of the prompt text, the image processing result corresponding to the visual position information is obtained, the accurate recognition of the local area of the to-be-processed image is realized, and the processing according to the features of multiple modalities is suitable for multi-modal processing tasks, which can realize the accurate recognition of complex image-text content, thereby ensuring the accuracy of the image processing result. BRIEF DESCRIPTION OF DRAWINGS
[0035] FIG. 1 is a schematic diagram of an application scenario of an image processing method according to one embodiment of the specification;
[0036] FIG. 2 is a flowchart of an image processing method according to one embodiment of the specification;
[0037] FIG. 3 is a flowchart of a processing process of an image processing method according to one embodiment of the specification;
[0038] FIG. 4 is a schematic diagram of the structure of an image processing apparatus according to one embodiment of the specification;
[0039] FIG. 5 is a flowchart of another image processing method according to one embodiment of the specification;
[0040] FIG. 6 is a schematic diagram of the structure of another image processing apparatus according to one embodiment of the specification;
[0041] FIG. 7 is a structural block diagram of a computing device according to one embodiment of the specification. DETAILED DESCRIPTION
[0042] In the following description, a lot of specific details are set forth in order to fully understand the specification. However, the specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the specification, so the specification is not limited by the specific implementation disclosed below.
[0043] The terminology used in this disclosure of one or more embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure of one or more embodiments. As used in this disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0044] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal sequence, but are used only to distinguish one piece of information from another. For example, without departing from the scope of the disclosure of one or more embodiments, first can be termed second, and similarly, second can be termed first. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining."
[0045] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the disclosure of one or more embodiments are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0046] In the disclosure of one or more embodiments, a large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, tens of billions, hundreds of billions, thousands of billions or even tens of billions of model parameters. The large model can also be called a foundation model. Through large-scale unlabeled corpus pre-training of the large model, a pre-trained model with hundreds of millions of parameters is produced. Such a model can adapt to a wide range of downstream tasks and has good generalization ability. For example, a large-scale language model (LLM) and a multi-modal pre-training model.
[0047] In practical applications, a large model can be widely applied in natural language processing (NLP) and computer vision fields, and can be applied to computer vision field tasks such as visual question answering (VQA), image captioning (IC), image generation, and natural language processing field tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the large model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, and the like.
[0048] In practical applications, there are a large number of image documents in industrial production and commercial activities, which are the main carriers and transmission media of information. Image document recognition and analysis aims to enable machines to understand text and visual content in image document scenes like humans, which is not only an important research topic in the field of machine learning, but also has great industry application value.
[0049] Deep learning small models have reached a bottleneck in processing image document analysis tasks, and small models cannot handle complex image-text parsing and understanding tasks alone. Typical complex image-texts, such as e-commerce detail image-texts, natural scene image-texts, various complex bills in the financial and medical industries, and complex forms in the transportation and circulation industries, are core objects for the corresponding industries to realize informatization. Therefore, an effective technical solution is urgently needed to solve the above problems.
[0050] First, the nomenclature related to one or more embodiments of the present specification is explained.
[0051] Text recognition: Optical Character Recognition (OCR), a technology used to convert printed or handwritten text into editable text.
[0052] Document parsing: a technology for detecting, recognizing, and extracting text, graphics, and image elements in an image.
[0053] Document understanding: understanding key information in a document, such as entity extraction, event extraction, sentiment analysis, and topic classification.
[0054] Multi-modal joint learning: a technology for jointly learning data from multiple modalities (such as text, images, audio, etc.).
[0055] Visual token: a feature vector obtained by passing an image through a visual encoder and then through a sampler.
[0056] Language token: A feature vector obtained by encoding text through a text encoder.
[0057] Large Language Model (LLM): Typically contains tens of billions or even hundreds of billions of parameters. These parameters are obtained by learning the complex patterns and rules of language, enabling high-quality contextual understanding and text generation capabilities.
[0058] VIT: Vision Transformer, a visual encoding neural network that can convert input images into a series of tokens.
[0059] ResNet: A visual encoding neural network.
[0060] Prompt: Refers to the initial input or instructions provided to the model, aiming to generate the desired output or response.
[0061] Instruction: Usually refers to the instructions or task descriptions directly provided to the model, indicating the specific operations or content generated by the model. Similar to "prompt", but "instruction" focuses more on explicitly instructing the model to perform a specific task.
[0062] Swin Transformers: A model based on the Transformer architecture, designed to handle computer vision tasks, by introducing a window mechanism to reduce computational complexity while maintaining local and global attention.
[0063] Backbone: Backbone network, which can be understood as the basic structure of the model, which can be understood as the network layer used to extract basic features in the model.
[0064] Resamplar: Resampling mechanism, usually used to describe the process of adjusting the sampling rate or resolution of data in digital signal processing, image processing and data science. Resampling can be applied to one-dimensional signals (such as audio signals), two-dimensional images or higher-dimensional data sets. It is mainly used to change the resolution of data to better adapt to different application scenarios or processing needs. In image processing, resampling is used to adjust the resolution of images, both reducing the image (downsampling) and enlarging the image (upsampling).
[0065] MLP: Multilayer Perceptron, a basic form of artificial neural network. MLP is usually used to solve supervised learning problems such as classification and regression tasks. It consists of multiple layers, including input layer, one or more hidden layers and output layer.
[0066] In the specification, two image processing methods are provided, and the specification also relates to two image processing apparatuses, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0067] Referring to FIG. 1, FIG. 1 is a schematic diagram of an application scenario of an image processing method according to an embodiment of the specification, and the image processing method includes the following steps.
[0068] Obtaining a to-be-processed image, prompt text associated with the to-be-processed image, and visual position information of the to-be-processed image;
[0069] Obtaining target text features corresponding to the prompt text, target image features corresponding to the to-be-processed image, and target position features corresponding to the visual position information;
[0070] According to the target text features, the target image features, and the target position features, obtaining an image processing result corresponding to the to-be-processed image, wherein the image processing result is a result of processing the to-be-processed image according to text content of the prompt text.
[0071] Specifically, FIG. 1 includes an end-side device 102 and a cloud-side device 104.
[0072] In a specific implementation, a user can send, through the end-side device 102, a to-be-processed image, prompt text associated with the to-be-processed image, and visual position information of the to-be-processed image to the cloud-side device 104, the cloud-side device 104 determines target text features corresponding to the prompt text, target image features corresponding to the to-be-processed image, and target position features corresponding to the visual position information, and determines an image processing result corresponding to the to-be-processed image according to the target text features, the target image features, and the target position features, and sends the image processing result to the end-side device 102, so as to display the image processing result to the user through a display interface of the end-side device 102.
[0073] The terminal-side device 102 can include a browser, an application (APP), or a web application such as a Hyper Text Markup Language 5 (H5) application, or a light application (also known as a small program, a lightweight application), or a cloud application, and the like, which can be developed based on a software development kit (SDK) of a corresponding service provided by the server-side, such as a real-time communication (RTC) SDK, and the like. The terminal-side device can be deployed in an electronic device, and needs to be run in dependence of the device or some APP in the device, and the like. The electronic device can have a display screen and support information browsing, and can be a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, and the like. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant communication tools, mailbox clients, social platform software, and the like.
[0074] The cloud-side device 104 can be understood as a server providing various services, including a physical server, a cloud server, for example, a server providing communication services for multiple clients, for example, a server for background training supporting a model used on a client, for example, a server processing data sent by a client, and the like. It should be noted that the cloud-side device 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The cloud-side device 104 can also be a server of a distributed system, or a server combined with a blockchain. The cloud-side device 104 can also be a cloud server of cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, and the like basic cloud computing services, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0075] It should be noted that the image processing method provided in the embodiments of the present specification can be executed by the cloud-side device 104, or can be executed by the terminal-side device 102; in other embodiments, the image processing method provided in the embodiments of the present specification can also be executed by the terminal-side device 102 and the cloud-side device 104 together.
[0076] Referring to FIG. 2, FIG. 2 is a flowchart of an image processing method according to an embodiment of the present specification, which specifically includes the following steps.
[0077] Step 202: obtaining a to-be-processed image, prompt text associated with the to-be-processed image, and visual position information of the to-be-processed image.
[0078] Specifically, the image processing method provided by the embodiments of the present specification can be applied to the recognition task of a document image. The recognition task of a document image is divided into three categories: visual tasks, language tasks, and multi-modal tasks. The visual task refers to a task related to the basic processing capability of a document image, and is usually a pre-link for subsequent language and multi-modal tasks. Specific tasks include: document image shadow removal, document image noise removal, and document image surface correction. The language task refers to a task that depends solely on text content. Language tasks include: document type, text question answering, text translation, and text generation. The multi-modal task refers to a task that depends on both images and semantics. These tasks need to be aligned within a multi-modal large model, and then a multi-modal output result is generated. Multi-modal tasks include: multi-language text detection and recognition, table structure recognition and reconstruction, chart structure analysis, image-text information extraction, and document image content generation.
[0079] The to-be-processed image can be understood as a document image that needs to be processed. The prompt text associated with the to-be-processed image can be understood as an instruction text (i.e., instruction) for indicating which processing is performed on the to-be-processed image. The prompt text can involve queries about image content, specific task requirements, etc. The prompt text can be, for example, “help me recognize the text in the target region of the image”, “help me remove the shadow in the target region of the image”, “help me recognize the table in the image”, etc. The visual position information of the to-be-processed image can be understood as the position information of the target region in the to-be-processed image that needs to be processed, which can be in the form of a point, a box, or a mask.
[0080] Based on this, the to-be-processed image that needs to be processed can be determined, the instruction text indicating which processing is performed on the to-be-processed image can be determined, and the position information of the target region in the to-be-processed image that needs to be processed can be determined.
[0081] In actual application, the to-be-processed image uploaded by the user through the client, the prompt text associated with the to-be-processed image input by the user through the client, and the visual position information obtained by the user by clicking or brushing the to-be-processed image can be received.
[0082] In specific implementation, the visual position information of the to-be-processed image is obtained, including:
[0083] A clicking instruction for the to-be-processed image is received.
[0084] The visual position information for the to-be-processed image is determined according to the clicking instruction.
[0085] The point selection instruction can include a click instruction, a box selection instruction, and a mask instruction for a to-be-processed region in the to-be-processed image.
[0086] Specifically, the point selection instruction sent by the user through the client for the to-be-processed region in the to-be-processed image can be received, and the visual position information of the target region in the to-be-processed image can be determined according to the point selection instruction.
[0087] In actual application, the visual position information can include a point (Point), a box (Box), and a mask (Mask). The point Prompt can be understood as that the user selects one or more key points in the image, and the model processes and understands the image region more accurately by identifying the positions and context information of the points. The box Prompt can be understood as that the user selects a specific region in the image by boxing, and provides the model with the range of the region of interest. The mask (Mask) Prompt can be understood as that the user provides a fine image mask to define the shape and boundary of a special region in the image. These prompts mainly correspond to positions, boundaries, and shapes, and the visual position information can be directly encoded as position tokens and input to the multi-modal processing model.
[0088] In summary, by determining the visual position information of the to-be-processed image, the specific target region in the image can be specified more finely for identification or detection segmentation, and accurate processing of the to-be-processed image can be achieved.
[0089] Specifically, the dynamic resolution adjustment mechanism can be used to intelligently select an image resolution suitable for the current task and hardware environment. By introducing an adaptive algorithm, the resolution of the to-be-processed image can be dynamically adjusted according to the complexity of the input to-be-processed image, the requirements of the image processing task, and the computing resources. The specific implementation is as follows.
[0090] After obtaining the to-be-processed image, the method further includes:
[0091] In response to the to-be-processed image containing text characters, the resolution of the to-be-processed image is adjusted according to the to-be-processed image and the text characters, and an adjusted to-be-processed image is obtained, wherein the resolution of the adjusted to-be-processed image satisfies a preset resolution condition.
[0092] Specifically, in the case where it is determined that the to-be-processed image contains text characters, the resolution of the to-be-processed image can be adjusted according to the image size of the to-be-processed image and the size of the text characters, and an adjusted to-be-processed image is obtained, so that the resolution of the adjusted to-be-processed image satisfies a preset resolution condition. The preset resolution condition can be understood as that the resolution satisfies a preset resolution threshold, or the resolution can also satisfy the computing resource requirements of the current processing device and the image processing task requirements.
[0093] In practical applications, scaling can be combined with the size of the text and the aspect ratio of the image to make the cut. In an embodiment of the present specification, the size of the text character can be calculated based on the edge detection algorithm and the connected domain algorithm applied to the image. Under the original image size, the text character is not scaled at 16-32 pixels, less than 16 pixels is enlarged *(16 / character height), and more than 32 pixels is reduced *(32 / character height). In addition, after scaling, the window is cut according to 256*256, and is divided into ceil(h / 256,w / 256) up to the integer, such as 1000*500 divided into 4*2 subgraphs. Wherein, ceil is the ceiling function, which rounds the given value up to the nearest integer.
[0094] In addition, in the case where the image to be processed contains or does not contain text characters, the resolution of the image to be processed can also be dynamically adjusted according to the complexity of the input image to be processed, the requirements of the image processing task, and the computing resources.
[0095] In summary, by dynamically adjusting the resolution of the image to be processed, the details of the image to be processed can be fully utilized, and the flexibility and adaptability of the subsequent multi-modal processing model can be improved, and the computing efficiency is improved on the premise of maximizing the use of image details.
[0096] Step 204: Obtain the target text feature corresponding to the prompt text, the target image feature corresponding to the image to be processed, and the target position feature corresponding to the visual position information.
[0097] Among them, the target text feature can be understood as the text feature vector of the prompt text, the target image feature can be understood as the image feature vector of the image to be processed, and the target position feature corresponding to the visual position information can be understood as the position feature encoding of the visual position information.
[0098] In specific implementation, obtaining the target text feature corresponding to the prompt text comprises:
[0099] The feature of the prompt text is extracted to obtain the target text feature corresponding to the prompt text.
[0100] In practical applications, the target text feature (i.e. text token) corresponding to the prompt text can be obtained by text encoder for feature extraction of the prompt text. The feature extraction of the prompt text can also be realized by a neural network model, which is not limited by the embodiments of the present specification.
[0101] In summary, by performing feature extraction on the prompt text, the encoding of the instruction input by the user through natural language is realized, which is convenient for subsequent combination with visual features to realize multi-modal processing tasks.
[0102] In specific implementation, obtaining the target image feature corresponding to the image to be processed comprises:
[0103] perform feature extraction on the to-be-processed image to obtain initial image features corresponding to the to-be-processed image;
[0104] perform sampling processing on the initial image features to obtain target image features corresponding to the to-be-processed image.
[0105] Specifically, the to-be-processed image can be subjected to feature extraction by a visual encoder to obtain initial image features corresponding to the to-be-processed image, and the initial image features are further subjected to sampling processing by a neural network to obtain target image features (i.e., visual tokens) corresponding to the to-be-processed image.
[0106] In the initial image features and the target image features, image spatial information can be retained.
[0107] In actual applications, deep learning architectures such as ResNet, Vision Transformers (ViT) or Swin Transformers are used. These backbones extract image features at multiple levels and scales, and have strong representation capabilities. After feature extraction by the encoder, the visual part is further encoded and processed by a few layers of transformer in the resamplar MLP to obtain visual tokens. This not only retains and compresses the image spatial information, but also aligns the pre-training task and the language features, so as to achieve more accurate cross-modal matching in the visual language model.
[0108] In summary, since the number of features of the initial image features is usually greater than that of the target text features, the alignment of the image and the text can be achieved by sampling.
[0109] In specific implementation, the visual position information includes position information of a to-be-processed part in the to-be-processed image.
[0110] obtaining target position features corresponding to the visual position information, including:
[0111] performing position encoding processing on the visual position information according to the position information of the to-be-processed part to obtain target position features corresponding to the visual position information.
[0112] The position information of the to-be-processed part can be understood as the region position information of the target region to be processed, and can include position coordinate information, boundary information and shape information, etc.
[0113] Accordingly, the visual position information can be subjected to position encoding processing according to the region position information of the target region to be processed in the to-be-processed image, and encoded into target position features (i.e., position tokens).
[0114] In summary, by encoding the visual position information, the position to be processed can be accurately determined during subsequent image processing.
[0115] At step 206, an image processing result corresponding to the to-be-processed image is obtained according to the target text feature, the target image feature, and the target position feature. The image processing result is a result of processing the to-be-processed image according to the text content of the prompt text.
[0116] Specifically, the image processing result of the to-be-processed image can be determined according to the target text feature, the target image feature, and the target position feature. The image processing result can be related to the prompt text. For example, if the prompt text is “help me identify the table in the image”, the image processing result is the table in the to-be-processed image. If the prompt text is “help me remove the shadow in the image”, the image processing result is the to-be-processed image with the shadow removed.
[0117] In an embodiment of the present specification, the image processing result corresponding to the to-be-processed image is determined according to the target text feature, the target image feature, and the target position feature, comprising:
[0118] The target text feature, the target image feature, and the target position feature are aligned to obtain the aligned target text feature.
[0119] The image processing result corresponding to the to-be-processed image is determined according to the aligned target text feature and the target image feature.
[0120] Specifically, the target text feature, the target image feature, and the target position feature can be aligned to obtain the aligned target text feature. The aligned target text feature is decoded to obtain text decoding information. The target image feature is visually decoded to obtain image decoding information. The image decoding information and the text decoding information are mapped to obtain the image processing result of the to-be-processed image.
[0121] In practical applications, aligning the target text feature, the target image feature, and the target position feature can be understood as a process of correctly corresponding the target text feature, the target image feature, and the target position feature in space or semantics. In specific implementation, space alignment can be achieved by using up-sampling, down-sampling, skip connection, and pyramid structure. Time alignment can be achieved by using time interpolation, recurrent neural network, and attention mechanism. Semantic alignment can be achieved by using joint embedding space, cross-modal attention mechanism, and multi-task learning, which are not limited in the embodiments of the present specification.
[0122] In summary, by aligning the image feature, the text feature, and the position feature, accurate alignment of language and vision is achieved, and the processing performance of the multi-modal task is ensured.
[0123] In another embodiment of the present specification, the image processing result corresponding to the to-be-processed image is determined according to the target text feature, the target image feature, and the target position feature, comprising:
[0124] The target text feature, the target image feature, and the target position feature are input into a multi-modal processing model to obtain the aligned target text feature output by the multi-modal processing model;
[0125] The target image feature is input into a visual decoder to obtain image decoding information of the target image feature;
[0126] The aligned target text feature is input into a text decoder to obtain text decoding information;
[0127] The image decoding information and the text decoding information are mapped to obtain the image processing result corresponding to the to-be-processed image.
[0128] Specifically, when the image decoding information and the text decoding information are mapped, the visual decoder and the text decoder can be mapped.
[0129] In actual application, in order to realize the conversion of the two cross-modal information of image information and text information, cross-modal mapping is needed. Specifically, after obtaining the target image feature and the aligned target text feature, when the image decoding information and the text decoding information are mapped, the encoder-decoder architecture can be used for mapping, wherein the encoder is used to encode the image feature, and the decoder is used to generate the text sequence. The attention mechanism can be introduced into the encoder-decoder architecture, which can enable the decoder to focus on different parts of the image when generating the text, and the target position feature is used to generate the description step by step according to the target image feature when the aligned target text feature is obtained, thereby improving the relevance of the generated image processing result and the image content.
[0130] In an embodiment of the present specification, the aligned target text feature is input into the text decoder to obtain the text decoding information, comprising:
[0131] The aligned target text feature is input into the text decoder, and the aligned target text feature is autoregressively decoded in the text decoder to obtain the text decoding information;
[0132] The image decoding information and the text decoding information are mapped to obtain the image processing result corresponding to the to-be-processed image, comprising:
[0133] The image decoding information and the text decoding information are mapped to obtain the text result corresponding to the to-be-processed image.
[0134] In this context, the text result corresponding to the image to be processed can be understood as the answer based on plain text output. In this case, the multimodal processing model can process language tasks, and the prompt text can be related to the language task.
[0135] In practical applications, during the processing of language tasks, autoregressive decoding can be used to provide answers based on plain text input by leveraging the capabilities of LLM itself.
[0136] In another embodiment of this specification, image decoding information and text decoding information are mapped to obtain the image processing result corresponding to the image to be processed, including:
[0137] The image decoding information and text decoding information are mapped to obtain the multimodal result at the visual location information corresponding to the image to be processed. The multimodal result includes at least text information and structural information.
[0138] In the case where the multimodal task in the multimodal processing model is a table recognition task, this structural information can be understood as a table structure.
[0139] In practical applications, convolutional image decoders can be used to efficiently process different image tasks in image-related tasks. For document denoising and shadow removal, generative networks are used with mean squared error loss (MSE Loss) to minimize the difference between the denoised and undenoised images; the input is a noisy image, and the output is the denoised image. In document correction tasks, deformable networks, such as Spatial Transformer Networks (STNs), are employed; the input is a tilted or deformed document image, and the output is the corrected image. In multimodal tasks, autoregressive decoding is also used. Unlike language tasks, multimodal tasks require the input to include corresponding positional and structural information at the output. For example, if the image to be processed contains a table containing ID, name, and region, then the image processing result output by the multimodal processing model after processing the image to be processed will be a JSON structure of table recognition: [{"ID":1,"Name":"XX","Age":30,"Region":"Region A"},{"ID":2,"Name":"XXX","Age":25,"Region":"Region B"},{"ID":3,"Name":"XXY","Age":35,"Region":"Region C"}].
[0140] Furthermore, the multimodal processing model provided in the embodiments of this specification can process multilingual tasks and can be connected to different large language models, thus enabling stronger language processing capabilities.
[0141] In summary, through the combination of the visual encoder, the multi-modal processing model and the visual decoder, accurate alignment of vision and language is achieved. In the visual task, the convolution-based visual decoder is used to process the characters, and the generation network and the deformation network are used to improve the authenticity and quality of the image. In the language task, through the autoregressive decoding mechanism, the large language model can give an answer based on pure text output. In the multi-modal task, the visual and location structure information are combined for multi-modal output, ensuring efficient processing of complex tasks.
[0142] In practical applications, the training step of the multi-modal processing model includes:
[0143] The training step of the multi-modal processing model includes:
[0144] Determine the training data corresponding to the image processing task, wherein the image processing task includes a language task and a multi-modal task;
[0145] According to the training data, the initial multi-modal processing model is trained until a multi-modal processing model satisfying a preset training stopping condition is obtained.
[0146] The initial multi-modal processing model can be understood as a machine learning model that has been pre-trained. In practical applications, the initial multi-modal processing model can be a large language model. Satisfying the preset training stopping condition can be understood as the model loss value reaching a preset loss value threshold and / or the number of training times reaching a preset number threshold.
[0147] In practical applications, the multi-modal two-stage multi-task pre-training is performed on the basis of the large language model. The language task includes document type, text question and answer, text translation and text generation. The multi-modal task can be understood as a task that depends on both images and semantics. These tasks need to be aligned within the multi-modal large model, and then output multi-modal output results.
[0148] In specific implementation, determining the training data corresponding to the image processing task includes:
[0149] Determine the image sample and the text label corresponding to the image sample;
[0150] According to the training data, the initial multi-modal processing model is trained until a multi-modal processing model satisfying a preset training stopping condition is obtained, including:
[0151] According to the image sample and the text label, the initial multi-modal processing model is trained until a multi-modal processing model satisfying a preset training stopping condition is obtained.
[0152] Specifically, in the basic pre-training, the machine learning model can be pre-trained by a large amount of data to obtain an initial multi-modal processing model. On the basis of the basic pre-training, the initial multi-modal processing model is trained for an image processing task. For example, in the case where the image processing task is a text recognition task in an image, an image sample corresponding to the task can be determined, and a text label corresponding to the image sample can be determined. The initial multi-modal processing model is trained according to the image sample and the text label to adapt to the text recognition task in the image until a multi-modal processing model satisfying a preset training stopping condition is obtained, so that the multi-modal processing model can process the text recognition task in the image.
[0153] In practical applications, the LLM can be pre-trained as the language model part of the document graph large model. This model is a pre-trained model that has undergone basic language understanding training. The processing capability of the model for document-related tasks can be improved using second-stage pre-training. The second-stage pre-training mainly involves tasks characteristic of documents, and the main tasks used include text detection and recognition tasks, table parsing and recognition tasks, chart parsing tasks, and image-text question answering tasks. Through the second-stage pre-training, the large model not only aligns the visual and language modalities, but also extends the processing capability of the LLM for image tokens and the representation capability of the LLM for image-text space.
[0154] In summary, based on the pre-trained large language model, multi-modal two-stage multi-task pre-training can be performed to extend the language and visual processing capabilities of the model.
[0155] Further, the training data includes text data of multiple language types;
[0156] According to the training data, the initial multi-modal processing model is trained until a multi-modal processing model satisfying a preset training stopping condition is obtained, including:
[0157] According to the training data, the initial multi-modal processing model is trained until a multi-modal processing model satisfying a preset training stopping condition is obtained, including:
[0158] The language types can be understood as languages, and the multiple language types can be Chinese, English, French, Spanish, Portuguese, Russian, Japanese, Chinese, Indonesian, Vietnamese, Thai, Hindi, Arabic, etc.
[0159] Specifically, when training the multi-modal processing model, a multi-language text task can be added, i.e., the initial multi-modal processing model is trained according to the text data of multiple language types and the image data corresponding to the specific image processing task until a multi-modal processing model satisfying a preset training stopping condition is obtained.
[0160] In summary, by adding a multi-language text task, the multi-modal processing model can process multiple language tasks.
[0161] One embodiment of the present specification determines the visual position in the to-be-processed image, the prompt text associated with the to-be-processed image, and the visual position information for the to-be-processed image, determines the target text feature corresponding to the prompt text, determines the target image feature corresponding to the to-be-processed image, and determines the target position feature corresponding to the visual position information, processes the to-be-processed image according to the text content of the prompt text based on the multi-modal features of the target text feature, the target image feature, and the target position feature, obtains the image processing result corresponding to the visual position information, accurately identifies the local area of the to-be-processed image, and can process multiple modalities while ensuring the accuracy of the image processing result.
[0162] The image processing method provided by the present specification will be further described below in combination with FIG. 3, taking the application of the image processing method in image recognition as an example. FIG. 3 is a flowchart of a processing process of an image processing method according to one embodiment of the present specification, which specifically includes the following steps.
[0163] Step 302: Determine the to-be-processed image, determine the prompt text associated with the to-be-processed image, and determine the visual position information for the to-be-processed image.
[0164] For example, the to-be-processed image uploaded by the user, the prompt text "help me identify the text content in the target area" associated with the to-be-processed image input by the user, and the visual position information of the target area obtained by the user by framing the to-be-processed image can be received. The visual position information is the coordinate information of the target area framed by the user.
[0165] Step 304: Feature extraction is performed on the prompt text to obtain the target text feature corresponding to the prompt text.
[0166] Specifically, the text encoder can be used to encode the prompt text to extract the features of the prompt text and obtain the target text feature (i.e., text token) corresponding to the prompt text.
[0167] Step 306: In response to the to-be-processed image containing text characters, the resolution of the to-be-processed image is adjusted based on the to-be-processed image and the text characters to obtain an adjusted to-be-processed image.
[0168] Specifically, in a case where it is determined that the to-be-processed image contains text characters, the text characters in the to-be-processed image can be scaled according to the size of the to-be-processed image and the size of the text characters, so as to adjust the resolution of the to-be-processed image and obtain an adjusted to-be-processed image.
[0169] Step 308: performing feature extraction on the adjusted to-be-processed image to obtain initial image features corresponding to the adjusted to-be-processed image, and performing sampling processing on the initial image features to obtain target image features corresponding to the adjusted to-be-processed image.
[0170] Specifically, the adjusted to-be-processed image can be subjected to multi-level and multi-scale feature extraction by a visual encoder to obtain the initial image features, and the initial image features can be subjected to sampling processing by a resamplar mechanism (i.e., an MLP network including multiple transformer network layers) to obtain the target image features (i.e., visual tokens).
[0171] Step 310: performing position encoding processing on the visual position information to obtain target position features corresponding to the visual position information.
[0172] Specifically, the visual position information can be subjected to position encoding processing to obtain the target position features (i.e., position tokens) corresponding to the visual position information.
[0173] Step 312: inputting the target text features, the target image features, and the target position features into a multi-modal processing model to obtain aligned target text features output by the multi-modal processing model.
[0174] Specifically, the target text features, the target image features, and the target position features can be input into the multi-modal processing model, and the target text features, the target image features, and the target position features can be aligned in the multi-modal processing model to obtain the aligned target text features.
[0175] Step 314: inputting the target image features into a visual decoder to obtain image decoding information output by the visual decoder.
[0176] Step 316: inputting the aligned target text features into a text decoder to obtain text decoding information output by the text decoder.
[0177] Step 318: mapping the image decoding information and the text decoding information to obtain an image processing result of the to-be-processed image.
[0178] The image processing result of the to-be-processed image can be determined according to a specific image processing task, such as an image processing result of a visual task, an image processing result of a language task, and an image processing result of a multi-modal task.
[0179] In the above example, the image processing result of the to-be-processed image is the text content in the recognized target region.
[0180] One embodiment of the present specification realizes the determination of the visual position in the to-be-processed image and the visual position corresponding to the visual position information by determining the to-be-processed image, the prompt text associated with the to-be-processed image, and the visual position information corresponding to the to-be-processed image. Subsequently, the target text feature corresponding to the prompt text is determined, the target image feature corresponding to the to-be-processed image is determined, and the target position feature corresponding to the visual position information is determined. According to the multi-modal features of the target text feature, the target image feature, and the target position feature, the to-be-processed image is processed according to the text content of the prompt text, the image processing result corresponding to the visual position information is obtained, the accurate recognition of the local region of the to-be-processed image is realized, and the accuracy of the image processing result can be ensured while being suitable for multi-modal tasks due to the processing according to the features of multiple modalities.
[0181] Corresponding to the method embodiments described above, the present specification also provides image processing device embodiments. FIG. 4 is a structural schematic diagram of an image processing device according to one embodiment of the present specification. As shown in FIG. 4, the device includes:
[0182] The first acquisition component 402 is configured to acquire a to-be-processed image, a prompt text associated with the to-be-processed image, and determine visual position information corresponding to the to-be-processed image;
[0183] The second acquisition component 404 is configured to acquire a target text feature corresponding to the prompt text, a target image feature corresponding to the to-be-processed image, and a target position feature corresponding to the visual position information;
[0184] The third acquisition component 406 is configured to acquire an image processing result corresponding to the to-be-processed image according to the target text feature, the target image feature, and the target position feature, wherein the image processing result is a result of processing the to-be-processed image according to the text content of the prompt text.
[0185] In an optional embodiment, the first acquisition component 402 is further configured to:
[0186] In response to the to-be-processed image containing text characters, adjusting the resolution of the to-be-processed image according to the to-be-processed image and the text characters to obtain an adjusted to-be-processed image, wherein the resolution of the adjusted to-be-processed image satisfies a preset resolution condition.
[0187] In an optional embodiment, the second acquisition component 404 is further configured to:
[0188] performing feature extraction on the to-be-processed image to obtain an initial image feature corresponding to the to-be-processed image;
[0189] The initial image features are sampled to obtain target image features corresponding to the image to be processed.
[0190] In an optional embodiment, the third obtaining component 406 is further configured to:
[0191] The target text features, the target image features, and the target position features are aligned to obtain aligned target text features;
[0192] The aligned target text features and the target image features are determined to obtain an image processing result corresponding to the image to be processed.
[0193] In an optional embodiment, the third obtaining component 406 is further configured to:
[0194] The target text features, the target image features, and the target position features are input into a multi-modal processing model to obtain aligned target text features output by the multi-modal processing model;
[0195] The target image features are input into a visual decoder to obtain image decoding information of the target image features;
[0196] The aligned target text features are input into a text decoder to obtain text decoding information;
[0197] The image decoding information and the text decoding information are mapped to obtain an image processing result corresponding to the image to be processed.
[0198] In an optional embodiment, the third obtaining component 406 is further configured to:
[0199] The aligned target text features are input into a text decoder, and the aligned target text features are autoregressively decoded in the text decoder to obtain text decoding information;
[0200] The image decoding information and the text decoding information are mapped to obtain an image processing result corresponding to the image to be processed, including:
[0201] The image decoding information and the text decoding information are mapped to obtain a text result corresponding to the image to be processed.
[0202] In an optional embodiment, the third obtaining component 406 is further configured to:
[0203] The image decoding information and the text decoding information are mapped to obtain a multi-modal result at a visual position corresponding to the image to be processed, wherein the multi-modal result at least includes text information and structure information.
[0204] In an optional embodiment, the apparatus further includes a training component configured to:
[0205] determine training data corresponding to the image processing task, wherein the image processing task comprises a language task and a multi-modal task;
[0206] train the initial multi-modal processing model according to the training data until a multi-modal processing model satisfying a preset training stopping condition is obtained.
[0207] In an optional embodiment, the training component is further configured to:
[0208] determine an image sample and determine a text label corresponding to the image sample;
[0209] train the initial multi-modal processing model according to the image sample and the text label until a multi-modal processing model satisfying a preset training stopping condition is obtained.
[0210] In an optional embodiment, the training data comprises text data of multiple language types;
[0211] The training component is further configured to:
[0212] train the initial multi-modal processing model according to the text data of multiple language types until a multi-modal processing model satisfying a preset training stopping condition is obtained.
[0213] In an optional embodiment, the first obtaining component 402 is further configured to:
[0214] receive a point selection instruction for the image to be processed;
[0215] determine visual position information of the image to be processed according to the point selection instruction.
[0216] In an optional embodiment, the visual position information comprises position information of a to-be-processed part in the image to be processed;
[0217] The second obtaining component 404 is further configured to:
[0218] perform position encoding processing on the visual position information according to the position information of the to-be-processed part to obtain target position features corresponding to the visual position information.
[0219] In an optional embodiment, the second obtaining component 404 is further configured to:
[0220] perform feature extraction on the prompt text to obtain target text features corresponding to the prompt text.
[0221] One embodiment of the present specification realizes the determination of the visual position in the to-be-processed image, the target text feature corresponding to the prompt text, the target image feature corresponding to the to-be-processed image, and the target position feature corresponding to the visual position information by determining the to-be-processed image, the prompt text associated with the to-be-processed image, and the visual position information corresponding to the to-be-processed image. According to the multi-modal features of the target text feature, the target image feature, and the target position feature, the to-be-processed image is processed according to the text content of the prompt text, the image processing result corresponding to the visual position information is obtained, the accurate recognition of the local area of the to-be-processed image is realized, and the accuracy of the image processing result can be ensured while being suitable for multi-modal tasks due to the processing according to the features of multiple modalities.
[0222] The above is a schematic scheme of an image processing device according to an embodiment of the present specification. It should be noted that the technical scheme of the image processing device belongs to the same concept as the technical scheme of the image processing method described above, and the details of the technical scheme of the image processing device that are not described in detail can be referred to the description of the technical scheme of the image processing method.
[0223] Referring to FIG. 5, FIG. 5 is a flowchart of another image processing method according to an embodiment of the present specification, which specifically includes the following steps.
[0224] Step 502: receiving a to-be-processed image and a prompt text associated with the to-be-processed image sent by a client, and receiving a point selection instruction for the to-be-processed image sent by the client, and determining visual position information of the to-be-processed image;
[0225] Step 504: obtaining a target text feature corresponding to the prompt text, a target image feature corresponding to the to-be-processed image, and a target position feature corresponding to the visual position information;
[0226] Step 506: obtaining an image processing result corresponding to the to-be-processed image according to the target text feature, the target image feature, and the target position feature, wherein the image processing result is a result of processing the to-be-processed image according to the text content of the prompt text;
[0227] Step 508: sending the image processing result to the client for display on a display interface of the client.
[0228] One embodiment of the present specification realizes the determination of the visual position in the to-be-processed image, the target text feature corresponding to the prompt text, the target image feature corresponding to the to-be-processed image, and the target position feature corresponding to the visual position information by determining the to-be-processed image, the prompt text associated with the to-be-processed image, and the visual position information corresponding to the to-be-processed image. According to the multi-modal features of the target text feature, the target image feature, and the target position feature, the to-be-processed image is processed according to the text content of the prompt text, the image processing result corresponding to the visual position information is obtained, the accurate recognition of the local area of the to-be-processed image is realized, and the accuracy of the image processing result can be ensured while being suitable for multi-modal tasks due to the processing according to the features of multiple modalities.
[0229] The above is a schematic scheme of an image processing method according to an embodiment of the present specification. It should be noted that the technical scheme of the image processing method belongs to the same concept as the technical scheme of the image processing method described above, and the details of the technical scheme of the image processing method that are not described in detail can be referred to the description of the technical scheme of the image processing method.
[0230] Corresponding to the above method embodiment, the present specification also provides an image processing device embodiment. FIG. 6 is a structural schematic diagram of another image processing device according to an embodiment of the present specification. As shown in FIG. 6, the device comprises:
[0231] The receiving component 602 is configured to receive the to-be-processed image and the prompt text associated with the to-be-processed image sent by the client, and receive the point selection instruction for the to-be-processed image sent by the client, and determine the visual position information of the to-be-processed image;
[0232] The fourth obtaining component 604 is configured to obtain the target text feature corresponding to the prompt text, the target image feature corresponding to the to-be-processed image, and the target position feature corresponding to the visual position information;
[0233] The fifth obtaining component 606 is configured to obtain the image processing result corresponding to the to-be-processed image according to the target text feature, the target image feature, and the target position feature, wherein the image processing result is the result of processing the to-be-processed image according to the text content of the prompt text;
[0234] The sending component 608 is configured to send the image processing result to the client and display it through the display interface of the client.
[0235] One embodiment of the present specification realizes the determination of the visual position in the to-be-processed image, the target text feature corresponding to the prompt text, the target image feature corresponding to the to-be-processed image, and the target position feature corresponding to the visual position information by determining the to-be-processed image, the prompt text associated with the to-be-processed image, and the visual position information of the to-be-processed image. According to the multi-modal features of the target text feature, the target image feature, and the target position feature, the to-be-processed image is processed according to the text content of the prompt text, the image processing result corresponding to the visual position information is obtained, the accurate recognition of the local area of the to-be-processed image is realized, and the accuracy of the image processing result can be ensured while being suitable for multi-modal tasks due to the processing according to the features of multiple modalities.
[0236] The above is a schematic scheme of an image processing device according to an embodiment of the present specification. It should be noted that the technical scheme of the image processing device belongs to the same concept as the technical scheme of the image processing method described above, and the details of the technical scheme of the image processing device that are not described in detail can be referred to the description of the technical scheme of the image processing method.
[0237] Corresponding to the method embodiments described above, the present specification also provides a model training platform embodiment. The model training platform comprises a request interface component, a model training component, and a response component.
[0238] The request interface component is configured to receive a model training request, wherein the model information of the machine learning model to be trained is included in the model training request.
[0239] The model training component is configured to determine the machine learning model to be trained according to the model information, and to perform model training on the machine learning model to be trained to obtain a trained multi-modal processing model. The multi-modal processing model is used to process the target text feature of the prompt text, the target image feature of the to-be-processed image, and the target position feature corresponding to the visual position information to obtain the image processing result of the to-be-processed image. The prompt text is the prompt text associated with the to-be-processed image, and the visual position information is the visual position information of the to-be-processed image.
[0240] The response component is configured to output the multi-modal processing model.
[0241] In an optional embodiment, the model training platform further comprises a data receiving component.
[0242] The data receiving component is configured to receive the training data input by a user and send the training data to the model training component.
[0243] The model training component is further configured to perform model training on the machine learning model to be trained according to the training data, to obtain the trained multi-modal processing model.
[0244] In an optional embodiment, the model training platform further comprises a model library, wherein the model library stores a plurality of machine learning models.
[0245] The model training component is further configured to determine the machine learning model to be trained from the model library according to the model information.
[0246] In summary, by training the multi-modal processing model by using the model training platform, the processing performance of the multi-modal processing model is ensured, so that the multi-modal processing model can process multi-modal tasks in subsequent applications.
[0247] FIG. 7 is a structural block diagram of a computing device according to an embodiment of the present specification. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to save data.
[0248] The computing device 700 further comprises an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 can include one or more of any type of network interface (for example, a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and the like.
[0249] In one embodiment of the present disclosure, the above-mentioned components of the computing device 700 and other components not shown in FIG. 7 can also be connected to each other, for example, through a bus. It should be understood that the computing device structure block diagram shown in FIG. 7 is merely for the purpose of example, and is not a limitation on the scope of the present disclosure. Other components can be added or replaced as needed by those skilled in the art.
[0250] The computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 700 can also be a mobile or stationary server.
[0251] The processor 720 is configured to execute computer programs / instructions that implement the steps of the above-mentioned image processing method when executed by the processor.
[0252] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, the computing device embodiment is described simply because it is basically similar to the image processing method embodiment, and the relevant part can be referred to the part of the image processing method embodiment.
[0253] An embodiment of the present specification also provides a computer-readable storage medium storing computer programs / instructions that implement the steps of the above-mentioned image processing method when executed by the processor.
[0254] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, the computer-readable storage medium embodiment is described simply because it is basically similar to the image processing method embodiment, and the relevant part can be referred to the part of the image processing method embodiment.
[0255] An embodiment of the present specification also provides a computer program product comprising computer programs / instructions that implement the steps of the above-mentioned image processing method when executed by the processor.
[0256] An embodiment of the present specification also provides a computer program, wherein the computer program implements the steps of the above-mentioned image processing method when executed by the processor.
[0257] The above is a schematic scheme of a computer program product of the embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the image processing method described above belong to the same concept, and the details of the technical scheme of the computer program product that are not described in detail can be referred to the description of the technical scheme of the image processing method.
[0258] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in which they are recited in the embodiments and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.
[0259] The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer-readable medium can include any entity or apparatus that can carry the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (Read-Only Memory, ROM for short), random access memory (Random Access Memory, RAM for short), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice, for example, in some regions, according to the patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0260] It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present specification are not limited by the order of the described actions, because according to the embodiments of the present specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the embodiments of the present specification.
[0261] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0262] The preferred embodiments of the present specification disclosed above are only used to help illustrate the present specification. Alternative embodiments do not describe all the details and limit the present application to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification. The present specification is limited only by the claims and their full scope and equivalents. Industrial applicability
[0263] The scheme provided by the embodiments of the present disclosure can be applied in the processing of images, obtaining a to-be-processed image, prompt text associated with the to-be-processed image, and determining visual position information for the to-be-processed image; obtaining target text features corresponding to the prompt text, target image features corresponding to the to-be-processed image, and target position features corresponding to the visual position information; and obtaining an image processing result corresponding to the to-be-processed image according to the target text features, the target image features, and the target position features, wherein the image processing result is a result of processing the to-be-processed image according to the text content of the prompt text, thereby solving the technical problem of low accuracy of the image processing result.
Claims
1. An image processing method, comprising: obtaining a to-be-processed image, prompt text associated with the to-be-processed image, and visual position information of the to-be-processed image; obtaining target text features corresponding to the prompt text, target image features corresponding to the to-be-processed image, and target position features corresponding to the visual position information; obtaining an image processing result corresponding to the to-be-processed image according to the target text features, the target image features, and the target position features, wherein the image processing result is a result of processing the to-be-processed image according to text content of the prompt text.
2. The image processing method of claim 1, after the obtaining the to-be-processed image, further comprising: in response to the to-be-processed image containing text characters, adjusting a resolution of the to-be-processed image according to the to-be-processed image and the text characters to obtain an adjusted to-be-processed image, wherein the resolution of the adjusted to-be-processed image satisfies a preset resolution condition.
3. The image processing method of claim 1 or 2, the obtaining the target image features corresponding to the to-be-processed image, comprising: performing feature extraction on the to-be-processed image to obtain initial image features corresponding to the to-be-processed image; performing sampling processing on the initial image features to obtain the target image features corresponding to the to-be-processed image.
4. The image processing method of claim 1, the obtaining the image processing result corresponding to the to-be-processed image according to the target text features, the target image features, and the target position features, comprising: performing alignment processing on the target text features, the target image features, and the target position features to obtain aligned target text features; determining the image processing result corresponding to the to-be-processed image according to the aligned target text features and the target image features.
5. The image processing method of claim 1, the obtaining the image processing result corresponding to the to-be-processed image according to the target text features, the target image features, and the target position features, comprising: inputting the target text features, the target image features, and the target position features into a multi-modal processing model to obtain aligned target text features output by the multi-modal processing model; inputting the target image features into a visual decoder to obtain image decoding information of the target image features; inputting the aligned target text features into a text decoder to obtain text decoding information; mapping the image decoding information and the text decoding information to obtain the image processing result corresponding to the to-be-processed image.
6. The image processing method of claim 5, the inputting the aligned target text features into a text decoder to obtain text decoding information, comprising: inputting the aligned target text features into the text decoder, and performing autoregressive decoding on the aligned target text features in the text decoder to obtain the text decoding information. The image decoding information and the text decoding information are mapped to obtain an image processing result corresponding to the image to be processed. The image decoding information and the text decoding information are mapped to obtain a text result corresponding to the image to be processed.
7. The image processing method of claim 5, wherein the image decoding information and the text decoding information are mapped to obtain an image processing result corresponding to the image to be processed. The image decoding information and the text decoding information are mapped to obtain a multi-modal result at the visual position information corresponding to the image to be processed, wherein the multi-modal result at least includes text information and structure information.
8. The image processing method of any one of claims 5-7, wherein the training step of the multi-modal processing model comprises: determining training data corresponding to an image processing task, wherein the image processing task includes a language task and a multi-modal task; training an initial multi-modal processing model according to the training data until the multi-modal processing model satisfying a preset training stop condition is obtained.
9. The image processing method of claim 8, wherein the determining training data corresponding to an image processing task comprises: determining an image sample and determining a text label corresponding to the image sample; the training an initial multi-modal processing model according to the training data until the multi-modal processing model satisfying a preset training stop condition is obtained comprises: training the initial multi-modal processing model according to the image sample and the text label until the multi-modal processing model satisfying the preset training stop condition is obtained.
10. The image processing method of claim 8, wherein the training data includes text data of multiple language types; the training an initial multi-modal processing model according to the training data until the multi-modal processing model satisfying a preset training stop condition is obtained comprises: training the initial multi-modal processing model according to the text data of the multiple language types until the multi-modal processing model satisfying the preset training stop condition is obtained.
11. The image processing method of claim 1, wherein the obtaining visual position information of the image to be processed comprises: receiving a point selection instruction for the image to be processed; determining the visual position information of the image to be processed according to the point selection instruction.
12. The image processing method of claim 1, wherein the visual position information includes position information of a to-be-processed part in the image to be processed; the obtaining a target position feature corresponding to the visual position information comprises: performing position encoding processing on the visual position information according to the position information of the to-be-processed part to obtain the target position feature corresponding to the visual position information.
13. An image processing method, comprising: receiving an image to be processed and prompt text associated with the image to be processed sent by a client, and receiving a point selection instruction for the image to be processed sent by the client to determine visual position information of the image to be processed. obtain a target text feature corresponding to the prompt text, a target image feature corresponding to the to-be-processed image, and a target position feature corresponding to the visual position information; obtain an image processing result corresponding to the to-be-processed image according to the target text feature, the target image feature, and the target position feature, wherein the image processing result is a result of processing the to-be-processed image according to text content of the prompt text; send the image processing result to the client, and display the image processing result on a display interface of the client.
14. A model training platform, comprising a request interface component, a model training component, and a response component; the request interface component is configured to receive a model training request, wherein the model information of a machine learning model to be trained is included in the model training request; the model training component is configured to determine the machine learning model to be trained according to the model information, and perform model training on the machine learning model to be trained to obtain a trained multi-modal processing model, wherein the multi-modal processing model is used to process a target text feature of a prompt text, a target image feature of a to-be-processed image, and a target position feature corresponding to visual position information, and obtain an image processing result of the to-be-processed image, the prompt text is a prompt text associated with the to-be-processed image, and the visual position information is visual position information of the to-be-processed image; the response component is configured to output the multi-modal processing model.
15. The model training platform of claim 14, further comprising a data receiving component, the data receiving component is configured to receive training data input by a user, and send the training data to the model training component; wherein, the model training component is further configured to perform model training on the machine learning model to be trained according to the training data to obtain the trained multi-modal processing model.
16. The model training platform of claim 14, further comprising a model library, wherein, the model library stores a plurality of machine learning models; the model training component is further configured to determine the machine learning model to be trained from the model library according to the model information.
17. A computing device, comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, and the computer programs / instructions, when executed by the processor, implement the steps of the method of any one of claims 1 to 13.
18. A computer-readable storage medium storing computer programs / instructions, wherein the computer programs / instructions, when executed by a processor, implement the steps of the method of any one of claims 1 to 13.
19. A computer program product comprising computer programs / instructions, wherein the computer programs / instructions, when executed by a processor, implement the steps of the method of any one of claims 1 to 13.
20. A computer program, wherein, the computer programs, when executed by a processor, implement the steps of the method of any one of claims 1 to 13.
Citation Information
Patent Citations
Cross-modal processing for vision and language
CN115017911A
Cross-modal-based image recognition model training method and apparatus, and electronic device
CN117635959A
Image processing method and device, equipment and medium
CN117711001A
Model training method and device, equipment, storage medium and program product
CN118378633A