Image processing method, image processing model training method and model training platform

By fusing sample image and text features into an image processing model, the model is trained to learn image and text knowledge, which solves the problem of insufficient accuracy and robustness of existing models in multimodal data processing and improves the accuracy of image processing.

CN121640209APending Publication Date: 2026-03-10ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to balance the accuracy, diversity, and robustness of models when processing complex multimodal data, leading to inaccurate image processing results.

Method used

By identifying the features of sample images and sample text, feature fusion processing is performed to train the image processing model, which learns image and text knowledge, thereby improving the model's accuracy and robustness.

Benefits of technology

The image processing model achieves accuracy and robustness in multimodal data processing, ensuring the accuracy of image processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640209A_ABST
    Figure CN121640209A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of computers, in particular to an image processing method, an image processing model training method and a model training platform, and the image processing method comprises the steps: determining a to-be-processed image; inputting the to-be-processed image into an image processing model to obtain a target image feature of the to-be-processed image; determining an image processing result of the to-be-processed image according to the target image feature; wherein the image processing model is obtained by training according to a sample image feature of a sample image, a first sample text feature of a first sample text and a fused sample feature, the fused sample feature is determined according to the sample image feature and a second sample text feature, and the second sample text feature is determined according to a second sample text; the first sample text and the second sample text are determined according to the sample image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of computer technology, and in particular to image processing methods, image processing model training methods, and model training platforms. Background Technology

[0002] With the development of computer technology, image-to-text (IPT) technology has made significant progress in computer vision and natural language processing. It typically relies on large pre-trained models to implement image processing techniques such as IPT. For example, contrastive learning can establish effective connections between images and text, thereby improving image processing performance. However, this single optimization method considers textual information, which can only describe a portion of the image content. When processing complex multimodal data, it is difficult to balance the accuracy, diversity, and robustness of the model, leading to inaccurate image processing results. Summary of the Invention

[0003] In view of this, the embodiments of this specification provide three image processing methods. One or more embodiments of this specification simultaneously relate to three image processing apparatuses, an image processing model training method, an image processing model training apparatus, a model training platform, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, an image processing method is provided, comprising:

[0005] Identify the image to be processed;

[0006] The image to be processed is input into the image processing model to obtain the target image features of the image to be processed;

[0007] Based on the target image features, determine the image processing result of the image to be processed;

[0008] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0009] According to a second aspect of the embodiments of this specification, an image processing apparatus is provided, comprising:

[0010] The first determining module is configured to determine the image to be processed;

[0011] The input module is configured to input the image to be processed into an image processing model to obtain the target image features of the image to be processed;

[0012] The second determining module is configured to determine the image processing result of the image to be processed based on the features of the target image;

[0013] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0014] According to a third aspect of the embodiments of this specification, an image processing model training method is provided, comprising:

[0015] The machine learning model to be trained is determined, the sample images associated with the image processing task are determined, and the first sample text and the second sample text corresponding to the sample images are determined.

[0016] Based on the machine learning model to be trained, determine the sample image features of the sample image;

[0017] Determine the first sample text features of the first sample text, and determine the second sample text features of the second sample text;

[0018] The sample image features and the second sample text features are fused to obtain fused sample features;

[0019] The machine learning model to be trained is trained based on the sample image features, the first sample text features, and the fused sample features to obtain a trained image processing model.

[0020] According to a fourth aspect of the embodiments of this specification, an image processing model training apparatus is provided, comprising:

[0021] The first determining module is configured to determine the machine learning model to be trained, determine the sample image associated with the image processing task, and determine the first sample text and the second sample text corresponding to the sample image.

[0022] The second determining module is configured to determine the sample image features of the sample image based on the machine learning model to be trained.

[0023] The third determining module is configured to determine a first sample text feature of the first sample text and a second sample text feature of the second sample text.

[0024] The fusion module is configured to perform fusion processing on the sample image features and the second sample text features to obtain fused sample features;

[0025] The training module is configured to train the machine learning model to be trained based on the sample image features, the first sample text features, and the fused sample features, to obtain a trained image processing model.

[0026] According to a fifth aspect of the embodiments of this specification, an image processing method is provided, comprising:

[0027] Receive product images of the target product sent by the client;

[0028] The product image is input into an image processing model to obtain the target image features of the product image;

[0029] Based on the features of the target image, other product images associated with the product image are determined;

[0030] Send the other product images to the client;

[0031] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0032] According to a sixth aspect of the embodiments of this specification, an image processing apparatus is provided, comprising:

[0033] The receiving module is configured to receive product images of the target product sent by the client;

[0034] The input module is configured to input the product image into an image processing model to obtain the target image features of the product image;

[0035] The determining module is configured to determine other product images associated with the product image based on the features of the target image;

[0036] The sending module is configured to send the other product images to the client.

[0037] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0038] According to a seventh aspect of the embodiments of this specification, an image processing method is provided, comprising:

[0039] Receive the image to be processed sent by the client;

[0040] The image to be processed is input into the image processing model to obtain the target image features of the image to be processed;

[0041] Based on the features of the target image, determine the descriptive text of the image to be processed;

[0042] Send the description text to the client;

[0043] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0044] According to an eighth aspect of the embodiments of this specification, an image processing apparatus is provided, comprising:

[0045] The receiving module is configured to receive images to be processed sent by the client;

[0046] The input module is configured to input the image to be processed into an image processing model to obtain the target image features of the image to be processed;

[0047] The determination module is configured to determine the descriptive text of the image to be processed based on the features of the target image;

[0048] The sending module is configured to send the description text to the client;

[0049] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0050] According to a ninth aspect of the embodiments of this specification, a model training platform is provided, including a request interface unit, a model training unit, and a response unit;

[0051] The request interface unit is used to receive a model training request, wherein the model training request includes model information of the machine learning model to be trained.

[0052] The model training unit is used to determine the machine learning model to be trained based on the model information, and to train the machine learning model to be trained to obtain a trained image processing model. The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0053] The response unit is used to output the image processing model.

[0054] According to a tenth aspect of the embodiments of this specification, a computing device is provided, comprising:

[0055] Memory and processor;

[0056] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which implement the steps of the above method when executed by the processor.

[0057] According to an eleventh aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0058] According to a twelfth aspect of an embodiment of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method described above.

[0059] One embodiment of this specification implements the following process during the training of an image processing model: determining sample images, corresponding first and second sample texts, and defining sample image features, first sample text features, and second sample text features. Based on these features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained using these features, enabling it to learn image knowledge, text knowledge, and image-text fusion knowledge. This knowledge guides the model's learning, improving its accuracy, diversity, and robustness. This facilitates the subsequent processing of images by ensuring the accuracy of the target image features extracted from the model, further guaranteeing the accuracy of the image processing results determined based on these target image features. Attached Figure Description

[0060] Figure 1 This is a schematic diagram illustrating an application scenario of an image processing method provided in one embodiment of this specification;

[0061] Figure 2 This is a flowchart illustrating an image processing method provided in one embodiment of this specification;

[0062] Figure 3 This is a flowchart illustrating an image processing model training method provided in one embodiment of this specification;

[0063] Figure 4 This is a flowchart illustrating the processing procedure of an image processing model training method provided in one embodiment of this specification.

[0064] Figure 5 This is a flowchart illustrating an image processing method provided in one embodiment of this specification;

[0065] Figure 6 This is a flowchart illustrating an image processing method provided in one embodiment of this specification;

[0066] Figure 7 This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification;

[0067] Figure 8 This is a schematic diagram of the structure of an image processing model training device provided in one embodiment of this specification;

[0068] Figure 9 This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification;

[0069] Figure 10This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification;

[0070] Figure 11 This is a schematic diagram of the structure of a model training platform provided in one embodiment of this specification;

[0071] Figure 12 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0072] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0073] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0074] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0075] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0076] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0077] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0078] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0079] Multimodal feature fusion: First, features of data of different modalities (such as images and text) are extracted using corresponding deep learning models, and then the features of different modalities are fused together.

[0080] Teacher model: In the field of model distillation, there is generally one teacher model and one student model. The teacher model usually performs better and can be used to guide the training of the student model.

[0081] K-means clustering: A typical clustering algorithm that clusters a batch of data, assigning the same category label to similar samples.

[0082] To address the aforementioned technical problems, this specification provides three image processing methods, as well as three image processing devices, an image processing model training method, an image processing model training device, a model training platform, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0083] See Figure 1 , Figure 1 A schematic diagram illustrating an application scenario of an image processing method provided according to an embodiment of this specification is shown.

[0084] Figure 1 The system includes a terminal device 102 and a cloud device 104. In a specific implementation, the user sends an image to be processed to the cloud device 104 through the terminal device 102. The cloud device 104 inputs the image to be processed into an image processing model to obtain the target image features of the image to be processed. Based on the target image features, the cloud device 104 determines the image processing result of the image to be processed and sends the image processing result back to the terminal device 102, which then displays it to the user through the display interface of the terminal device 102.

[0085] The edge device 102 may include a browser, an app (application), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. The edge device can be developed based on a software development kit (SDK) provided by the server, such as a real-time communication (RTC) SDK. The edge device can be deployed in an electronic device and depends on the device's operation or certain apps within the device to run. The electronic device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0086] Cloud-side device 104 can be understood as a server providing various services, including physical servers and cloud servers. Examples include servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that cloud-side device 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. Cloud-side device 104 can also be a server for a distributed system, or a server integrated with blockchain. Cloud-side device 104 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0087] It is worth noting that the image processing method provided in the embodiments of this specification can be executed by the cloud-side device 104. In other embodiments of this specification, the image processing model can be deployed in the edge device 102, so that the edge device 102 can also have similar functions to the cloud-side device 104, thereby executing the image processing method provided in the embodiments of this specification. In other embodiments, the image processing method provided in the embodiments of this specification can also be jointly executed by the edge device 102 and the cloud-side device 104.

[0088] See Figure 2 , Figure 2 A flowchart of an image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0089] Step 202: Determine the image to be processed.

[0090] Specifically, the image processing methods provided in the embodiments of this specification can be applied to the fields of image retrieval and image-to-text (IPT) processing. For example, in the field of image retrieval, the image processing methods provided in the embodiments of this specification can be used to retrieve other images similar to the image to be processed. In the field of IPT processing, the image processing methods provided in the embodiments of this specification can be used to generate image description text corresponding to the image to be processed.

[0091] Here, the image to be processed can be understood as the image that needs to be processed.

[0092] In practical applications, the image to be processed can be uploaded by the user through the client, or it can be selected by the user through the client.

[0093] For example, on an e-commerce platform, if a user uploads an image of a mobile phone of brand XX through the client, then that image of the mobile phone is the image to be processed.

[0094] Step 204: Input the image to be processed into the image processing model to obtain the target image features of the image to be processed.

[0095] In this context, the image processing model can be understood as a trained neural network model. In practical applications, the image processing model can be a transformer neural network model. Preferably, the image processing model can include multiple network layers, such as 12 or 24 layers. The target image features can be understood as the image feature vector of the image to be processed.

[0096] Based on this, the image to be processed can be input into a trained neural network model to obtain the image feature vector of the image to be processed, which is output by the model.

[0097] In one embodiment of this specification, the image processing model can be deployed on a client, and the client can directly input the image to be processed into the image processing model to obtain the target image features of the image to be processed.

[0098] In another embodiment of this specification, the image processing model can be deployed on a server. The client can send the image to be processed to the server, and the server can input the image to be processed into the image processing model to obtain the target image features of the image to be processed.

[0099] In another embodiment of this specification, the image processing service provided by the image processing model can be deployed on the server through the model call interface. The client can send the image to be processed to the server, and the server can call the image processing model through the model call interface, input the image to be processed into the image processing model, and obtain the target image features of the image to be processed.

[0100] Using the previous example, we can input the image of a mobile phone of brand XX into the image processing model to obtain the target image features of the mobile phone.

[0101] Step 206: Determine the image processing result of the image to be processed based on the target image features.

[0102] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0103] In one embodiment of this specification, determining the image processing result of the image to be processed based on the target image features includes:

[0104] Obtain features from multiple reference texts;

[0105] Based on the similarity between the target image features and the plurality of reference text features, the target text features corresponding to the target image features are determined from the plurality of reference text features;

[0106] The target text corresponding to the target text feature is determined, and the target text is determined as the image processing result of the image to be processed.

[0107] In this context, the reference text features can be understood as the text features used for reference in the image-to-text task. These reference text features can be multiple text features stored in a database, each corresponding to a text. A text feature can be understood as a feature vector of the text. The target text can be understood as the image description text of the image to be processed.

[0108] Based on this, when processing image-to-text tasks, multiple reference text features can be obtained from the database. The similarity between the target image features and each of the multiple reference text features can be calculated. The reference text feature with the highest similarity to the target image features can be identified as the target text feature. The target text corresponding to the target text feature can be identified, and the target text can be identified as the image processing result of the image to be processed.

[0109] Understandably, the database can be determined based on the task type of the image processing task. For example, if the image processing task is a text-to-image task of landscape images, then the database can store reference text features of image description text about landscape images.

[0110] Using the previous example, multiple reference text features can be obtained from the database. The similarity between the target image feature of the mobile phone image and each of the multiple reference text features is calculated. The reference text feature with the highest similarity to the target image feature is determined as the target text feature, and the target text corresponding to the target text feature is determined to be "XX brand mobile phone".

[0111] In summary, by determining the target text features based on similarity, the corresponding image description text for the image to be processed can be determined.

[0112] In another embodiment of this specification, determining the image processing result of the image to be processed based on the target image features includes:

[0113] Acquire features from multiple reference images;

[0114] Based on the similarity between the target image features and the plurality of reference image features, the target reference image features corresponding to the target image features are determined from the plurality of reference image features;

[0115] The target image corresponding to the features of the target reference image is determined, and the target image is determined as the image processing result of the image to be processed.

[0116] In this context, reference image features can be understood as image features used for reference in similar image retrieval tasks. These reference image features can be multiple image features stored in a database, with each feature corresponding to a single image. Image features can be understood as feature vectors of an image. The target image can be understood as an image similar to the image to be processed.

[0117] Based on this, when processing image retrieval tasks, multiple reference image features can be obtained from the database, the similarity between the target image feature and each of the multiple reference image features can be calculated, the reference image feature with the highest similarity to the target image feature can be determined as the target reference image feature, the target image corresponding to the target reference image feature can be determined, and the target image can be determined as the image processing result of the image to be processed.

[0118] Using the previous example, multiple reference image features can be obtained from the database. The similarity between the target image feature of the mobile phone image and each of the multiple reference image features is calculated. The reference image feature with the highest similarity to the target image feature is determined as the target reference image feature. The target image corresponding to the target reference image feature can be other mobile phone images of the XX brand, or it can be a mobile phone image of the YY brand, which is similar to the XX brand.

[0119] In summary, by determining the features of the target reference image based on similarity, other images similar to the image to be processed can be identified.

[0120] In practical applications, determining the image to be processed includes:

[0121] Receive the image to be processed sent by the client;

[0122] After determining the image processing result of the image to be processed, the method further includes:

[0123] The image processing results are sent to the client and displayed through the client's interface.

[0124] Specifically, the image processing method provided in the embodiments of this specification can be applied to the server. The server can receive the image to be processed sent by the client, input the image to be processed into the image processing model, obtain the target image features corresponding to the image to be processed, determine the image processing result of the image to be processed based on the target image features, and send the image processing result to the client for display through the client's display interface.

[0125] In summary, interaction with the user is achieved by receiving the image to be processed sent by the client and returning the image processing result of the image to the client.

[0126] In practical applications, the training steps of the image processing model include:

[0127] The machine learning model to be trained is determined, the sample images associated with the image processing task are determined, and the first sample text and the second sample text corresponding to the sample images are determined.

[0128] Based on the machine learning model to be trained, determine the sample image features of the sample image;

[0129] Determine the first sample text features of the first sample text, and determine the second sample text features of the second sample text;

[0130] The sample image features and the second sample text features are fused to obtain fused sample features;

[0131] The machine learning model to be trained is trained based on the sample image features, the first sample text features, and the fused sample features to obtain a trained image processing model.

[0132] Specifically, the training steps of the image processing model are similar to the image processing model training method provided in the embodiments of this specification. For details, please refer to the following description of the image processing model training method, which will not be repeated here.

[0133] See Figure 3 , Figure 3 A flowchart of an image processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0134] Step 302: Determine the machine learning model to be trained, determine the sample images associated with the image processing task, and determine the first sample text and the second sample text corresponding to the sample images.

[0135] In this context, the machine learning model to be trained can be understood as the neural network model that needs to be trained. The first sample text and the second sample text can be understood as image description text for the sample images. The first sample text and the second sample text can be the same or different. These two sample texts can be used for comparative learning and distillation learning of the subsequent model, respectively, to improve the model's performance.

[0136] Based on this, the machine learning model to be trained is determined, the sample image associated with the image processing task is determined, and the first sample text and the second sample text describing the sample image are determined based on the sample image.

[0137] Understandably, sample images, the corresponding first sample text, and the second sample text can be used as a set of sample data to train the machine learning model. Usually, multiple sets of sample data can be used to train the machine learning model to be trained.

[0138] Step 304: Determine the sample image features of the sample image based on the machine learning model to be trained.

[0139] The sample image features can be understood as the image feature vector of the sample image.

[0140] Specifically, taking one set of sample data as an example, the sample image can be input into the machine learning model to be trained to obtain the sample image features.

[0141] Step 306: Determine the first sample text features of the first sample text, and determine the second sample text features of the second sample text.

[0142] In this context, the first sample text feature can be understood as the feature vector of the first sample text, and the second sample text feature can be understood as the feature vector of the second sample text.

[0143] Specifically, the first sample text and the second sample text can be input into the text processing model to obtain the first sample text features of the first sample text and the second sample text features of the second sample text output by the text processing model.

[0144] In practical applications, text processing models can be neural network models or any model used to extract text features.

[0145] Step 308: Perform fusion processing on the sample image features and the second sample text features to obtain fused sample features.

[0146] Specifically, the sample image features and the second sample text features can be input into the fusion model to obtain fused sample features.

[0147] In practical applications, the fusion model can be a neural network model or any model used to fuse features.

[0148] In specific implementation, the step of fusing the sample image features and the second sample text features to obtain fused sample features includes:

[0149] The sample image features and the second sample text features are added together to obtain fused sample features; or

[0150] The sample image features and the second sample text features are weighted and summed to obtain fused sample features.

[0151] In one embodiment of this specification, the sample image features and the second sample text features can be added together to achieve the fusion of the sample image features and the second sample text features.

[0152] In another embodiment of this specification, the sample image features and the second sample text features can be weighted and summed to achieve the fusion of the sample image features and the second sample text features.

[0153] It is understandable that other fusion methods can also be used to fuse the features of the sample image and the features of the second sample text, such as feature concatenation, feature multiplication, etc., but the embodiments in this specification do not limit this.

[0154] In summary, by fusing features, the fused sample features become multimodal features that combine image and text knowledge, which facilitates the enhancement of the model's multimodal data processing capabilities during subsequent model training.

[0155] Step 310: Train the machine learning model to be trained based on the sample image features, the first sample text features, and the fused sample features to obtain the trained image processing model.

[0156] In specific implementation, the step of training the machine learning model to be trained based on the sample image features, the first sample text features, and the fused sample features to obtain the trained image processing model includes:

[0157] Based on the features of the sample image and the features of the first sample text, the machine learning model to be trained is trained to obtain the initial image processing model after training.

[0158] The initial image processing model is trained based on the sample image features and the fused sample features to obtain a trained image processing model.

[0159] Specifically, the machine learning model to be trained can be trained by comparative learning based on the features of the sample images and the features of the first sample text to obtain the initial image processing model. Then, the initial image processing model can be trained by distillation learning based on the features of the sample images and the features of the fused samples to obtain the final image processing model.

[0160] In summary, by combining contrastive learning and distillation learning, the image processing model can learn image processing capabilities and the ability to use the fusion model as a teacher model.

[0161] Further, the step of training the initial image processing model based on the sample image features and the fused sample features to obtain the trained image processing model includes:

[0162] Determine the category information corresponding to the features of the fused samples;

[0163] The sample image features are used as training samples, and the category information is used as training labels to train the initial image processing model, thereby obtaining the trained image processing model.

[0164] The category information corresponding to the fused sample features can be defined category information, such as category information 1 corresponding to fused sample feature F1 and category information 2 corresponding to fused sample feature F2.

[0165] Specifically, we can define the category information corresponding to the fused sample features, use the sample image features as training samples, use the category information as pseudo-classification labels, train the initial image processing model, and obtain the trained image processing model.

[0166] In summary, combining classification training with image processing model training improves model performance.

[0167] In practical applications, there are multiple sample images, and both the sample image features and the fused sample features are multiple.

[0168] Determining the category information corresponding to the fused sample features includes:

[0169] Clustering is performed on multiple fused sample features to obtain clustering results, and the category information of each fused sample feature is determined based on the clustering results.

[0170] The step of using the sample image features as training samples and the category information as training labels to train the initial image processing model to obtain the trained image processing model includes:

[0171] Multiple sample image features are used as training samples, and the category information of each fused sample feature is used as training labels to train the initial image processing model, thereby obtaining the trained image processing model.

[0172] Specifically, clustering algorithms can be used to cluster multiple fused sample features to obtain clustering results. Based on the clustering results, the category information of each fused sample feature can be determined. The multiple sample image features are used as training samples, and the category information of each fused sample feature is used as training labels to train the initial image processing model, thus obtaining the trained image processing model.

[0173] In practical applications, the clustering algorithm can be the K-means clustering algorithm, which can cluster multiple fused sample features into 4096 classes using K-means clustering and calculate the category information corresponding to each fused sample feature.

[0174] In summary, clustering algorithms can be used to cluster the features of multiple fused samples, which facilitates the subsequent determination of the corresponding category information and reduces the computation time and computational resource consumption of clustering multiple fused sample features.

[0175] Further, the step of training the machine learning model to be trained based on the sample image features and the first sample text features to obtain a trained initial image processing model includes:

[0176] Calculate the loss function based on the sample image features and the first sample text features;

[0177] The machine learning model to be trained is trained according to the loss function to obtain the initial image processing model after training.

[0178] In practical applications, the loss function is shown in the following formula.

[0179]

[0180] Where N is the total number of samples used for model training, and i and j are the sample labels, i.e., the i-th sample and the j-th sample, Let V be the loss function. i Let be the sample image features of the i-th sample image, where 'a' is the first sample text and 'b' is the second sample text. For the first sample text feature of the i-th first sample text, Let τ1 be the second sample text feature of the j-th second sample text, τ1 be a preset parameter value (which can be 0.07 in practical applications), and σ be the model parameter.

[0181] In summary, by conducting comparative learning training on the image processing model, the image processing model learns image processing capabilities.

[0182] In summary, one embodiment of this specification implements the following process during the training of an image processing model: determining sample images, corresponding first and second sample texts, and defining sample image features, first sample text features, and second sample text features. Based on these features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained using these features, enabling it to learn image knowledge, text knowledge, and image-text fusion knowledge during training. This knowledge guides the model's learning, improving its accuracy, diversity, and robustness. Furthermore, it ensures the accuracy of the image processing results determined based on these target image features during subsequent image processing, leveraging the accuracy of the target image features extracted by the model.

[0183] The following is in conjunction with the appendix Figure 4 Taking the image processing model training method provided in this specification as an example in model training, the image processing model training method will be further explained. Among them, Figure 4 The present specification illustrates a flowchart of an image processing model training method according to an embodiment, which includes the following steps.

[0184] Step 402: Determine the sample image, and the first sample text and the second sample text corresponding to the sample image.

[0185] For example, sample images can be like Figure 4 As shown, the first sample text could be "Two cats are snuggling together", and the second sample text could be "The cat is in the bamboo basket".

[0186] Step 404: Perform comparative learning on the image processing model EI based on the sample image and the first sample text.

[0187] Specifically, the text processing model ET is used to extract text features from the first sample text to obtain the first sample text features. The image processing model EI is used to extract image features from the sample image to obtain the sample image features. Based on the first sample text features and the sample image features, a loss function is calculated. Based on the loss function, the image processing model EI is trained.

[0188] Step 406: Perform distillation learning on the image processing model EI based on the sample images and the second sample text.

[0189] Specifically, the text processing model ET is used to extract features from the second sample text to obtain its corresponding second sample text features. These second sample text features and sample image features are then input into the fusion model EF (i.e., the teacher model) to obtain fused sample features. A similar process is performed on each sample image to obtain multiple fused sample features. These fused sample features are then clustered to define the category information corresponding to each fused sample feature. The sample image features are used as training samples, and the category information is used as training labels to train the image processing model EI.

[0190] Furthermore, during distillation learning using a fusion model, the gradient uppropagation operation can be stopped to prevent model parameters from being updated in certain situations.

[0191] See Figure 5 , Figure 5 A flowchart of an image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0192] Step 502: Receive the product image of the target product sent by the client;

[0193] Step 504: Input the product image into the image processing model to obtain the target image features of the product image;

[0194] Step 506: Based on the features of the target image, determine other product images associated with the product image;

[0195] Step 508: Send the other product images to the client;

[0196] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0197] The product image of the target product can be understood as the image to be processed, and other product images can be understood as the image processing results of the image to be processed. Other product images associated with the target product image can be understood as other product images similar to the target product image.

[0198] Specifically, the image processing method provided in the embodiments of this specification can be applied to the field of image retrieval. In specific implementation, the server can receive the product image of the target product sent by the client, input the product image into the image processing model, obtain the target image features of the product image, determine other product images similar to the product image based on the target image features, and send the other product images to the client for display through the client's display interface.

[0199] Understandably, the training process of this image processing model is similar to that of the aforementioned image processing model, and will not be repeated here. During the training process, sample images, corresponding first and second sample texts are determined. The sample image features, first sample text features, and second sample text features are also determined. Based on the sample image features and second sample text features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained based on the sample image features, first sample text features, and fused sample features. This allows the image processing model to learn image knowledge, text knowledge, and image-text fusion knowledge during training. This knowledge guides the learning of the image processing model, achieving accuracy, diversity, and robustness. This facilitates the subsequent processing of product images of target goods, ensuring the accuracy of the target image features extracted by the image processing model, and further guaranteeing the accuracy of other product images determined based on these target image features.

[0200] See Figure 6 , Figure 6 A flowchart of an image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0201] Step 602: Receive the image to be processed sent by the client;

[0202] Step 604: Input the image to be processed into the image processing model to obtain the target image features of the image to be processed;

[0203] Step 606: Determine the descriptive text of the image to be processed based on the features of the target image;

[0204] Step 608: Send the description text to the client;

[0205] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0206] The descriptive text of the image to be processed can be understood as the image processing result of the image to be processed.

[0207] Specifically, the image processing method provided in the embodiments of this specification can be applied to the field of image-to-text. In specific implementation, the server can receive the image to be processed sent by the client, input the image to be processed into the image processing model, obtain the target image features of the image to be processed, determine the descriptive text of the image to be processed based on the target image features, and send the descriptive text to the client for display through the client's display interface.

[0208] Understandably, the training process of this image processing model is similar to that of the aforementioned image processing model, and will not be repeated here. During the training process, sample images, first sample text, and second sample text corresponding to the sample images are determined. The sample image features of the sample images, the first sample text features of the first sample text, and the second sample text features of the second sample text are also determined. Based on the sample image features and the second sample text features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained based on the sample image features, the first sample text features, and the fused sample features. This allows the image processing model to learn image knowledge, text knowledge, and image-text fusion knowledge during training. This knowledge guides the learning of the image processing model, achieving accuracy, diversity, and robustness. This facilitates the subsequent processing of images by utilizing the accuracy of the target image features extracted by the image processing model, further ensuring the accuracy of the descriptive text determined based on the target image features.

[0209] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 7 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. Figure 7 As shown, the device includes:

[0210] The first determining module 702 is configured to determine the image to be processed;

[0211] Input module 704 is configured to input the image to be processed into an image processing model to obtain the target image features of the image to be processed;

[0212] The second determining module 706 is configured to determine the image processing result of the image to be processed based on the features of the target image;

[0213] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0214] In an optional embodiment, the second determining module 706 is further configured to:

[0215] Obtain features from multiple reference texts;

[0216] Based on the similarity between the target image features and the plurality of reference text features, the target text features corresponding to the target image features are determined from the plurality of reference text features;

[0217] The target text corresponding to the target text feature is determined, and the target text is determined as the image processing result of the image to be processed.

[0218] In an optional embodiment, the second determining module 706 is further configured to:

[0219] Acquire features from multiple reference images;

[0220] Based on the similarity between the target image features and the plurality of reference image features, the target reference image features corresponding to the target image features are determined from the plurality of reference image features;

[0221] The target image corresponding to the features of the target reference image is determined, and the target image is determined as the image processing result of the image to be processed.

[0222] In an optional embodiment, the first determining module 702 is further configured to:

[0223] Receive the image to be processed sent by the client;

[0224] The image processing results are sent to the client and displayed through the client's interface.

[0225] In an optional embodiment, the device further includes a training module configured to:

[0226] The machine learning model to be trained is determined, the sample images associated with the image processing task are determined, and the first sample text and the second sample text corresponding to the sample images are determined;

[0227] Based on the machine learning model to be trained, determine the sample image features of the sample image;

[0228] Determine the first sample text features of the first sample text, and determine the second sample text features of the second sample text;

[0229] The sample image features and the second sample text features are fused to obtain fused sample features;

[0230] The machine learning model to be trained is trained based on the sample image features, the first sample text features, and the fused sample features to obtain a trained image processing model.

[0231] One embodiment of this specification implements the following process during the training of an image processing model: determining sample images, corresponding first and second sample texts, and defining sample image features, first sample text features, and second sample text features. Based on these features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained using these features, enabling it to learn image knowledge, text knowledge, and image-text fusion knowledge. This knowledge guides the model's learning, improving its accuracy, diversity, and robustness. This facilitates the subsequent processing of images by ensuring the accuracy of the target image features extracted from the model, further guaranteeing the accuracy of the image processing results determined based on these target image features.

[0232] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0233] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing model training device. Figure 8 A schematic diagram of an image processing model training apparatus according to one embodiment of this specification is shown. Figure 8 As shown, the device includes:

[0234] The first determining module 802 is configured to determine the machine learning model to be trained, determine the sample image associated with the image processing task, and determine the first sample text and the second sample text corresponding to the sample image.

[0235] The second determining module 804 is configured to determine the sample image features of the sample image based on the machine learning model to be trained.

[0236] The third determining module 806 is configured to determine the first sample text features of the first sample text and the second sample text features of the second sample text.

[0237] The fusion module 808 is configured to perform fusion processing on the sample image features and the second sample text features to obtain fused sample features;

[0238] The training module 810 is configured to train the machine learning model to be trained based on the sample image features, the first sample text features, and the fused sample features, to obtain a trained image processing model.

[0239] In an optional embodiment, the fusion module 808 is further configured to:

[0240] The sample image features and the second sample text features are added together to obtain fused sample features; or

[0241] The sample image features and the second sample text features are weighted and summed to obtain fused sample features.

[0242] In an optional embodiment, the training module 810 is further configured to:

[0243] Based on the features of the sample image and the features of the first sample text, the machine learning model to be trained is trained to obtain the initial image processing model after training.

[0244] The initial image processing model is trained based on the sample image features and the fused sample features to obtain a trained image processing model.

[0245] In an optional embodiment, the training module 810 is further configured to:

[0246] Determine the category information corresponding to the features of the fused samples;

[0247] The sample image features are used as training samples, and the category information is used as training labels to train the initial image processing model, thereby obtaining the trained image processing model.

[0248] In one optional embodiment, there are multiple sample images, and both the sample image features and the fused sample features are multiple.

[0249] The training module 810 is further configured as follows:

[0250] Clustering is performed on multiple fused sample features to obtain clustering results, and the category information of each fused sample feature is determined based on the clustering results.

[0251] The step of using the sample image features as training samples and the category information as training labels to train the initial image processing model to obtain the trained image processing model includes:

[0252] Multiple sample image features are used as training samples, and the category information of each fused sample feature is used as training labels to train the initial image processing model, thereby obtaining the trained image processing model.

[0253] In an optional embodiment, the training module 810 is further configured to:

[0254] Calculate the loss function based on the sample image features and the first sample text features;

[0255] The machine learning model to be trained is trained according to the loss function to obtain the initial image processing model after training.

[0256] One embodiment of this specification implements the following process during the training of an image processing model: determining sample images, corresponding first and second sample texts, and defining sample image features, first sample text features, and second sample text features. Based on these features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained using these features, enabling it to learn image knowledge, text knowledge, and image-text fusion knowledge. This knowledge guides the model's learning, improving its accuracy, diversity, and robustness. This facilitates the subsequent processing of images by ensuring the accuracy of the target image features extracted from the model, further guaranteeing the accuracy of the image processing results determined based on these target image features.

[0257] The above is a schematic scheme of an image processing model training device according to this embodiment. It should be noted that the technical solution of this image processing model training device and the technical solution of the image processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the image processing model training device, please refer to the description of the technical solution of the image processing model training method described above.

[0258] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 9 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. Figure 9 As shown, the device includes:

[0259] The receiving module 902 is configured to receive a product image of the target product sent by the client;

[0260] Input module 904 is configured to input the product image into an image processing model to obtain the target image features of the product image;

[0261] The determining module 906 is configured to determine other product images associated with the product image based on the features of the target image;

[0262] The sending module 908 is configured to send the other product images to the client.

[0263] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0264] One embodiment of this specification implements the following process during the training of an image processing model: determining sample images, corresponding first and second sample texts, and defining sample image features, first sample text features, and second sample text features. Based on these features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained using these features, enabling it to learn image knowledge, text knowledge, and image-text fusion knowledge. This knowledge guides the model's learning, improving its accuracy, diversity, and robustness. This facilitates the subsequent processing of product images using the accuracy of the target image features extracted by the model, further ensuring the accuracy of other product images determined based on these target image features.

[0265] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0266] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 10 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. Figure 10 As shown, the device includes:

[0267] The receiving module 1002 is configured to receive the image to be processed sent by the client;

[0268] Input module 1004 is configured to input the image to be processed into an image processing model to obtain the target image features of the image to be processed;

[0269] The determining module 1006 is configured to determine the descriptive text of the image to be processed based on the features of the target image;

[0270] Sending module 1008 is configured to send the description text to the client;

[0271] The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0272] One embodiment of this specification implements the following process during the training of an image processing model: determining sample images, corresponding first and second sample texts, and determining sample image features, first sample text features, and second sample text features. Based on these features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained using these features, enabling it to learn image knowledge, text knowledge, and image-text fusion knowledge during training. This knowledge guides the model's learning, improving its accuracy, diversity, and robustness. This facilitates the subsequent processing of images by ensuring the accuracy of the target image features extracted from the image, further guaranteeing the accuracy of the descriptive text determined based on these target image features.

[0273] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0274] Corresponding to the above method embodiments, this specification also provides embodiments of a model training platform. Figure 11 A schematic diagram of the structure of a model training platform provided in one embodiment of this specification is shown. Figure 11 As shown, the model training platform includes:

[0275] The request interface unit 1102, the model training unit 1104, and the response unit 1106;

[0276] The request interface unit 1102 is used to receive a model training request, wherein the model training request includes model information of the machine learning model to be trained.

[0277] The model training unit 1104 is used to determine the machine learning model to be trained based on the model information, and to train the machine learning model to be trained to obtain a trained image processing model. The image processing model is trained based on the sample image features of the sample image, the first sample text features of the first sample text, and the fused sample features. The fused sample features are determined based on the sample image features and the second sample text features. The second sample text features are determined based on the second sample text. The first sample text and the second sample text are determined based on the sample image.

[0278] The response unit 1106 is used to output the image processing model.

[0279] Optionally, the model training platform also includes a data receiving unit.

[0280] The data receiving unit is used to receive training data input by the user and send the training data to the model training unit;

[0281] The model training unit is further configured to train the machine learning model to be trained based on the training data to obtain a trained image processing model.

[0282] Optionally, the model training platform also includes a model library, wherein the model library stores multiple machine learning models;

[0283] The model training unit is further configured to determine the machine learning model to be trained from the model library based on the model information.

[0284] One embodiment of this specification implements the following process during the training of an image processing model: determining sample images, corresponding first and second sample texts, and defining sample image features, first sample text features, and second sample text features. Based on these features, fused sample features are determined to obtain stronger multimodal features. The image processing model is then trained using these features, enabling it to learn image knowledge, text knowledge, and image-text fusion knowledge. This knowledge guides the model's learning, improving its accuracy, diversity, and robustness. This facilitates the subsequent processing of images by ensuring the accuracy of the target image features extracted from the model, further guaranteeing the accuracy of the image processing results determined based on these target image features.

[0285] The above is an illustrative scheme of a model training platform according to this embodiment. It should be noted that the technical solution of this model training platform and the technical solution of the image processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the model training platform, please refer to the description of the technical solution of the image processing model training method described above.

[0286] Figure 12A structural block diagram of a computing device 1200 according to an embodiment of this specification is shown. The components of the computing device 1200 include, but are not limited to, a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and a database 1250 is used to store data.

[0287] The computing device 1200 also includes an access device 1240, which enables the computing device 1200 to communicate via one or more networks 1260. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1240 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0288] In one embodiment of this application, the aforementioned components of the computing device 1200 and Figure 12 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 12 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0289] The computing device 1200 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1200 can also be a mobile or stationary server.

[0290] The memory 1210 is used to store computer programs / instructions, and the processor 1220 is used to execute the computer programs / instructions stored in the memory 1210. When the computer programs / instructions are executed by the processor, they implement the steps of the above method.

[0291] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computing device can be referred to the description of the technical solution of the above method.

[0292] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0293] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the method described above belong to the same concept, and all details not described in detail in the technical solution of the storage medium can be referred to the description of the technical solution of the method described above.

[0294] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0295] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of the above method.

[0296] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0297] The computer program / instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0298] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0299] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0300] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An image processing model training method, comprising: determining a machine learning model to be trained, determining a sample image associated with an image processing task, and first sample text and second sample text corresponding to the sample image; determining a sample image feature of the sample image according to the machine learning model to be trained; determining a first sample text feature of the first sample text and a second sample text feature of the second sample text; performing fusion processing on the sample image feature and the second sample text feature to obtain a fusion sample feature; training an image processing model according to the sample image feature, the first sample text feature, and the fusion sample feature.

2. The image processing model training method of claim 1, wherein the fusion processing on the sample image feature and the second sample text feature to obtain a fusion sample feature comprises: performing feature addition on the sample image feature and the second sample text feature to obtain a fusion sample feature; or performing weighted summation on the sample image feature and the second sample text feature to obtain a fusion sample feature.

3. The image processing model training method of claim 1 or 2, wherein the training of the machine learning model to be trained according to the sample image feature, the first sample text feature, and the fusion sample feature to obtain a trained image processing model comprises: training the machine learning model to be trained according to the sample image feature and the first sample text feature to obtain a trained initial image processing model; training the initial image processing model according to the sample image feature and the fusion sample feature to obtain a trained image processing model.

4. The image processing model training method of claim 3, wherein the training of the initial image processing model according to the sample image feature and the fusion sample feature to obtain a trained image processing model comprises: determining category information corresponding to the fusion sample feature; training the initial image processing model using the sample image feature as a training sample and the category information as a training label to obtain a trained image processing model.

5. The image processing model training method of claim 4, wherein the sample image is a plurality of sample images, and the sample image feature and the fusion sample feature are a plurality of sample image features and a plurality of fusion sample features, respectively. The determination of the category information of each fusion sample feature from the plurality of fusion sample features comprises: performing clustering processing on the plurality of fusion sample features to obtain a clustering result, and determining the category information of each fusion sample feature from the plurality of fusion sample features according to the clustering result. The training of the initial image processing model using the sample image feature as a training sample and the category information as a training label to obtain a trained image processing model comprises: training the initial image processing model using the plurality of sample image features as training samples and the category information of each fusion sample feature as a training label to obtain a trained image processing model.

6. The image processing model training method of claim 3, wherein the training of the machine learning model to be trained according to the sample image feature and the first sample text feature to obtain an initial image processing model after training comprises: calculating a loss function according to the sample image feature and the first sample text feature; and training the machine learning model to be trained according to the loss function to obtain the initial image processing model after training.

7. An image processing method, comprising: determining an image to be processed; inputting the image to be processed into an image processing model to obtain a target image feature of the image to be processed; and determining an image processing result of the image to be processed according to the target image feature; wherein the image processing model is trained according to a sample image feature of a sample image, a first sample text feature of a first sample text, and a fusion sample feature, the fusion sample feature is determined according to the sample image feature and a second sample text feature, the second sample text feature is determined according to a second sample text, and the first sample text and the second sample text are determined according to the sample image.

8. The image processing method of claim 7, wherein the determining of the image processing result of the image to be processed according to the target image feature comprises: obtaining a plurality of reference text features; determining a target text feature corresponding to the target image feature from the plurality of reference text features according to a similarity between the target image feature and the plurality of reference text features; and determining a target text corresponding to the target text feature as the image processing result of the image to be processed.

9. The image processing method of claim 7, wherein the determining of the image processing result of the image to be processed according to the target image feature comprises: obtaining a plurality of reference image features; determining a target reference image feature corresponding to the target image feature from the plurality of reference image features according to a similarity between the target image feature and the plurality of reference image features; and determining a target image corresponding to the target reference image feature as the image processing result of the image to be processed.

10. The image processing method of any one of claims 7-9, wherein the determining of the image to be processed comprises: receiving the image to be processed sent by a client; and after the determining of the image processing result of the image to be processed, the method further comprises: sending the image processing result to the client for display on a display interface of the client.

11. The image processing method of any one of claims 7-9, wherein the training of the image processing model comprises: determining a machine learning model to be trained, determining a sample image associated with an image processing task, and determining a first sample text and a second sample text corresponding to the sample image; determining a sample image feature of the sample image according to the machine learning model to be trained; determining a first sample text feature of the first sample text and a second sample text feature of the second sample text. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ fusing processing is performed on the sample image feature and the second sample text feature to obtain a fused sample feature; training is performed on the machine learning model to be trained according to the sample image feature, the first sample text feature and the fused sample feature, to obtain a trained image processing model.

12. An image processing method, comprising: receiving a product image of a target product sent by a client; inputting the product image into an image processing model to obtain a target image feature of the product image; determining other product images associated with the product image according to the target image feature; sending the other product images to the client; wherein the image processing model is trained according to a sample image feature of a sample image, a first sample text feature of a first sample text and a fused sample feature, the fused sample feature is determined according to the sample image feature and a second sample text feature, the second sample text feature is determined according to a second sample text, and the first sample text and the second sample text are determined according to the sample image.

13. An image processing method, comprising: receiving an image to be processed sent by a client; inputting the image to be processed into an image processing model to obtain a target image feature of the image to be processed; determining a description text of the image to be processed according to the target image feature; sending the description text to the client; wherein the image processing model is trained according to a sample image feature of a sample image, a first sample text feature of a first sample text and a fused sample feature, the fused sample feature is determined according to the sample image feature and a second sample text feature, the second sample text feature is determined according to a second sample text, and the first sample text and the second sample text are determined according to the sample image.

14. A model training platform, comprising a request interface unit, a model training unit and a response unit; the request interface unit is configured to receive a model training request, wherein the model training request comprises model information of a machine learning model to be trained; the model training unit is configured to determine the machine learning model to be trained according to the model information, and to train the machine learning model to be trained to obtain a trained image processing model, wherein the image processing model is trained according to a sample image feature of a sample image, a first sample text feature of a first sample text and a fused sample feature, the fused sample feature is determined according to the sample image feature and a second sample text feature, the second sample text feature is determined according to a second sample text, and the first sample text and the second sample text are determined according to the sample image; the response unit is configured to output the image processing model.

15. The model training platform of claim 14, further comprising a data receiving unit, the data receiving unit is configured to receive training data input by a user and send the training data to the model training unit. The model training unit is further configured to perform model training on the machine learning model to be trained according to the training data, and obtain a trained image processing model.

16. The model training platform of claim 14, further comprising a model library, wherein, The model library stores a plurality of machine learning models. The model training unit is further configured to determine the machine learning model to be trained from the model library according to the model information.

17. A computing device comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method of any one of claims 1 to 13.

18. A computer-readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the method of any one of claims 1 to 13.

19. A computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the method of any one of claims 1 to 13.