Data processing method and device based on machine learning model, medium and equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2026-06-22
- Publication Date
- 2026-08-07
AI Technical Summary
然而,目前基于机器学习模型进行视觉数据处理时存在准确性低的问题
[0009]通过上述技术方案,在接收到输入信息的情况下,将该输入信息发送给第一机器学习模型,以通过第一机器学习模型自动生成输入信息对应的视觉数据,由于第一机器学习模型通过训练第二机器学习模型获取,且第二机器学习包括第一网络与第二网络,故通过第一机器学习模型能够准确生成输入信息对应的视觉数据,在此过程中,由于第一网络通过图像生成相关的第一损失训练,其对应的是生成分支,以及第二网络通过文本生成相关的第二损失训练,其对应的是推理支路,故通过将两个分支解耦,能够避免二者相互干扰,进而提高视觉数据处理的准确性。
Smart Images

Figure CN122530337A_ABST
Abstract
Description
Technical Field
[0001] This content relates to the field of computer vision technology, specifically to a data processing method, apparatus, medium, and device based on a machine learning model. Background Technology
[0002] With the rapid development of artificial intelligence technology, the application of machine learning models for automatic processing of visual data is becoming increasingly widespread. Examples include applications that use machine learning models to generate text-to-image representations, or applications that perform image editing, image localization, and image segmentation. However, current methods for processing visual data using machine learning models suffer from low accuracy. Summary of the Invention
[0003] This content section is provided to briefly introduce the concepts, which will be described in detail in the examples section later. This content section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] Firstly, a data processing method based on a machine learning model is provided, the method comprising: Receive input information, wherein the input information includes at least a first natural language description; The input information is sent to a first machine learning model, which is a multimodal large language model; The system receives visual data, which is generated by a first machine learning model based on the input information. The first machine learning model is obtained by training a second machine learning model, which includes a first network and a second network. The first network is trained using a first loss, and the second network is trained using a second loss. The first loss is related to image generation, and the second loss is related to text generation.
[0005] Secondly, a data processing apparatus based on a machine learning model is provided, the apparatus comprising: The receiving module is configured to receive input information, which includes at least first natural language description information; The sending module is configured to send the input information to a first machine learning model, wherein the first machine learning model is a multimodal large language model; The receiving module is further configured to receive visual data generated by the first machine learning model based on the input information. The first machine learning model is obtained by training a second machine learning model, which includes a first network and a second network. The first network is trained using a first loss, and the second network is trained using a second loss. The first loss is related to image generation, and the second loss is related to text generation.
[0006] Thirdly, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0007] Fourthly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0008] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0009] The above technical solution involves sending the input information to a first machine learning model upon receipt. This model automatically generates visual data corresponding to the input information. Since the first machine learning model is trained by a second machine learning model, which includes both a first and a second network, the first machine learning model can accurately generate the visual data corresponding to the input information. In this process, the first network is trained using a first loss related to image generation (corresponding to the generation branch), and the second network is trained using a second loss related to text generation (corresponding to the inference branch). By decoupling the two branches, mutual interference can be avoided, thereby improving the accuracy of visual data processing.
[0010] Other features and advantages of this article will be described in detail in the following examples section. Attached Figure Description
[0011] The above and other features, advantages, and aspects of this document will become more apparent when viewed in conjunction with the accompanying drawings and the following examples. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a schematic diagram illustrating an implementation environment according to an exemplary embodiment.
[0012] Figure 2This is a flowchart illustrating a data processing method based on a machine learning model according to an exemplary embodiment.
[0013] Figure 3 This is an example diagram illustrating the architecture of a second machine learning model in a data processing method based on a machine learning model, according to an exemplary embodiment.
[0014] Figure 4 This is an example diagram illustrating the automatic generation of training samples in a data processing method based on a machine learning model, according to an exemplary embodiment.
[0015] Figure 5 This is an example diagram illustrating task type classification in a data processing method based on a machine learning model, according to an exemplary embodiment.
[0016] Figure 6 This is an example diagram illustrating training samples for a text-to-image generation task in a data processing method based on a machine learning model, according to an exemplary embodiment.
[0017] Figure 7 This is an example diagram of training samples corresponding to a visual localization task in a data processing method based on a machine learning model, according to an exemplary embodiment.
[0018] Figure 8 This is an example diagram of training samples corresponding to a visual chain task in a data processing method based on a machine learning model, according to an exemplary embodiment.
[0019] Figure 9 This is a block diagram illustrating a data processing apparatus based on a machine learning model according to an exemplary embodiment.
[0020] Figure 10 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation
[0021] The present invention will now be described in more detail with reference to the accompanying drawings. While certain scenarios are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the scenarios set forth herein; rather, these scenarios are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and the scenarios depicted are for illustrative purposes only and are not intended to limit the scope of this document.
[0022] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0023] The term "comprising" and its variations as used herein can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.
[0024] It should be noted that the concepts of "first" and "second" mentioned here are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0025] It should be noted that the terms "one" and "more" used here are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0026] The names of messages or information exchanged between the multiple devices in the implementation are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0027] It is understandable that before using the technical solutions provided here, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in accordance with relevant laws and regulations, and their authorization should be obtained through appropriate means.
[0028] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations described herein.
[0029] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0030] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described herein. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the technical solution.
[0031] At the same time, it is understood that the data involved in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0032] Multimodal Large Language Models (MLLMs) have made significant progress in tasks such as image understanding, visual question answering, and optical character recognition. The trend is towards unifying visual understanding and visual generation within a single model architecture, demonstrating that generated data can, under certain conditions, improve the performance of understanding tasks. However, MLLMs still have many shortcomings when facing highly complex visual understanding tasks where text is difficult to express precisely.
[0033] For example, in related technologies, multimodal large language models are mainly trained to understand visual concepts through text descriptions. However, most visual tasks inherently require pixel-level or region-level precise understanding, and simple text supervision signals are usually insufficient to provide sufficiently fine-grained visual perception training for multimodal large language models. This results in multimodal large language models performing poorly when answering questions related to image spatial relationships, precise target locations, and subtle visual differences.
[0034] Furthermore, multimodal large language models mainly employ a hybrid architecture combining autoregression (AR) and diffusion models to process visual data. However, this architecture suffers from high inference costs and high architectural complexity. Moreover, these works focus on the image generation quality of the model, and the generation and understanding tasks are trained in the same model. This coupled training often leads to interference between the two capabilities, making it difficult to steadily improve the understanding ability.
[0035] Furthermore, training multimodal large language models typically employs a simple approach of mixing general generated data such as text-to-image conversion, lacking a systematic analysis of the impact of different types of generation tasks on comprehension capabilities. Moreover, during the reasoning process, multimodal large language models suffer from the gradual loss of visual information within the deeper layers of the model, leading to poor performance on specific tasks, which can be complex reasoning tasks requiring precise visual evidence.
[0036] To address the aforementioned issues, this paper proposes a data processing method, apparatus, medium, and device based on a machine learning model. This method, through training a first machine learning model, can accurately process visual data, thereby comprehensively enhancing the complex visual understanding capabilities of multimodal large language models without increasing costs.
[0037] The data processing methods based on machine learning models provided in this content can be applied in scenarios such as... Figure 1 As shown, Figure 1 As shown, the application scenario can include terminal 101 and server 102. Terminal 101 and server 102 are connected in communication.
[0038] Terminal 101 can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and laptop computer. Terminal 101 has communication capabilities and can access wired or wireless networks. Terminal 101 can refer to one of multiple terminals; those skilled in the art will understand that the number of such terminals can be more or less.
[0039] Server 102 can be a standalone physical server, a server cluster consisting of multiple physical servers, or a distributed file system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0040] For example, server 102 and terminal 101 can be connected directly or indirectly through wired or wireless communication, without limitation.
[0041] The terminal 101 can receive input information, which can be entered by the user, and send the input information to the server 102. When the server 102 receives the input information, it can generate visual data related to the input information, and then send the visual data to the terminal 101 and instruct the terminal 101 to display the visual data.
[0042] Figure 2 This is a flowchart illustrating a data processing method based on a machine learning model according to an exemplary embodiment. The apparatus can be implemented in software and / or hardware. Figure 2 As shown, the method may include the following steps.
[0043] In step S210, input information is received.
[0044] In some implementations, the input information may include at least a first natural language description, based on which a first machine learning model can be instructed to perform visual data processing operations. Here, the input information may be requirement information, based on which the type of task to be performed can be determined; different user requirements result in different corresponding task types.
[0045] As an example, input information can be used to instruct the generation of a specified image. In this case, the input information includes natural language description information (text information), and the corresponding task type is a text-to-image generation task (generation-type task). This text-to-image generation task can be a text-to-image task. Here, the image generation task can use long text and complex prompts with details to generate images. For example, the input information related to the image generation task could be "Generate a fantasy epic movie poster full of fighting scenes".
[0046] As another example, input information can be used to instruct the editing of a specified image. In this case, the task type corresponding to the input information is an image editing task, which can be used to locally add or modify the subject, task, clothing, or scene details. Here, the input information related to the image editing task can include natural language description information and image information. For example, the input information related to the image editing task may include the Chinese text information "Generate a standing target object and put the clothes in the image onto the generated target object," and the image information is the image mentioned in the text information.
[0047] As another example, the input information is used to instruct the segmentation of a specified image. In this case, the task type corresponding to the input information is an image segmentation task. This image segmentation task is used to peel a specified object from the background pixel by pixel according to instructions to generate a black and white mask image, accurately delineating the boundaries of different overlapping objects and the background. Here, the input information related to the image segmentation task can include natural language description information and image information. For example, the text information included in the input information could be "find which items are commonly used to make ice cubes and chilled drinks," and the image information included in the input information could be an image including a refrigerator.
[0048] As another example, the input information is used to instruct the localization of a specified image. In this case, the task type corresponding to the input information is an image localization task, which can also be called a target localization task. It involves circling a specified target object in an image with a rectangular border based on a natural language question. The input information related to the image localization task can include natural language descriptions and image information. For example, the text information included in the input information could be "Find the 'Add Product' button in the image," and the image information included could include the "Add Product" button and the product itself.
[0049] As another example, the input information is used to instruct step-by-step additions, deletions, replacements, background removal, style changes, etc., of an image, i.e., modifying the image based on chained reasoning. In this case, the task type corresponding to the input information is a visual chained reasoning task. The input information related to visual chained reasoning tasks can include natural language descriptions and image information. For example, the text input information could be "Add a closed hardcover book next to the glasses; replace the hardcover book with a slim smartphone; remove all backgrounds, leaving only the glasses in the image, and change the image to a screen printing style." The image input information could include an image of the glasses. Here, the input information can be entered by the user through multiple dialogues based on their needs.
[0050] It is evident that different input information leads to different types of tasks performed by the subsequent machine learning model, and consequently, different types of output visual data. Furthermore, the input information can be multimodal data, which can at least include natural language descriptions, enabling the processing of visual data based on these descriptions.
[0051] In step S220, the input information is sent to the first machine learning model.
[0052] As an alternative approach, upon receiving input information, the client can send the input information to the first machine learning model. For example, the client can input the user's input information as a prompt word into the first machine learning model, or the client can process the user's input information first to obtain the target prompt information and then input the target prompt information into the first machine learning model.
[0053] In other words, the information received by the first machine learning model can be the raw input information received by the client, or it can be information obtained by secondary processing or manipulation of the raw input information received by the client. The specific information that the input to the first machine learning model refers to is not explicitly limited here; it can be selected based on the actual situation.
[0054] In some implementations, the first machine learning model can be a multimodal large language model, which can be an artificial intelligence model extended from a large language model. It can simultaneously receive and process multiple types of input data, such as text and images. In other words, the multimodal large language model is based on cross-modal semantic association capabilities acquired through large-scale pre-training. It can not only complete various tasks such as visual understanding, graphic-text interaction, and content generation, but also achieve the fusion, understanding, and conversion of information from different modalities.
[0055] In summary, the first machine learning model is based on a large language model architecture, integrates an image encoding module, supports multimodal inputs such as text and images, and has the capabilities of visual perception, semantic reasoning, and multimodal content generation. Based on the first machine learning model, complex visual tasks such as target localization (image localization), image segmentation, and image editing can be achieved.
[0056] Understandably, after receiving input information, the first machine learning model can generate visual data based on that input information and output that visual data. Here, the first machine learning model can be deployed on a server or a client; there are no specific restrictions on which device it is deployed on, and the choice can be made based on the actual situation.
[0057] In step S230, visual data is received.
[0058] As described above, the first machine learning model can generate visual data based on the input information received. This first machine learning model can be obtained by training a second machine learning model; that is, after obtaining a training dataset, inputting the training dataset into the second machine learning model allows for model training, resulting in the first machine learning model.
[0059] In one scenario, the second machine learning model can be a multimodal large language model, which may include a first network and a second network. The first network can be trained using a first loss, and the second network can be trained using a second loss. Here, the first loss may be related to image generation, and the second loss may be related to text generation.
[0060] The first network can be a generative branch of the second machine learning model. It can be an auxiliary processing branch dedicated to the model training phase, which can optimize the model's visual features based on the first loss to improve the accuracy of visual representation. It can be removed from the model after training. The first loss can be a visual class supervised loss, which can be a cosine similarity loss.
[0061] For example, the first network can be connected to a shared layer (the underlying shared feature processing layer), which can operate during model training. During model training, the first network can perform visual representation prediction based on the input visual features and update parameters according to the first loss (image generation loss). Therefore, the first network is mainly used to enhance the model's visual perception and detail extraction capabilities.
[0062] The second network can be the understanding branch of the second machine learning model, or it can be the main processing branch of the model. This second network can continue to run during both the training and inference phases. Based on this second network, multimodal features can be fused to complete semantic processing and sequence generation, so as to realize core functions such as image and text understanding and content output. Here, the second loss can be the text cross-entropy loss.
[0063] For example, similar to the first network, the second network can also be connected to the shared layer. However, unlike the first network, the second network can function normally during both the model training and inference phases. During model training, the second network can perform high-level semantic parsing on the fused multimodal features and iteratively optimize parameters using a second loss (text generation loss) to achieve business functions such as visual understanding and multimodal content generation. Therefore, the second network is the core branch that provides services to the model.
[0064] To better illustrate the structure of the first machine learning model, this content provides the following... Figure 3 The example diagram shown is as follows. Figure 3 The generated branch (a) can correspond to the first network, and the understood branch (b) can correspond to the second network. Before training the second machine learning model, a training dataset can be obtained. This training dataset can include multiple training samples, each of which can include image information and second natural language description information. The second natural language description information can be training text information, and the image information can be training image information.
[0065] As described above, the first machine learning model can perform multiple task types. To ensure this capability, the training dataset must include data related to at least one task type.
[0066] After obtaining the training dataset, the task type can be obtained, which is associated with the training samples. Then, the task complexity can be determined based on the task type. The training dataset is then input into the second machine learning model based on the task complexity.
[0067] In other words, after obtaining a training dataset containing multiple training samples, this paper can determine the task type corresponding to each training sample in the training dataset, then determine the task complexity corresponding to that task type, and based on this, input the training dataset into the second machine learning model in sequence according to the task complexity. Here, the task type can include at least one of the following: text-to-image generation task (Text2Image), image editing task (Image Edit), image segmentation task (Segmentation), visual localization task (Grounding), and visual chained reasoning task (Visual-CoT).
[0068] For example, the training dataset can be divided into five categories according to the task types mentioned above. Then, the task complexity of these five categories can be determined. Based on this, the corresponding training samples are input into the second machine learning model according to the order of task complexity for training. The task complexity of the five categories, from low to high, is as follows: text-to-image generation task, visual localization task, image segmentation task, image editing task, and visual chain reasoning task.
[0069] As can be seen, the text-to-image generation task has the lowest task complexity, while the visual chain task has the highest. Therefore, we can first input the training samples related to the image generation task into the second machine learning model to train the model using these training samples. Then, we can sequentially input the training samples corresponding to subsequent task types into the second machine learning model in an increasing order of complexity. In other words, this paper can introduce generation tasks of different complexities in a manner that progresses from easy to difficult.
[0070] Optionally, during the training of the second machine learning model, progressive loss weights can be used, such as automatically adjusting the weights based on the ratio of the first loss (generation loss) to the second loss (understanding loss) during training.
[0071] For example, in the first training phase, training samples related to a first complexity can be used to train the model. For instance, the task corresponding to the first complexity could be a text-to-image generation task. Here, the weight of the first loss in the first training phase could be 0.6, and the weight of the second loss could be 0.4. In the second training phase, training samples related to a second complexity can be used to train the model. For instance, the task corresponding to the second complexity could be an image segmentation task or an object localization task. The second complexity can be greater than the first complexity. Here, the weight of the first loss in the second training phase could be 0.4, and the weight of the second loss could be 0.6. In the third training phase, training samples related to a third complexity can be used to train the model. For instance, the task corresponding to the third complexity could be an image editing task or a visual chain reasoning task. The third complexity can be greater than the second complexity. Here, the weight of the first loss in the third training phase could be 0.2, and the weight of the second loss could be 0.8.
[0072] In summary, after obtaining the training dataset, the training samples can be input into the second machine learning model in a specified order to train the model. By inputting the training dataset into the second machine learning model in this way, it can be ensured that the supervision signal for the generation task can play an effective role throughout the entire training process, that is, it can be ensured that it is neither overwhelmed nor excessively interfered with the generation ability.
[0073] As an alternative approach, the training dataset can be automatically generated as follows: acquire a first image; generate a first label, wherein the first label can be obtained by recognizing the first image; detect a first bounding box based on the first label, wherein the first bounding box can belong to the first image; normalize the first bounding box to obtain normalized coordinate information; generate training question-answering samples based on the first image, the first bounding box, and the normalized coordinate information; and generate a training dataset from the training question-answering samples and the first image.
[0074] Here, the process of generating the training dataset can be applied to generating training samples for visual localization tasks. That is, upon receiving an instruction to generate training data related to a visual localization task, training question-and-answer samples can be automatically generated based on the above method, where the input data can be images.
[0075] To better illustrate the process of obtaining training samples, the following is given: Figure 4 The example diagram shown is based on Figure 4As can be seen, after acquiring the first image (original image), the first image can be identified based on a label recognition model to obtain multiple first labels included in the first image. For example, an open-domain recognition model can be used to automatically generate all noteworthy target labels in the image. For instance, by recognizing the first image, the generated first labels may include dress, target object, lantern, sofa, poster, and bow tie.
[0076] Based on this, an open vocabulary detection operation can be performed, which involves inputting the first image and automatically generated labels into an open vocabulary object detection model to obtain the bounding boxes of the targets that match the labels. The object detection model can be an open vocabulary object detection model. For example... Figure 4 As shown, the object detection model can obtain the bounding boxes corresponding to the target object, the bow tie, and the lantern. The combination of these bounding boxes can be a set of target candidate sets. Each bounding box can include a label (category), coordinates, and area.
[0077] Optionally, after obtaining the bounding box set through detection, the bounding boxes can be normalized to obtain normalized coordinate information. Before this, a sampling operation can be performed on the bounding box set. During the sampling process, the corresponding sampling logic can be determined based on the attributes of the bounding boxes (the size of the bounding boxes). The attributes of the bounding boxes can include their size; different bounding box sizes require different sampling logic. For example, a first sampling logic can be used for simple samples, while a second sampling logic can be used for difficult samples.
[0078] The first sampling logic can be based on area ratio, ensuring that larger targets have a higher probability of being sampled. The second sampling logic can use a uniform sampling method to ensure that the target bounding box has an equal probability of being sampled, thus including data on small targets or occluded objects in the training samples. Additionally, locally magnified and cropped images can be generated.
[0079] For example, the bounding box labeled "target object" has the largest area and the highest probability of being sampled. The normalized bounding box coordinates of "target object" can be output later. The bounding box labeled "bow tie" can be sampled uniformly, meaning that small targets like bow ties and large targets like target objects have the same probability of being selected. The normalized bounding box coordinates of the bow tie and its corresponding zoomed crop can be output later.
[0080] Based on the aforementioned bounding box attributes, the sampling difficulty can be determined. For labels of large targets, probability sampling can be performed according to area ratio, while for small targets, uniform sampling can be used to ensure that their probability of selection is the same as that of large targets. By controlling the target bounding box sampling strategy, the difficulty of the synthesized data can be adjusted. In easy mode, the sampling probability can be proportional to the bounding box area to prioritize large targets; in hard mode, uniform sampling can be performed to ensure that the subsequently generated training samples contain more small targets and complex samples.
[0081] After obtaining the normalized coordinate information, training question-answering samples can be generated based on the first image, the first bounding box, and the normalized coordinate information. For example, a large language model can be used, inputting the first image and a magnified view of the target area, to generate localization commands closely related to the defined target, ensuring the uniqueness and scene relevance of the commands. Here, the first image may be labeled with a sampled bounding box.
[0082] After obtaining the normalized coordinate information, this information, along with the first image with labeled bounding boxes and a locally magnified cropped image, can be input into the base model to generate training question-and-answer samples. Here, the base model can be a multimodal model used for automatically writing training instructions. For example, for the first image, the automatically generated training document sample could be "Find the red lantern placed above the right shoulder of the target object," with the corresponding answer being [0.78, 0.18, 0.92, 0.34], where these coordinates can be the target bounding box coordinates.
[0083] As described above, the task types in this paper can include at least one of the following: text-to-image generation, image editing, image segmentation, visual localization, and visual chain reasoning. The above method can automatically construct multiple training samples that cover the five task types mentioned above, and each task type can be further subdivided, such as... Figure 5 The 15 seed task types are shown.
[0084] The text-to-image generation task can include three sub-task types: sketch-based interactive iterative editing, general scene text-to-image generation, and world-sense-based text-to-image generation. The image segmentation task can include three sub-task types: combined segmentation control, reasoning followed by mask generation, and multi-object pixel-level resolution. The image editing task can include three sub-task types: domain-specific scene editing, subject clothing transfer editing, and style and atmosphere editing. The target localization task (visual localization task) can include three sub-task types: logical reasoning-based target localization, graphical user interface element localization, and structural transformation localization. The visual chain reasoning task can include three sub-task types: cropping and mask-assisted reasoning, image-text interleaved distribution reasoning, and multi-object spatial localization.
[0085] During the training of the second machine learning model, the model can be trained according to the complexity of the aforementioned 15 task types, in a specified order and with corresponding weights. By introducing generated task data corresponding to these various types, the overall gain of different types of generated tasks on the model's understanding ability can be maximized.
[0086] As an example, for the text-to-image (T2I) generation task, this paper proposes a model that can predict the corresponding image embedding sequence based on a given text description. This task can serve as a general foundational task for generation. Training samples of this type enable the model to learn the alignment relationship between text and visual space, allowing it to retain more visual information in later stages and effectively mitigating the loss of visual information in deep reasoning. For instance, training samples for the text-to-image generation task can be as follows: Figure 6 As shown.
[0087] As another example, for the image editing task (Edit), this paper provides an input image and editing instructions, such as editing quality settings like "change the background to blue" or "remove the target object in the image," and uses a machine learning model to predict the edited target image. Training the model with training samples related to this type of task enables it to understand the operability of image content, enhancing the model's deep understanding of image components, spatial layout, and semantic relationships.
[0088] As another example, for image segmentation tasks, this paper transforms semantic segmentation into a generative paradigm. Given an input image and segmentation instructions, the model is instructed to predict the embedding representation of the segmentation result in an autoregressive manner to obtain the segmented image. Training the model with training samples relevant to this type of task enables it to learn precise pixel-level visual perception capabilities, enhancing the model's fine-grained understanding of target boundaries and spatial extent in images.
[0089] As another example, for visual grounding tasks, this paper proposes a model to predict the bounding box coordinates of target regions, given an image and a natural language description. Visual grounding tasks can be divided into two steps: automatic annotation and data synthesis. First, an open-vocabulary detection model can be used to automatically generate bounding box annotations for all valuable targets in the image. Then, target bounding boxes are sampled, and a large language model is used to generate location instructions closely related to the target based on the defined region and a magnified local image, ensuring that each instruction corresponds to only one target. Training the model with training samples relevant to this type of task enables it to learn the ability to associate natural language descriptions with precise spatial locations. For example, training samples for visual grounding tasks can be as follows: Figure 7 As shown.
[0090] As another example, for visual chain reasoning tasks, this paper provides a visual question requiring complex reasoning. During the answering process, the model can generate intermediate visual representations, such as intermediate images annotated with key regions, simulating the "thinking with image" reasoning process. Training the model with training samples relevant to this type of task enables it to learn to actively utilize and retain visual information during reasoning, rather than relying solely on text chains. For example, training samples for visual chain reasoning tasks could be like... Figure 8 As shown.
[0091] It should be noted that the task types in this paper are not limited to text-to-image generation, image editing, image segmentation, visual localization, and visual chain reasoning. They can also include depth estimation, normal prediction, optical flow estimation, image colorization, and super-resolution tasks. By introducing generation tasks that require precise pixel / region-level visual perception, fine-grained visual supervision signals can be provided to the model.
[0092] Based on the above method, high-quality training data can be automatically constructed from any business image data. In particular, it can trigger the automatic construction of high-quality localization training data from any business data without manual annotation, which can reduce the cost of generating training datasets to a certain extent.
[0093] As an alternative approach, after obtaining the training dataset, the image information from the training samples can be input into a visual encoder (Vision Transformer, VIT) to obtain a visual feature sequence. Based on this, a second machine learning model is trained using the visual feature sequence and / or a text word sequence to obtain a first machine learning model. The text word sequence can be obtained based on second natural language description information.
[0094] As described above, the training dataset includes training samples that can include image information and second natural language description information (text information). During the training of the second machine learning model, training information for different modalities can be input into the corresponding encoders to obtain feature sequences of the associated modalities. For example, image information can be input into a visual encoder to extract visual feature sequences, and second natural language description information can be input into a text tokenizer to obtain text word sequences (text token sequences).
[0095] During the training of the second machine learning model based on visual feature sequences and / or text word sequences, the visual feature sequences and text word sequences can be concatenated first to obtain a concatenated sequence. This concatenated sequence is then sent to the shared network (shared layer) of the second machine learning model to obtain intermediate hidden features. Here, the shared network can be... Figure 3 The numbers 0 to I shown are shown. split The first and second networks are then connected in a hierarchical manner. The intermediate hidden features are then input into the first and second networks respectively to obtain predicted image features and predicted text features. Here, the predicted image features can be obtained using an autoregressive approach. Finally, the second machine learning model is trained based on the predicted image features and predicted text features.
[0096] For example, the visual feature sequence and the text word sequence can be concatenated and input into a multimodal large language model. For the generation task, the multimodal large language model can predict the predicted image features (embedding sequence) of the target image one by one through an autoregressive method. Based on the predicted image features and standard image features, a first loss can be calculated, which can be used as a supervision signal on the generation side. The standard image features can be the embeddings extracted from the target image by the visual encoder (VIT).
[0097] Understandably, during the training of the second machine learning model based on image features, standard image features can be acquired, and a first loss can be determined based on these standard image features and the predicted image features. Based on this first loss, the first network and the shared network are then trained. Therefore, the supervision signal for the generation task (first network) in this paper can be applied to the hidden representation space of the large language model, thus guiding the model to learn more accurate visual representations. Furthermore, the visual representation space can simultaneously serve the understanding task (second network), thereby achieving a positive transfer of understanding ability from generative training.
[0098] Furthermore, to further improve the quality of generative supervision, this paper introduces a Reconstruction-Aware Encoder (RAE) generative representation supervision mechanism. For example, high-quality visual representations (standard image features) provided by an encoder with image reconstruction capabilities can be used as the supervision target of the first network (the generative side), enabling the embedding predicted by the Next-Event Prediction (NEP) to not only contain semantic information but also richer visual detail information.
[0099] like Figure 3 As shown, the information output by the visual encoder can be used as a supervision target, i.e., as standard image features, which can reduce the representation gap caused by resolution transformation. Here, the visual encoder can be a dedicated image reconstruction decoder, thus enabling end-to-end image generation.
[0100] Furthermore, upon receiving image information from training samples, the visual encoder can extract and encode features from the image information to convert it into a visual feature sequence, which can be an image embedding vector. Based on this, the visual feature sequence can be input into the first feature projection layer (Projector Pφ) and the second feature projection layer (Projector Pφ').
[0101] The first feature projection layer can be the image feature projection unit of the backbone of the second machine learning model, based on... Figure 3 As can be seen, the first feature projection layer can be located between the visual encoder and the model backbone. The first feature projection layer can be used to map image features to the text embedding space to generate an image embedding sequence that matches the text token dimension.
[0102] Optionally, the second feature projection layer can be a generative branch projection unit dedicated to the training phase, whose parameters can be initialized by copying the standard projector (first feature projection layer). Figure 3It can be seen that the second feature projection layer can be located between the image encoder (ViT) and the Vision Head of the first network.
[0103] It should be noted that the first feature projection layer remains operational during both the model training and inference phases, and is primarily supervised by text cross-entropy loss, without any direct visual supervision signal. The second feature projection layer can be enabled during the model training phase, and can provide the model with direct visual supervision signals through cosine similarity loss, optimizing the fine-grained expression of visual features. After training is completed, the second feature projection layer can be removed, meaning it does not participate in the model inference process.
[0104] After obtaining the first loss, not only can the first network and the shared network be trained and optimized based on the first loss, but the first feature projection layer can also be trained and optimized based on the first loss. The first network may include a vision head for generating an image prediction head and a first high-level semantic processing layer (I... split+1 ~Lth), where the image generation prediction head can be a dedicated feature processing head for the generation branch, and the first high-level semantic processing layer can be used to extract visual features.
[0105] It should be noted that the second feature projection layer can be updated using EMA (Exponential Moving Average). EMA can be used to maintain the smoothing parameters of the second feature projection layer, preventing training fluctuations from causing drift in the standard image features (standard label features), thus ensuring the stability of the supervision signal for the first network. The first feature projection layer is trainable and participates in gradient backpropagation, while the second feature projection layer is frozen by EMA and does not participate in gradient backpropagation.
[0106] By employing the image generation training paradigm under the aforementioned native autoregressive architecture, machine learning models can simultaneously learn image understanding and image generation tasks in a unified autoregressive manner. In other words, this paper demonstrates that a pure autoregressive architecture can be used for image generation training. For example, the visual representation (embedding) of an image can be used as the prediction target, progressively predicting the next image embedding token (lexical embedding sequence) during the autoregressive decoding process, thus entering the next event prediction (NEP).
[0107] Optionally, during the training of the second machine learning model based on text features, label text features can be extracted. These label text features can be obtained based on the second natural language description information. On this basis, a second loss is determined based on the label text features and the predicted text features, and the second network and the shared network are trained according to the second loss.
[0108] Upon receiving the second natural language information from the training samples, feature extraction can be performed on this second natural language description information to obtain the label text features. The concatenated sequence is then input into the second network, and prediction can be used to obtain the predicted image features. Based on these, a second loss can be derived using the predicted image features and the label text features.
[0109] After obtaining the second loss, not only can the second network and the shared network be trained and optimized based on this second loss, but the first feature projection layer can also be trained and optimized based on this second loss. The second network may include a language prediction output head (LM Head) and a second high-level semantic processing layer (I... split+1 ~Lth), where the language prediction output head can be used to predict the probability distribution of the next word (autoregressive generation), and the second high-level semantic processing layer can be used to extract multimodal semantic features.
[0110] By adding a set of model parameter branches specifically for generation tasks to the standard multimodal large language model, the ability of the large language model to understand visual data can be improved without increasing inference costs, thus improving the accuracy of visual data processing. Here, the model parameter branches specifically for generation tasks can be MoT (Mixture-of-Transformers) branches.
[0111] based on Figure 3 As can be seen, the feature processing layer (Transformer layer) in the backbone network of the Large Language Models (LLM backbone) in this paper can be extended into a two-branch structure. The understanding branch (second network) can retain the original Transformer parameters and is responsible for handling the forward propagation and loss calculation for the understanding task; the generation branch (first network) can include additional Transformer layer parameters, such as MoT layers, which can be specifically responsible for handling the forward propagation for the generation task. Although the two branches share the visual encoder (ViT) and feature projection layer at the input, they can be computed using independent parameters within the backbone network of the Large Language Model.
[0112] As described above, during the training of the second machine learning model, the second network (understanding branch) receives the second loss (text cross-entropy loss) for the understanding task, while the first network (generation branch) receives the first loss (vision loss) of the NEP paradigm. The first and second loss signals propagate back through their respective independent branches, which can, to some extent, allow the generation loss to directly interfere with the optimization of the understanding-side parameters. Therefore, this paper demonstrates that by decoupling the generation and understanding branches through MoT, it can improve understanding capabilities through generative training.
[0113] Understandably, although the large language model parameters of the first and second networks are independent, they share the same visual encoder and feature projection layer. The training of the first network optimizes the shared visual encoder parameters through gradient backpropagation, enabling the visual encoder to learn richer and more accurate visual representations. These improved visual representations, when input into the second network, can provide higher-quality visual information for the understanding task, thereby achieving an indirect positive transfer from generation training to understanding ability.
[0114] By employing a NEP-based image training paradigm and a MoT decoupling architecture, along with a systematic organization strategy incorporating multi-type generation task data, this paper comprehensively enhances the complex visual understanding capabilities of machine learning models without increasing inference costs. Here, the MoT decoupling architecture is merely an example; it can be replaced with other architectures. For instance, a LoRA (Low-Rank Adaptation Adapter) branch can be used to replace the MoT architecture, allowing for the addition of an independent LoRA adapter for generation tasks. This adapter can be used to avoid direct conflicts between the two types of losses on shared parameters. No specific architecture is explicitly restricted here.
[0115] It's worth noting that when passing the hidden features of a large language model to the first network (the generation branch) for NEP prediction, features from the intermediate or shallower layers of the large language model can be extracted for the generation task. Extracting shallower features for the generation task can improve understanding performance. The main reason is that shallower features are closer to the original visual representation and more consistent with the goal of the generation task. This can, to some extent, avoid the dilution of the generation supervision signal by excessive semantic abstraction in deep features, thereby improving the understanding ability of the machine learning model.
[0116] For example, features from approximately half of the total number of layers in the model can be extracted as hidden features and passed to the first network. For instance, in a large language model with 28 layers, hidden features from layers 14 onwards can be passed to the first network. Figure 3 The I shown splitThe first network can have 14 layers, and based on this, hidden features from layers 15 onwards can be passed to the first network.
[0117] Alternatively, the aforementioned NEP can be replaced with discrete token prediction. For example, a discrete image tokenizer can be used to encode the image into a discrete token sequence, which can then be trained using the standard next token prediction method. As can be seen, this paper utilizes an autoregressive paradigm to incorporate image generation into the training. The specific image representation method is not explicitly limited here and can be adjusted according to the actual scenario.
[0118] Alternatively, this paper can fuse the multi-layer features of the large language model and use the fused multi-layer features to generate predictions, or it can adaptively select target features and pass the selected target features to the first network. For example, the optimal layer can be automatically selected through a learnable gating mechanism. By selecting the optimal hidden representation of the large language model, the transmission of generative supervision can be better realized.
[0119] The generated training dataset enables the training of a second machine learning model. After training, the first network of the second machine learning model can be removed to obtain the first machine learning model. In other words, the first machine learning model can include both a second network and a shared network. Upon receiving input information, features can be extracted using this first machine learning model to obtain initial features. Based on these initial features, the shared network is input to obtain hidden features. These hidden features can then be used to predict the visual data corresponding to the input information. In other words, the first machine learning model, obtained through training, can predict the corresponding visual data based on the input information.
[0120] As can be seen, during the inference phase, the first machine learning model can perform forward computation using only the second network (understanding branch), while the parameters of the first network (generation branch) do not participate in inference. Therefore, this paper not only achieves low computational cost, low latency, and low memory usage when performing inference operations, but also realizes the goal of "multi-task learning during training and zero additional cost during inference".
[0121] Upon receiving input information, this content sends the input information to a first machine learning model, which automatically generates the corresponding visual data. Since the first machine learning model is trained using a second machine learning model, which includes both the first and second networks, the first model can accurately generate the visual data corresponding to the input information. In this process, the first network is trained using a first loss related to image generation (corresponding to the generation branch), and the second network is trained using a second loss related to text generation (corresponding to the inference branch). Decoupling these two branches avoids mutual interference, thereby improving the accuracy of visual data processing. Furthermore, by incorporating generation-guided machine learning models when handling comprehension tasks, attention can be focused on the visual regions relevant to the problem, rather than being evenly distributed across the entire image.
[0122] Figure 9 This is a schematic diagram illustrating the structure of a data processing apparatus based on a machine learning model, according to an exemplary embodiment. Figure 9 As shown, this paper provides a data processing device 900 based on a machine learning model, which may include a receiving module 910 and a transmitting module 920.
[0123] The receiving module 910 is configured to receive input information, which includes at least first natural language description information; The sending module 920 is configured to send the input information to a first machine learning model, wherein the first machine learning model is a multimodal large language model; The receiving module 930 is further configured to receive visual data generated by the first machine learning model based on the input information. The first machine learning model is obtained by training a second machine learning model. The second machine learning model includes a first network and a second network. The first network is trained using a first loss, and the second network is trained using a second loss. The first loss is related to image generation, and the second loss is related to text generation.
[0124] In some embodiments, the data processing apparatus 900 based on a machine learning model further includes: The training module is configured to acquire a training dataset, which includes multiple training samples, including image information and second natural language description information; input the image information to a visual encoder to obtain a visual feature sequence; and train a second machine learning model based on the visual feature sequence and / or a text word sequence to obtain a first machine learning model, wherein the text word sequence is obtained based on the second natural language description information.
[0125] In some embodiments, the second machine learning model further includes a shared network, and the training module is further configured to concatenate the visual feature sequence and the text word sequence to obtain a concatenated sequence; send the concatenated sequence to the shared network to obtain intermediate hidden features; input the intermediate hidden features into the first network and the second network respectively to obtain predicted image features and predicted text features, wherein the predicted image features are obtained based on an autoregressive method; and train the second machine learning model based on the predicted image features and the predicted text features.
[0126] In some implementations, the training module is further configured to acquire standard image features; determine the first loss based on the standard image features and the predicted image features; and train the first network and the shared network based on the first loss.
[0127] In some implementations, the training module is further configured to extract labeled text features, which are obtained based on the second natural language description information; determine the second loss based on the labeled text features and the predicted text features; and train the second network and the shared network according to the second loss.
[0128] In some implementations, the training dataset includes training data related to at least one task type, and the training module is further configured to acquire the task type associated with the training samples; determine the task complexity based on the task type; and input the training dataset into the second machine learning model according to the task complexity.
[0129] In some implementations, the task type includes at least one of text-to-image generation, image editing, image segmentation, visual localization, and visual chain reasoning.
[0130] In some implementations, the training module is further configured to: acquire a first image; generate a first label, the first label being acquired by recognizing the first image; detect a first bounding box based on the first label, the first bounding box belonging to the first image; normalize the first bounding box to obtain normalized coordinate information; generate training question-answering samples based on the first image, the first bounding box, and the normalized coordinate information; and generate the training dataset from the training question-answering samples and the first image.
[0131] In some embodiments, the first machine learning model includes the second network and a shared network, and the data processing device 900 based on the machine learning model further includes: The processing module is configured to extract the input information to obtain initial features; input the initial features into the shared network to obtain hidden features; and predict the visual data, which is obtained based on the second network and the hidden features.
[0132] Upon receiving input information, this content sends the input information to a first machine learning model, which automatically generates visual data corresponding to the input information. Since the first machine learning model is obtained by training a second machine learning model, and the second machine learning model includes both the first and second networks, the first machine learning model can accurately generate visual data corresponding to the input information. In this process, the first network is trained using a first loss related to image generation, which corresponds to the generation branch, and the second network is trained using a second loss related to text generation, which corresponds to the inference branch. Therefore, by decoupling the two branches, mutual interference between them can be avoided, thereby improving the accuracy of visual data processing.
[0133] The following is for reference. Figure 10 The diagram illustrates a structural schematic of an electronic device 1000 suitable for implementing the above-described technical solution. Terminal devices may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 10 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.
[0134] like Figure 10 As shown, the electronic device 1000 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for the operation of the electronic device 1000. The processing unit 1001, the ROM 1002, and the RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0135] Typically, the following devices can be connected to the input / output interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic device 1000 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 1000 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0136] In particular, depending on certain circumstances, the processes described in the flowchart above can be implemented as computer software programs. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1008, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the above-described methods.
[0137] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0138] In some implementations, the client and server can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the internet (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0139] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0140] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: receive input information, the input information including at least first natural language description information; send the input information to a first machine learning model, the first machine learning model being a multimodal large language model; and receive visual data generated by the first machine learning model based on the input information, the first machine learning model being acquired by training a second machine learning model, the second machine learning model including a first network and a second network, the first network being trained using a first loss, the second network being trained using a second loss, the first loss being related to image generation, and the second loss being related to text generation.
[0141] Computer program code for performing the above operations can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0142] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0143] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself; for example, a receiving module can also be described as a "module for receiving visual data".
[0144] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0145] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0146] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of this document is not limited to technical solutions formed by specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features provided herein that have similar functions.
[0147] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limitations on the scope of the technical solution. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.
[0148] Although the technical solution has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the aforementioned apparatus, the specific manner in which each module performs its operation has already been described in detail in the section concerning the method, and will not be elaborated upon here.
Claims
1. A data processing method based on a machine learning model, comprising: Receive input information, wherein the input information includes at least a first natural language description; The input information is sent to a first machine learning model, which is a multimodal large language model; The system receives visual data, which is generated by a first machine learning model based on the input information. The first machine learning model is obtained by training a second machine learning model, which includes a first network and a second network. The first network is trained using a first loss, and the second network is trained using a second loss. The first loss is related to image generation, and the second loss is related to text generation.
2. The data processing method based on a machine learning model according to claim 1, further comprising: Obtain a training dataset, which includes multiple training samples, including image information and second natural language description information; The image information is input into the visual encoder to obtain a visual feature sequence; The second machine learning model is trained based on the visual feature sequence and / or text word sequence to obtain the first machine learning model, wherein the text word sequence is obtained based on the second natural language description information.
3. The data processing method based on a machine learning model according to claim 2, wherein the second machine learning model further includes a shared network, and the step of training the second machine learning model based on the visual feature sequence and / or text word sequence includes: The visual feature sequence and the text word sequence are concatenated to obtain the concatenated sequence; Send the spliced sequence to the shared network to obtain the intermediate hidden features; The intermediate hidden features are input into the first network and the second network respectively to obtain predicted image features and predicted text features, wherein the predicted image features are obtained based on an autoregressive method. The second machine learning model is trained based on the predicted image features and the predicted text features.
4. The data processing method based on a machine learning model according to claim 3, wherein training the second machine learning model based on the visual feature sequence and / or text word sequence comprises: Obtain standard image features; The first loss is determined based on the standard image features and the predicted image features; The first network and the shared network are trained based on the first loss.
5. The data processing method based on a machine learning model according to claim 3, wherein training the second machine learning model based on the visual feature sequence and / or text word sequence comprises: Extracting label text features, which are obtained based on the second natural language description information; The second loss is determined based on the labeled text features and the predicted text features; The second network and the shared network are trained based on the second loss.
6. The data processing method based on a machine learning model according to claim 2, wherein the training dataset includes at least one type of task-related training data, and the method further includes: Obtain the task type, which is associated with the training sample; Determine the task complexity based on the task type; The training dataset is input into the second machine learning model according to the task complexity.
7. The data processing method based on a machine learning model according to claim 6, wherein the task type includes at least one of text-to-image generation task, image editing task, image segmentation task, visual localization task, and visual chain reasoning task.
8. The data processing method based on a machine learning model according to claim 2, wherein obtaining the training dataset includes: Get the first image; A first label is generated, which is obtained by recognizing the first image; A first bounding box is detected based on the first label, and the first bounding box belongs to the first image; The first bounding box is normalized to obtain normalized coordinate information; Training question-answering samples are generated based on the first image, the first bounding box, and the normalized coordinate information; The training dataset is generated from the training question-and-answer samples and the first image.
9. The data processing method based on a machine learning model according to claim 1, wherein the first machine learning model includes the second network and a shared network, and the generation of the visual data includes: Extract the input information to obtain initial features; The initial features are input into the shared network to obtain hidden features; The visual data is predicted based on the second network and the hidden features.
10. A data processing device based on a machine learning model, comprising: The receiving module is configured to receive input information, which includes at least first natural language description information; The sending module is configured to send the input information to a first machine learning model, wherein the first machine learning model is a multimodal large language model; The receiving module is further configured to receive visual data generated by the first machine learning model based on the input information. The first machine learning model is obtained by training a second machine learning model, which includes a first network and a second network. The first network is trained using a first loss, and the second network is trained using a second loss. The first loss is related to image generation, and the second loss is related to text generation.
11. A computer-readable medium having a computer program stored thereon, wherein, When executed by a processing device, the computer program performs the steps of the method according to any one of claims 1-9.
12. An electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-9.
13. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-9.