Method and device for model training, equipment, storage medium and program product
By dynamically selecting the loss function and combining sample input data and thought chain information to optimize the training of the multimodal language model, the problems of poor generalization ability and excessive resource consumption of traditional models are solved, and efficient and accurate multimodal data processing is achieved.
Patent Information
- Application Number
- CN202511197481.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional multimodal language models have poor generalization capabilities when processing multimodal data, and the training process consumes excessive resources, making it difficult to maintain efficiency and accuracy in cross-modal interactions.
By dynamically selecting the loss function, utilizing sample input data and thought chain information, dynamically determining the target loss of the training sample, and combining the first loss and the second loss, the training process of the machine learning model is optimized, computing overhead is reduced, and model accuracy is improved.
It improves the training efficiency and accuracy of multimodal language models, reduces training costs, and enhances the model's generalization and cross-modal interaction capabilities.
Smart Images

Figure CN120804718A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, electronic device, computer-readable storage medium and computer program product for model training. BACKGROUND
[0002] With the development of machine learning technology, it has been possible to utilize machine learning models to perform tasks in a variety of application environments. A multimodal keyphrase prediction (MMKP) task is a task aiming to determine keyphrases (e.g., labels) of multimodal media data, which can include data of any modality such as images, texts, videos, etc. A multimodal language model (MLM) is usually utilized to perform the MMKP task. The multimodal language model is capable of processing and fusing data inputs of multiple modalities, supporting applications across various domains, including image captioning, visual question answering, and video analysis, etc. SUMMARY
[0003] In a first aspect of the present disclosure, a method for model training is provided. The method comprises: obtaining a training sample, the training sample comprising sample input data, a sample keyphrase, and sample thinking chain information, the sample thinking chain information indicating an inference process from the sample input data to the sample keyphrase; determining, by utilizing a machine learning model, a first predicted keyphrase corresponding to the training sample based at least on the sample input data, the machine learning model being configured to process a model input to generate a keyphrase and / or thinking chain information corresponding to the model input; determining a first loss based on a difference between the first predicted keyphrase and the sample keyphrase; determining a target loss to be used for the training sample based on a comparison between the first loss and a threshold loss, wherein the target loss is selected from the first loss or a second loss, and the second loss is based on a difference between predicted thinking chain information corresponding to the training sample and the sample thinking chain information, and a difference between a second predicted keyphrase corresponding to the training sample and the sample keyphrase; and training the machine learning model based at least on the target loss to be used for the training sample.
[0004] In a second aspect of the disclosure, an apparatus for model training is provided. The apparatus includes: a sample obtaining module configured to obtain a training sample, the training sample including sample input data, a sample keyword, and sample thinking chain information, the sample thinking chain information indicating an inference process from the sample input data to the sample keyword; a keyword determining module configured to determine, based on at least the sample input data, a first predicted keyword corresponding to the training sample by using a machine learning model, the machine learning model being configured to process a model input to generate a keyword and / or thinking chain information corresponding to the model input; a first loss determining module configured to determine a first loss based on a difference between the first predicted keyword and the sample keyword; a target loss determining module configured to determine, based on a comparison between the first loss and a threshold loss, a target loss to be used for the training sample, wherein the target loss is selected from the first loss or a second loss, and the second loss is based on a difference between predicted thinking chain information corresponding to the training sample and the sample thinking chain information, and a difference between a second predicted keyword corresponding to the training sample and the sample keyword; and a model training module configured to train the machine learning model based on at least the target loss to be used for the training sample.
[0005] In a third aspect of the disclosure, an electronic device is provided. The electronic device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to the first aspect of the disclosure.
[0006] In a fourth aspect of the disclosure, a computer-readable storage medium is provided, having stored thereon computer-executable instructions that, when executed by a processor, cause the processor to perform the method according to the first aspect of the disclosure.
[0007] In a fifth aspect of the disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0008] It should be understood that the content described in this section is not intended to limit key features or important features of the embodiments of the disclosure, nor is it intended to limit the scope of the disclosure. Other features of the disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, advantages, and aspects of embodiments of the disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0010] Figure 1 A schematic diagram showing an example environment for model training and application is shown in accordance with some embodiments of the present disclosure;
[0011] Figure 2A and Figure 2B A schematic diagram showing an example architecture for model training is shown in accordance with some embodiments of the present disclosure;
[0012] Figure 3 A flowchart showing a method for model training is shown in accordance with some embodiments of the present disclosure;
[0013] Figure 4 An example block diagram of an apparatus for model training is shown in accordance with some embodiments of the present disclosure; and
[0014] Figure 5 A block diagram of an electronic device that can implement one or more embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0016] In the description of embodiments of the present disclosure, the term "includes" and its derivatives, shall be understood to be open-ended, i.e., to mean "includes but is not limited to". The term "based on" means "based, at least in part, on". The term "one embodiment" or "an embodiment" shall mean "at least one embodiment". The term "some embodiments" shall mean "at least some embodiments". The terms "a" or "an" shall mean "one or more". The terms "first", "second", and the like can refer to different or same objects. Other explicit or implicit definitions can also be included below.
[0017] In this document, unless explicitly stated otherwise, performing a step "in response to" A does not mean that the step is performed immediately after A, but can include one or more intermediate steps.
[0018] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the obtaining, use, storage or deletion of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0019] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0020] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using personal information of the user, so that the user can voluntarily choose whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0021] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0022] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0023] As used herein, the term "model" can learn an association between respective inputs and outputs from training samples (which can also be referred to as training data, sample data, etc.), so that after training, a corresponding output can be generated for a given input. The generation of a model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this document, a "model" can also be referred to as a "machine learning model", a "learning model", a "machine learning network" or a "learning network", which are used interchangeably herein.
[0024] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, which generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications generally include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence, so that the output of a previous layer is provided as the input of a subsequent layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also referred to as processing nodes or neurons), each of which processes the input from the previous layer.
[0025] Generally, machine learning can include three stages, i.e., a training stage, a testing stage, and an application stage (also referred to as an inference stage). In the training stage, a given model can be trained using a large number of training samples, and the parameter values are iteratively updated until the model can obtain consistent inferences from the training samples that meet the expected target. Through training, the model can be considered to be able to learn the association (also referred to as the mapping) between the input and the output from the training samples. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be merged into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values obtained through training to determine the corresponding model output.
[0026] As mentioned above, the multi-modal keyword prediction task is a task aiming to determine the keyword of multi-modal media data, and a multi-modal language model is usually used to perform the task. The multi-modal media data usually includes text and visual data (such as images, videos, etc.) matching the text. The conventional multi-modal language model is limited by the model capacity and multi-task ability, and usually relies on an external optical character recognition (OCR) system and a visual feature extraction module to process the visual data, and uses the processing result of the visual data to enhance the understanding ability of the text. The generalization ability and integration ability of such a multi-modal language model are usually poor, and the cross-modal interaction ability of processing different modal data is also poor.
[0027] In addition, since the data of different modalities usually includes different information and corresponds to different features, the training process of the multi-modal language model usually needs to use a very large amount of multi-modal data to train the model in multiple stages, so that the model can fully learn the association pattern between different modalities. Traditionally, zero-shot, supervised fine-tuning (SFT), etc. are usually used to train the multi-modal language model.
[0028] Although some model training schemes that help improve the performance of the model are traditionally proposed, the traditional scheme of increasing the amount of training samples usually has the problem of excessive consumption of resources. The traditional supervised fine-tuning scheme usually results in poor generalization ability of the model, and the model relies heavily on the training samples, and may not be able to output accurate results when facing unfamiliar data. Currently, in order to overcome the disadvantages of supervised fine-tuning training, multi-modal Chain-of-Thought (COT) information is proposed, aiming to activate and supplement the world knowledge of the model and strengthen its reasoning ability.
[0029] However, reality shows that directly introducing COT information does not immediately improve model performance, and COT will increase additional computational overhead during the inference phase. In addition, models that introduce COT often suffer from problems of overthinking and content redundancy. Overthinking means that the model generates overly general model outputs during inference, which affects the accuracy of the model output. Content redundancy means that for similar model inputs, the model outputs generated by the model will also be highly similar, for example, they will include almost identical reasoning paths, which will lead to serious model redundancy and reduce the effectiveness of the model. It is hoped that the accuracy and effectiveness of the trained model can be guaranteed while conveniently and efficiently training the multimodal language model.
[0030] In view of this, according to an embodiment of the present disclosure, an improved scheme for model training is provided. According to the scheme of an embodiment of the present disclosure, a training sample is obtained, and the training sample includes sample input data, sample keywords and sample thought chain information, and the sample thought chain information indicates the reasoning process from the sample input data to the sample keyword. Using a machine learning model, a first predicted keyword corresponding to the training sample is determined based on at least the sample input data, and the machine learning model is configured to process the model input to generate keywords and / or thought chain information corresponding to the model input. A first loss is determined based on the difference between the first predicted keyword and the sample keyword. A target loss to be used for the training sample is determined based on a comparison between the first loss and the threshold loss, wherein the target loss is selected from the first loss or the second loss, and the second loss is based on the difference between the predicted thought chain information corresponding to the training sample and the sample thought chain information, and the difference between the second predicted keyword corresponding to the training sample and the sample keyword. The machine learning model is trained based at least on the target loss to be used for the training sample.
[0031] In this way, for different training samples, the loss to be used for model training can be dynamically determined from different types of losses (i.e., the first loss and the second loss) based on the comparison results of the first loss of the training sample and the threshold loss. In this way, depending on the current capabilities of the model, the training sample can be used to continuously update the model according to different types of training losses. This not only improves the model's capabilities, but also speeds up model training and reduces training costs.
[0032] Figure 1 Schematic diagram of an example environment 100 for model training and application according to some embodiments of the present disclosure is shown. Figure 1 The example environment 100 of FIG. 1 shows three different stages of the model, including a pre-training stage 102, a fine-tuning stage 104, and an application stage 106. After the pre-training stage 102 or the fine-tuning stage 104 is completed, there may also be a testing stage, which is not shown in the figure.
[0033] The training sample set construction system 110 can be configured to generate a training sample set for training the machine learning model 125, which can include at least one of the pre-training sample set 131 and the training sample set 141. The training sample set construction system 110 can employ any suitable manner to generate the training sample set. In some embodiments, the training sample set construction system 110 can utilize the generative model 115 to generate the training sample set.
[0034] Both the generative model 115 and the machine learning model 125 can be based on any suitable model structure, including but not limited to a Transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), and the like. In some embodiments, the generative model 115 and the machine learning model 125 can also be based on a language model (LM) or other any suitable model. By way of example only, the generative model 115 and the machine learning model 125 can be based on a multi-modal language model, a multi-modal large language model (MLLM), a visual language model (VLM), and the like.
[0035] In the pre-training stage 102, the model pre-training system 130 is configured to perform pre-training of the machine learning model 125 with the pre-training sample set 131. At the beginning of the pre-training, the individual components in the machine learning model 125 can have initial parameter values. The pre-training process is to update the parameter values of the machine learning model 125 to desired values based on the data in the pre-training sample set 131. The pre-training task is to help the parameter update of the machine learning model 125.
[0036] In the pre-training stage 102, the machine learning model 125 can learn strong generalization capability through the pre-training sample set 131 that includes a large amount of data. After the pre-training is completed, the parameter values of the machine learning model 125 have been updated, with pre-trained parameter values. The pre-trained machine learning model 125, for example, can perform simple tasks with relatively high accuracy.
[0037] The pre-trained machine learning model 125 can be provided to the fine-tuning stage 104 for fine-tuning for different downstream tasks by the model fine-tuning system 140. In the fine-tuning stage 104, the parameter values of the machine learning model 125 are further adjusted with the training sample set 141. The parameters of the machine learning model 125 are also updated and adjusted with the corresponding training algorithm during the fine-tuning. Since the model has learned a lot of knowledge from the training samples in the pre-training stage, a desired downstream task model can be obtained with a small amount of training samples in the fine-tuning stage 104. The fine-tuned machine learning model 125, for example, can perform a specific downstream task with relatively high accuracy.
[0038] In some embodiments, a testing phase can be further included after the fine-tuning phase 104, and the model performance of the machine learning model 125 can be tested with a test dataset. The dataset used in the testing phase can be of the same type as the fine-tuning phase. With the testing phase, only the machine learning model that passes the test can be provided to the subsequent application phase 106, and the machine learning model that fails the test can be retrained.
[0039] In the application phase 106, the obtained machine learning model 125 with trained parameter values can be provided to a model application system 150 for use. In the application phase 106, the machine learning model 125 can be utilized to process corresponding input data 151 in actual scenarios to obtain corresponding output data 152. The specific content of the input data 151 and the output data 152 can be flexibly set based on the actual scenario.
[0040] In Figure 1 In some embodiments, the training sample set construction system 110, the model pre-training system 130, the model fine-tuning system 140, and the model application system 150 can be deployed at any appropriate electronic device. The electronic device can be any type of device with computing capability, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof.
[0041] The server device can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network, and big data and artificial intelligence platform. The server device may, for example, include a computing system / server such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.
[0042] It can be appreciated that the training sample set construction system 110, the model pre-training system 130, the model fine-tuning system 140, and the model application system 150 can be deployed on the same electronic device, or can be deployed on different electronic devices, which are not limited in the present disclosure. For example only, the training sample set construction system 110, the model pre-training system 130, the model fine-tuning system 140, and the model application system 150 can all be deployed at the electronic device 120 as shown in FIG. 1. Figure 1
[0043] It should be appreciated that Figure 1 The components and arrangements shown in the example environment 100 shown are for example only and a computing system suitable for implementing the example implementations described in the present disclosure can include one or more different components, other components, and / or different arrangements of components. For example, while shown as separate, the training sample set construction system 110, the model pre-training system 130, the model fine-tuning system 140, and the model application system 150 can be integrated in the same system or device. For example, at least the model pre-training system 130 and the model fine-tuning system 140 can be integrated in a model training system or device. Implementations of the present disclosure are not limited in this regard.
[0044] Some example embodiments of the present disclosure will be hereinafter described with continued reference to the drawings. In this document, the model training method involved in the present disclosure is implemented at the electronic device 120, and the electronic device 120 is deployed with at least the model pre-training system 130, the model fine-tuning system 140, and the model application system 150 as an example to be described exemplarily.
[0045] Reference Figure 2A , Figure 2A An example architecture 200A for model training according to some embodiments of the present disclosure is shown. The example 200A can be implemented at the electronic device 120 and specifically can be implemented at the model pre-training system 130 and / or the model fine-tuning system 140 in the Figure 1 The training process shown in the example 200A can be a training process of the machine learning model 125, which can be configured to process a model input to generate keyword and / or thought chain information corresponding to the model input. As mentioned previously, the machine learning model 125 can be based on any suitable model structure, and for example only, the machine learning model 125 can be based on an MLM.
[0046] In embodiments of the present disclosure, the electronic device 120 obtains training samples 210 for training the machine learning model 125, the training samples 210 including sample input data 216, sample keywords 213, and sample Chain-of-Thought (COT) information 215. In some embodiments, the sample input data 216 can be multi-modal data. By way of example only, the sample input data 216 can include sample text 212 and sample visual content 211 that matches the sample text. The sample visual content 211 can include images and / or video (e.g., video frames extracted from a video). It can be appreciated that the sample input data 216 can also include other any suitable content, such as audio, file, etc., without limitation. The sample keywords 213 can be keywords corresponding to the text in the sample input data 216, for example, which can include one or more sample keywords. The sample Chain-of-Thought information 215 can indicate an inference process from the sample input data 216 to the sample keywords 213. The sample Chain-of-Thought information 215 can be determined manually or by means of predetermined rules, algorithms, models, etc.
[0047] The training samples 210 can be manually determined and provided to the electronic device 120 by a user, or can be generated and provided to the electronic device 120 by the training sample set construction system 110 based on predetermined rules, algorithms, etc. The training sample set construction system 110 can be deployed at the electronic device 120 or at other devices. In some embodiments, the training sample set construction system 110 can obtain original samples, which can include sample input data 216 and sample keywords. The training sample set construction system 110 can process the original samples to determine sample Chain-of-Thought information corresponding to the original samples, and then generate training samples 210 based on the original samples (i.e., the sample input data 216 and the sample keywords) and the sample Chain-of-Thought information. The electronic device 120 can train the machine learning model 125 using such multiple training samples 210.
[0048] The training sample set construction system 110 can determine the sample Chain-of-Thought information by means of the generative model 115, for example. Specifically, referring to Figure 2B , Figure 2B An example 200B for model training is shown in accordance with some embodiments of the present disclosure. As shown in FIG. 2B, the training sample set construction system 110 can obtain original samples 220, which can include sample input data 216 and sample keywords 213. The training sample set construction system 110 can determine sample Chain-of-Thought information 215 for the original samples 220 by means of the generative model 115, and then generate training samples 210 based on the original samples 220 (i.e., the sample input data 216 and the sample keywords 213) and the sample Chain-of-Thought information 215. Figure 2BAs shown, the training sample set construction system 110 can determine the prompt word information 261, the sample input data 216, and the sample keyword 213 as a model input, and provide the model input to the generative model 115. The prompt word information 261 can be used to instruct the generative model 115 to generate the sample thought chain information 215 indicating an inference process from the sample input data 216 to the sample keyword 213. The prompt word information 261 may, for example, also indicate a format of the sample thought chain information 215, a way to generate the thought chain information, etc. The training sample set construction system 110 can obtain a model output generated by the generative model 115 for the above model input, which can indicate the sample thought chain information 215 for the above model input.
[0049] During the training of the machine learning model 125, the electronic device 120 can perform training of the machine learning model 125 with a large number of training samples 210. For each training sample 210, the electronic device 120 can determine, based at least on the sample input data 216, a first predicted keyword corresponding to the training sample 210, by utilizing the machine learning model 125. As an example, the electronic device 120 can generate the first predicted keyword based on the sample input data 216 and the sample keyword 213, by utilizing the machine learning model 125. The first predicted keyword can include one or more predicted keywords.
[0050] In some embodiments, the electronic device 120 can also obtain prompt word information 214 for the machine learning model 125, which can be used at least to instruct the machine learning model 125 how to generate the predicted keyword. The electronic device 120 can provide the prompt word information 214 to the machine learning model 125 together with other input data. For example, the electronic device 120 can provide the sample visual content 211, the sample text 212, the sample keyword 213, and the prompt word information 214 to the machine learning model 125 to instruct the machine learning model 125 to generate the first predicted keyword corresponding to the training sample 210. As an example only, the prompt word information 214 can include at least the text “You are a very helpful assistant, please analyze the keywords of the article based on the given image and article”.
[0051] As to the specific way that the machine learning model 125 generates the first predicted keyword, in some embodiments, the machine learning model 125 comprises a visual encoder 220, a text encoder 230, and a decoder 240. The visual encoder 220 is configured to perform encoding on data of the visual modality (e.g., the sample visual content 211) to determine a sample visual feature sequence 225 corresponding to the sample visual content 211. The text encoder 230 is configured to perform encoding on data of the text modality to determine a corresponding text feature sequence, e.g., to perform encoding on the sample text 212 to determine a sample text feature sequence corresponding to the sample text 212. The decoder 240 is configured to determine a predicted text feature sequence 245 based at least on the sample text feature sequence and the sample visual feature sequence 225.
[0052] As to the specific way that the decoder 240 determines the predicted text feature sequence 245, in some embodiments, an input text feature sequence 235 for the decoder 240 can be determined. In the scenario of generating the first predicted keyword, the input data of the machine learning model 125 can comprise at least the sample visual content 211, the sample text 212, and the sample keyword 213, where the sample text 212 and the sample keyword 213 can each be data of the text modality, and the text encoder 230 can determine a sample text feature sequence and a sample keyword feature sequence corresponding to the sample text 212 and the sample keyword 213, respectively.
[0053] The input text feature sequence 235 for the decoder 240 can comprise the sample text feature sequence and the sample keyword feature sequence. It is noted that if the input data of the machine learning model 125 further comprises the prompt word information 214 of the text modality, the text encoder 230 can further determine a prompt word feature sequence corresponding to the prompt word information 214, and the input text feature sequence 235 for the decoder 240 can further comprise the prompt word feature sequence.
[0054] The decoder 240 can determine the predicted text feature sequence 245 in a sequence-to-sequence manner, for example. Specifically, for a first input text feature in the input text feature sequence 235 (which can be any appropriate input text feature), the decoder 240 can determine a first predicted text feature corresponding to the first input text feature based on a portion of the input text feature sequence 235 before the first input text feature, a portion of the predicted text output feature sequence corresponding to the portion of the input text feature sequence 235, and the sample visual feature sequence 225.
[0055] The decoder 240 can determine the predicted text features corresponding to the input text feature sequence 235 in turn based on such a manner, and further determine a predicted text feature sequence 245 corresponding to the input text feature sequence 235. In the scenario of determining the first predicted keyword, the predicted text feature sequence 245 can be further decoded to obtain the first predicted keyword. The electronic device 120 can determine a first loss 252 based on the difference between the first predicted keyword and the sample keyword 213. The first loss 252 can also be referred to as a supervised fine-tuning (SFT) loss.
[0056] The electronic device 120 can obtain a threshold loss, and determine a comparison result between the first loss 252 and the threshold loss. The threshold loss can be flexibly configured to be any appropriate numerical value based on the visual scene. The electronic device 120 can determine whether it is necessary to determine the second predicted keyword and the predicted thought chain information corresponding to the training sample 210 based on the comparison result between the first loss 252 and the threshold loss. For example only, the electronic device 120 can determine whether it is necessary to determine the second predicted keyword and the predicted thought chain information corresponding to the training sample 210 based on the following formula:
[0057]
[0058] wherein indicates the first loss 252 corresponding to the training sample 210, y s indicates the model output of the model in the process of determining the first predicted keyword, y c indicates the model output of the model in the process of determining the second predicted keyword and the predicted thought chain information, y d indicates the current model output of the machine learning model 125. y indicates the threshold loss, which can be any appropriate numerical value. According to the above formula (1), in the case where the first loss reaches (i.e., is greater than or equal to) the threshold loss, the model output is the predicted keyword based on SFT inference, and in the case where the first loss is less than the threshold loss, the model output is the predicted keyword based on thought chain (COT) inference and the thought chain information.
[0059] Thus, the electronic device 120 can determine that the machine learning model 125 can be trained directly based on the first loss obtained based on the first predicted keyword in response to the first loss 252 reaching the threshold loss. In this case, the electronic device 120 does not need to determine the second predicted keyword and the predicted train of thought information. The electronic device 120 can determine that the second predicted keyword and the predicted train of thought information corresponding to the training sample 210 need to be determined in response to the first loss 252 not reaching the threshold loss, to further train the machine learning model 125 based on the second loss corresponding to the second predicted keyword and the predicted train of thought information. Thus, the electronic device 120 does not need to determine the second predicted keyword and the predicted train of thought information of all training samples by means of the model, which can reduce the computational amount of the model.
[0060] In the case of determining that the second predicted keyword and the predicted train of thought information need to be determined, the electronic device 120 may, for example, also determine the second predicted keyword and the predicted train of thought information corresponding to the training sample 210 based on at least the sample input data 216 by means of the machine learning model 125. The input data used to generate the first predicted keyword and the second predicted keyword can be different. The electronic device 120 may, for example, also need to determine the second predicted keyword and the predicted train of thought information based on the sample train of thought information 215. As an example, the electronic device 120 may, for example, generate the second predicted keyword and the predicted train of thought information based on the sample input data 216, the sample keyword 213 and the sample train of thought information 215 by means of the machine learning model 125. Similarly, the electronic device 120 may, for example, also obtain prompt word information for the machine learning model 125 for determining the second predicted keyword and the predicted train of thought information, which can be used to inform the machine learning model 125 how to generate the predicted keyword and the predicted train of thought information.
[0061] It should be noted that the prompt word information for the first predicted keyword and the prompt word information for the second predicted keyword can be the same or different, and herein is only exemplarily described by taking the two as the same prompt word information as an example.
[0062] The electronic device 120 may, for example, provide the sample visual content 211, the sample text 212, the sample keyword 213, the prompt word information 214 and the sample train of thought information 215 to the machine learning model 125 together to instruct the machine learning model 125 to generate the second predicted keyword and the predicted train of thought information corresponding to the training sample 210.
[0063] The sample thought chain information 215 may, for example, be text modality data. The text encoder 230 of the machine learning model 125 may, for example, determine a sample thought chain feature sequence corresponding to the sample thought chain information 215. The input text feature sequence 235 for the decoder 240 of the machine learning model 125 may, for example, include the sample text feature sequence, the sample keyword feature sequence, the prompt word feature sequence, and the sample thought chain feature sequence. The decoder 240 may, again, employ a sequence-to-sequence manner to determine a predicted text feature sequence 245 corresponding to the input text feature sequence 235. In scenarios where a second predicted keyword and predicted thought chain information are determined, the predicted text feature sequence 245 may, subsequently, be further decoded to obtain the second predicted keyword and the predicted thought chain information.
[0064] The electronic device 120 may, for example, determine a second loss 254 based on a difference between the predicted thought chain information corresponding to the training sample 210 and the sample thought chain information 215 (which may, for example, be referred to as a thought chain difference), and a difference between the second predicted keyword corresponding to the training sample 210 and the sample keyword 213 (which may, for example, be referred to as a keyword difference). The second loss 254 may, again, be referred to as a COT loss. By way of example only, the thought chain difference and the keyword difference may be weighted, and the second loss 254 may be determined based on a result of the weighting. The weights corresponding to the thought chain difference and the keyword difference may, for example, be default, user-specified, or determined by the model.
[0065] In some embodiments, the electronic device 120 may, for example, determine a loss for the machine learning model 125 based on the following formula:
[0066]
[0067] where T indicates a length of the input text feature sequence 235, i.e., a number of input text features included in the input text feature sequence 235. t indicates an ordering of the first input text feature in the input text feature sequence 235, i.e., the first input text feature to be predicted is the t-th feature in the input text feature sequence 235. i.e., y d is a concatenation of y p and i.e., y p represents a partial input text feature sequence before the first input text feature, represents a partial predicted text output feature sequence corresponding to the partial input text feature sequence. v indicates the sample visual feature sequence, and θ represents the current model parameters of the machine learning model 125. represents the loss of the machine learning model 125. The electronic device 120 can determine the cross-entropy loss corresponding to the current training sample 210 based on formula (2) and the current input text feature sequence 235.
[0068] It can be understood that the model input and the model output corresponding to the first loss 252 and the second loss 254 are different. If the input data of the model is the input data used only for generating the first predicted keyword, the electronic device 120 can determine the current y d = y s , That is, the loss currently determined by the model is the first loss determined based on the model output in the process of determining the first predicted keyword. If the input data of the model is the input data used for generating the second predicted keyword and the predicted thought chain information, the electronic device 120 can determine the current y d = y c ,, wherein represents the second loss. That is, the loss currently determined by the model is the second loss determined based on the model output in the process of determining the second predicted keyword and the predicted thought chain information.
[0069] For example, the numerical value corresponding to each of the first loss 252 and the second loss 254 can belong to the set [0, 1], and the threshold loss γ in the foregoing can be 0.4, for example. The goal of model training is to constantly reduce the loss value of the loss function by updating the model parameters, so that the loss value is minimized or reduced to a predetermined target, thereby completing the training of the model.
[0070] Thus, in the case where the first loss 252 reaches the threshold loss, the training sample 210 can be determined as a difficult sample, and at this time, training the model using the first loss 252 can improve the training effect of the model. In the case where the first loss 252 does not reach the threshold loss, the training sample 210 can be determined as a simple sample, and at this time, training the model using the second loss 254 can avoid overfitting and also help improve the training effect of the model. Thus, for different training samples, it can be dynamically determined whether to train the model based on the first loss 252 (SFT loss) or the second loss 254 (COT loss), which can determine the loss that is more matched with the training sample, and help improve the training effect. Through such training, the machine learning model can be trained to have the ability to automatically determine whether to generate thought chain information for the model input.
[0071] As to the specific manner of model training, in some embodiments, the electronic device 120 can synchronize the training of the encoder (e.g., the visual encoder 220 and the text encoder 230) and the decoder (e.g., the decoder 240) of the machine learning model 125, that is, the model parameters of the encoder and the model parameters of the decoder are both changed during the training process. In some embodiments, the encoder in the machine learning model 125 can also be a trained encoder, in which case the electronic device 120 can only train the decoder, that is, the model parameters of the encoder remain unchanged, and only the model parameters of the decoder are changed during the training process.
[0072] In some embodiments, in the case where the machine learning model 125 includes multiple encoders, some of the encoders can also be trained, and the electronic device 120 can train the other encoders and the decoder in addition to the trained encoders, that is, the model parameters of the trained encoders remain unchanged, and the model parameters of the other encoders and the model parameters of the decoder are both changed during the training process. For example, for the machine learning model 125 shown in the figure, the visual encoder 220 can be a trained encoder, and during the model training process, the model parameters of the visual encoder 220 remain unchanged, and the model parameters of the decoder 240 are changed during the training process. In this way, the trained encoder can be used to extract features, which can reduce the amount of parameters that need to be updated during the training process and improve the efficiency of the training. Figure 2A
[0073] The electronic device 110 may, for example, determine that the training of the machine learning model 125 is complete in response to determining that the training termination requirement is met. The training termination requirement may, for example, indicate that the first loss and / or the second loss corresponding to the training sample 210 is less than a threshold value or is 0. After the training of the machine learning model 125 is complete, the electronic device 110 (specifically, the model application system 150) can input target input data in an actual scenario to the trained machine learning model 125, and the target input data 125 may, for example, include visual content and text (e.g., an image-text pair, a video-text pair, a set of images and texts, etc.).
[0074] The electronic device 120 can use the trained machine learning model 125 to generate a target keyword of the target input data, or a target keyword and a thinking chain process from the target input data to the target keyword. The target keyword can include one or more keywords. For example only, the electronic device 120 can determine the difficulty of the target input data. The electronic device 120 may, for example, use the machine learning model 125 to generate a target keyword and a thinking chain process corresponding to the target input data in the case where it is determined that the target input data is relatively simple, and can use the machine learning model 125 to generate a target keyword corresponding to the target input data in the case where it is determined that the target input data is relatively difficult.
[0075] Of course, here is only an example, the electronic device 120 can also generate the target keyword corresponding to the target input data by using the machine learning model 125 in the case of determining that the target input data is relatively simple, and can generate the target keyword and the thinking chain process corresponding to the target input data by using the machine learning model 125 in the case of determining that the target input data is relatively difficult. The present disclosure does not limit this. Thus, the output of the machine learning model can be determined based on the actual scene, which can improve the generalization ability of the model.
[0076] Figure 3 A flowchart of a method 300 for model training is shown, according to some embodiments of the present disclosure. The method 300 may, for example, be implemented at the electronic device 120 of Figure 1 The method 300 will be described with reference to the environment 100 of Figure 1 The method 300 will be described with reference to the environment 100 of
[0077] At block 310, the electronic device 120 obtains a training sample, the training sample including sample input data, a sample keyword, and sample thinking chain information, the sample thinking chain information indicating an inference process from the sample input data to the sample keyword.
[0078] At block 320, the electronic device 120 determines, based at least on the sample input data, a first predicted keyword corresponding to the training sample by using a machine learning model, the machine learning model being configured to process a model input to generate a keyword and / or thinking chain information corresponding to the model input.
[0079] At block 330, the electronic device 120 determines a first loss based on a difference between the first predicted keyword and the sample keyword.
[0080] At block 340, the electronic device 120 determines, based on a comparison between the first loss and a threshold loss, a target loss to be used for the training sample, wherein the target loss is selected from the first loss or a second loss, and the second loss is based on a difference between predicted thinking chain information corresponding to the training sample and the sample thinking chain information, and a difference between a second predicted keyword corresponding to the training sample and the sample keyword.
[0081] At block 350, the electronic device 120 trains the machine learning model based at least on the target loss to be used for the training sample.
[0082] In some embodiments, determining, based on the comparison between the first loss and the threshold loss, the target loss to be used for the training sample includes: in response to determining that the first loss reaches the threshold loss, determining the first loss as the target loss to be used for the training sample; and in response to determining that the first loss does not reach the threshold loss, determining the second loss as the target loss to be used for the training sample.
[0083] In some embodiments, the first predicted keyword is determined by generating, based on the sample input data and the sample keyword, the first predicted keyword using the machine learning model, and the second predicted keyword and the predicted thinking chain information are determined by generating, based on the sample input data, the sample keyword and the sample thinking chain information, the second predicted keyword and the predicted thinking chain information using the machine learning model.
[0084] In some embodiments, the sample input data comprises sample text and sample visual content, the machine learning model comprises a text encoder, a visual encoder and a decoder, and wherein determining, based on at least the sample input data, the first predicted keyword corresponding to the training sample using the machine learning model comprises: performing encoding on the sample text using the text encoder to obtain a sample text feature sequence corresponding to the sample text; performing encoding on the sample visual content using the visual encoder to obtain a sample visual feature sequence corresponding to the sample visual content; and determining, based on at least the sample text feature sequence and the sample visual feature sequence, a predicted text feature sequence using the decoder, the predicted text feature sequence being decoded to obtain the first predicted keyword.
[0085] In some embodiments, during the training process of the machine learning model, the model parameters of the visual encoder remain unchanged, and the model parameters of the decoder change with the training process.
[0086] In some embodiments, obtaining the training sample comprises: obtaining the sample input data and the sample keyword; generating, based on the sample input data and the sample keyword, the sample thinking chain information using a generative model, the generative model being configured to process a model input to generate corresponding thinking chain information; and determining the training sample based on the sample input data, the sample keyword and the sample thinking chain information.
[0087] In some embodiments, the method 300 further comprises: inputting target input data to the trained machine learning model; and generating, based on the target input data, a target keyword of the target input data, or the target keyword and a thinking chain process from the target input data to the target keyword using the machine learning model.
[0088] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. Figure 4 An exemplary structural block diagram of an apparatus 400 for model training according to some embodiments of the present disclosure is shown. The apparatus 400 can be implemented as or included in the electronic device 120. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware or any combination thereof.
[0089] As Figure 4As shown, the apparatus 400 includes a sample obtaining module 410 configured to obtain a training sample, the training sample including a sample input data, a sample keyword, and sample thinking chain information, the sample thinking chain information indicating an inference process from the sample input data to the sample keyword. The apparatus 400 further includes a keyword determining module 420 configured to determine, based on at least the sample input data, a first predicted keyword corresponding to the training sample by utilizing a machine learning model, the machine learning model being configured to process a model input to generate a keyword and / or thinking chain information corresponding to the model input. The apparatus 400 further includes a first loss determining module 430 configured to determine a first loss based on a difference between the first predicted keyword and the sample keyword. The apparatus 400 further includes a target loss determining module 440 configured to determine, based on a comparison between the first loss and a threshold loss, a target loss to be used for the training sample, wherein the target loss is selected from the first loss or a second loss, and the second loss is based on a difference between predicted thinking chain information corresponding to the training sample and the sample thinking chain information, and a difference between a second predicted keyword corresponding to the training sample and the sample keyword. The apparatus 400 further includes a model training module 450 configured to train the machine learning model based on at least the target loss to be used for the training sample.
[0090] In some embodiments, the target loss determining module 440 is further configured to: in response to determining that the first loss reaches the threshold loss, determine the first loss as the target loss to be used for the training sample; and in response to determining that the first loss does not reach the threshold loss, determine the second loss as the target loss to be used for the training sample.
[0091] In some embodiments, the first predicted keyword is determined by: generating, based on the sample input data and the sample keyword, the first predicted keyword by utilizing the machine learning model, and the second predicted keyword and the predicted thinking chain information are determined by: generating, based on the sample input data, the sample keyword, and the sample thinking chain information, the second predicted keyword and the predicted thinking chain information by utilizing the machine learning model.
[0092] In some embodiments, the sample input data includes a sample text and a sample visual content, the machine learning model includes a text encoder, a visual encoder, and a decoder, and wherein the keyword determining module 420 is further configured to: perform encoding on the sample text by utilizing the text encoder to obtain a sample text feature sequence corresponding to the sample text; perform encoding on the sample visual content by utilizing the visual encoder to obtain a sample visual feature sequence corresponding to the sample visual content; and determine, based on at least the sample text feature sequence and the sample visual feature sequence, a predicted text feature sequence by utilizing the decoder, the predicted text feature sequence being decoded to obtain the first predicted keyword.
[0093] In some embodiments, during the training process of the machine learning model, the model parameters of the visual encoder remain unchanged, and the model parameters of the decoder change as the training process.
[0094] In some embodiments, the sample obtaining module 410 is further configured to: obtain sample input data and a sample keyword; generate, by using the generative model, sample thinking chain information based on the sample input data and the sample keyword, the generative model being configured to process a model input to generate corresponding thinking chain information; and determine a training sample based on the sample input data, the sample keyword, and the sample thinking chain information.
[0095] In some embodiments, the apparatus 400 further includes: a data input module configured to input target input data to the trained machine learning model; and a generation module configured to generate, by using the machine learning model, a target keyword of the target input data, or the target keyword and a thinking chain process from the target input data to the target keyword.
[0096] The units and / or modules included in the apparatus 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine-executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 400 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0097] It should be understood that one or more steps in the above methods can be performed by an appropriate electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices may, for example, include the electronic device 110 of Figure 1 .
[0098] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the disclosure can be implemented is shown. It should be understood that Figure 5 The electronic device 500 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to implement the electronic device 120 of Figure 1 or the apparatus 400 of Figure 4 .
[0099] As described above, the electronic device 500 can be used to implement the electronic device 110 of Figure 5As shown, the electronic device 500 is in the form of a general electronic device. Components of the electronic device 500 can include, but are not limited to, one or more processing units or processors 510, memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processor 510 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in the memory 520. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve parallel processing capability of the electronic device 500.
[0100] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any available media that is accessible by the electronic device 500 and includes both volatile and non-volatile media, removable and non-removable media. The memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable media and can include machine-readable media such as a flash drive, a magnetic disk drive, or any other media that can be used to store information and / or data and that can be accessed by the electronic device 500.
[0101] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 disk drives for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0102] The communication unit 540 enables communication with other electronic devices over communication media. Additionally, functionality of the components of the electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. As such, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in a networked environment.
[0103] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown), such as storage devices, display devices, etc., one or more devices that enable a user to interact with the electronic device 500, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 500 to communicate with one or more other electronic devices, as desired via the communication unit 540. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0104] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0105] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0106] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0107] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0108] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The machine-readable medium can be a single medium, or multiple media, of the same or different type. The computer program product can be one or more computer program components embodied in medium and / or transmission signals. The computer program product can have one or more computer program components embodied in medium and / or transmission signals.
[0109] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although the implementations of the disclosure have been described with regard to one or more implementations, it will be recognized that a variety of modifications and changes can be made to these implementations without departing from the broader spirit and scope of the implementations as set forth in the preceding disclosure. For example, certain aspects of the implementations can be performed using hardware, software, and / or firmware, or any combination thereof. The above-described implementations should therefore be regarded as merely illustrative, and not as narrowing the scope of the disclosure, which is defined by the appended claims and their equivalents.
Claims
1. A method for model training, comprising: Obtaining a training sample, wherein the training sample includes sample input data, sample keywords, and sample thought chain information, wherein the sample thought chain information indicates a reasoning process from the sample input data to the sample keywords; Determining, using a machine learning model, a first predicted keyword corresponding to the training sample based at least on the sample input data, the machine learning model being configured to process a model input to generate keywords and / or thought chain information corresponding to the model input; determining a first loss based on a difference between the first predicted keyword and the sample keyword; determining a target loss to be used for the training sample based on a comparison between the first loss and a threshold loss, wherein the target loss is selected from the first loss or the second loss, and the second loss is based on a difference between predicted thought chain information corresponding to the training sample and the sample thought chain information, and a difference between a second predicted keyword corresponding to the training sample and the sample keyword; as well as The machine learning model is trained based at least on a target loss to be used for the training samples.
2. The method of claim 1 , wherein determining a target loss to be used for the training sample based on a comparison between the first loss and a threshold loss comprises: In response to determining that the first loss reaches the threshold loss, determining the first loss as the target loss to be used for the training sample; as well as In response to determining that the first loss does not reach the threshold loss, the second loss is determined as the target loss to be used for the training sample.
3. The method of claim 1 , wherein the first predicted keyword is determined by: Using the machine learning model, generating the first predicted keyword based on the sample input data and the sample keyword, And the second predicted keyword and the predicted thought chain information are determined as follows: The machine learning model is used to generate the second predicted keywords and the predicted thought chain information based on the sample input data, the sample keywords and the sample thought chain information.
4. The method of claim 1 , wherein the sample input data comprises sample text and sample visual content, the machine learning model comprises a text encoder, a visual encoder, and a decoder, and wherein determining, using the machine learning model, at least based on the sample input data, a first predicted keyword corresponding to the training sample comprises: Encoding the sample text using the text encoder to obtain a sample text feature sequence corresponding to the sample text; Encoding the sample visual content using the visual encoder to obtain a sample visual feature sequence corresponding to the sample visual content; as well as The decoder is used to determine a predicted text feature sequence based at least on the sample text feature sequence and the sample visual feature sequence, and the predicted text feature sequence is decoded to obtain the first predicted keyword.
5. The method according to claim 4, wherein during the training process of the machine learning model, the model parameters of the visual encoder remain unchanged, and the model parameters of the decoder change with the training process.
6. The method according to claim 1, wherein obtaining the training sample comprises: Obtaining sample input data and the sample keywords; generating the sample thought chain information based on the sample input data and the sample keywords using a generative model, wherein the generative model is configured to process the model input to generate corresponding thought chain information; as well as The training sample is determined based on the sample input data, the sample keywords and the sample thought chain information.
7. The method according to claim 1, further comprising: Inputting target input data into the trained machine learning model; as well as Utilizing the machine learning model, a target keyword of the target input data, or the target keyword and a thought chain process from the target input data to the target keyword is generated.
8. A device for model training, comprising: a sample acquisition module configured to obtain a training sample, wherein the training sample includes sample input data, sample keywords, and sample thought chain information, wherein the sample thought chain information indicates a reasoning process from the sample input data to the sample keywords; a keyword determination module, configured to determine a first predicted keyword corresponding to the training sample based on at least the sample input data using a machine learning model, wherein the machine learning model is configured to process a model input to generate keywords and / or thought chain information corresponding to the model input; a first loss determining module configured to determine a first loss based on a difference between the first predicted keyword and the sample keyword; a target loss determination module configured to determine a target loss to be used for the training sample based on a comparison between the first loss and a threshold loss, wherein the target loss is selected from the first loss or the second loss, and the second loss is based on a difference between predicted thought chain information corresponding to the training sample and the sample thought chain information, and a difference between a second predicted keyword corresponding to the training sample and the sample keyword; as well as A model training module is configured to train the machine learning model based at least on a target loss to be used for the training samples.
9. An electronic device comprising: at least one processor; as well as At least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 7 when executed by the at least one processor.
10. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 7.
11. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 7.