Pre-training method and device for multi-modal feature extraction model, equipment and medium

By applying the recognition difficulty coefficient and sample balance coefficient in the pre-training of the multimodal feature extraction model and optimizing the network parameters, the problem of the difficulty of balancing sample recognition difficulty during training of the multimodal feature extraction model is solved, and the performance of the model in downstream tasks is improved.

CN120611338APending Publication Date: 2025-09-09BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410256843.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing pre-training schemes for multimodal feature extraction models find it difficult to balance difficult-to-identify and easy-to-identify samples in multimodal information during training, resulting in insufficient quality of extracted multimodal fusion features and, consequently, insignificant performance improvement in downstream tasks.

Method used

By applying the recognition difficulty coefficient to the predicted logarithmic term of the cross entropy loss function, the weight of difficult-to-recognize samples is increased, and the sample balancing coefficient and sample discarding coefficient are applied to the predicted logarithmic term to balance the influence of positive and negative samples and optimize the network parameters of the visual encoder, text encoder and fusion layer.

Benefits of technology

It improves the pre-training effect of the multimodal feature extraction model and enhances the performance of the model in downstream tasks, especially in image description and visual text retrieval tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611338A_ABST
    Figure CN120611338A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a pre-training method and device for a multi-modal feature extraction model, a medium and equipment. A specific embodiment comprises the following steps: determining a plurality of first fusion features with classification labels as positive samples and a plurality of second fusion features with classification labels as negative samples according to visual information and text information in a plurality of pieces of multi-modal information, a visual encoder, a text encoder and a fusion layer; inputting each first fusion feature and each second fusion feature into a classifier, calculating a corrected cross entropy loss according to an output predicted value, and applying an identification difficulty coefficient which is positively correlated to the predicted value and approaches to a classification boundary value on a predicted logarithm item; and network parameters of the visual encoder, the text encoder and the fusion layer are updated with the goal of decreasing the loss of the cross entropy. Through the method, the characterization capability of the multi-modal fusion features extracted by the pre-trained model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of software testing technology, and in particular to a pre-training method, apparatus, device, and medium for a multimodal feature extraction model. Background Art

[0002] In recent years, model pre-training has been rapidly developed and widely used in single-modal processing fields such as computer vision and natural language processing. By pre-training the feature extraction model, the performance of downstream tasks can be improved, such as improving the performance of image recognition tasks. Furthermore, the model pre-training method has been gradually applied to multimodal processing tasks. By pre-training the multimodal feature extraction model, the performance of multimodal downstream tasks can also be improved, such as image description (Image Captioning) tasks, vision-language retrieval (Vision-Language Retrieval) tasks and other tasks. However, the existing multimodal feature extraction model pre-training scheme still has the problem that the quality of the multimodal fusion features extracted by the pre-trained model is insufficient, and when the multimodal fusion features are used for downstream tasks, the improvement of task performance is not significant enough.

[0003] Therefore, a new pre-training method for multimodal feature extraction models is needed. Summary of the Invention

[0004] The embodiments of the present disclosure describe a pre-training method, apparatus, device, and medium for a multimodal feature extraction model.

[0005] According to a first aspect, a pre-training method for a multimodal feature extraction model is provided, comprising:

[0006] Determining, based on the visual information and textual information in the plurality of multimodal information, as well as the visual encoder, the textual encoder, and the fusion layer, a plurality of first fused features whose classification labels are positive samples, and a plurality of second fused features whose classification labels are negative samples;

[0007] Each first fusion feature and the second fusion feature are input into the first classifier to obtain a prediction value of whether each first fusion feature and the second fusion feature is a positive sample; a prediction logarithmic term is determined according to the logarithm of the prediction value, and a recognition difficulty coefficient is applied to the prediction logarithmic term to obtain a modified cross entropy loss, wherein the recognition difficulty coefficient is positively correlated with the prediction value approaching the classification cutoff value; with the goal of reducing the cross entropy loss, the network parameters of the visual encoder, the text encoder and the fusion layer are updated.

[0008] According to a second aspect, a testing device for a multimodal feature extraction model is provided, comprising:

[0009] A feature fusion unit is configured to determine, based on the visual information and textual information in the plurality of multimodal information, the visual encoder, the textual encoder, and the fusion layer, a plurality of first fused features for which the classification labels are positive samples, and a plurality of second fused features for which the classification labels are negative samples;

[0010] The training unit is configured to input each first fusion feature and the second fusion feature into the first classifier to obtain a prediction value for whether each first fusion feature and the second fusion feature is a positive sample; determine a prediction logarithmic term according to the logarithm of the prediction value, apply a recognition difficulty coefficient to the prediction logarithmic term, and obtain a modified cross entropy loss, wherein the recognition difficulty coefficient is positively correlated with the prediction value approaching the classification cutoff value; and update the network parameters of the visual encoder, the text encoder, and the fusion layer with the goal of reducing the cross entropy loss.

[0011] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.

[0012] According to a fourth aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.

[0013] According to an embodiment of the present disclosure, a method, apparatus, device and medium for a multimodal feature extraction model are provided. First, based on the visual information and textual information in a plurality of multimodal information, as well as the visual encoder, the text encoder and the fusion layer, a plurality of first fusion features with classification labels as positive samples and a plurality of second fusion features with classification labels as negative samples can be determined. Then, each of the first fusion features and the second fusion features is input into the first classifier to obtain a prediction value for whether each of the first fusion features and the second fusion features is a positive sample; a prediction logarithm term is determined according to the logarithm of the prediction value, and a recognition difficulty coefficient is applied to the prediction logarithm term to obtain a modified cross entropy loss, wherein the recognition difficulty coefficient is positively correlated with the predicted value approaching the classification cutoff value; with the goal of minimizing the cross entropy loss, the network parameters of the visual encoder, the text encoder and the fusion layer are updated. Through this method, the pre-training effect of the multimodal feature extraction model can be enhanced, and the characterization ability of the multimodal fusion features extracted by the pre-trained model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A schematic diagram of a pre-training method for a multimodal feature extraction model according to an embodiment of the present disclosure is shown;

[0015] Figure 2A schematic diagram of a process for pre-training a multimodal feature extraction model according to an embodiment of the present disclosure is shown;

[0016] Figure 3 A schematic diagram of a pre-training loss function of a multimodal feature extraction model according to an embodiment of the present disclosure is shown;

[0017] Figure 4 A schematic diagram of a pre-training method for a multimodal feature extraction model according to another embodiment of the present disclosure is shown;

[0018] Figure 5 A schematic block diagram of a pre-training device for a multimodal feature extraction model according to an embodiment of the present disclosure is shown;

[0019] Figure 6 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown;

[0020] Figure 7 A schematic diagram of the structure of a storage medium suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0021] The technical solutions provided in this specification are further described in detail below in conjunction with the accompanying drawings and embodiments. It will be understood that the specific embodiments described herein are merely for explaining the relevant inventions and are not intended to limit the inventions. It should also be noted that, for ease of description, only the portions relevant to the relevant inventions are shown in the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features therein may be combined with each other.

[0022] In the description of the implementations of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one / an implementation" or "the implementation" should be understood as "at least one / an implementation." The term "some implementations" should be understood as "at least some implementations." Other explicit and implicit definitions may be included below.

[0023] As mentioned above, model pre-training has been rapidly developed and widely used in single-modal processing fields such as computer vision and natural language processing. By pre-training the feature extraction model, the performance of downstream tasks can be improved, such as improving the performance of image recognition tasks. Furthermore, the model pre-training method has gradually been applied to multimodal processing tasks. By pre-training the multimodal feature extraction model, the performance of multimodal downstream tasks can also be improved, such as image captioning tasks and vision-language retrieval tasks. However, the existing pre-training schemes for multimodal feature extraction models still have the problem that the quality of the multimodal fusion features extracted by the pre-trained model is insufficient, and then when the multimodal fusion features are used for downstream tasks, the improvement in task performance is not significant enough. Specifically, the reasons are as follows: On the one hand, the existing pre-training schemes for multimodal processing usually do not have or find it difficult to balance the difficult-to-identify samples and easy-to-identify samples in the multimodal information used for training during training, resulting in insufficient representational ability of the multimodal fusion features extracted by the pre-trained model. Secondly, because it is usually easier to obtain a large number of negative samples, existing multimodal processing pre-training schemes usually use more negative samples than positive samples, which can cause an imbalance in the impact of positive and negative samples on network parameter updates during training, leading to insufficient representation capabilities of the multimodal fusion features extracted by the pre-trained model. Thirdly, existing multimodal processing pre-training schemes often use large samples of multimodal information, and the quality of these samples is often uneven. For example, some of the modal information is missing or damaged, and the quality of the multimodal information is poor. After being used for pre-training, this can reduce the training effect of the pre-training, resulting in insufficient representation capabilities of the multimodal fusion features extracted by the pre-trained model.

[0024] In order to solve the above technical problems, the embodiments of the present disclosure provide a pre-training method for a multimodal feature extraction model. Figure 1 FIG. 1 is a schematic diagram showing a pre-training method of a multimodal feature extraction model according to an embodiment of the present disclosure. Figure 1As shown, in some embodiments, for example, multiple multimodal information can be obtained, such as multimodal information A and multimodal information B. The visual information (for example, video information) and text information included in each multimodal information are input into the video encoder and the text encoder respectively to obtain the video features and text features corresponding to each visual information and text information. For example, the visual information AV1 included in the multimodal information A and the video information BV2 included in the multimodal information B are input into the visual encoder to obtain the visual feature V1 (corresponding to the visual information AV1) and the video feature V2 (corresponding to the visual information AV2). The text information AT1 included in the multimodal information A and the text information BT2 included in the multimodal information B are input into the text encoder to obtain the text feature T1 (corresponding to the visual information AT1) and the visual feature T2 (corresponding to the visual information AT2). Then, the visual features and text features derived from the same multimodal information are input into the fusion layer to obtain the fusion features of the positive sample. The visual features and text features derived from different multimodal information are input into the fusion layer to obtain the fusion features of the positive sample. For example, the visual feature V1 and the text feature T1 are input into the fusion layer to obtain the fusion feature C1 of the positive sample. The visual feature V1 and the text feature T2 are input into the fusion layer to obtain the fusion feature C2 of the negative sample. The visual feature V2 and the text feature T1 are input into the fusion layer to obtain the fusion feature C3 of the negative sample. Next, the fusion features of the positive sample and the negative sample are input into the classifier (for example, the first classifier), and the modified cross entropy loss is calculated based on the classification result of whether the input fusion feature is a positive sample. The difference between the loss function for calculating the modified cross entropy loss (or called the modified cross entropy function) and the loss function for calculating the conventional cross entropy loss (or called the conventional cross entropy function) is that a recognition difficulty coefficient is applied to the predicted logarithmic term contained in the conventional cross entropy function, and the recognition difficulty coefficient is positively correlated to the predicted value approaching the classification cutoff value. Thereafter, the network parameters of the visual encoder, text encoder and fusion layer are updated based on the obtained modified cross entropy loss.

[0025] The advantages of this method are as follows: First, the conventional cross-entropy loss in existing pre-training schemes is typically calculated based on a predicted logarithmic term determined based on a logarithmic function of the predicted value. However, based solely on the calculation of the predicted logarithmic term, it is difficult to distinguish between samples with greater recognition difficulty (or simply hard-to-identify samples) and samples with lower recognition difficulty (or simply easy-to-identify samples) in the input samples. Consequently, when updating network parameters based on the cross-entropy loss, the difference between hard-to-identify and easy-to-identify samples is not obvious. In contrast, this method provides a method for determining the difficulty of sample recognition by determining a recognition difficulty coefficient that is positively correlated with the predicted value approaching the classification cutoff. Furthermore, by applying the sample recognition difficulty to the predicted logarithmic term of the cross-entropy function, the predicted logarithmic term corresponding to the hard-to-identify sample is weighted. This improves the pre-training effect of the hard-to-identify samples on the network parameters during training, specifically improving the representational power of the features extracted by the pre-trained visual encoder, text encoder, and fusion layer. Second, in some embodiments, a sample balancing coefficient can also be applied to the predicted logarithmic term to increase the weight of positive samples. Under the premise that the number of positive samples is significantly lower than the number of negative samples, increasing the weight of positive samples can balance the impact of positive and negative samples on network parameter updates without reducing the number of negative samples used for training. This further improves the feature extraction capability of the pre-trained model while maintaining the generalization performance of the pre-trained model. Thirdly, in some embodiments, a sample discard coefficient can also be applied to the predicted logarithmic term to determine whether to abandon the update of network parameters based on samples with poor information quality, thereby further improving the effect of pre-training.

[0026] The detailed process of this method is further described below.

[0027] Figure 2 FIG. 1 is a flow chart showing a method for pre-training a multimodal feature extraction model according to an embodiment of the present disclosure. Figure 2 As shown, the method comprises at least the following steps:

[0028] Step S201, determining a plurality of first fused features with classification labels of positive samples and a plurality of second fused features with classification labels of negative samples based on the visual information and textual information in the plurality of multimodal information, as well as the visual encoder, the textual encoder, and the fusion layer;

[0029] Step S203: Input each first fusion feature and the second fusion feature into the first classifier to obtain a prediction value of whether each first fusion feature and the second fusion feature is a positive sample; determine the prediction logarithm term according to the logarithm of the prediction value, apply the recognition difficulty coefficient to the prediction logarithm term, and obtain a modified cross entropy loss, wherein the recognition difficulty coefficient is positively correlated with the prediction value approaching the classification cutoff value; with the goal of minimizing the cross entropy loss, update the network parameters of the visual encoder, the text encoder, and the fusion layer.

[0030] First, in step S201, based on the visual information and textual information in the multiple multimodal information, as well as the visual encoder, the textual encoder and the fusion layer, determine a plurality of first fusion features with the classification label as positive samples and a plurality of second fusion features with the classification label as negative samples. In one embodiment, specifically, the visual information in the multiple multimodal information can be input into the visual encoder respectively to obtain a plurality of visual features, and the textual information in the multiple multimodal information can be input into the text encoder respectively to obtain a plurality of textual features. In this step, the visual information and textual information in the multiple multimodal information can be input into the visual encoder and the textual encoder respectively to obtain a plurality of visual features and a plurality of textual features corresponding to the multiple multimodal information respectively. In other words, a plurality of visual features corresponding to the visual information in the multiple multimodal information and a plurality of textual features corresponding to the textual information in the multiple multimodal information are obtained. In Figure 2 In the example shown, for example, visual information AV1 and video information BV2 can be input into a visual encoder to obtain visual features V1 and video features V2. Text information AT1 and text information BT2 can be input into a text encoder to obtain text features T1 and visual features T2.

[0031] In different embodiments, the specific types or network structures of the visual encoder and the text encoder may be different, and this specification does not limit this. In one embodiment, the visual encoder may be, for example, a convolutional neural network (CNN) or a recurrent neural network (RNN). In another embodiment, the text encoder may be, for example, a fully connected neural network (FCN), a recurrent neural network (RNN), a convolutional neural network (CNN), and a recursive neural network (RNN).

[0032] In different embodiments, the specific methods for obtaining multiple multimodal information may vary. In one embodiment, multiple multimodal information may be obtained from a preset multimodal information set. In a specific embodiment, before obtaining the multiple multimodal information from the multimodal sample set, defective multimodal information is pre-eliminated from the multimodal information set, where the visual information or textual information in the defective multimodal information is damaged or missing. In this way, the quality of the multimodal information used for pre-training can be improved, thereby enhancing the pre-training effect.

[0033] In different embodiments, the specific content of the visual information and text information included in the multimodal information may be different, and this specification does not limit this. In one embodiment, the visual information may be video information or image information, and the text information may be one or more of the title, subtitles, or description information of the video information or image information. In a specific embodiment, the video information may be a video frame included in the video. In different specific embodiments, the specific method of extracting the video frame may be different. In a specific embodiment, for example, the video may be divided into a predetermined number of segments, and at least one video frame may be extracted from each segment.

[0034] After obtaining multiple text features, the visual features and text features of the same multimodal information can be input into the fusion layer to obtain multiple first fusion features with classification labels as positive samples; the visual features and text features of different multimodal information can be input into the fusion layer to obtain multiple second fusion features with classification labels as negative samples. Figure 2 In the example shown, the visual feature V1 and the text feature T1 can be input into the fusion layer to obtain the fused feature C1 (first fused feature) whose classification label is a positive sample. The visual feature V1 and the text feature T2 can be input into the fusion layer to obtain the fused feature C2 (second fused feature) whose classification label is a negative sample. The visual feature V2 and the text feature T1 can be input into the fusion layer to obtain the fused feature C3 (second fused feature) whose classification label is a negative sample.

[0035] Then, in step S203, each of the first and second fused features is input into the first classifier to obtain a predicted value for whether each of the first and second fused features is a positive sample. A predicted logarithmic term is determined based on the logarithm of the predicted value, and a recognition difficulty coefficient is applied to the predicted logarithmic term to obtain a modified cross-entropy loss. The recognition difficulty coefficient can be positively correlated with the predicted value approaching the classification cutoff value. In other words, the recognition difficulty coefficient can be negatively correlated with the predicted value approaching the upper or lower classification threshold.

[0036] In different embodiments, the specific confirmation method of the recognition difficulty coefficient may be different. In one embodiment, the first prediction logarithmic term can be determined based on whether the first fusion feature is the first prediction value of the positive sample, and the first recognition difficulty coefficient is applied to the first prediction logarithmic term to obtain a modified first cross entropy loss, and the first recognition difficulty coefficient is the first hyperparameter power of the difference between the upper threshold and the first prediction value; the second prediction logarithmic term is determined based on whether the second fusion feature is the second prediction value of the positive sample, and the second recognition difficulty coefficient is applied to the second prediction logarithmic term to obtain a modified second cross entropy loss, and the recognition difficulty coefficient is the first hyperparameter power of the second prediction value. In different specific embodiments, the value of the first hyperparameter may be different. In a specific embodiment, for example, it can be 2.

[0037] In another embodiment, an identification difficulty coefficient and a sample balance coefficient determined based on the predicted value may also be applied to the prediction logarithmic term. In a specific embodiment, the first prediction logarithmic term can be determined based on whether the first fused feature is the first predicted value of a positive sample, and the identification difficulty coefficient and the first sample balance coefficient can be applied to the first prediction logarithmic term to obtain a modified first cross entropy loss, wherein the first sample balance coefficient is a second hyperparameter. Furthermore, the second prediction logarithmic term can be determined based on whether the second fused feature is the second predicted value of a positive sample, and the identification difficulty coefficient and the second sample balance coefficient can be applied to the second prediction logarithmic term to obtain a modified second cross entropy loss, wherein the second sample balance coefficient is the difference between one and the second hyperparameter. In different specific embodiments, the value of the second hyperparameter may also be different. In a specific embodiment, for example, it can be 0.75. By applying the sample balance coefficient, the weight of positive samples in training can be increased, thereby improving the feature extraction capability of the model when the number of negative samples is greater than that of positive samples.

[0038] In another embodiment, a recognition difficulty coefficient, a sample balance coefficient, and a sample discard coefficient may also be applied to the prediction logarithmic term. The sample discard coefficient is used to indicate whether to discard the fused features targeted by the cross entropy loss. In order to reduce the interference of poor quality positive samples on the pre-training effect, in a specific embodiment, the prediction logarithmic term applied by the sample discard coefficient can be the prediction logarithmic term corresponding to the fused features whose classification label is a positive sample.

[0039] In a specific embodiment, Figure 3 As shown, the corrected cross entropy loss can be expressed by the following formula:

[0040]

[0041] Where L is the modified cross entropy loss, p is the predicted value, -log(p) is the predicted logarithmic term for positive samples, -log(1-p) is the predicted logarithmic term for negative samples, (1-p) γ is the recognition difficulty coefficient imposed on the predicted logarithmic term for the positive sample, p γ is the recognition difficulty coefficient applied to the predicted logarithmic item for negative samples, γ is the first hyperparameter, the sample balance coefficient applied to the predicted logarithmic item for positive samples is α, the sample balance coefficient applied to the predicted logarithmic item for negative samples is (1-α), and α is the second hyperparameter.

[0042] After determining the corrected cross-entropy loss, the network parameters of the visual encoder, text encoder, and fusion layer can be updated with the goal of minimizing the cross-entropy loss. The visual encoder, text encoder, and fusion layer can constitute a multimodal feature extraction model. Specifically, in one embodiment, the network parameters of the visual encoder, text encoder, and fusion layer can be updated, for example, using a BP (back propagation) algorithm.

[0043] Thereafter, in one embodiment, the updated visual encoder, text encoder, fusion layer, and second classifier may be combined to perform the classification task corresponding to the second classifier. Thus, the pre-trained multimodal feature extraction model is utilized to improve the classification performance of the downstream classification task (i.e., the classification task corresponding to the second classifier). In different specific embodiments, the specific classification task corresponding to the second classifier may be different, and this specification does not limit this. In a specific embodiment, for example, it may be a multimodal classification task of vision and text. Figure 4 FIG2 shows a schematic diagram of a pre-training method for a multimodal feature extraction model according to another embodiment of the present disclosure. Figure 4 As shown, for example, the visual information AV3 included in the multimodal information C can be input into the visual encoder to obtain visual features V3. The text information AT3 included in the multimodal information C can be input into the text encoder to obtain text features T3. Then, the visual features V3 and text features T3 are input into the fusion layer to obtain the fused features C4 of the positive sample. Next, the fused features C4 are input into, for example, a second classifier based on the classification results of the multimodal information C.

[0044] Figure 5 A schematic block diagram of a device for a pre-training method of a multimodal feature extraction model according to an embodiment of the present disclosure is shown. The device is used to perform the following steps: Figure 2 As shown in the method. Figure 5 As shown, the apparatus 500 includes:

[0045] The feature fusion unit 501 is configured to determine, based on the visual information and textual information in the plurality of multimodal information, the visual encoder, the textual encoder, and the fusion layer, a plurality of first fused features for which the classification label is a positive sample, and a plurality of second fused features for which the classification label is a negative sample;

[0046] The training unit 502 is configured to input each first fusion feature and the second fusion feature into the first classifier to obtain a prediction value for whether each first fusion feature and the second fusion feature is a positive sample; determine a prediction logarithm term based on the logarithm of the prediction value, apply a recognition difficulty coefficient to the prediction logarithm term, and obtain a modified cross entropy loss, wherein the recognition difficulty coefficient is positively correlated to the prediction value approaching the classification cutoff value; update the network parameters of the visual encoder, the text encoder, and the fusion layer with the goal of reducing the cross entropy loss.

[0047] The embodiment of the present disclosure also provides an electronic device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the following is achieved: Figure 2 The method shown.

[0048] You can also refer to the following Figure 6 , which shows a structural diagram of an electronic device 600 suitable for implementing an embodiment of the present application. Figure 6 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0049] like Figure 6 As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601. The above-mentioned processing device 601 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, ROM 602 and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0050] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 7 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0051] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the pre-training method of a multimodal feature extraction model provided in the embodiment of the present application are executed.

[0052] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the following steps provided in the present disclosure: Figure 2 A pre-training method for a multimodal feature extraction model as shown in . Figure 7 A schematic diagram of a storage medium for implementing an embodiment of the present application. Figure 7As shown, the storage medium 700 may be a non-transitory computer-readable storage medium for storing non-transitory computer-executable instructions 701. When the non-transitory computer-executable instructions 701 are executed by the processor, a pre-training method for a multimodal feature extraction model provided in an embodiment of the present application may be implemented. For example, when the non-transitory computer-executable instructions 701 are executed by the processor, one or more steps in a pre-training method for a multimodal feature extraction model provided in an embodiment of the present application may be executed. For example, the storage medium 700 may be applied to the above-mentioned electronic device. For example, the storage medium 700 may include a memory in the electronic device. For the description of the storage medium 700, reference may be made to the description of the memory in the embodiment of the electronic device, and the repeated parts will not be repeated here. The specific functions and technical effects of the storage medium 700 may refer to the description of the pre-training method for a multimodal feature extraction model provided in an embodiment of the present application, and will not be repeated here.

[0053] It should be noted that the computer-readable medium of the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a memory card of a smart phone, a storage component of a tablet computer, a portable computer disk, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. In the embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0054] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the server, the one or more programs enable the electronic device to implement a pre-training method for a multimodal feature extraction model provided in an embodiment of the present application.

[0055] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0056] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions. The units involved in the embodiments described in the present disclosure may be implemented using software or hardware. The name of the unit does not, in some cases, constitute a limitation on the unit itself. The functions described above in this document may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0057] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences between the other embodiments. In particular, the storage medium and computing device embodiments are described briefly because they are generally similar to the method embodiments. For relevant portions, refer to the description of the method embodiments.

[0058] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned disclosed concepts. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the present disclosure to form a technical solution. In addition, although the operations are described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination.

[0059] The above specific implementation methods further describe in detail the purpose, technical solutions and beneficial effects of the embodiments of the present invention. Although the subject matter has been described in a language specific to structural features and / or method logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. It should be understood that the above are only specific implementation methods of the embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. A pre-training method for a multimodal feature extraction model, comprising: Determining, based on the visual information and textual information in the plurality of multimodal information, as well as the visual encoder, the textual encoder, and the fusion layer, a plurality of first fused features whose classification labels are positive samples, and a plurality of second fused features whose classification labels are negative samples; Input each first fusion feature and the second fusion feature into the first classifier to obtain a prediction value of whether each first fusion feature and the second fusion feature is a positive sample; determine the prediction logarithm term according to the logarithm of the prediction value, apply the recognition difficulty coefficient to the prediction logarithm term, and obtain the modified cross entropy loss, wherein, The recognition difficulty coefficient is positively correlated with the predicted value approaching the classification cutoff value; With the goal of minimizing the cross entropy loss, the network parameters of the visual encoder, text encoder, and fusion layer are updated.

2. The method according to claim 1, wherein Based on the visual information and text information in the multiple multimodal information, a plurality of first fused features with classification labels as positive samples and a plurality of second fused features with classification labels as negative samples are determined through a visual encoder, a text encoder, and a fusion layer, including: The visual information in multiple multimodal information is respectively input into the visual encoder to obtain multiple visual features, and the text information in the multiple multimodal information is respectively input into the text encoder to obtain multiple text features; the visual features and text features of the same multimodal information are input into the fusion layer to obtain multiple first fusion features with classification labels as positive samples; the visual features and text features of different multimodal information are input into the fusion layer to obtain multiple second fusion features with classification labels as negative samples.

3. The method according to claim 1, wherein The recognition difficulty coefficient is positively correlated with the predicted value approaching the classification cutoff value, including: The recognition difficulty coefficient is positively correlated with the predicted value approaching the classification cutoff value, and negatively correlated with the predicted value approaching the upper threshold or lower threshold of the classification; A predicted logarithmic term is determined according to the logarithm of the predicted value, and a recognition difficulty coefficient is applied to the predicted logarithmic term to obtain a modified cross entropy loss, including: Determining a first prediction logarithmic term based on whether the first fused feature is a first prediction value of a positive sample, applying a first recognition difficulty coefficient to the first prediction logarithmic term to obtain a modified first cross entropy loss, where the first recognition difficulty coefficient is a first hyperparameter power of a difference between the upper threshold and the first prediction value; A second prediction logarithmic term is determined based on whether the second fusion feature is a second prediction value of a positive sample, and a second recognition difficulty coefficient is applied to the second prediction logarithmic term to obtain a modified second cross entropy loss, where the recognition difficulty coefficient is a first hyperparameter power of the second prediction value.

4. The method according to claim 1, wherein Apply identification difficulty coefficients to the predicted logarithmic terms, including: Apply the recognition difficulty coefficient and sample balance coefficient to the prediction logarithm.

5. The method according to claim 4, wherein A predicted logarithmic term is determined according to the logarithm of the predicted value, and a recognition difficulty coefficient is applied to the predicted logarithmic term to obtain a modified cross entropy loss, including: Determining a first prediction logarithmic term based on whether the first fused feature is a first prediction value of a positive sample, applying a recognition difficulty coefficient and a first sample balance coefficient to the first prediction logarithmic term to obtain a modified first cross entropy loss, wherein the first sample balance coefficient is a second hyperparameter; A second prediction logarithmic term is determined based on whether the second fusion feature is a second prediction value of a positive sample, and a recognition difficulty coefficient and a second sample balance coefficient are applied to the second prediction logarithmic term to obtain a modified second cross entropy loss, wherein the second sample balance coefficient is the difference between one and a second hyperparameter.

6. The method according to claim 4, wherein: Apply the recognition difficulty coefficient and sample balance coefficient to the prediction logarithm, including: The recognition difficulty coefficient, the sample balance coefficient and the sample discarding coefficient are applied to the prediction logarithmic term, and the sample discarding coefficient is used to indicate whether to discard the prediction logarithmic term applied by the sample discarding coefficient.

7. The method according to claim 6, wherein: The predicted logarithmic term imposed by the sample discarding coefficient is the predicted logarithmic term corresponding to the fusion feature whose classification label is a positive sample.

8. The method according to claim 1, wherein The plurality of multimodal information are obtained from a preset multimodal information set; The method further comprises: Before acquiring the plurality of multimodal information from the multimodal sample set, defective multimodal information is pre-eliminat ed from the multimodal information set, wherein visual information or text information in the defective multimodal information is damaged or missing.

9. The method according to claim 1, further comprising: The updated visual encoder, text encoder, fusion layer, and second classifier are combined to perform a classification task corresponding to the second classifier.

10. The method according to claim 1, wherein The visual information is video information or image information, and the text information is one or more of the title, subtitle, or description information of the video information or image information.

11. A training device for a multimodal feature extraction model, comprising: a feature fusion unit configured to determine, based on the visual information and textual information in the plurality of multimodal information, the visual encoder, the textual encoder, and the fusion layer, a plurality of first fused features for which the classification labels are positive samples, and a plurality of second fused features for which the classification labels are negative samples; The training unit is configured to input each of the first fused features and the second fused features into a first classifier to obtain a prediction value of whether each of the first fused features and the second fused features is a positive sample; determine a prediction logarithm term based on the logarithm of the prediction value, and apply a recognition difficulty coefficient to the prediction logarithm term to obtain a modified cross entropy loss, wherein the recognition difficulty coefficient is positively correlated with the prediction value approaching a classification cutoff value; With the goal of minimizing the cross entropy loss, the network parameters of the visual encoder, text encoder, and fusion layer are updated.

12. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 10.

13. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Image classification model optimization method based on confusion samples

    CN115908927A

  • Multi-modal fusion named entity recognition method and device based on debiasing contrast learning, medium and product

    CN117010389A

  • Transformer fault diagnosis method and system based on knowledge constraint neural network

    CN117574264A