Method and apparatus for generating multimodal classification model
By generating a multimodal classification model and using a multimodal fusion network to train multimodal data such as images and text, the problem of insufficient classification accuracy of multimodal data in the existing technology is solved, and more efficient multimodal data understanding and classification is achieved.
Patent Information
- Application Number
- CN202110394335.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-04-13
AI Technical Summary
The prior art is difficult to effectively process multimodal data, such as image and text data, resulting in insufficient understanding and classification accuracy of complex data.
A multimodal classification model generation method is proposed. By obtaining the preset sample set and a pre-established multimodal fusion network, the multimodal fusion network is trained to generate a multimodal classification model. The method includes inputting subsamples of different modalities into corresponding classification models, extracting eigenvectors and threshold vectors, and performing fusion training through the modal fusion module.
Through the training of multimodal fusion network, mutual complementation and interpretation can be achieved between multimodal data, improving the accuracy of multimodal target classification.
Smart Images

Figure CN113762321B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, specifically to the field of artificial intelligence technologies, and particularly to a method and apparatus for generating a multi-modal classification model, a method and apparatus for multi-modal target classification, an electronic device, a computer-readable medium, and a computer program product. Background Art
[0002] Data information has multiple modalities, such as images, texts, videos, audios, etc.; due to the large differences in different types of algorithms, fields, principles, application scopes, etc., most traditional models separately process data of only one modality. However, in reality, many data exist in the form of two or more modalities simultaneously. For example, the products in e-commerce exist in the forms of images and texts, and the data of the two forms are in a relationship of mutual explanation and complementation. Losing any one piece of information may lead to a deviation in the understanding of the product. Summary of the Invention
[0003] Embodiments of the present disclosure provide a method and apparatus for generating a multi-modal classification model, an electronic device, a computer-readable medium, and a computer program product.
[0004] In a first aspect, an embodiment of the present disclosure provides a method for generating a multi-modal classification model, the method including: obtaining a preset sample set, the sample set including at least one sample, and the sample including at least two sub-samples of different modalities; obtaining a pre-established multi-modal fusion network, the multi-modal fusion network including: a threshold module, a modality fusion module, and at least two classification models, each classification model classifying data of different modalities; performing the following training steps: selecting a sample from the sample set; respectively inputting the sub-samples of different modalities of the sample into the classification models corresponding to the respective modalities to obtain feature vectors output by each classification model, extracting a threshold vector of all the feature vectors through the threshold module, and inputting all the feature vectors and the threshold vector into the modality fusion module, and in response to determining that the multi-modal fusion network satisfies the training completion condition, taking the multi-modal fusion network as the multi-modal classification model.
[0005] In some embodiments, the training completion condition of the multi-modal fusion network includes: sequentially training each classification model, the threshold module, and the modality fusion module in the multi-modal fusion network; after each classification model, the threshold module, and the modality fusion module are all trained, simultaneously training all the modules in the multi-modal fusion network until the multi-modal fusion network is trained.
[0006] In some embodiments, training each classification model, threshold module, and modality fusion module in the above-mentioned multi-modal fusion network in sequence includes: cutting off the gradient backpropagation of the threshold module, and for each classification model in the multi-modal fusion network, setting the threshold vector of the threshold module to a constant value corresponding to this classification model; training this classification model to obtain the trained classification model; after all classification models are trained, turning on the gradient backpropagation of the threshold module, fixing the parameters of all trained classification models, and training the threshold module and the modality fusion module to obtain the trained threshold module and modality fusion module.
[0007] In some embodiments, the above-mentioned classification models include: an image classification model, a text classification model; for each classification model in the multi-modal fusion network, setting the threshold vector of the threshold module to a constant value corresponding to this classification model; training this classification model to obtain the trained classification model, including: setting the threshold vector assigned by the threshold module to the image classification model as a vector of all ones with a set dimension, setting the threshold vector assigned by the threshold model to the text classification model as a vector of all zeros with a set dimension; training the image classification model to obtain the trained image classification model; setting the threshold vector assigned by the threshold module to the text classification model as a vector of all ones with a set dimension, setting the threshold vector assigned by the threshold model to the image classification model as a vector of all zeros with a set dimension; training the text classification model to obtain the trained text classification model.
[0008] In some embodiments, the above-mentioned modality fusion module fuses all feature vectors and threshold vectors using the following formula:
[0009]
[0010] where i takes an integer value between 0 and 511, i is the index of a 512-dimensional vector, represents the i-th value of the feature vector of the image classification model, represents the i-th value of the feature vector of the text classification model, g i represents the i-th value of the threshold vector of the threshold module, represents the i-th value of the vector after fusion by the modality fusion module.
[0011] In some embodiments, the above-mentioned threshold module includes: two first threshold sub-modules connected in series and a second threshold sub-module; the first threshold sub-module includes: a fully connected layer, a batch normalization layer, and a first activation layer; the second threshold sub-module includes: a fully connected layer and a second activation layer, and the activation function of the first activation layer is different from the activation function of the second activation layer.
[0012] Second aspect, embodiments of the present disclosure provide a multi-modal target classification method, which includes: obtaining a target to be classified, where the target has at least two different types of modal data; inputting the target into a multi-modal classification model generated by using the method described in any implementation manner of the first aspect, and obtaining a classification result output by the multi-modal classification model.
[0013] Third aspect, embodiments of the present disclosure provide a multi-modal classification model generation device, which includes: a sample acquisition unit configured to obtain a preset sample set, where the sample set includes at least one sample, and the sample includes at least two different modal sub-samples; a network acquisition unit configured to obtain a pre-established multi-modal fusion network, where the multi-modal fusion network includes: a threshold module, a modal fusion module, and at least two classification models, and each classification model classifies different modal data; a selection unit configured to select a sample from the sample set; an input unit configured to respectively input the different modal sub-samples of the sample into the classification models corresponding to each modality, and obtain feature vectors output by each classification model; an extraction unit configured to extract threshold vectors of all the feature vectors through the threshold module; a fusion unit configured to input all the feature vectors and the threshold vectors into the modal fusion module; an output unit configured to, in response to determining that the multi-modal fusion network meets the training completion condition, use the multi-modal fusion network as the multi-modal classification model.
[0014] In some embodiments, the above output unit includes: a single-training sub-unit configured to sequentially train each classification model, the threshold module, and the modal fusion module in the multi-modal fusion network; a combined-training sub-unit configured to, after each classification model, the threshold module, and the modal fusion module are all trained, train all the modules in the multi-modal fusion network simultaneously until the multi-modal fusion network is trained.
[0015] In some embodiments, the above single-training sub-unit is further configured to cut off the gradient backpropagation of the threshold module, set the threshold vector of the threshold module to a constant value corresponding to the classification model for each classification model in the multi-modal fusion network; train the classification model to obtain the trained classification model; after all the classification models are trained, turn on the gradient backpropagation of the threshold module, fix the parameters of all the trained classification models, and train the threshold module and the modal fusion module to obtain the trained threshold module and modal fusion module.
[0016] In some embodiments, the classification model includes: an image classification model and a text classification model; the single-training subunit is further configured to: set the threshold vector assigned by the threshold module to the image classification model as a vector of all ones with a set dimension, and set the threshold vector assigned by the threshold model to the text classification model as a vector of all zeros with a set dimension; train the image classification model to obtain a trained image classification model; set the threshold vector assigned by the threshold module to the text classification model as a vector of all ones with a set dimension, and set the threshold vector assigned by the threshold model to the image classification model as a vector of all zeros with a set dimension; train the text classification model to obtain a trained text classification model.
[0017] In some embodiments, the above-mentioned modality fusion module fuses all feature vectors and threshold vectors using the following formula:
[0018]
[0019] where i takes an integer value between 0 and 511, and i is the index of a 512-dimensional vector. represents the i-th value of the feature vector of the image classification model. represents the i-th value of the feature vector of the text classification model, and g i represents the i-th value of the threshold vector of the threshold module. represents the i-th value of the vector after fusion by the modality fusion module.
[0020] In some embodiments, the above-mentioned threshold module includes: two first threshold sub-modules connected in series and a second threshold sub-module; the first threshold sub-module includes: a fully connected layer, a batch normalization layer, and a first activation layer; the second threshold sub-module includes: a fully connected layer and a second activation layer, and the activation function of the first activation layer is different from that of the second activation layer.
[0021] Fourthly, an embodiment of the present disclosure provides a multi-modal target classification device, which includes: an acquisition unit configured to acquire a target to be classified, where the target has at least two different types of modality data; a classification unit configured to input the target into the multi-modal classification model generated by the method described in any implementation manner of the first aspect to obtain a classification result output by the multi-modal classification model.
[0022] Fifthly, an embodiment of the present disclosure provides an electronic device, which includes: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0023] Sixthly, an embodiment of the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0024] Seventhly, an embodiment of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0025] The multi-modal classification model generation method and device provided by the embodiment of the present disclosure first obtain a preset sample set; secondly, obtain a pre-established multi-modal fusion network; thirdly, select a sample from the sample set; fourthly, use the sample in the sample set to train the multi-modal fusion network to obtain a multi-modal classification model, and the training step is: input the sub-samples of different modalities of the sample into the classification models corresponding to each modality respectively to obtain the feature vectors output by each classification model; extract the threshold vectors of all the feature vectors through a threshold module, and input all the feature vectors and the threshold vectors into a modality fusion module, and in response to determining that the multi-modal fusion network meets the training completion condition, use the multi-modal fusion network as the multi-modal classification model. Thus, the provided multi-modal classification model can enhance the mutual complementarity and mutual explanation between multi-modal data during the training process, and improve the accuracy of multi-modal target classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present disclosure will become more obvious:
[0027] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;
[0028] Figure 2 is a flowchart of an embodiment of the multi-modal classification model generation method according to the present disclosure;
[0029] Figure 3 is a schematic structural diagram of a multi-modal fusion network provided by the present disclosure;
[0030] Figure 4 is another schematic structural diagram of the multi-modal fusion network provided by the present disclosure;
[0031] Figure 5 is a flowchart of an embodiment of the multi-modal target classification method according to the present disclosure;
[0032] Figure 6 is a schematic structural diagram of an embodiment of the multi-modal classification model generation device provided by the present disclosure;
[0033] Figure 7 is a schematic structural diagram of an embodiment of a multi-modal object classification device according to the present disclosure;
[0034] Figure 8 is a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners
[0035] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that, for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0036] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and embodiments.
[0037] Figure 1 An exemplary system architecture 100 to which the multi-modal classification model generation method, multi-modal classification model generation device, multi-modal object classification method, and multi-modal object classification device of the embodiments of the present disclosure can be applied is shown.
[0038] As Figure 1 shown, the system architecture 100 may include terminals 101, 102, a network 103, a database server 104, and a server 105. The network 103 is used to provide a medium for communication links between the terminals 101, 102, the database server 104, and the server 105. The network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0039] The user 110 may use the terminals 101, 102 to interact with the server 105 through the network 103 to receive or send messages, etc. Various client applications may be installed on the terminals 101, 102, such as model training applications, image conversion applications, shopping applications, payment applications, web browsers, and instant messaging tools, etc.
[0040] The terminals 101, 102 here may be hardware or software. When the terminals 101, 102 are hardware, they may be various electronic devices with a display screen, including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, and desktop computers, etc. When the terminals 101, 102 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (for example, used to provide distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0041] When terminals 101 and 102 are hardware, a target information collection device can also be installed thereon. The target information collection device can be various devices capable of implementing the function of collecting multimodal information (images, texts, videos, audio), such as cameras, sensors, etc. User 110 can use the target information collection devices on terminals 101 and 102 to collect multimodal targets to be classified.
[0042] The database server 104 can be a database server providing various services. For example, a sample set can be stored in the database server. The sample set contains a large number of samples. Among them, the samples can include at least two subsamples of different modalities. A subsample is a type of data, and subsamples of different modalities are embodied as data of different modalities. For example, a subsample is one of text data, image data, video data, audio data, etc. In this way, user 110 can also select samples from the sample set stored in database server 104 through terminals 101 and 102.
[0043] The server 105 can also be a server providing various services, such as a background server that supports various applications displayed on terminals 101 and 102. The background server can use the samples in the sample set sent by terminals 101 and 102 to train the initial model and can send the training results (such as the generated multimodal classification model) to terminals 101 and 102. In this way, the user can apply the generated multimodal classification model for target classification.
[0044] Here, the database server 104 and the server 105 can also be hardware or software. When they are hardware, they can be implemented as a distributed server cluster composed of multiple servers or as a single server. When they are software, they can be implemented as multiple software or software modules (for example, used to provide distributed services) or as a single software or software module. No specific limitation is made here.
[0045] It should be noted that the multimodal classification model generation method or the multimodal target classification method provided by the embodiments of the present disclosure is generally executed by the server 105. Correspondingly, the multimodal classification model generation device or the multimodal target classification device is generally also set in the server 105.
[0046] It should be pointed out that in the case where the server 105 can implement the related functions of the database server 104, the database server 104 may not be provided in the system architecture 100.
[0047] It should be understood that Figure 1 the numbers of terminals, networks, database servers, and servers in are merely illustrative. According to the implementation requirements, there can be any number of terminals, networks, database servers, and servers.
[0048] As Figure 2 , Flow 200 of an embodiment of a method for generating a multimodal classification model according to the present disclosure is shown. The method for generating a multimodal classification model includes the following steps:
[0049] Step 201, obtaining a preset sample set.
[0050] In this embodiment, the execution subject of the method for generating a multimodal classification model (such as Figure 1 the server 105 shown) can obtain the sample set through various means. For example, the execution subject can obtain the existing sample set stored therein from a database server (such as Figure 1 the database server 104 shown) through a wired connection method or a wireless connection method. For another example, a user can collect samples through a terminal (such as Figure 1 the terminals 101, 102 shown). In this way, the execution subject can receive the samples collected by the terminal and store these samples locally, thereby generating a sample set.
[0051] Here, the sample set may include at least one sample. Among them, a sample may include at least two subsamples of different modalities. A subsample is a type of data, and a modality refers to the way in which something occurs or exists. Subsamples of different modalities are embodied as data of different modalities. For example, a subsample is one of data such as text data, image data, video data, audio data, etc. A sample is a combination of two or more subsamples of different modalities. For example, a sample includes text data and image data. For another example, a sample includes text data, image data, and audio data.
[0052] The modality types of the subsamples in the sample are not limited herein and can be any combination. The specific implementation process can refer to the sample selection step of step 203.
[0053] Step 202, obtaining a pre-established multimodal fusion network.
[0054] In this embodiment, the structure of obtaining a pre-established multimodal fusion network is as Figure 3 shown. The multimodal fusion network includes: a threshold module, a modality fusion module, and at least two classification models, and each classification model classifies different modality data. In Figure 3 , the at least two classification models include a first classification model... an Nth classification model, where N≥2. The first classification model to the Nth classification model are used to classify N types of modality data to obtain their respective classification results.
[0055] In this embodiment, the data characteristics of different modalities are different, and the structures of each classification model in at least two classification models are different. For example, text data is a type of sequential data with a logical relationship between words before and after. The classification model can adopt a recurrent neural network. For example, by using a recurrent neural network, "reasoning" is performed on the text information to obtain relevant semantic information. Image data has completely different characteristics from text data. When processing image data, a convolutional neural network can be adopted. For example, multiple convolutional kernels are used to capture local information of the image data, and then semantic information of the picture is extracted layer by layer.
[0056] It should be noted that at least two classification models can be classification models trained by using the above-mentioned preset sample set. For example, the classification model classifies an image (unknown type) of a product and outputs a feature vector belonging to types such as clothes, shoes, and socks; optionally, at least two classification models can also be models that have not been trained with samples.
[0057] In this embodiment, the threshold module is used to obtain a threshold vector of all classification models according to the output vectors of all classification models. The threshold vector can be used to represent the weight of the feature vector output by the classification model among all feature vectors. The output vectors of all classification models are concatenated and then enter the threshold module to obtain a threshold vector. This threshold vector is input into the modality fusion module to fuse the output vectors of all classification models.
[0058] In some optional implementation manners of this embodiment, the threshold module includes: two first threshold sub-modules connected in series and a second threshold sub-module; the first threshold sub-module includes: a fully connected layer, a batch normalization layer, and a first activation layer; the second threshold sub-module includes: a fully connected layer and a second activation layer, and the activation function of the first activation layer is different from the activation function of the second activation layer. For example, the activation function of the first activation layer adopts the relu function, and the second activation layer uses the sigmoid function. Here, if the number of modalities is greater than 2, the structure of the fully connected layer in the second threshold sub-module is adjusted so that a vector with an output dimension of (512, N) is obtained, where N is the number of modalities, and softmax is used as the activation function of the second activation layer. Then, one column of this vector is multiplied by the corresponding position of the feature vector (with a length of 512) of a corresponding modality to obtain a vector with a length of 512. Then, all N vectors with a length of 512 are added to obtain a vector with a length of 512, and then subsequent calculations are performed. This threshold module improves the reliability of the threshold vector.
[0059] In this embodiment, the modality fusion module fuses the output vectors of all classification modules based on the threshold vectors of each classification model to obtain a fusion result, and this fusion result is the final classification result after fusing the inputs of all classification models. As Figure 3As shown, the component has multiple inputs, including the feature vectors output by the first to the Nth classification models, and the threshold vectors output by the threshold model corresponding to each classification model (i.e., one threshold vector corresponds to one classification model). It should be noted that the modality fusion module needs to fuse vectors with the same dimension. Therefore, there is a dimension conversion component in the modality fusion module. For example, Figure 4 in Figure 4 , the FC+BN+Relu between the image classification model and the modality fusion component is a kind of dimension conversion component, which includes a fully connected layer, a normalization layer, and a Relu activation function layer (the activation layer uses Relu as the activation function).
[0060] Step 203: Select samples from the sample set.
[0061] In this embodiment, the execution subject can select samples from the sample set obtained in step 201 and perform the training steps of steps 203 to 207. Among them, the selection method and the number of selected samples are not limited in this disclosure. For example, at least one sample can be randomly selected, or samples with better clarity (i.e., higher pixels) and clear text content can be selected from them.
[0062] According to the input requirements of the classification model, each sample can include at least two subsamples of different modalities, and each subsample has a corresponding classification model. For example, in this embodiment, at least two classification models include an image classification model and a text classification model. Then the selected samples include an image modality subsample and a text modality subsample. The image modality subsample is input into the image classification model, and the text modality subsample is input into the text classification model.
[0063] Step 204: Input the subsamples of different modalities of the sample into the classification models corresponding to each modality respectively, and obtain the feature vectors output by each classification model.
[0064] In this embodiment, the classification model is used to classify data of each modality. For example, the classification model includes an image classification model. An image containing a coat, trousers, and socks is input into the image classification model, and the classification model obtains the labels for classifying the coat, trousers, and socks in the image.
[0065] Step 205: Extract the threshold vectors of all the feature vectors through the threshold module.
[0066] In this embodiment, the feature vectors output by at least two classification models are very different from each other. When fusing at least two feature vectors, in order to achieve mutual complementarity and reference of various information, the learnable threshold module determines how to perform information fusion according to the characteristics of the actual various information, and the threshold vectors output by the threshold module are the specific values at which at least two feature vectors can be fused.
[0067] Step 206: Input all the feature vectors and threshold vectors into the modality fusion module.
[0068] In this embodiment, the threshold module, the modality fusion module, and at least two classification models can be jointly trained. During the joint training, while the threshold module updates the threshold vectors of each classification model according to the gradient backpropagation algorithm, the threshold vectors of at least two classification models in the modality fusion module can also be updated simultaneously.
[0069] Step 207: In response to determining that the multi-modal fusion network meets the training completion condition, use the multi-modal fusion network as the multi-modal classification model.
[0070] In this embodiment, after the samples are input into the multi-modal fusion network, a loss function is calculated. The loss function is used to represent the gap between the predicted value and the known answer. When training the multi-modal fusion network, by continuously changing all the parameters in the multi-modal fusion network, the loss function is continuously reduced, thereby training a multi-modal classification model with higher accuracy.
[0071] Specifically, the gradient backpropagation method can be used to calculate the loss value of the loss function. The gradient backpropagation method is a method that combines backpropagation (abbreviated as BP) with the gradient descent method. This method calculates the gradient of the loss function for all the weights in the network. This gradient is fed back to the gradient descent method to update the weights to minimize the loss function.
[0072] In this embodiment, optionally, the training completion condition includes at least one of the following: the number of training iterations of the multi-modal fusion network reaches a predetermined iteration threshold, and the loss value of the loss function is less than a predetermined loss value threshold. For example, the training iteration reaches 5,000 times and the loss value is less than 0.05. After training is completed, use the multi-modal fusion network as the multi-modal classification model. Setting the training completion condition can accelerate the model convergence speed.
[0073] If the multi-modal fusion network does not meet the training completion condition, adjust the relevant parameters in the multi-modal fusion network to make the loss value converge. Based on the adjusted multi-modal fusion network, continue to execute steps 203-207 until the multi-modal fusion network meets the training completion condition.
[0074] The method for generating a multimodal classification model provided by the embodiments of the present disclosure first obtains a preset sample set; secondly, obtains a pre-established multimodal fusion network; thirdly, selects samples from the sample set; and then trains the multimodal fusion network with the samples in the sample set to obtain a multimodal classification model. The training step is as follows: input the sub-samples of different modalities of the sample into the classification models corresponding to each modality respectively to obtain the feature vectors output by each classification model; extract the threshold vectors of all the feature vectors through a threshold module, and input all the feature vectors and the threshold vectors into a modality fusion module. In response to determining that the multimodal fusion network meets the training completion condition, the multimodal fusion network is used as the multimodal classification model. Thus, the provided multimodal classification model can make the multimodal targets complement and explain each other during the training process, improving the accuracy of multimodal target classification.
[0075] Since there are significant differences between multimodal data, the training processes of classification models corresponding to different modality data are also different. Therefore, the training process of the multimodal fusion network needs to have the function of separate individual training. That is to say, after the multimodal fusion network is built, the individual training of each classification model in at least two classification models needs to be carried out separately first. After each classification model is fully trained, the fusion training of all modules in the multimodal fusion network is carried out.
[0076] In some optional implementation manners of this embodiment, the training completion condition of the multimodal fusion network includes: training each classification model, threshold module, and modality fusion module in the multimodal fusion network in sequence; after all the classification models, threshold module, and modality fusion module are trained, training all the modules in the multimodal fusion network simultaneously until the multimodal fusion network is trained.
[0077] In this optional implementation manner, the individual training of each module can be controlled by a switch. When all the modules in the multimodal fusion network are jointly trained, the gradient backpropagation of all the models and all the modules can also be turned on simultaneously through the switch for training. The specific control process of the switch refers to Figure 4 the detailed description of the illustrated embodiment.
[0078] Specifically, training each classification model, threshold module, and modality fusion module in the multimodal fusion network in sequence means: first, training each classification module separately; secondly, after all the classification models are trained, training the threshold module and modality fusion module. During the training process of the threshold module and modality fusion module, all the classification models need to be frozen. After this stage of training is completed, all the classification models are unfrozen, and then the overall multimodal fusion network is jointly trained.
[0079] The following takes the classification models as an image classification model and a text classification model to introduce the training process of the multimodal fusion network. The training completion conditions of the multimodal fusion network training are divided into four stages:
[0080] ① Turn the switch to the image classification model and train the image classification model;
[0081] ② Turn the switch to the text classification model and train the text classification model;
[0082] ③ Turn on all the switches, freeze the image classification model and the text classification model (that is, fix the parameters of the trained image classification model and text classification model), and then train the threshold module and the modality fusion module;
[0083] ④ Unfreeze the image classification model and the text classification model (the parameters of the trained image classification model and text classification model are no longer fixed values, but are updated with the update of the parameters of the multimodal fusion network), and start training all the modules and models of the multimodal fusion network.
[0084] In this optional implementation, by independently training each classification model, threshold module, and modality fusion module in the multimodal fusion network in sequence, and then jointly training the entire multimodal fusion network, the various classification models, threshold module, and modality fusion module can achieve the effect of sufficient training, improving the reliability of the multimodal fusion network training.
[0085] In some optional implementations of this embodiment, sequentially training each classification model, threshold module, and modality fusion module in the multimodal fusion network includes: cutting off the gradient backpropagation of the threshold module, setting the threshold vector of the threshold module to a constant value corresponding to the classification model for each classification model in the multimodal fusion network; training the classification model to obtain the trained classification model; after all the classification models are trained, turning on the gradient backpropagation of the threshold module, fixing the parameters of all the trained classification models, and training the threshold module and the modality fusion module to obtain the trained threshold module and modality fusion module.
[0086] The above-mentioned cutting off and turning on the gradient backpropagation of the threshold module can be implemented by a switch, and the specific control process of the switch can be found in the Figure 4 detailed description of the illustrated embodiment.
[0087] In this optional implementation, by cutting off the gradient back propagation of the threshold module, the threshold vector of the threshold module is set for each classification model in the network to complete the independent training of each classification model; and after all the classification models are trained, the gradient back propagation of the threshold module is turned on, and the threshold module and the modal fusion module are trained to obtain the trained threshold module and modal fusion module, thereby simply and conveniently achieving full training of each module in the multimodal fusion network.
[0088] In some optional implementations of this embodiment, such as Figure 4 As shown, the at least two classification models include: an image classification model and a text classification model.
[0089] Specifically, the model structure of the image classification model and the text classification model can be selected according to the scenario. For example, tensorflow2.0 is used as the basic architecture of the image classification model, and transfer learning is used to fine-tune the original weights. On this basis, the image data enhancement effect is improved, and the model output is used as a feature vector, that is, the output of the model is used as a picture encoding process.
[0090] Text data is sequence data, and a recurrent neural network or a Bert (Bidirectional Encoder Representations from Transformers) type network based on attention can be used. For example, the text classification model uses the Bi-LSTM model, which is suitable for bidirectional reasoning. This allows for bidirectional reasoning of text modal data and better mining of the relationship between targets (referring to targets with text modal data, such as titles) and categories.
[0091] In practice, when the target is the title of a product, the title is usually a stack of multiple words. The text classification model segments the title and then partially shuffles the words. It should be noted that the local shuffling here is a method of text data enhancement. For example, the words ranked 3-6 are shuffled internally. The reason for this is that although the title is disordered, it contains a certain overall order preference. For example, the product brand word usually appears at the front of the title, and the product size and color are more likely to appear at the end.
[0092] In this optional implementation, the input of the modality fusion module has three parts. One is the feature vector output by the image classification model, which is converted into a dimension conversion component ( Figure 4The FC + BN + Relu) between the Chinese image classification model and the modality fusion component yields a 512-dimensional feature vector. The second is the feature vector of the text classification model, which passes through a maintenance transformation component ([ Figure 4 The FC + BN + Tanh) between the Chinese text classification model and the modality fusion component yields a 512-dimensional feature vector. The third is the threshold vector output by the gate component, which is a 512-dimensional vector. The modality fusion component fuses the feature vectors of the two classification models, namely the image classification model and the text classification model, based on the value of the threshold vector. The fused vector (512-dimensional) then passes through a fully connected layer, a normalization layer, and a Relu activation layer ([ Figure 4 The FC + BN + Relu) connected to the output end of the modality fusion component in [ Figure 4 The FC) in [ becomes a vector with the number of dimensions equal to the number of categories for classification, obtaining the final output of the modality fusion module.
[0093] In this alternative implementation, for each classification model in the multi-modal fusion network, the threshold vector of the threshold module is set to a constant value corresponding to the classification model; the classification model is trained to obtain the trained classification model, including: setting the threshold vector assigned by the threshold module to the image classification model as a vector of all ones with a set dimension, setting the threshold vector assigned by the threshold model to the text classification model as a vector of all zeros with a set dimension; training the image classification model to obtain the trained image classification model; setting the threshold vector assigned by the threshold module to the text classification model as a vector of all ones with a set dimension, setting the threshold vector assigned by the threshold model to the image classification model as a vector of all zeros with a set dimension; training the text classification model to obtain the trained text classification model.
[0094] In this alternative implementation, the above control process can be adopted Figure 4The switch control shown is, in practice, that the switch directly adds an assignable interface to the threshold vector and sets the opening or closing of the gradient backpropagation channel. For example, when only training the image classification model part, the threshold vector is set as a 512-dimensional all-ones vector and set as a constant value, cutting off the propagation of the gradient to the threshold module part. At this time, a threshold vector of all-ones is directly assigned to the image classification model (the output value of the modality fusion component is a threshold vector of all-ones), and a threshold vector of all-zeros is assigned to the text classification model (the output value of the modality fusion component is a threshold vector of all-zeros). Conversely, if the threshold vector is set as a 512-dimensional all-zeros vector, only the text classification model is trained (equivalent to switching the switch to the text classification model). When the threshold vector is not assigned, both classification models are open. At this time, the outputs of the two classification models are fused through the threshold vector output by the threshold module, and the threshold vector at this time is obtained through the learning of the threshold module. During the training process, the gradient backpropagation will reach all parts of the modality fusion network and update all weights.
[0095] In this embodiment, Figure 4 the switch in is used to control the training process of the multi-modal fusion network and can also be used to observe the trained multi-modal classification model after the model training is completed. For example, after the training is completed, turn off one of the classification models and observe the difference between the output result and the output of the model before the joint training, etc.
[0096] In this alternative implementation, by separately assigning an all-ones threshold vector to the image classification model or the text classification model, the image classification model or the text classification model is independently trained, achieving the purpose of separately training each classification model in at least two classification models at different times and improving the reliability of the separate training of the classification models.
[0097] In some alternative implementations of this embodiment, the modality fusion component in the modality fusion module can fuse all feature vectors and threshold vectors using the following formula (1):
[0098]
[0099] where i takes an integer between 0 and 511, and i is the index of the 512-dimensional vector. represents the i-th value of the feature vector of the image classification model. represents the i-th value of the feature vector of the text classification model, and g i represents the i-th value of the threshold vector of the threshold module. represents the i-th value of the vector after the modality fusion module fuses.
[0100] In this optional implementation, the above formula is used to simply and conveniently realize the fusion of feature vectors output by at least two different classification models, thereby improving the effect of fusion after classification of different modal data.
[0101] See also Figure 5 , which shows a process 500 of an embodiment of a multimodal target classification method provided by the present disclosure. The multimodal target classification method may include the following steps:
[0102] Step 501: Acquire a target to be classified, where the target has at least two different modal data.
[0103] In this embodiment, the execution subject of the multimodal target classification method (for example Figure 1 The server 105 shown in FIG. 104 may obtain the target to be classified in a variety of ways. For example, the execution subject may obtain the target from a database server (e.g. Figure 1 The target to be classified is obtained from the database server 104 shown in the figure, and the target has at least two different modal data. For another example, the execution subject can also receive a terminal (for example Figure 1 The targets to be classified are collected by the terminals 101, 102) or other devices shown.
[0104] In this embodiment, the target to be classified can be a product information including at least two different modal data, and the product information can specifically include information such as images, text, video, and voice, wherein the image can be a color image and / or a grayscale image, etc., and the format of the image is not limited in this disclosure. For example, products on the Internet are targets that exist in the form of images and texts. The pictures of the products are image data, and the product titles, product details, filled-in attributes, etc. exist in the form of text data.
[0105] Step 502: input the target into the multimodal classification model to obtain the classification result output by the multimodal classification model.
[0106] In this embodiment, the execution subject can input the target obtained in step 501 into the multimodal classification model to generate a classification result. The multimodal classification model trained and generated in steps 201-207 can determine the type of the input target. For example, the input target includes image data (picture of the clothes) and text data (title, details, attributes, etc. of the clothes), and the multimodal classification model outputs "T-shirt".
[0107] In this embodiment, the multimodal classification model can be adopted as described above. Figure 2 The specific generation process can be found in Figure 2 The relevant description of the embodiments will not be repeated here.
[0108] It should be noted that the multimodal target classification method of this embodiment can be used to test the multimodal classification model generated by the above embodiments. Then, the multimodal classification model can be continuously optimized according to the conversion result. This method can also be a practical application method of the multimodal classification model generated by the above embodiments. Using the multimodal classification model generated by the above embodiments to perform target classification helps to improve the accuracy of target classification.
[0109] When processing multimodal products (for example, the product includes a picture and a title. The title describes the product in the picture and provides information such as size and color, while the picture intuitively gives the appearance of the product) (for example, obtaining product attributes, classifying products, making personalized recommendations, etc.), it is necessary to have a fine-grained understanding and recognition of the product. If only one of the modal data of the product is used to identify the product, the amount of information is not enough. For example, a piece of clothing may contain attributes such as size, color, style, collar type, version, material, whether it is a suit or a set of several pieces, and a laptop computer may contain attributes such as appearance, CPU model, memory, hard disk, etc. The recognition results obtained by only using one modal data of pictures or text are not accurate, and the recognition results are not accurate when simply superimposing two modal data without fusion learning.
[0110] In this embodiment, the multimodal product is input into the multimodal classification model generated by the multimodal classification model generation method, and the classification result obtained after the multimodal classification model accurately identifies and classifies the multimodal product can be obtained. For example, the multimodal product is a picture of a T-shirt and a title of a T-shirt, wherein the keyword "dress" is added to the title of the T-shirt in order to attract the traffic of the dress. If the classification result obtained by a unimodal data classification model (such as a text classification model or an image classification model) is: T-shirt (type) 60% (confidence). The classification result of the product by the multimodal classification model is: T-shirt (type) 90% (confidence). Therefore, the multimodal classification model can accurately understand the product information on the basis of the fusion of multimodal data information, and accurately classify the product on the basis of accurately understanding the product information.
[0111] Continue to see Figure 6 , as a response to the above Figure 2 The present disclosure provides an embodiment of a multimodal classification model generation device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0112] like Figure 6As shown in the figure, the multi-modal classification model generation device 600 of this embodiment may include: a sample acquisition unit 601, a network acquisition unit 602, a selection unit 603, an input unit 604, an extraction unit 605, a fusion unit 606, and an output unit 607. Among them, the sample acquisition unit 601 is configured to acquire a preset sample set, the sample set includes at least one sample, and the sample includes at least two sub-samples of different modalities. The network acquisition unit 602 is configured to acquire a pre-established multi-modal fusion network, and the multi-modal fusion network includes: a threshold module, a modality fusion module, and at least two classification models, and each classification model classifies different modality data. The selection unit 603 is configured to select a sample from the sample set. The input unit 604 is configured to respectively input the sub-samples of different modalities of the sample into the classification models corresponding to each modality to obtain the feature vectors output by each classification model. The extraction unit 605 is configured to extract the threshold vectors of all the feature vectors through the threshold module. The fusion unit 606 is configured to input all the feature vectors and threshold vectors into the modality fusion module. The output unit 607 is configured to, in response to determining that the multi-modal fusion network meets the training completion condition, use the multi-modal fusion network as the multi-modal classification model.
[0113] In some optional implementation manners of this embodiment, the above output unit 607 includes: a single-training sub-unit (not shown in the figure), a combined-training sub-unit (not shown in the figure). Among them, the single-training sub-unit is configured to sequentially train each classification model, the threshold module, and the modality fusion module in the multi-modal fusion network; the combined-training sub-unit is configured to, after all the classification models, the threshold module, and the modality fusion module are trained, simultaneously train all the modules in the multi-modal fusion network until the multi-modal fusion network is trained.
[0114] In some optional implementation manners of this embodiment, the above single-training sub-unit is further configured to cut off the gradient backpropagation of the threshold module, set the threshold vector of the threshold module to a constant value corresponding to the classification model for each classification model in the multi-modal fusion network; train the classification model to obtain the trained classification model; after all the classification models are trained, turn on the gradient backpropagation of the threshold module, fix the parameters of all the trained classification models, and train the threshold module and the modality fusion module to obtain the trained threshold module and modality fusion module.
[0115] In some alternative implementation manners of this embodiment, the classification model includes: an image classification model and a text classification model; the single-training subunit is further configured to: set the threshold vector assigned by the threshold module to the image classification model as a vector of all ones with a set dimension, and set the threshold vector assigned by the threshold model to the text classification model as a vector of all zeros with a set dimension; train the image classification model to obtain a trained image classification model; set the threshold vector assigned by the threshold module to the text classification model as a vector of all ones with a set dimension, and set the threshold vector assigned by the threshold model to the image classification model as a vector of all zeros with a set dimension; train the text classification model to obtain a trained text classification model.
[0116] In some alternative implementation manners of this embodiment, the modality fusion module fuses all the feature vectors and threshold vectors by using the following formula:
[0117]
[0118] where i takes an integer between 0 and 511, and i is the index of a 512-dimensional vector. represents the i-th value of the feature vector of the image classification model. represents the i-th value of the feature vector of the text classification model, and g i represents the i-th value of the threshold vector of the threshold module. represents the i-th value of the vector after fusion by the modality fusion module.
[0119] In some alternative implementation manners of this embodiment, the above-mentioned threshold module includes: two first threshold sub-modules connected in series and a second threshold sub-module; the first threshold sub-module includes: a fully connected layer, a batch normalization layer, and a first activation layer; the second threshold sub-module includes: a fully connected layer and a second activation layer, and the activation function of the first activation layer is different from the activation function of the second activation layer.
[0120] It can be understood that the units described in the apparatus 600 correspond to the respective steps in the method described with reference to Figure 2 Therefore, the operations, features, and beneficial effects described above for the method also apply to the apparatus 600 and the units included therein, and will not be elaborated herein.
[0121] Continue to refer to Figure 7 , as an implementation of the method shown above Figure 5 The present disclosure provides an embodiment of a multi-modal target classification apparatus. This apparatus embodiment corresponds to the method embodiment shown in Figure 5 and this apparatus can be specifically applied to various electronic devices.
[0122] As shown in Figure 7As shown in the figure, the multimodal target classification device 700 of this embodiment may include: an acquisition unit 701 configured to acquire a target to be classified, where the target has at least two different types of modal data. A classification unit 702 configured to input the target into the multimodal classification model generated by the method described in the above Figure 2 or Figure 5 embodiment to obtain the classification result output by the multimodal classification model.
[0123] It can be understood that the various units described in the device 700 correspond to the respective steps in the method described with reference to Figure 5 Accordingly, the operations, features, and beneficial effects described above for the method also apply to the device 700 and the units included therein, and will not be elaborated herein.
[0124] Next, with reference to Figure 8 , a schematic structural diagram of an electronic device 800 suitable for implementing the embodiments of the present disclosure is shown.
[0125] As Figure 8 shown, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0126] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 shows an electronic device 800 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 8 Each block shown in
[0127] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-described functions defined in the method of the embodiment of the present disclosure are performed.
[0128] It should be noted that the computer-readable medium in the embodiment of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiment of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiment of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0129] The above computer-readable medium may be included in the above server; or it may exist independently without being assembled into the server. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the server, the server is caused to: obtain a preset sample set, the sample set including at least one sample, and the sample including at least two sub-samples of different modalities; obtain a pre-established multi-modal fusion network, the multi-modal fusion network including: a threshold module, a modality fusion module, and at least two classification models, each classification model classifying data of different modalities; perform the following training steps: select a sample from the sample set; respectively input the sub-samples of different modalities of the sample into the classification models corresponding to the respective modalities to obtain the feature vectors output by each classification model, extract the threshold vectors of all the feature vectors through the threshold module, and input all the feature vectors and the threshold vectors into the modality fusion module, and in response to determining that the multi-modal fusion network meets the training completion condition, use the multi-modal fusion network as a multi-modal classification model.
[0130] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0132] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as a processor including a sample acquisition unit, a network acquisition unit, a selection unit, an input unit, an extraction unit, a fusion unit, and an output unit. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself. For example, the sample extraction unit can also be described as a unit "configured to acquire a preset sample set".
[0133] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.
Claims
1. A method for generating a multi-modal classification model, the method comprising: Obtain a pre-set sample set, where the sample set includes at least one sample, and the sample includes at least two sub-samples of different modalities; wherein, the at least two sub-samples of different modalities include at least two of text data, image data, video data, and audio data; Obtain a pre-established multi-modal fusion network, where the multi-modal fusion network includes: a threshold module, a modality fusion module, and at least two classification models, and each classification model classifies data of different modalities; Execute the following training steps: Select a sample from the sample set; input the sub-samples of different modalities of the sample into the classification models corresponding to each modality respectively to obtain the feature vectors output by each classification model, extract the threshold vector of all the feature vectors through the threshold module, the threshold vector is used to represent the weight of the feature vector output by the classification model among all the feature vectors, and input all the feature vectors and the threshold vector into the modality fusion module, where the modality fusion module includes a dimension conversion component for converting the dimensions of all the feature vectors to the same dimension as the threshold vector, and in response to determining that the multi-modal fusion network meets the training completion condition, use the multi-modal fusion network as a multi-modal classification model.
2. The method according to claim 1, wherein, The training completion conditions of the multi-modal fusion network include: Train each classification model, the threshold module, and the modality fusion module in the multi-modal fusion network in sequence; After each classification model, the threshold module, and the modality fusion module are all trained, train all the modules in the multi-modal fusion network simultaneously until the multi-modal fusion network is trained.
3. The method according to claim 2, wherein, The training each classification model, the threshold module, and the modality fusion module in the multi-modal fusion network in sequence includes: Cut off the gradient backpropagation of the threshold module, for each classification model in the multi-modal fusion network, set the threshold vector of the threshold module to a constant value corresponding to the classification model; train the classification model to obtain the trained classification model; After all the classification models are trained, turn on the gradient backpropagation of the threshold module, fix the parameters of all the trained classification models, and train the threshold module and the modality fusion module to obtain the trained threshold module and modality fusion module.
4. The method according to claim 3, wherein, The classification models include: an image classification model, a text classification model; the for each classification model in the multi-modal fusion network, set the threshold vector of the threshold module to a constant value corresponding to the classification model; train the classification model to obtain the trained classification model, includes: Set the threshold vector assigned by the threshold module to the image classification model as a vector of all ones with a set dimension, and set the threshold vector assigned by the threshold module to the text classification model as a vector of all zeros with a set dimension; Train the image classification model to obtain the trained image classification model; Set the threshold vector assigned by the threshold module to the text classification model as a vector of all ones with a set dimension, and set the threshold vector assigned by the threshold module to the image classification model as a vector of all zeros with a set dimension; Train the text classification model to obtain a trained text classification model.
5. The method according to claim 3, wherein, The modality fusion module fuses all the feature vectors and the threshold vectors using the following formula: Among them, i takes integer values between 0 and 511, and i is the index of a 512-dimensional vector. represents the i-th value of the feature vector of the image classification model. represents the i-th value of the feature vector of the text classification model, g i represents the i-th value of the threshold vector of the threshold module. represents the i-th value of the vector after fusion by the modality fusion module.
6. The method according to any one of claims 1-5, wherein, The threshold module includes: Two first threshold sub-modules and one second threshold sub-module connected in series; The first threshold sub-module includes: a fully connected layer, a batch normalization layer, and a first activation layer; The second threshold sub-module includes: a fully connected layer and a second activation layer, and the activation function of the first activation layer is different from the activation function of the second activation layer.
7. A multi-modal target classification method, the method comprising: Obtain a target to be classified, where the target has at least two different modality data; Input the target into the multi-modal classification model generated by the method according to any one of claims 1-6 to obtain the classification result output by the multi-modal classification model.
8. A multi-modal classification model generation device, the device comprising: A sample acquisition unit configured to acquire a preset sample set, where the sample set includes at least one sample, and the sample includes at least two sub-samples of different modalities; wherein, the at least two sub-samples of different modalities include at least two of text data, image data, video data, and audio data; A network acquisition unit configured to acquire a pre-established multi-modal fusion network, where the multi-modal fusion network includes: a threshold module, a modality fusion module, and at least two classification models that respectively classify different modality data; A selection unit configured to select a sample from the sample set; An input unit configured to respectively input the sub-samples of different modalities of the sample into the classification models corresponding to each modality to obtain the feature vectors output by each classification model; An extraction unit configured to extract the threshold vectors of all the feature vectors through the threshold module, where the threshold vectors are used to represent the weights of the feature vectors output by the classification models among all the feature vectors; A fusion unit configured to input all the feature vectors and the threshold vectors into the modality fusion module, where the modality fusion module includes a dimension conversion component for converting the dimensions of all the feature vectors to be the same as the dimensions of the threshold vectors; An output unit configured to, in response to determining that the multi-modal fusion network meets the training completion condition, use the multi-modal fusion network as a multi-modal classification model.
9. A multimodal target classification device, the device comprising: An acquisition unit configured to acquire a target to be classified, where the target has at least two different modality data; A classification unit configured to input the target into the multi-modal classification model generated by the method according to any one of claims 1-6 to obtain the classification result output by the multi-modal classification model.
10. An electronic device, comprising: One or more processors; A storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
11. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method according to any one of claims 1-7.
12. A computer program product comprising a computer program, the computer program implementing the method according to any one of claims 1-7 when executed by a processor.
Citation Information
Patent Citations
Multi-mode fusion commodity classification system
CN106909946A
Cross-modal information retrieval method and device and storage medium
CN109816039A