Classification method and device based on multi-modal data, equipment and medium

By using multimodal classification models in multimodal data classification tasks and combining technologies such as graph neural networks to effectively integrate text, visual and audio modal data, the problem of low classification accuracy of multimodal data is solved, and higher classification results are achieved.

CN120011893APending Publication Date: 2025-05-16XIAN JIAOTONG LIVERPOOL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510380675.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The effective fusion of multimodal data remains a key challenge, resulting in low accuracy of the classification results of the model in the multimodal data classification task.

Method used

By obtaining the data of the object to be classified in text, visual and audio modes and inputting it into a multimodal classification model based on sample training, data processing and model training are used for graph neural networks to improve the accuracy of classification results.

Benefits of technology

The accuracy of the classification results of multimodal classification models in multimodal data is improved, and by effectively fusion of information from different modal data, the model's understanding and classification ability of complex data is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011893A_ABST
    Figure CN120011893A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a classification method and device based on multi-modal data, equipment and a medium. The method comprises the steps of obtaining to-be-classified data of a to-be-classified object in a candidate mode; wherein the candidate modes comprise a text mode, a visual mode and an audio mode; inputting the to-be-classified data into the trained multi-modal classification model to obtain an object classification result; wherein the multi-modal classification model is obtained based on joint training of text sample data, visual sample data and audio sample data of the sample training object. And the accuracy of the classification result determined by the multi-modal classification model based on the multi-modal data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a classification method, device, equipment and medium based on multimodal data. Background Art

[0002] The multimodal classification task aims to improve the model's ability to understand and classify complex data by fusing information from different data modalities (such as text, images, sounds, etc.). This task is widely used in fields such as emotion recognition, medical diagnosis, and behavior recognition. However, due to the differences in feature distribution, representation, and information content of different modal data, the effective fusion of multimodal data remains a key challenge. Therefore, it is crucial to improve the accuracy of the model's classification results based on multimodal data. Summary of the invention

[0003] The present invention provides a classification method, device, equipment and medium based on multimodal data to improve the accuracy of the classification results of the model based on multimodal data.

[0004] According to one aspect of the present invention, a classification method based on multimodal data is provided, comprising:

[0005] Acquire data to be classified of the object to be classified in a candidate modality; wherein the candidate modality includes a text modality, a visual modality, and an audio modality;

[0006] Inputting the data to be classified into a trained multimodal classification model to obtain an object classification result;

[0007] The multimodal classification model is obtained by joint training based on text sample data, visual sample data and audio sample data of the sample training object.

[0008] According to another aspect of the present invention, there is provided a classification device based on multimodal data, comprising:

[0009] A module for acquiring data to be classified, used to acquire data to be classified of an object to be classified in a candidate mode; wherein the candidate mode includes a text mode, a visual mode and an audio mode;

[0010] A classification result determination module, used for inputting the data to be classified into a trained multimodal classification model to obtain an object classification result;

[0011] The multimodal classification model is obtained by joint training based on text sample data, visual sample data and audio sample data of the sample training object.

[0012] According to another aspect of the present invention, there is provided an electronic device, comprising:

[0013] one or more processors;

[0014] A memory for storing one or more programs;

[0015] When one or more programs are executed by one or more processors, the one or more processors can execute any one of the classification methods based on multimodal data provided by the embodiments of the present invention.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions, and the computer instructions are used to enable a processor to implement any classification method based on multimodal data provided by an embodiment of the present invention when executed.

[0017] The embodiment of the present invention provides a classification scheme based on multimodal data, by obtaining the data to be classified of the object to be classified in the candidate modality; wherein the candidate modality includes text modality, visual modality and audio modality; inputting the data to be classified into a trained multimodal classification model to obtain the object classification result; wherein the multimodal classification model is trained based on the text sample data, visual sample data and audio sample data of the sample training object. The above scheme trains the multimodal classification model according to the text sample data, visual sample data and audio sample data of the sample training object, so that the multimodal classification model can better process the data to be classified of the object to be classified in the candidate modality, thereby improving the accuracy of the object classification result determined by the trained multimodal classification model, that is, improving the accuracy of the classification result determined by the multimodal classification model based on the multimodal data.

[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 is a flow chart of a classification method based on multimodal data provided in Example 1 of the present invention;

[0021] Figure 2 is a flow chart of a classification method based on multimodal data provided in Embodiment 2 of the present invention;

[0022] Figure 3 This is an architecture diagram of a classification process based on multimodal data provided in Embodiment 3 of the present invention;

[0023] Figure 4 is a structural schematic diagram of a classification device based on multimodal data provided by Embodiment 4 of the present invention;

[0024] Figure 5 It is a structural schematic diagram of an electronic device for implementing a classification method based on multimodal data provided by Embodiment 5 of the present invention. DETAILED DESCRIPTION

[0025] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0026] Embodiment 1

[0027] Figure 1 It is a flowchart of a classification method based on multimodal data provided in Example 1 of the present invention. This embodiment can be applied to the situation where the category of the object to be classified is determined by inputting the multimodal data of the object to be classified into a multimodal classification model. The method can be executed by a classification device based on multimodal data, which can be implemented in software and / or hardware and can be configured in an electronic device that carries a classification function based on multimodal data.

[0028] See also Figure 1 The classification method based on multimodal data shown includes:

[0029] S110, obtaining data to be classified of the object to be classified in the candidate mode.

[0030] Among them, the object to be classified refers to an object that needs to be classified. The candidate modality refers to a pre-set data modality. Exemplarily, the candidate modality may include text modality, visual modality and audio modality. Text modality refers to a data modality displayed through text. Visual modality refers to a data modality displayed through vision. Audio modality refers to a data modality displayed through audio.

[0031] The data to be classified refers to data used for category determination. Exemplarily, the data to be classified may include text data to be classified, visual data to be classified, and audio data to be classified. The text data to be classified is data in text mode, the visual data to be classified is data in visual mode, and the audio data to be classified is data in audio mode.

[0032] The text data to be classified refers to the text data of the object to be classified. The visual data to be classified refers to the visual data of the object to be classified. The audio data to be classified refers to the audio data of the object to be classified.

[0033] It should be noted that the embodiments of the present invention do not make any settings for the text data to be classified, the visual data to be classified, and the audio data to be classified, and the technicians can obtain them as needed. For example, the text data to be classified may include natural language descriptions and conversation content, etc.; the visual data to be classified may include images and videos, etc.; the audio data to be classified may include voice, music, and ambient sound, etc.

[0034] Specifically, the text data to be classified in the text mode, the visual data to be classified in the visual mode, and the audio data to be classified in the audio mode of the object to be classified are obtained.

[0035] S120, inputting the data to be classified into the trained multimodal classification model to obtain the object classification result.

[0036] Among them, the multimodal classification model can be used to determine the category of the object to be classified through the data to be classified under the candidate modality. The embodiment of the present invention does not specifically limit the network structure of the multimodal classification model, which can be set by the technician according to experience or needs. Exemplarily, the multimodal classification model can adopt a graph neural network. The object classification result refers to the category to which the object to be classified belongs.

[0037] Exemplarily, the multimodal classification model is trained based on the text sample data, visual sample data and audio sample data of the sample training object. The sample training object refers to an object that provides data required for training the multimodal classification model. The text sample data refers to the text data of the sample training object. The visual sample data refers to the visual data of the sample training object. The audio sample data refers to the audio data of the sample training object.

[0038] It should be noted that an object to be classified has a multimodal data, that is, the multimodal data of the object to be classified may include text data to be classified, visual data to be classified and audio data to be classified; correspondingly, a sample training object also has a multimodal data, that is, the multimodal data of the sample training object may include text sample data, visual sample data and audio sample data.

[0039] The embodiment of the present invention provides a classification scheme based on multimodal data, by obtaining the data to be classified of the object to be classified in the candidate modality; wherein the candidate modality includes text modality, visual modality and audio modality; inputting the data to be classified into a trained multimodal classification model to obtain the object classification result; wherein the multimodal classification model is trained based on the text sample data, visual sample data and audio sample data of the sample training object. The above scheme trains the multimodal classification model according to the text sample data, visual sample data and audio sample data of the sample training object, so that the multimodal classification model can better process the data to be classified of the object to be classified in the candidate modality, thereby improving the accuracy of the object classification result determined by the trained multimodal classification model, that is, improving the accuracy of the classification result determined by the multimodal classification model based on the multimodal data.

[0040] Embodiment 2

[0041] Figure 2 is a flowchart of a classification method based on multimodal data provided by Example 2 of the present invention. Based on the above embodiments, this embodiment further adds "inputting text sample data, visual sample data and audio sample data of the sample training object into a pre-built multimodal classification model to obtain candidate sample vectors; wherein the candidate sample vectors include candidate text vectors, candidate audio vectors and candidate visual vectors; determining a spatial reference vector and a spatial alignment vector in the candidate sample vectors, and aligning the spatial alignment vector with the spatial reference vector in terms of spatial distribution to obtain a reference sample vector corresponding to the spatial alignment vector; determining a sample relationship image of the sample training object according to the spatial reference vector and the reference sample vector, and According to the image attribute matrix corresponding to the sample relationship image, the image edges in the sample relationship image are aligned with the image nodes in terms of attributes to obtain an updated edge degree matrix; wherein the image attribute matrix includes a point-edge relationship matrix, a node weight matrix and an edge weight matrix; according to the updated edge degree matrix and the image attribute matrix, the target sample vector of the sample training object is determined, and according to the target sample vector, the predicted classification result of the multimodal classification model is determined; according to the predicted classification result, the actual classification result of the corresponding sample training object, the spatial loss value and the attribute loss value, the model loss value of the multimodal classification model is determined, and according to the model loss value, the multimodal classification model is trained "operation to improve the training mechanism of the multimodal classification model. It should be noted that, for the part not described in detail in the embodiment of the present invention, reference may be made to the description of other embodiments.

[0042] See also Figure 2 The classification method based on multimodal data shown includes:

[0043] S210: Input text sample data, visual sample data, and audio sample data of the sample training object into a pre-built multimodal classification model to obtain a candidate sample vector.

[0044] The sample training object refers to an object that provides data required for training a multimodal classification model. The text sample data refers to the text data of the sample training object. The visual sample data refers to the visual data of the sample training object. The audio sample data refers to the audio data of the sample training object.

[0045] The candidate sample vector refers to a vector obtained by performing vector transformation on the sample training data of the sample training object. Exemplarily, the candidate sample vector includes a candidate text vector, a candidate audio vector and a candidate visual vector.

[0046] Specifically, feature extraction is performed on text sample data to obtain candidate text vectors, feature extraction is performed on audio sample data to obtain candidate audio vectors, and feature extraction is performed on visual sample data to obtain candidate visual vectors.

[0047] For example, for text feature extraction, the input text sample data is extracted using a pre-trained text feature extraction model, and the average of the vectors of the last four layers is used as a candidate text vector through a Bi-GRU (Bidirectional Gated Recurrent Unit). The candidate text vector is a 1024-dimensional vector:

[0048]

[0049] in, represents the candidate text vector; Represents text sample data; RoBERTa represents the model used for text feature extraction, that is, the text feature extraction model.

[0050] For example, for visual feature extraction, the input visual sample data is extracted using a visual feature extraction model, and a 342-dimensional or 300-dimensional candidate visual vector is obtained through a fully connected layer. The visual feature extraction model can be a densely connected convolutional network or a three-dimensional convolutional neural network model.

[0051]

[0052] in, represents a candidate visual vector; Represents visual sample data; DenseNet represents densely connected convolutional networks; 3D-CNN represents three-dimensional convolutional neural networks. The visual feature extraction model can be used to extract visual features.

[0053] For example, for audio feature extraction, the audio sample data is input and extracted using the audio feature extraction model, and then a 1582-dimensional or 342-dimensional candidate audio vector is obtained through a fully connected layer:

[0054]

[0055] in, represents a candidate audio vector; Represents audio sample data; openSMILE represents the model used for audio feature extraction, that is, the audio feature extraction model.

[0056] S220 , determining a spatial reference vector and a spatial alignment vector in the candidate sample vector, and aligning the spatial alignment vector with the spatial reference vector in terms of spatial distribution, to obtain a reference sample vector corresponding to the spatial alignment vector.

[0057] The spatial reference vector can be used as a standard for the spatial distribution of the spatial alignment vector, that is, the spatial reference vector is a reference standard for the spatial distribution of the spatial alignment vector. The spatial alignment vector refers to a candidate sample vector that needs to be spatially aligned to the spatial reference vector. The reference sample vector refers to a vector obtained after aligning the spatial alignment vector at the spatial distribution level.

[0058] In an optional embodiment, determining a spatial reference vector and a spatial alignment vector in a candidate sample vector includes: grouping the candidate sample vectors by modality to obtain a sample vector group, and determining a modality group contribution of the sample vector group; wherein the sample vector group includes candidate sample vectors under at least two candidate modalities; determining a target modality from the candidate modalities based on the modality group contribution, and using the candidate sample vector under the target modality as a spatial reference vector, and using the candidate sample vectors under other candidate modalities except the target modality as a spatial alignment vector.

[0059] The sample vector group refers to a vector combination obtained by randomly grouping candidate sample vectors under different candidate modalities. Exemplarily, the sample vector group may include a first sample vector group, a second sample vector group, a third sample vector group, and a fourth sample vector group. The first sample vector group may be a combination of a candidate text vector and a candidate visual vector, the second sample vector group may be a combination of a candidate text vector and a candidate audio vector, the third sample vector group may be a combination of a candidate audio vector and a candidate visual vector, and the fourth sample vector group may be a combination of a candidate text vector, a candidate visual vector, and a candidate audio vector.

[0060] Among them, the modal group contribution can be used to quantify the importance of the sample vector group in determining the classification result. Exemplarily, the modal group contribution can be determined based on a preset modal evaluation index. The modal evaluation index can be used to evaluate the contribution of multimodal data. The embodiment of the present invention does not impose any restrictions on the setting of the modal evaluation index, which can be set by a technician based on experience or needs. The target modality refers to the candidate modality with the highest single modality contribution.

[0061] Exemplarily, a sample vector group including candidate sample vectors under three candidate modalities is taken as a sample vector global group, and a sample vector group including candidate sample vectors under two candidate modalities is taken as a sample vector local group; for any sample vector local group, according to the sample vector global group, a candidate modality that does not exist in the sample vector local group is taken as a reference modality; the difference between the modality group contribution corresponding to the sample vector global group and the modality group contribution corresponding to the sample vector local group is taken as the single modality contribution of the reference modality; the target modality is determined according to the single modality contribution of each reference modality. Preferably, the reference modality with the highest single modality contribution is taken as the target modality.

[0062] Among them, the global group of sample vectors refers to a combination of candidate text vectors, candidate visual vectors and candidate audio vectors. The local group of sample vectors may include a combination of candidate text vectors and candidate audio vectors, a combination of candidate audio vectors and candidate visual vectors, and a combination of candidate text vectors and candidate visual vectors. The reference modality refers to a candidate modality that does not exist in any local group of sample vectors. The single modality contribution can be used to quantify the importance of the candidate sample vector under any candidate modality in determining the classification result.

[0063] For example, if the global group of sample vectors is A, the local group of sample vectors including candidate text vectors and candidate audio vectors is B, the local group of sample vectors including candidate audio vectors and candidate visual vectors is C, and the local group of sample vectors including candidate text vectors and candidate visual vectors is D, then the difference between the modal group contribution corresponding to A and the modal group contribution corresponding to B is taken as the unimodal contribution a of the visual modality; the difference between the modal group contribution corresponding to A and the modal group contribution corresponding to C is taken as the unimodal contribution b of the text modality; the difference between the modal group contribution corresponding to A and the modal group contribution corresponding to D is taken as the unimodal contribution c of the audio modality; the unimodal contribution y with the largest value is determined from the unimodal contribution a, the unimodal contribution b and the unimodal contribution c, and the candidate modality corresponding to the unimodal contribution y is taken as the target modality.

[0064] Exemplarily, after obtaining the representation information of the three modalities (i.e., obtaining the candidate text vector, the candidate audio vector, and the candidate visual vector), the contribution of different modalities to the task is measured according to the evaluation indicators of the downstream task (i.e., the modality evaluation indicator) to help identify the main modality (i.e., the target modality) that can be used as the cross-modal alignment module.

[0065] Exemplarily, the single-mode contribution can be determined by the following formula:

[0066] Δπ(x)=q(Jπ(x)∪x)-q(Jπ(x)),x∈{A,V,T};

[0067] Among them, Jπ(x) represents all modal combinations excluding the candidate modality x in a specific combination π, that is, the local group of sample vectors, q(·) can be used as an evaluation function to measure the gain, and the evaluation indicators of downstream tasks, such as accuracy and weighted F-Socre, are used to measure the prediction performance of each sample vector group to obtain the marginal contribution Δπ(x), which represents the performance improvement brought by the candidate modality x when combined with other modalities. If the local group of sample vectors Jπ(x) is set to {A, V} and a candidate modality x=T is added, then the single modal contribution of the candidate modality T can be obtained. Because the candidate modality T does not appear in the modal combination Jπ(x), the modality that brings the greatest performance improvement is the modality that contributes the most to the task.

[0068] It can be understood that by introducing the contribution of sample vector groups and modal groups, determining the target mode, and determining the spatial reference vector and spatial alignment vector based on the target mode, it is possible to consider the contribution of each mode, determine the target mode that does not require spatial optimization, and the candidate mode that requires spatial optimization, thereby improving the accuracy of the determined spatial reference vector and spatial alignment vector.

[0069] For example, the target modality (such as text modality) with the highest assessed unimodal contribution is used as the main modality (i.e., spatial reference modality), and the visual modality and audio modality (i.e., spatial alignment modality) are aligned to the main modality. This alignment operation is based on the OT distribution optimization algorithm, in which the cost from the original distribution to the target distribution (i.e., the distribution of the main modality) is evaluated by the Wasserstein Distance (WD). Specifically, WD A→T and WD V→T They represent the alignment of the distributions of the audio and visual modalities toward the distribution of the textual modality using the Wilstein distance, respectively.

[0070]

[0071] Among them, A represents audio mode; T represents text mode; V represents visual mode; WDA→T It indicates that the spatial distribution of the candidate audio vectors in the audio mode is aligned with the spatial distribution of the candidate text vectors in the text mode; WD V→T Indicates that the spatial distribution of candidate visual vectors in the visual mode is aligned with the spatial distribution of candidate text vectors in the text mode; inf represents the infimum, which is used to seek the transmission plan with the minimum transmission cost, similar to min; Represents the search for the optimal transmission plan between the audio modality and the text modality γ A , so that the transmission cost is minimized; X A represents the space of audio mode; Y represents the space of text mode; c A represents the integral part of the cost function, which is used to measure the difference between the two spatial distributions of the text modality and the audio modality; A candidate audio vector representing the i-th sample training object; Represents the candidate text vector of the i-th sample training object; represents the search for the optimal transfer plan between the visual modality and the textual modality γ V , so that the transmission cost is minimized; X V represents the space of visual modalities; c V represents the integral part of the cost function, which is used to measure the difference between the two spatial distributions of the textual modality and the visual modality; Represents the candidate visual vector of the i-th sample training object.

[0072] Continuing the above example, we get the reference sample vector and And the space reference vector

[0073] S230. Determine a sample relationship image of the sample training object according to the spatial reference vector and the reference sample vector, and align the image edges in the sample relationship image to the image nodes according to the image attribute matrix corresponding to the sample relationship image to obtain an updated edge degree matrix.

[0074] Among them, the sample relationship image refers to a hypergraph that shows the relationship between each modality. The image attribute matrix refers to a matrix determined based on the basic attributes in the sample relationship image. Exemplarily, the image attribute matrix includes a point-edge relationship matrix, a node weight matrix, and an edge weight matrix, wherein the point-edge relationship matrix can be used to record the connection relationship between the image nodes and the image edges in the sample relationship image. The node weight matrix can be used to record the weights corresponding to each image node in the sample relationship image. The edge weight matrix can be used to record the weights corresponding to each image edge (i.e., hyperedge) in the sample relationship image.

[0075] It should be noted that the image nodes and image edges in the sample relationship image are preset with corresponding weights. The embodiment of the present invention does not impose any restrictions on the weights of the image nodes and the weights of the image edges. The technicians can set them according to experience or needs, or determine them repeatedly through a large number of experiments. In the point-edge relationship matrix, the relationship between the image node and the image edge can be represented by 0 and 1, where 1 can represent that the image node and the image edge have a connection relationship, and 0 can represent that the image node and the image edge do not have a connection relationship.

[0076] The edge degree matrix can be used to record the degree information of the image edge. For example, the edge degree matrix can record the number of image nodes connected by the image edge, or the connection strength between the image edge and the image node.

[0077] In an optional embodiment, the image edges in the sample relationship image are attribute-aligned to the image nodes according to the image attribute matrix corresponding to the sample relationship image to obtain an updated edge degree matrix, including: determining the point-edge weight matrix according to the point-edge relationship matrix, the node weight matrix and the edge weight matrix; determining the node degree matrix and the edge degree matrix according to the point-edge relationship matrix or the point-edge weight matrix, and determining the updated edge degree matrix according to the node degree matrix and the edge degree matrix.

[0078] The node-edge weight matrix can be used to record the connection strength between image nodes and image edges. The node degree matrix can be used to record the degree information of image nodes. Exemplarily, the node degree matrix can record the number of image edges connected to the image nodes, or the connection strength between the image nodes and the image edges.

[0079] Exemplarily, in order to determine the updated edge degree matrix, the concept of slicing is introduced, that is, the edge degree information of the hyperedge node pairs in the edge degree matrix is ​​updated one by one, and finally the updated edge degree matrix is ​​obtained. For example, the edge degree information of each hyperedge node pair can be updated by the following formula:

[0080]

[0081] Among them, WD e→n represents the OT optimization of aligning image edges to image nodes, that is, the image edges are aligned to the image nodes in terms of attributes; e m represents the mth image edge, i.e. the mth hyperedge; n m represents the mth image node; (e m ,n m ) represents a hyperedge node pair; ε represents all hyperedge node pairs in the sample relationship image; Represents the degree information of the mth image edge, that is, the image edge e m The number of connected image nodes or the strength of the connections; Represents the degree information of the mth image node, that is, image node n m The number of connected image edges or the strength of the connection; WD can be used to measure the Wasserstein distance between image nodes and image edges; ∫() is the integral term, which can be expressed as the calculation of Sliced ​​Wasserstein Distance (SWD) on the unit sphere. This projection to a low-dimensional space is to reduce the computational cost; θ∈S d-1 represents the calculation performed on the unit sphere; P θ Represents the projection matrix, used to reduce complexity.

[0082] Furthermore, after minimum optimization, the optimal hyper-edge degree matrix, that is, the updated edge degree matrix, is obtained.

[0083] It can be understood that by introducing the point-edge weight matrix, determining the node degree matrix and edge degree matrix according to the point-edge weight or the point-edge relationship matrix, and updating the edge degree matrix, the diversity of the determined node degree matrix and edge degree matrix is ​​improved, and the accuracy of the updated edge degree matrix is ​​improved.

[0084] In an optional embodiment, a point-edge weight matrix is ​​determined based on a point-edge relationship matrix, a node weight matrix, and an edge weight matrix, including: for any image node in a sample relationship image, determining the point-edge association relationship between the image node and each image edge in the sample relationship image based on the point-edge relationship matrix; determining a matching result between the image node and any image edge with which there is an association, and determining a degree of point-edge association between the image node and the image edge based on the matching result, the node weight matrix, and the edge weight matrix; determining a point-edge weight matrix based on the degree of point-edge association between each image node and each image edge in the sample relationship image.

[0085] The point-edge association relationship may characterize the connection relationship between the image nodes and the image edges, that is, whether there is a connection between any image node and any image edge. Exemplarily, the point-edge association relationship may be that the point-edge is associated or that the point-edge is not associated. The point-edge association degree may be used to quantify the closeness between the image nodes and image edges that are associated with the point-edge.

[0086] The matching result can indicate whether the image node and the image edge associated with the point edge belong to the same sample training object. Specifically, for any image node p, determine another image node s connected to the image edge associated with the image node p, and judge whether the image node p and the image node s belong to the same sample training object.

[0087] Exemplarily, for an image node n and an image edge e associated with the image node n, if the matching result is a successful match, it indicates that the image node n and the image edge e both belong to the same sample training object. At this time, the product of the weight corresponding to the image node n and the weight corresponding to the image edge e can be used as the point-edge association degree between the image node n and the image edge e; if the matching result is a failed match, it indicates that the image node n and the image edge e do not belong to the same sample training object. At this time, the preset association default value can be used as the point-edge association degree between the image node n and the image edge e. The embodiment of the present invention does not impose any limitation on the size of the preset association default value, which can be set by a technician based on experience, or determined repeatedly through a large number of experiments. For example, the preset association default value can be 0.

[0088] Specifically, the point-edge weight matrix is ​​determined according to the point-edge association degree between the image nodes and the image edges in which the point-edge association exists in the sample relationship image.

[0089] For example, construct a hypergraph H g =(V H ,E H ,ω,γ) to carry out information modeling, where V H Represents the set of image nodes in the sample relationship image, representing a multimodal data unit, each image node corresponds to a single-modal sample vector; in a multimodal scenario, each data sample will be decomposed into three heterogeneous modal nodes. H Represents the set of image edges in the sample relation image, including both intra-modality and inter-modality image edges.

[0090] Exemplary, the point-edge relationship matrix of the sample relationship image Among them, H can represent the connection relationship between image nodes and image edges in the sample relationship image, that is, whether the record is associated. 1 in H can indicate that there is an association between the image edge and the image node, and 0 in H indicates that there is no association between the image edge and the image node.

[0091] Furthermore, in order to strengthen the association between image nodes and image edges, the weighted association matrix (i.e., point-edge weight matrix) is calculated. For example, the point-edge weight matrix can be determined by the following formula:

[0092]

[0093] in, Indicates the degree of point-edge association between image nodes and image edges; γ e(n) represents the weight corresponding to the image node; ω(e) represents the weight corresponding to the image edge; n∈e indicates that the matching result is a successful match. Further, according to the point-edge association degree between each image node and the image edge, the point-edge weight matrix is ​​determined

[0094] It can be understood that by determining the point-edge correlation degree between the image nodes and image edges associated with the point-edge according to the point-edge correlation relationship, and constructing a point-edge weight matrix, the closeness between the image nodes and image edges associated with the point-edge is taken into consideration, thereby improving the accuracy of the determined point-edge weight matrix.

[0095] S240: Determine a target sample vector of the sample training object according to the updated edge degree matrix and image attribute matrix, and determine a predicted classification result of the multimodal classification model according to the target sample vector.

[0096] Among them, the target sample vector refers to the vector after the candidate sample vector is updated. Exemplarily, the target sample vector may include a target text vector, a target visual vector and a target audio vector. The target text vector refers to the text vector obtained after the candidate text vector is updated. The target visual vector refers to the visual vector obtained after the candidate visual vector is updated. The target audio vector refers to the audio vector obtained after the candidate audio vector is updated. The predicted classification result refers to the category result of the sample training object output by the multimodal classification model.

[0097] In an optional embodiment, the target sample vector of the sample training object is determined based on the updated edge degree matrix and image attribute matrix, and the predicted classification result of the multimodal classification model is determined based on the target sample vector, including: updating the candidate sample vector of the sample training object based on the updated edge degree matrix, edge weight matrix, point-edge relationship matrix, node degree matrix, and point-edge weight matrix to obtain the target sample vector; determining the sample modal vector of each sample training object based on the target sample vector, and determining the predicted classification result output by the multimodal classification model based on the sample modal vector.

[0098] The sample modal vector can be understood as a vector obtained by concatenating the multimodal candidate sample vectors after spatial distribution alignment and attribute alignment. The sample modal vector can be used to predict the classification results of the sample training object.

[0099] For example, the GNN (Graph Neural Network) network propagation update, the following formula is to explain how the hypergraph (i.e., sample relationship image) is iterated so that the multimodal classification model can better learn multimodal associations, v can represent the candidate vector matrix before the update, and v' represents the target vector matrix after the update. Among them, the candidate vector matrix can be used to record the candidate sample vectors of the sample training object; the target vector matrix can be used to record the target sample vectors of the sample training object.

[0100]

[0101] Where v' represents the target vector matrix; σ represents the nonlinear activation function; D -1 represents the inverse of the node degree matrix; H represents the point-edge relationship matrix; W e represents the edge weight matrix, which can be used to adjust the importance of different image edges; represents the inverse of the updated edge degree matrix; represents the transpose of the point-edge weight matrix; v represents the candidate vector matrix. It should be noted that is a submatrix extracted from D, that is, D represents the degree matrix of all image nodes, Represents image node n m The degree matrix of .

[0102] Furthermore, the target sample vector of each sample training object is determined according to the target vector matrix.

[0103] Exemplarily, the sample modal vector may be determined by the following formula:

[0104]

[0105] Among them, s i Represents the sample modal vector of the i-th sample training object; Represents the target text vector of the i-th sample training object; represents the target audio vector of the i-th sample training object; Represents the target visual vector of the i-th sample training object; Indicates splicing.

[0106] Furthermore, the final category can be predicted by applying the Softmax layer, that is, the predicted classification result is determined by the following formula:

[0107]

[0108] in, represents the predicted classification result of the i-th sample training object; softmax represents the normalized exponential function; ReLU stands for Rectified Linear Unit, an activation function widely used in artificial neural networks; b τ is a bias term, which can be used to adjust the result after linear transformation; τ represents a constant term; W represents a trainable weight matrix, which is used to adjust the ReLU (s i ) for linear transformation. Specifically, after activation, the score of each class label of the sample training object is calculated, and finally the highest score is selected as the final predicted class of the sample training object through the argmax (argument of the maximum, a large value of the independent variable point set or parameter set) operation, that is, the class label with the highest prediction result probability is used as the predicted classification result.

[0109] It can be understood that by updating the candidate sample vectors to obtain the target sample vectors, and then determining the sample modal vectors based on the target sample vectors, the sample modal vectors are optimized, thereby improving the accuracy of the predicted classification results of the multimodal classification model based on the sample modal vector output.

[0110] S250, determining the model loss value of the multimodal classification model according to the predicted classification result, the actual classification result of the corresponding sample training object, the spatial loss value and the attribute loss value, and training the multimodal classification model according to the model loss value.

[0111] The actual classification result refers to the real category label of the sample training object. The spatial loss value refers to the loss value of the multimodal classification model when the spatial distribution is aligned. The attribute loss value refers to the loss value of the multimodal classification model when the attribute is aligned. The model loss value refers to the loss value of the multimodal classification model.

[0112] In an optional embodiment, the model loss value of the multimodal classification model is determined based on the predicted classification result, the actual classification result of the corresponding sample training object, the spatial loss value and the attribute loss value, including: determining the predicted result probability of the predicted classification result, and encoding the actual classification result to obtain the actual result probability; determining the basic loss value of the multimodal classification model based on the predicted result probability, the actual result probability and the number of samples and the number of category labels of the sample training object; obtaining the spatial loss value and the attribute loss value, and determining the model loss value of the multimodal classification model based on the basic loss value, the spatial loss value and the attribute loss value.

[0113] Among them, the prediction result probability can be used to quantify the accuracy of the predicted classification result. Specifically, the prediction result probability can be used to evaluate the possibility that the sample training object is the corresponding predicted classification result. The actual result probability refers to the probability of the actual classification result of the sample training object. Exemplarily, the actual result probability can be 1. The number of samples refers to the number of sample training objects. The number of category labels refers to the number of category labels that can be predicted by the multimodal classification model. The basic loss value refers to the basic loss value of the multimodal classification model. Exemplarily, the basic loss value can be determined using an existing loss function. For example, the basic loss value can be determined using the L2 regularized category cross entropy loss.

[0114] Exemplarily, the basic loss value can be determined by the following formula:

[0115]

[0116] in, represents the basic loss value; N represents the number of samples; i represents the i-th sample training object; C represents the number of category labels; j represents the j-th category label; y i,j Represents the unique hot encoding of the actual classification result, that is, the actual result probability; Represents the probability of the predicted result; represents the L2 norm of all available training parameters; λ represents the weight of L2 regularization.

[0117] Furthermore, the model loss value can be determined by the following formula:

[0118]

[0119] in, Represents the model loss value; Indicates the space loss value; Represents the attribute loss value.

[0120] It can be understood that by determining the model loss value of the multimodal classification model based on the basic loss value, spatial loss value and attribute loss value, the comprehensiveness and accuracy of the determined model loss value are improved, thereby improving the accuracy of subsequent training of the multimodal classification model based on the model loss value, and improving the performance of the multimodal classification model.

[0121] S260: Obtain the data to be classified of the object to be classified in the candidate mode.

[0122] Among them, candidate modalities include text modality, visual modality and audio modality.

[0123] S270: Input the data to be classified into the trained multimodal classification model to obtain the object classification result.

[0124] Among them, the multimodal classification model is trained based on the text sample data, visual sample data and audio sample data of the sample training object.

[0125] The embodiment of the present invention provides a classification scheme based on multimodal data, by adding text sample data, visual sample data and audio sample data of a sample training object into a pre-built multimodal classification model to obtain a candidate sample vector; wherein the candidate sample vector includes a candidate text vector, a candidate audio vector and a candidate visual vector; a spatial reference vector and a spatial alignment vector in the candidate sample vector are determined, and the spatial alignment vector is spatially aligned with the spatial reference vector to obtain a reference sample vector corresponding to the spatial alignment vector; a sample relationship image of the sample training object is determined according to the spatial reference vector and the reference sample vector, and a sample relationship image corresponding to the sample relationship image is obtained according to the image corresponding to the sample relationship image. The image edges in the sample relationship image are aligned to the image nodes in terms of attributes to obtain an updated edge degree matrix; wherein the image attribute matrix includes a point-edge relationship matrix, a node weight matrix and an edge weight matrix; according to the updated edge degree matrix and the image attribute matrix, the target sample vector of the sample training object is determined, and according to the target sample vector, the predicted classification result of the multimodal classification model is determined; according to the predicted classification result, the actual classification result of the corresponding sample training object, the spatial loss value and the attribute loss value, the model loss value of the multimodal classification model is determined, and according to the model loss value, the multimodal classification model is trained, thereby improving the training mechanism of the multimodal classification model. The above scheme shortens the difference in spatial distribution between different candidate modalities by aligning the spatial alignment vector to the spatial reference vector, that is, reduces the semantic difference between different candidate modalities, which is beneficial to the subsequent learning of the multimodal classification model; at the same time, the image edges in the sample relationship image are aligned to the image nodes in terms of attributes, eliminating the attribute differences between the image nodes and the image edges, that is, eliminating the differences at the network level, improving the information fusion capability, and facilitating the iterative process of the sample relationship image; and, in this scheme, by introducing spatial distribution alignment and attribute alignment, the semantic differences and attribute differences between different candidate modalities are further shortened, the accuracy of the multimodal classification model trained based on multimodal data is improved, and the accuracy of the multimodal classification model prediction is improved; and, the embodiment of the present invention also introduces spatial loss values ​​and attribute loss values, determines the model loss value based on the spatial loss value and the attribute loss value, and improves the comprehensiveness and accuracy of the determined model loss value.

[0126] Embodiment 3

[0127] The embodiment of the present invention provides an optional example based on the above embodiment. It should be noted that for the part not described in detail in the embodiment of the present invention, reference can be made to the description of other embodiments.

[0128] Current multimodal classification methods mainly rely on feature-level fusion, decision-level fusion, and multimodal modeling under the deep learning framework (such as the Transformer architecture). However, these methods usually face the following problems: (1) The information contribution of different modalities is uneven, and they are easily disturbed by low-contribution modalities, thus affecting classification performance; (2) There are large differences in the data distribution of different modalities, which makes it difficult to align features in the same semantic space, affecting the complementarity of cross-modal information; (3) Traditional fusion methods mostly rely on fixed structures and lack the ability to model the dynamic characteristics of different modalities, making it difficult to fully explore the potential correlation of multimodal data.

[0129] In the existing technology, graph neural networks (GNNs) have received extensive attention in multimodal learning tasks due to their advantages in modeling complex relationships and structured data. GNNs model the dependencies within and between modalities through graph structures, allowing information to propagate between nodes and edges, thereby effectively integrating multimodal information. However, existing GNN solutions often ignore the differences in contribution between modalities and the influence of different node and edge attributes during information propagation, resulting in a large room for improvement in the performance of the model in multimodal fusion.

[0130] In view of the above problems, the embodiments of the present invention propose an advanced multimodal data distribution alignment and multimodal fusion method under the GNN framework, including: (1) a modal contribution evaluation module, which uses the evaluation indicators of downstream tasks to dynamically evaluate the contribution of each candidate modality to the classification task and serves the subsequent cross-modal distribution alignment; (2) a cross-modal distribution alignment module: based on the contribution evaluated by the previous module, the optimal transport (OT) algorithm is used to optimize the data distribution of different candidate modalities so that they are aligned to a unified feature space and the semantic differences between modalities are reduced; (3) a fusion mechanism based on a hypergraph. In the hypergraph structure constructed by the GNN network, the attribute edges between different candidate modalities are different. This module improves the information fusion capability by aligning the edges and nodes with different attributes; (4) a more flexible cross entropy loss. Adding the OT optimization loss at the GNN network level on the basis of the traditional cross entropy loss can more efficiently realize multimodal fusion information transmission.

[0131] Exemplarily, the modal contribution evaluation module in the embodiment of the present invention dynamically evaluates the single-modal contribution of each candidate modality through the evaluation index of the downstream task itself, thereby selecting the target modality with the largest single-modal contribution; the cross-modal distribution alignment module is based on the target modality evaluated by the modal contribution evaluation module, and uses the optimal transmission algorithm to align the remaining candidate modalities to the target modality. This data preprocessing at the representation level can well eliminate modal heterogeneity, thereby obtaining a more unified representation space for subsequent GNN network-based learning; the fusion mechanism based on the hypergraph is to optimize the differences in attribute edges between different candidate modalities during the fusion and information transmission stage of the GNN network, and improve the efficiency of multimodal fusion and information transmission by aligning attribute edges to nodes. The advantage is that it can not only effectively eliminate the adverse effects of modal heterogeneity on subsequent network learning, but also significantly improve the ability of multimodal information transmission and fusion; in addition, the method has shown performance improvement in the multimodal dialogue emotion recognition task, further proving its broad application prospects and robustness.

[0132] The classification method based on multimodal data proposed in the embodiment of the present invention has greatly promoted the depth and breadth of application of multimodal classification technology. Through advanced multimodal data preprocessing and graph neural network technology, it has promoted the application of analysis based on multimodal classification in all aspects of modern society, enabling users to manage and utilize multimodal information more efficiently and accurately.

[0133] For example, see Figure 3 The architecture diagram of the classification process based on multimodal data is shown in Figure 1. Specifically, the input of the three modalities of visual modality, text modality and audio modality will first pass through the modal contribution evaluation module, select the main modality with the highest single modal contribution (i.e., the target modality), and then enter the cross-modal distribution alignment module for representation level preprocessing, and then enter the hypergraph-based GNN network for fusion and learning, and finally predict the final category (i.e., the predicted category result) through the classifier. The training process is to add the loss of representation and network level alignment on the basis of cross entropy loss.

[0134] The modality contribution evaluation module in the embodiment of the present invention first performs feature extraction, and then uses the extracted representation information (i.e., candidate sample vectors) to evaluate the modality contribution. The construction of the hypergraph (i.e., sample relationship image) in the GNN network includes node construction, hyperedge construction, and the weight mechanism between hyperedges and nodes. g =(V H ,E H ,ω,γ), the node set V HRepresents a multimodal data unit. Each image node corresponds to a single-modal data sample representation, that is, a candidate sample vector. In a multimodal scenario, each data sample will be decomposed into three heterogeneous modal nodes, representing: text modal node Audio Mode Node and the visual modal node Therefore, the total number of image nodes in the sample relationship image is |V H |=3K, K represents the total number of sample training objects, and the coefficient 3 corresponds to the three candidate modalities of each sample training object.

[0135] Among them, the hyperedge set E H It aims to model the relationship between modalities and modalities to enhance the interaction of cross-modal information. H |=3+K image edges, which are divided into the following two categories: intra-modal image edges and inter-modal image edges. Intra-modal image edges: For each candidate modality x∈{A,V,T}, all image nodes of candidate modalities x will be connected to the same image edge throughout the conversation. The above connection method is used to capture the feature dynamics and potential associations within the same candidate modality. Inter-modal image edges: The image nodes of the three candidate modalities of each sample training object are connected by an independent image edge. This structure can effectively fuse cross-modal information and enhance the complementarity of features of different modalities, thereby building a more complete representation.

[0136] Among them, the weight mechanism between hyperedges and nodes: the connection strength between each hyperedge (i.e., image edge) e and image node n is composed of the hyperedge weight ω(e) and the node weight γ e (n) is characterized to determine the point-edge association matrix and the point-edge weight matrix respectively, which are used to control the information propagation process in the hypergraph. The benefits are: strengthening the influence of key modes and samples through learnable weight parameters (giving higher weights to important connections); adjusting the propagation path and intensity of information in the sample relationship image, optimizing cross-modal feature fusion; and providing an interpretable modal contribution quantification basis for downstream tasks.

[0137] Exemplary, fusion mechanism based on sample relational images: To further enhance information aggregation, the OT method is introduced into the GNN framework to optimize information transfer by aligning the image node distribution with the corresponding image edge distribution. The attributes of the image edges between the candidate modalities in the original sample relational images are different. This alignment process ensures the consistency between the image node representation and the image edge representation, thereby enhancing the ability of GNN to capture the interaction of image edge information with different attributes. The goal of alignment is to minimize the Wilstein distance between the image node and image edge distributions.

[0138] The classification method based on multimodal data provided by the embodiment of the present invention has an efficient multimodal classification effect, that is, the embodiment of the present invention shows good performance in the multimodal classification dialogue emotion recognition task. The accuracy and stability of classification are improved by combining the information of the three modes of text, audio and vision. The combination of the modal contribution evaluation module and the cross-modal alignment module eliminates the modal heterogeneity well and obtains a more unified and similar modal analysis, which is conducive to the fusion and information transfer in the subsequent GNN network. In the GNN network, by aligning the image nodes and image edges in the hypergraph (sample relationship image), this alignment process ensures the consistency between the image node representation and the image edge representation, thereby improving the efficiency of fusion and information transfer at the network level. In this way, the optimization at the network level and the representation level effectively improves the accuracy of multimodal classification; it has an innovative hypergraph paradigm, that is, after the traditional hypergraph is constructed, considering that the hyperedges (i.e., image edges) between modalities have attribute differences, the hyperedges are innovatively aligned to the image nodes. This optimization method can help multimodal fusion to be more efficient, so that the network can learn the fused multimodal information more easily and accurately. Such optimization can serve as a new paradigm for hypergraph-based GNN networks and play a significant role in a variety of downstream tasks.

[0139] The classification method based on multimodal data proposed in the embodiment of the present invention provides an efficient modal contribution evaluation. The embodiment of the present invention discovered and paid attention to the problem of different single-modal contributions in multimodal classification tasks, and calculated the marginal contribution (i.e., single-modal contribution) through the evaluation index of the task itself, thereby effectively selecting the target modality with the highest contribution to the task to serve the subsequent data preprocessing at the representation level. This evaluation mechanism has low computational complexity and can be adapted to a variety of different evaluation indicators, so it has good practicality and generalization capabilities.

[0140] The classification method based on multimodal data proposed in the embodiment of the present invention provides an innovative cross-modal alignment mechanism. The embodiment of the present invention pays attention to the impact of modal distribution and heterogeneity on subsequent network learning, and designs cross-modal alignment based on optimal transmission. By aligning other candidate modalities to the target modality with the highest contribution to the selected single modality, a unified and similar representation distribution is obtained. This mechanism can not only effectively eliminate the semantic gap, but also significantly improve the efficiency of subsequent model learning in the network.

[0141] The classification method based on multimodal data proposed in the embodiment of the present invention provides a fusion mechanism based on a hypergraph. Based on the traditional hypergraph construction, the embodiment of the present invention pays attention to the differences in the hyperedges between the candidate modalities, so the alignment between the hyperedges and the image nodes is extracted, and this alignment depends on the optimal transmission algorithm. This alignment process ensures the consistency between the image node representation and the image edge representation, thereby enhancing the ability of GNN to capture the interaction of hyperedge information of different attributes.

[0142] The classification method based on multimodal data proposed in the embodiment of the present invention provides a flexible cross entropy loss. On the basis of the traditional cross entropy loss function, the embodiment of the present invention also adds the loss function in the spatial alignment process and the loss function in the attribute alignment process to the cross entropy loss function. This way of constructing the loss function can help the model to adjust dynamically, so as to give full play to the advantages of optimal transmission at the representation and network levels.

[0143] For example, take the multimodal emotion recognition dataset I as an example. If the multimodal emotion recognition dataset I includes 151 two-person dialogues and 7433 sentences, and six emotion category labels are preset, namely sadness, happiness, anger, neutral, frustration and excitement. Each sentence is labeled with one of the six category labels. The training and verification process used 120 dialogues and 5810 sentences, and the testing process used 31 dialogues and 1623 sentences. Based on the multimodal emotion recognition dataset I, an accuracy of 74.00% and a weighted F1-Score of 73.86% were obtained, and the F1-Scores of 83.27%, 64.31%, 66.88%, 70.85%, 71.32% and 81.83% were obtained for sadness, happiness, anger, neutral, frustration and excitement respectively.

[0144] For example, take the multimodal emotional dialogue dataset M as an example. If the multimodal emotional dialogue dataset M includes 1433 rounds of dialogue and 13708 sentences, and seven emotional category labels are preset, namely disgust, fear, neutral, joy, anger, surprise and sadness. Each sentence is labeled with one of the seven category labels. The training and verification process used 1153 dialogues, a total of 11098 sentences, and the remaining 280 dialogues, a total of 2610 sentences, for testing. Based on the multimodal emotional dialogue dataset M, an accuracy of 68.35% and a weighted F1-Score of 67.03% were obtained, and F1-Scores of 28.00%, 28.17%, 80.25%, 63.73%, 53.29%, 64.12% and 41.86% were obtained for disgust, fear, neutral, joy, anger, surprise and sadness, respectively.

[0145] Embodiment 4

[0146] Figure 4 : is a structural diagram of a classification device based on multimodal data provided by Embodiment 4 of the present invention. This embodiment is applicable to the case where the category of an object to be classified is determined by inputting the multimodal data of the object to be classified into a multimodal classification model. The method can be executed by a classification device based on multimodal data, which can be implemented in software and / or hardware, and can be configured in an electronic device that carries a classification function based on multimodal data.

[0147] like Figure 4 As shown, the device includes: a module 410 for acquiring data to be classified and a module 420 for determining classification results.

[0148] The to-be-classified data acquisition module 410 is used to acquire the to-be-classified data of the to-be-classified object in the candidate modality; wherein the candidate modality includes text modality, visual modality and audio modality;

[0149] A classification result determination module 420 is used to input the data to be classified into a trained multimodal classification model to obtain an object classification result;

[0150] The multimodal classification model is obtained by joint training based on text sample data, visual sample data and audio sample data of the sample training object.

[0151] The embodiment of the present invention provides a classification scheme based on multimodal data, by obtaining the data to be classified of the object to be classified in the candidate modality; wherein the candidate modality includes text modality, visual modality and audio modality; inputting the data to be classified into a trained multimodal classification model to obtain the object classification result; wherein the multimodal classification model is trained based on the text sample data, visual sample data and audio sample data of the sample training object. The above scheme trains the multimodal classification model according to the text sample data, visual sample data and audio sample data of the sample training object, so that the multimodal classification model can better process the data to be classified of the object to be classified in the candidate modality, thereby improving the accuracy of the object classification result determined by the trained multimodal classification model, that is, improving the accuracy of the classification result determined by the multimodal classification model based on the multimodal data.

[0152] Optionally, the multimodal classification model is trained based on the following devices:

[0153] A candidate sample vector determination module is used to input the text sample data, visual sample data and audio sample data of the sample training object into a pre-built multimodal classification model to obtain a candidate sample vector; wherein the candidate sample vector includes a candidate text vector, a candidate audio vector and a candidate visual vector;

[0154] A reference sample vector determination module, used to determine a spatial reference vector and a spatial alignment vector in the candidate sample vector, and align the spatial alignment vector with the spatial reference vector in terms of spatial distribution to obtain a reference sample vector corresponding to the spatial alignment vector;

[0155] An edge degree matrix updating module is used to determine the sample relationship image of the sample training object according to the spatial reference vector and the reference sample vector, and to align the image edges in the sample relationship image to the image nodes in terms of attributes according to the image attribute matrix corresponding to the sample relationship image, so as to obtain an updated edge degree matrix; wherein the image attribute matrix includes a point-edge relationship matrix, a node weight matrix and an edge weight matrix;

[0156] A prediction classification result determination module, used to determine the target sample vector of the sample training object according to the updated edge degree matrix and the image attribute matrix, and determine the prediction classification result of the multimodal classification model according to the target sample vector;

[0157] The model training module is used to determine the model loss value of the multimodal classification model according to the predicted classification result, the actual classification result of the corresponding sample training object, the spatial loss value and the attribute loss value, and train the multimodal classification model according to the model loss value.

[0158] Optionally, the edge degree matrix updating module includes:

[0159] A point-edge weight matrix determining unit, configured to determine a point-edge weight matrix according to the point-edge relationship matrix, the node weight matrix and the edge weight matrix;

[0160] The edge degree matrix updating unit is used to determine the node degree matrix and the edge degree matrix according to the point-edge relationship matrix or the point-edge weight matrix, and to determine the updated edge degree matrix according to the node degree matrix and the edge degree matrix.

[0161] Optionally, a vertex-edge weight matrix determination unit is used to:

[0162] For any image node in the sample relationship image, determine the point-edge association relationship between the image node and each image edge in the sample relationship image according to the point-edge relationship matrix;

[0163] Determine a matching result between the image node and any image edge associated with the image node, and determine the degree of point-edge association between the image node and the image edge according to the matching result, the node weight matrix and the edge weight matrix;

[0164] A point-edge weight matrix is ​​determined according to the point-edge association degree between each image node and each image edge in the sample relationship image.

[0165] Optionally, the prediction and classification result determination module is specifically used to:

[0166] According to the updated edge degree matrix, the edge weight matrix, the point-edge relationship matrix, the node degree matrix, and the point-edge weight matrix, the candidate sample vector of the sample training object is updated to obtain a target sample vector;

[0167] According to the target sample vector, the sample modality vector of each sample training object is determined, and according to the sample modality vector, the predicted classification result output by the multimodal classification model is determined.

[0168] Optional, model training module, specifically used for:

[0169] Determining a predicted result probability of the predicted classification result, and encoding the actual classification result to obtain an actual result probability;

[0170] Determining a basic loss value of the multimodal classification model according to the predicted result probability, the actual result probability, and the number of samples and the number of category labels of the sample training objects;

[0171] A spatial loss value and an attribute loss value are obtained, and a model loss value of the multimodal classification model is determined according to the basic loss value, the spatial loss value and the attribute loss value.

[0172] Optionally, the reference sample vector determination module is specifically used to:

[0173] Performing modality grouping on the candidate sample vectors to obtain a sample vector group, and determining a modality group contribution of the sample vector group; wherein the sample vector group includes candidate sample vectors under at least two candidate modalities;

[0174] According to the contribution of the modality group, a target modality is determined from the candidate modalities, and the candidate sample vector under the target modality is used as a spatial reference vector, and the candidate sample vectors under other candidate modalities except the target modality are used as spatial alignment vectors.

[0175] The multimodal data-based classification device provided in the embodiment of the present invention can execute the multimodal data-based classification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each multimodal data-based classification method.

[0176] In the technical solution of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of the data to be classified, text sample data, visual sample data and audio sample data involved all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0177] Embodiment 5

[0178] Figure 5 It is a structural diagram of an electronic device for implementing a classification method based on multimodal data provided by Embodiment 5 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0179] like Figure 5 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0180] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0181] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a classification method based on multimodal data.

[0182] In some embodiments, the classification method based on multimodal data can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the classification method based on multimodal data described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the classification method based on multimodal data in any other appropriate manner (for example, by means of firmware).

[0183] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0184] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0185] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0186] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0187] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0188] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0189] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0190] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A classification method based on multimodal data, characterized in that: include: Acquire data to be classified of the object to be classified in a candidate modality; wherein the candidate modality includes a text modality, a visual modality, and an audio modality; Inputting the data to be classified into a trained multimodal classification model to obtain an object classification result; The multimodal classification model is obtained by joint training based on text sample data, visual sample data and audio sample data of the sample training object.

2. The method according to claim 1, characterized in that: The multimodal classification model is trained based on the following method: Inputting text sample data, visual sample data and audio sample data of the sample training object into a pre-built multimodal classification model to obtain candidate sample vectors; wherein the candidate sample vectors include candidate text vectors, candidate audio vectors and candidate visual vectors; Determine a spatial reference vector and a spatial alignment vector in the candidate sample vector, and align the spatial alignment vector with the spatial reference vector in terms of spatial distribution to obtain a reference sample vector corresponding to the spatial alignment vector; Determine a sample relationship image of the sample training object according to the spatial reference vector and the reference sample vector, and align the image edges in the sample relationship image to the image nodes according to the image attribute matrix corresponding to the sample relationship image, so as to obtain an updated edge degree matrix; wherein the image attribute matrix includes a point-edge relationship matrix, a node weight matrix and an edge weight matrix; Determining a target sample vector of the sample training object according to the updated edge degree matrix and the image attribute matrix, and determining a predicted classification result of the multimodal classification model according to the target sample vector; According to the predicted classification result, the actual classification result of the corresponding sample training object, the spatial loss value and the attribute loss value, the model loss value of the multimodal classification model is determined, and the multimodal classification model is trained according to the model loss value.

3. The method according to claim 2, characterized in that The step of aligning the image edges in the sample relationship image to the image nodes in terms of attributes according to the image attribute matrix corresponding to the sample relationship image to obtain an updated edge degree matrix includes: Determine a point-edge weight matrix according to the point-edge relationship matrix, the node weight matrix and the edge weight matrix; A node degree matrix and an edge degree matrix are determined according to the point-edge relationship matrix or the point-edge weight matrix, and an updated edge degree matrix is ​​determined according to the node degree matrix and the edge degree matrix.

4. The method according to claim 3, characterized in that: The step of determining a point-edge weight matrix according to the point-edge relationship matrix, the node weight matrix and the edge weight matrix comprises: For any image node in the sample relationship image, determine the point-edge association relationship between the image node and each image edge in the sample relationship image according to the point-edge relationship matrix; Determine a matching result between the image node and any image edge associated with the image node, and determine the degree of point-edge association between the image node and the image edge according to the matching result, the node weight matrix and the edge weight matrix; A point-edge weight matrix is ​​determined according to the point-edge association degree between each image node and each image edge in the sample relationship image.

5. The method according to claim 2, characterized in that: The step of determining a target sample vector of the sample training object according to the updated edge degree matrix and the image attribute matrix, and determining a predicted classification result of the multimodal classification model according to the target sample vector, comprises: According to the updated edge degree matrix, the edge weight matrix, the point-edge relationship matrix, the node degree matrix, and the point-edge weight matrix, the candidate sample vector of the sample training object is updated to obtain a target sample vector; According to the target sample vector, the sample modality vector of each sample training object is determined, and according to the sample modality vector, the predicted classification result output by the multimodal classification model is determined.

6. The method according to claim 2, characterized in that Determining the model loss value of the multimodal classification model according to the predicted classification result, the actual classification result of the corresponding sample training object, the space loss value and the attribute loss value includes: Determining a predicted result probability of the predicted classification result, and encoding the actual classification result to obtain an actual result probability; Determining a basic loss value of the multimodal classification model according to the predicted result probability, the actual result probability, and the number of samples and the number of category labels of the sample training objects; A spatial loss value and an attribute loss value are obtained, and a model loss value of the multimodal classification model is determined according to the basic loss value, the spatial loss value and the attribute loss value.

7. The method according to claim 2, characterized in that The determining of the spatial reference vector and the spatial alignment vector in the candidate sample vector comprises: Performing modality grouping on the candidate sample vectors to obtain a sample vector group, and determining a modality group contribution of the sample vector group; wherein the sample vector group includes candidate sample vectors under at least two candidate modalities; According to the contribution of the modality group, a target modality is determined from the candidate modalities, and the candidate sample vector under the target modality is used as a spatial reference vector, and the candidate sample vectors under other candidate modalities except the target modality are used as spatial alignment vectors.

8. A classification device based on multimodal data, characterized in that: include: A module for acquiring data to be classified, used to acquire data to be classified of an object to be classified in a candidate mode; wherein the candidate mode includes a text mode, a visual mode and an audio mode; A classification result determination module, used for inputting the data to be classified into a trained multimodal classification model to obtain an object classification result; The multimodal classification model is obtained by joint training based on text sample data, visual sample data and audio sample data of the sample training object.

9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement a classification method based on multimodal data as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a classification method based on multimodal data as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Intelligent classification and grading method for multi-modal data set

    CN120850124A