A multi-modal feature fusion object classification method and system based on a two-tower structure
Through the multimodal feature fusion method based on the double-tower structure, the problems of insufficient expression capabilities and lack of modality are solved, and the effective fusion of multimodal data and the accuracy of target classification are achieved.
Patent Information
- Application Number
- CN202510001260.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing target classification technology has the problem of model mismatch due to insufficient single modal expression ability and lack of modality of the target object.
The multimodal feature fusion method based on the double-tower structure is adopted. By acquiring multimodal data, unsupervised and supervised data sets are constructed, the features of different modalities are extracted, and unsupervised and supervised training is carried out through the double-tower multimodal feature fusion model to achieve the fusion and query of multimodal features.
It effectively improves the generalization ability of the model in different scenarios, reduces the impact of missing details and differences in the characteristics of similar targets, and achieves accurate target classification.
Smart Images

Figure CN119397487B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target recognition, and in particular, to a multi-modal feature fusion target classification method and system based on a two-tower structure. Background Art
[0002] Image target classification refers to automatically determining the predetermined category to which the original image belongs through a computer program. Traditional image classification techniques mainly rely on manually designed features and classical machine learning algorithms. In recent years, the introduction of deep learning models, especially convolutional neural networks, has greatly promoted the development of image classification techniques. Deep learning models can automatically learn useful feature representations from the original image, greatly improving the accuracy of classification. In certain specific fields, target classification is also a crucial task, which involves identifying and distinguishing different types of targets for effective supervision, decision-making, and action. However, due to the complex and variable target environment, it is difficult to obtain complete information of a single modality for some target objects. For example, only a partial image or some description materials of the target object can be obtained, and the local target information is easy to be confused and difficult to represent the global information of the target object. Summary of the Invention
[0003] The present invention provides a multi-modal feature fusion target classification method and system based on a two-tower structure, which solves the problem that the existing target classification technology is not suitable for the model due to the insufficient single-modal expression ability of the target object and the lack of modality.
[0004] In a first aspect, an embodiment of the present invention provides a multi-modal feature fusion target classification method based on a two-tower structure, and the method includes the following processes:
[0005] Obtain multi-modal data, and construct an unsupervised data set and a supervised data set of multiple modalities based on the multi-modal data;
[0006] Extract data features of different modalities from the unsupervised data set and the supervised data set of multiple modalities to construct an unsupervised feature training set and a supervised feature training set;
[0007] Perform unsupervised training and supervised training on the two-tower multi-modal feature fusion model based on the unsupervised feature training set and the supervised feature training set;
[0008] Use the two-tower multi-modal feature fusion model after unsupervised training and supervised training to perform multi-modal feature query on the target object, and implement multi-modal target classification based on similarity retrieval.
[0009] In the above embodiments, the present invention constructs a dual - tower multi - modal feature fusion model to enable the dynamic fusion of data features of multi - modal data, so as to achieve the complementarity between different modal features, and solve the problems of insufficient expression ability of single - modal data and model mismatch caused by modal loss.
[0010] As some alternative embodiments of the present application, the multi - modal data includes text data, image data, and signal data.
[0011] In the above embodiments, the multi - modal data involved in the present invention has stronger expression ability and can achieve the complementarity between different modal data.
[0012] As some alternative embodiments of the present application, the process of constructing a multi - modal unsupervised dataset and a supervised dataset based on multi - modal data is as follows:
[0013] Obtain the original data of the multi - modal data; wherein, the original data of the multi - modal data includes the original data without class labels and the original data with class labels;
[0014] For the original data without class labels, construct enhanced data by deleting part of the information of the original data, and construct data pairs based on the original data and the enhanced data to obtain a multi - modal unsupervised dataset;
[0015] For the original data with class labels, construct data pairs based on the original data of two target objects to obtain a multi - modal supervised dataset.
[0016] In the above embodiments, through the construction of the unsupervised dataset and the supervised dataset, the present invention enables the dual - tower multi - modal feature fusion model to effectively achieve unsupervised learning and supervised learning.
[0017] As some alternative embodiments of the present application, the process of extracting data features of different modalities from the multi - modal unsupervised dataset and the supervised dataset is as follows:
[0018] If the modality of the data is image data, use a pre - trained image feature extraction model to extract image features from the image data to obtain image features;
[0019] If the modality of the data is text data, use a pre - trained text feature extraction model to extract text features from the text features to obtain text features;
[0020] If the modality of the data is signal data, use a pre - trained residual neural network to extract signal features from the signal data to obtain signal features.
[0021] In some alternative embodiments of the present application, the dual - tower multi - modal feature fusion model includes a feature input layer, a feature mapping layer, a feature normalization layer, a feature splicing layer, a feature fusion layer, and a similarity calculation layer.
[0022] In the above - mentioned embodiments, the present invention constructs a dual - tower multi - modal feature fusion model to input the data feature pairs of multi - modal data into the dual towers, and the model is trained and discriminated based on the similarity between the two object features, solving the problem that the existing models have poor adaptability to data of different scales.
[0023] In some alternative embodiments of the present application, the processes of unsupervised training and supervised training of the dual - tower multi - modal feature fusion model based on an unsupervised feature training set and a supervised feature training set are as follows:
[0024] Perform unsupervised training on the dual - tower multi - modal feature fusion model based on the data features of the unsupervised feature training set;
[0025] Perform supervised training on the dual - tower multi - modal feature fusion model after unsupervised training based on the data features of the supervised feature training set;
[0026] Wherein, the data features include image features, text features, and signal features.
[0027] In the above - mentioned embodiments, the present invention constructs a contrast - learning - based dual - tower multi - modal feature fusion model that is compatible with unsupervised learning and supervised learning, and realizes the training of the model on unlabeled data and labeled data through contrast learning, solving the problem of difficult acquisition of labeled data.
[0028] In some alternative embodiments of the present application, the process of performing unsupervised training on the dual - tower multi - modal feature fusion model based on the data features of the unsupervised feature training set is as follows:
[0029] Input the data features of the unsupervised training set into the dual - tower multi - modal feature fusion model;
[0030] Perform feature mapping, normalization, and feature splicing and fusion processing on the data features through the dual - tower multi - modal feature fusion model to obtain the multi - modal features of the target object;
[0031] Calculate the similarity of the multi - modal features of the target object through the dual - tower multi - modal feature fusion model to obtain the multi - modal feature similarity between the target objects, and construct a cross - entropy loss function according to the multi - modal feature similarity. Through iterative training of the cross - entropy loss function, the unsupervised training of the dual - tower multi - modal feature fusion model is realized.
[0032] In the above embodiments, the present invention adopts an unsupervised learning reinforcement model to enhance the robustness of the missing information, and adopts a dynamic feature extraction method to handle the missing modality cases, effectively improving the generalization ability of the model in different scenarios.
[0033] As some alternative embodiments of the present application, the process of supervised training for the dual-tower multi-modal feature fusion model after unsupervised training based on the data features of the supervised feature training set is as follows:
[0034] Input the data features of the supervised training set into the dual-tower multi-modal feature fusion model after unsupervised training;
[0035] Through the dual-tower multi-modal feature fusion model, perform feature mapping, normalization, and feature splicing and fusion processing on the data features to obtain the multi-modal features of the target object;
[0036] Through the dual-tower multi-modal feature fusion model, calculate the similarity of the multi-modal features of the target object to obtain the multi-modal feature similarity between the target objects, and construct a cross-entropy loss function based on the multi-modal feature similarity. Through the iterative training of the cross-entropy loss function, the supervised training of the dual-tower multi-modal feature fusion model is realized.
[0037] In the above embodiments, the present invention discriminates the target category based on the feature similarity between the target objects, and through the dual-tower structure, combines the supervised learning reinforcement model to enhance the perception ability of the relevant features of the same-class target features and the sensitivity to the differences of different-class target features, reducing the influence caused by the missing details and the differences of the same-class target features.
[0038] As some alternative embodiments of the present application, the process of using the dual-tower multi-modal feature fusion model after unsupervised training and supervised training to perform multi-modal feature query on the target object and realizing multi-modal target classification based on similarity retrieval is as follows:
[0039] Obtain the multi-modal data of the target object to be classified, and perform data feature extraction on the multi-modal data in different modalities to obtain the data features of the target object to be classified;
[0040] Input the data features of the target object to be classified into the dual-tower multi-modal feature fusion model after unsupervised training and supervised training. Through the dual-tower multi-modal feature fusion model, perform feature mapping, normalization, and feature splicing and fusion processing on the data features to obtain the multi-modal features of the target object to be classified;
[0041] Based on the multi-modal features of the target object, perform feature query to obtain the multi-modal features of several known categories, and select the category corresponding to the multi-modal feature with the highest similarity as the multi-modal target classification result through similarity calculation.
[0042] In the above embodiments, the present invention constructs a dual-tower architecture model, which can simultaneously achieve feature learning in supervised and unsupervised modes, and based on a large amount of unlabeled data, improve the basic learning ability of the model for domain data, and combine a small amount of labeled data to improve the model's perception ability of feature differences between different categories, breaking through the problem of difficult acquisition of data in certain specific application scenarios and achieving accurate target classification.
[0043] In a second aspect, the present invention provides a multi-modal feature fusion target classification system based on a dual-tower structure, and the system includes:
[0044] An original data acquisition unit, which is used to acquire multi-modal data and construct a multi-modal unsupervised data set and a supervised data set based on the multi-modal data;
[0045] A data feature construction unit, which is used to extract data features of different modalities from the multi-modal unsupervised data set and the supervised data set to construct an unsupervised feature training set and a supervised feature training set;
[0046] A model joint training unit, which performs unsupervised training and supervised training on the dual-tower multi-modal feature fusion model based on the unsupervised feature training set and the supervised feature training set;
[0047] A multi-modal target classification unit, which uses the dual-tower multi-modal feature fusion model after unsupervised training and supervised training to perform multi-modal feature query on the target object, and realizes multi-modal target classification based on similarity retrieval.
[0048] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the multi-modal feature fusion target classification method based on a dual-tower structure.
[0049] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the multi-modal feature fusion target classification method based on a dual-tower structure.
[0050] The beneficial effects of the present invention are as follows:
[0051] 1. The present invention constructs a dual-tower multi-modal feature fusion model, which uses unsupervised learning to strengthen the robustness of the model to missing information, and adopts a dynamic feature extraction method to cope with the missing situation of modalities, effectively improving the generalization ability of the model in different scenarios.
[0052] 2. The present invention discriminates the target category based on the feature similarity between target objects, and through a two - tower structure combined with supervised learning, it enhances the model's perception ability of relevant features of the same - type target features and the sensitivity to the differences in different - type target features, reducing the influence caused by missing details and differences in the same - type target features.
[0053] 3. By constructing a two - tower multi - modal feature fusion model, the present invention can simultaneously achieve feature learning in both supervised and unsupervised modes, and based on a large amount of unlabeled data, it enhances the model's basic learning ability for domain data, and combined with a small amount of labeled data, it enhances the model's perception ability of feature differences between different categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0055] Figure 1 It is a schematic structural diagram of a computer device for the hardware operating environment of the embodiments of the present invention;
[0056] Figure 2 It is a flowchart of the multi - modal feature fusion target classification method of the embodiments of the present invention;
[0057] Figure 3 It is a schematic structural diagram of the two - tower multi - modal feature fusion model of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0059] To solve the problem that the existing target classification technology has model mismatch due to the insufficient single - modal expression ability and modal missing of target objects. The embodiments of the present invention provide a multi - modal feature fusion target classification method and system based on a two - tower structure. Before introducing the specific technical solutions of the present application, the hardware operating environment involved in the embodiments of the present application will be introduced first.
[0060] Please refer to Figure 1 , Figure 1 It is a schematic structural diagram of a computer device for the hardware operating environment involved in the embodiments of the present invention.
[0061] As Figure 1As shown in the figure, the computer device may include: a processor, such as a Central Processing Unit (CPU), a communication bus, a user interface, a network interface, and a memory. Among them, the communication bus is used to realize the connection and communication between these components. The user interface may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface may also include a standard wired interface and a wireless interface. The network interface may be an optional wired interface or a wireless interface (such as a Wi-Fi interface). The memory may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory may also be a storage device independent of the aforementioned processor.
[0062] Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0063] As Figure 1 shown, the memory as a storage medium may include an operating system, a network communication module, a user interface module, and an electronic program module.
[0064] In Figure 1 the computer device shown, the network interface is mainly used for data communication with a network server; the user interface is mainly used for data interaction with a user; the processor and memory in the computer device of the present application may be arranged in the computer device. The computer device calls the electronic program stored in the data storage module through the processor and executes the multi-modal feature fusion target classification method based on the dual-tower structure provided in the embodiments of the present application.
[0065] Based on the hardware environment of the foregoing embodiments, an embodiment of the present application provides a multi-modal feature fusion target classification method based on a dual-tower structure. Please refer to Figure 2 , Figure 2 which is a flowchart of the multi-modal feature fusion target classification method. The method process is as follows:
[0066] (1) Obtain multi-modal data, and construct an unsupervised data set and a supervised data set based on the multi-modal data.
[0067] In the embodiments of the present invention, the multi-modal data includes, but is not limited to, text data, image data, signal data, etc.
[0068] In the embodiments of the present invention, the process of constructing a multi-modal unsupervised dataset and a supervised dataset based on multi-modal data is as follows:
[0069] (1.1) Obtain the original data of the multi-modal data; wherein, the original data of the multi-modal data includes the original data without class labels and the original data with class labels.
[0070] (1.2) For the original data without class labels, construct augmented data by deleting part of the information of the original data, and construct data pairs based on the original data and the augmented data to obtain a multi-modal unsupervised dataset; for example, construct augmented data by means of data augmentation such as deleting some pixels of the image, deleting some text description segments, etc., and construct data pairs according to the original data and the augmented data; wherein, the samples of the unsupervised dataset are represented as follows: ;
[0071] Wherein, represents the original data of the target object in modalities 1 to n of modality a, represents the augmented data of the target object in modalities 1 to n of modality b. If the target object lacks the data of a certain modality, NULL is used to replace the data of that modality.
[0072] (1.3) For the original data with class labels, construct data pairs based on the original data of two target objects to obtain a multi-modal supervised dataset; wherein, the samples of the supervised dataset are represented as follows: ;
[0073] Wherein, represents the original data of the target object in modalities 1 to n of modality a, represents the original data of the target object in modalities 1 to n of modality b.
[0074] Specifically, for the sample data of the multi-modal supervised dataset, mark the label of the sample data when storing the sample data. If the target object a and the target object b belong to the same category, mark the label of the sample data as True; otherwise, if the target object a and the target object b do not belong to the same category, mark the label of the sample data as False.
[0075] (2) Extract data features of different modalities from the multi-modal unsupervised dataset and the supervised dataset to construct an unsupervised feature training set and a supervised feature training set.
[0076] In the embodiments of the present invention, the process of extracting data features of different modalities is as follows:
[0077] (2.1) If the modality of the data If the data is image data, a pre-trained image feature extraction model ( ) is used to extract image features from the image data to obtain image features. Specifically, the image feature extraction model can also be a residual neural network ( ), a convolutional neural network ( ), etc.
[0078] Among them, extracting image features using a pre-trained image feature extraction model ( ) is expressed as:
[0079] ;
[0080] ;
[0081] (2.2) If the modality of the data is text data, a pre-trained text feature extraction model (ERNIE) is used to extract text features from the text data to obtain text features. Specifically, the text feature extraction model can also be a BERT model, etc.
[0082] Among them, extracting text features using a pre-trained text feature extraction model (ERNIE) is expressed as:
[0083] ;
[0084] ;
[0085] (2.3) If the modality of the data is signal data, a pre-trained residual neural network ( ) is used to extract signal features from the signal data to obtain signal features.
[0086] Among them, extracting signal features using a pre-trained residual neural network is expressed as:
[0087] In the embodiments of the present invention, for other modality data, a general feature extraction model in that modality can be used to extract relevant features. If the corresponding modality data of the object is missing, i.e., NULL, its corresponding feature vector is set to a zero vector.
[0088] (3) Unsupervised training and supervised training are performed on the two-tower multi-modal feature fusion model based on the unsupervised feature training set and the supervised feature training set.
[0089] In the embodiments of the present invention, the two-tower multi-modal feature fusion model includes a feature input layer, a feature mapping layer, a feature normalization layer, a feature splicing layer, a feature fusion layer, and a similarity calculation layer. Please refer to Figure 3 , Figure 3It is a schematic structural diagram of the dual - tower multi - modal feature fusion model.
[0090] In the embodiments of the present invention, the process of unsupervised training and supervised training of the dual - tower multi - modal feature fusion model based on the unsupervised feature training set and the supervised feature training set is as follows:
[0091] (3.1) Perform unsupervised training on the dual - tower multi - modal feature fusion model based on the data features of the unsupervised feature training set, where the data features include image features, text features, and signal features.
[0092] Specifically, the process of unsupervised training is as follows:
[0093] (3.11) Input the data features of the unsupervised training set into the dual - tower multi - modal feature fusion model; that is, access the data features through the feature input layer of the dual - tower multi - modal feature fusion model.
[0094] (3.12) Perform feature mapping, normalization, and feature splicing and fusion processing on the data features through the dual - tower multi - modal feature fusion model to obtain the multi - modal features of the target object;
[0095] Specifically, since the dimensions of data features of different modalities are often different, a single - layer feed - forward neural network (FFN) is used as the feature mapping layer to map the data features of different modalities into a vector space of the same dimension. That is, the original data and augmented data features of the samples in the unsupervised dataset are regarded as data features from different target objects. For the data features of the same modality of different target objects, a single - layer feed - forward neural network is used to achieve feature mapping:
[0096] ;
[0097] ;
[0098] Specifically, in order to avoid the influence of too large or too small features on the fusion during the fusion process, a feature normalization layer is used to adjust the feature scale of the data features to ensure the consistency of the feature vector scales between different modalities:
[0099] ;
[0100] ;
[0101] Specifically, a feature splicing layer is used to splice the data features of the same modality, and a multi - layer perceptron (MLP) is used as the feature fusion layer to fuse the spliced data features to construct the multi - modal features of the target object , :
[0102] ;
[0103] ;
[0104] (3.13) Calculate the similarity of the multimodal features of the target objects through the dual - tower multimodal feature fusion model to obtain the multimodal feature similarity between the target objects, that is, use the cosine similarity algorithm of the similarity calculation layer to calculate a single sample of the unsupervised data set of the multimodal feature similarity between two target objects:
[0105] ;
[0106] Specifically, after obtaining the multimodal feature similarity, construct a cross - entropy loss function according to the multimodal feature similarity:
[0107] ;
[0108] Specifically, through the iterative training of the cross - entropy loss function, the unsupervised training of the dual - tower multimodal feature fusion model is realized, so that the dual - tower multimodal feature fusion model can still judge the category according to the similarity of the existing features when some features of the target objects are lost, reduce the sensitivity of the dual - tower multimodal feature fusion model to feature loss, and improve the robustness of the dual - tower multimodal feature fusion model.
[0109] (3.2) Perform supervised training on the dual - tower multimodal feature fusion model after unsupervised training based on the data features of the supervised feature training set, where the data features include image features, text features, and signal features.
[0110] Specifically, the process of supervised training is as follows:
[0111] (3.21) Input the data features of the supervised training set into the dual - tower multimodal feature fusion model after unsupervised training.
[0112] (3.22) Through the dual - tower multimodal feature fusion model, perform feature mapping, normalization, and feature splicing and fusion processing on the data features to obtain the multimodal features of the target objects.
[0113] Specifically, the processes of feature mapping, normalization, and feature splicing and fusion processing in supervised training are the same as those in unsupervised training. For detailed content, refer to steps (2.11)~step (2.12).
[0114] (3.23) Through the dual - tower multimodal feature fusion model, calculate the similarity of the multimodal features of the target objects to obtain the multimodal feature similarity between two target objects of a single sample of the supervised data set and construct a cross - entropy loss function according to the multimodal feature similarity:
[0115] ;
[0116] Among them, when the sample is marked as True, ; when marked as False, ; Through iterative training based on the loss function, make the mapping features of objects in the same category as similar as possible, and as different as possible between objects in different categories, so as to achieve the supervised training of the two-tower multi-modal feature fusion model.
[0117] (4) Use the two-tower multi-modal feature fusion model after unsupervised training and supervised training to perform multi-modal feature queries on the target object, and achieve multi-modal target classification based on similarity retrieval.
[0118] Specifically, in the actual implementation process, use the two-tower multi-modal feature fusion model to construct multi-modal features for the target object of the category to be queried , retrieve the feature vectors of known category objects based on the multi-modal features, and select the object feature with the highest correlation with the feature of the target object of the category to be queried , and use the category of this object as the category of the target object to be queried.
[0119] ;
[0120] ;
[0121] In summary, the present invention constructs a two-tower multi-modal feature fusion model, and combines multi-modal data to perform unsupervised learning and supervised learning on the model, improves the basic learning ability of the model for domain data based on a large amount of unlabeled data, and combines a small amount of labeled data to improve the ability of the model to perceive feature differences between different categories, breaking through the problem of difficult data acquisition in some specific application scenarios.
[0122] In addition, in one embodiment, based on the same inventive concept as the foregoing embodiment, the embodiment of the present invention provides a multi-modal feature fusion target classification system based on a two-tower structure. The system corresponds one-to-one with the method of Embodiment 1. The system includes:
[0123] An original data acquisition unit, which is used to acquire multi-modal data and construct an unsupervised data set and a supervised data set in multi-modal based on the multi-modal data;
[0124] A data feature construction unit, which is used to extract data features of different modalities from the unsupervised data set and the supervised data set in multi-modal to construct an unsupervised feature training set and a supervised feature training set;
[0125] A model joint training unit, which performs unsupervised training and supervised training on the dual - tower multi - modal feature fusion model based on an unsupervised feature training set and a supervised feature training set;
[0126] A multi - modal target classification unit, which uses the dual - tower multi - modal feature fusion model after unsupervised training and supervised training to perform multi - modal feature query on the target object, and realizes multi - modal target classification based on similarity retrieval.
[0127] It should be noted that in this embodiment, each unit in the multi - modal feature fusion target classification system based on the dual - tower structure corresponds one by one to each step in the multi - modal feature fusion target classification method based on the dual - tower structure in the foregoing embodiment. Therefore, the specific implementation manners and achieved technical effects of this embodiment can refer to the implementation manners of the foregoing multi - modal feature fusion target classification method based on the dual - tower structure, and will not be elaborated here.
[0128] In addition, in one embodiment, the present application further provides a computer device, which includes a processor, a memory, and a computer program stored in the memory. When the computer program is run by the processor, it implements the method in the foregoing embodiment.
[0129] In addition, in one embodiment, the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is run by the processor, it implements the method in the foregoing embodiment.
[0130] In some embodiments, the computer - readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD - ROM; or it can be various devices including one or any combination of the above - mentioned memories. The computer can be various computing devices including intelligent terminals and servers.
[0131] In some embodiments, the executable instructions can be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as an independent program or being deployed as a module, component, sub - routine, or other unit suitable for use in a computing environment.
[0132] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0133] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0134] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0135] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0136] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory / random access memory, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a multimedia terminal device (which can be a mobile phone, a computer, a television receiver, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0137] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A multimodal feature fusion target classification method based on a dual-tower structure, characterized in that: The method comprises the following steps: Acquire multimodal data, and construct a multimodal unsupervised data set and a supervised data set based on the multimodal data; the multimodal data includes text data, image data, and signal data; The training samples of unsupervised and supervised datasets are represented as: s su,i =[M 1,a ,…,M n,a ;M 1,b ,…,M n,b ] Among them, M 1,a ,…,M n,a ;M 1,b , …, M n,b Represents multimodal data of target objects a and b of a single training sample; Extract data features of different modalities from multimodal unsupervised data sets and supervised data sets to construct unsupervised feature training sets and supervised feature training sets; Based on the unsupervised feature training set and the supervised feature training set, the dual-tower multi-mode feature fusion model is trained unsupervised and supervised; The dual-tower multi-mode feature fusion model includes a feature input layer, a feature mapping layer, a feature normalization layer, a feature splicing layer, a feature fusion layer and a similarity calculation layer; The process of unsupervised training of the dual-tower multi-mode feature fusion model based on the data features of the unsupervised feature training set is as follows: Input the data features of the unsupervised training set into the dual-tower multi-mode feature fusion model; The data features are mapped, normalized, and concatenated through a dual-tower multi-mode feature fusion model to obtain the multi-mode features of the target object. The multimodal features of the target objects are similarly calculated through the dual-tower multimodal feature fusion model to obtain the multimodal feature similarity between the target objects, and a cross-entropy loss function is constructed based on the multimodal feature similarity. The unsupervised training of the dual-tower multimodal feature fusion model is achieved through iterative training of the cross-entropy loss function. In unsupervised feature training, the formula for constructing the cross entropy loss function based on multimodal feature similarity is: Among them, sim i (a, b) represents the multimodal feature similarity between two target objects of a single sample; The process of supervised training of the dual-tower multi-mode feature fusion model after unsupervised training based on the data features of the supervised feature training set is as follows: Input the data features of the supervised training set into the dual-tower multi-mode feature fusion model after unsupervised training; The data features are mapped, normalized, and concatenated through a dual-tower multi-mode feature fusion model to obtain the multi-mode features of the target object. The multimodal features of the target objects are similarly calculated through the dual-tower multimodal feature fusion model to obtain the multimodal feature similarity between the target objects, and a cross-entropy loss function is constructed based on the multimodal feature similarity. The supervised training of the dual-tower multimodal feature fusion model is achieved through iterative training of the cross-entropy loss function. Among them, in supervised feature training, the formula for constructing the cross entropy loss function based on the multi-modal feature similarity is: Among them, sim i (a, b) represents the multimodal feature similarity between two target objects a and b in a single training sample. If target object a and target object b belong to the same category, the training sample s is marked. su,i True; otherwise, if the target object a and the target object b do not belong to the same category, the training sample s is marked su,i is False, when the training sample s su,i When marked as True, y i =1, when the training sample s su,i When the mark is False, y i =0; The dual-tower multimodal feature fusion model after unsupervised training and supervised training is used to perform multimodal feature query on the target object, and multimodal target classification is achieved based on similarity retrieval.
2. According to the multimodal feature fusion target classification method based on the double-tower structure of claim 1, it is characterized in that: The process of constructing multimodal unsupervised and supervised datasets based on multimodal data is as follows: Acquire original data of multimodal data; wherein the original data of the multimodal data includes original data without category labels and original data with category labels; For the original data without category labels, the enhanced data is constructed by deleting part of the original data information, and a data pair is constructed based on the original data and the enhanced data to obtain a multimodal unsupervised dataset; For the original data with category labels, data pairs are constructed based on the original data of two target objects to obtain a multimodal supervised dataset.
3. According to the multimodal feature fusion target classification method based on the double-tower structure of claim 1, it is characterized in that: The process of extracting data features of different modalities for multimodal unsupervised data sets and supervised data sets is as follows: If the modality of the data is image data, a pre-trained image feature extraction model is used to extract image features from the image data to obtain image features; If the modality of the data is text data, a pre-trained text feature extraction model is used to extract text features to obtain text features; If the mode of the data is signal data, a pre-trained residual neural network is used to extract signal features from the signal data to obtain signal features.
4. According to the multimodal feature fusion target classification method based on the double-tower structure of claim 3, it is characterized in that: The process of unsupervised training and supervised training of the dual-tower multi-mode feature fusion model based on the unsupervised feature training set and the supervised feature training set is as follows: Based on the data features of the unsupervised feature training set, the dual-tower multi-mode feature fusion model is unsupervisedly trained; Based on the data features of the supervised feature training set, the dual-tower multi-mode feature fusion model after unsupervised training is supervised trained; The data features include image features, text features and signal features.
5. According to the multimodal feature fusion target classification method based on the double-tower structure of claim 1, it is characterized in that: The dual-tower multimodal feature fusion model after unsupervised training and supervised training is used to perform multimodal feature query on the target object, and the process of multimodal target classification based on similarity retrieval is as follows: Acquire multimodal data of the target object to be classified, and extract data features of different modes from the multimodal data to obtain data features of the target object to be classified; The data features of the target object to be classified are input into the dual-tower multi-mode feature fusion model after unsupervised training and supervised training, and the data features are subjected to feature mapping, normalization, feature splicing and fusion processing through the dual-tower multi-mode feature fusion model to obtain the multi-mode features of the target object to be classified; A feature query is performed based on the multimodal features of the target object to obtain multimodal features of several known categories, and the category corresponding to the multimodal feature with the highest similarity is selected as the multimodal target classification result through similarity calculation.
6. A multimodal feature fusion target classification system based on a dual-tower structure, characterized in that: The system comprises: An original data acquisition unit, the original data acquisition unit is used to acquire multimodal data and construct a multimodal unsupervised data set and a supervised data set based on the multimodal data; the multimodal data includes text data, image data and signal data; The training samples of unsupervised and supervised datasets are represented as: s su,i =[M 1,a ,…,M n,a ;M 1,b ,…,M n,b ] Among them, M 1,a ,…,M n,a ;M 1,b , …, M n,b Represents multimodal data of target objects a and b of a single training sample; A data feature construction unit, wherein the data feature construction unit is used to extract data features of different modalities from a multimodal unsupervised data set and a supervised data set to construct an unsupervised feature training set and a supervised feature training set; A model joint training unit, wherein the model joint training unit performs unsupervised training and supervised training on the dual-tower multi-mode feature fusion model based on an unsupervised feature training set and a supervised feature training set; The dual-tower multi-mode feature fusion model includes a feature input layer, a feature mapping layer, a feature normalization layer, a feature splicing layer, a feature fusion layer and a similarity calculation layer; The process of unsupervised training of the dual-tower multi-mode feature fusion model based on the data features of the unsupervised feature training set is as follows: Input the data features of the unsupervised training set into the dual-tower multi-mode feature fusion model; The data features are mapped, normalized, and concatenated through a dual-tower multi-mode feature fusion model to obtain the multi-mode features of the target object. The multimodal features of the target objects are similarly calculated through the dual-tower multimodal feature fusion model to obtain the multimodal feature similarity between the target objects, and a cross-entropy loss function is constructed based on the multimodal feature similarity. The unsupervised training of the dual-tower multimodal feature fusion model is achieved through iterative training of the cross-entropy loss function. In unsupervised feature training, the formula for constructing the cross entropy loss function based on multimodal feature similarity is: Among them, sim i (a, b) represents the multimodal feature similarity between two target objects of a single sample; The process of supervised training of the dual-tower multi-mode feature fusion model after unsupervised training based on the data features of the supervised feature training set is as follows: Input the data features of the supervised training set into the dual-tower multi-mode feature fusion model after unsupervised training; The data features are mapped, normalized, and concatenated through a dual-tower multi-mode feature fusion model to obtain the multi-mode features of the target object. The multimodal features of the target objects are similarly calculated through the dual-tower multimodal feature fusion model to obtain the multimodal feature similarity between the target objects, and a cross-entropy loss function is constructed based on the multimodal feature similarity. The supervised training of the dual-tower multimodal feature fusion model is achieved through iterative training of the cross-entropy loss function. Among them, in supervised feature training, the formula for constructing the cross entropy loss function based on the multi-modal feature similarity is: Among them, sim i (a, b) represents the multimodal feature similarity between two target objects a and b in a single training sample. If target object a and target object b belong to the same category, the training sample s is marked. su,i True; otherwise, if the target object a and the target object b do not belong to the same category, the training sample s is marked su,i is False, when the training sample s su,i When marked as True, y i =1, when the training sample s su,i When the mark is False, y i =0; A multimodal target classification unit uses a dual-tower multimodal feature fusion model after unsupervised training and supervised training to perform multimodal feature query on the target object, and realizes multimodal target classification based on similarity retrieval.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the multimodal feature fusion target classification method based on the dual-tower structure as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method for multimodal feature fusion target classification based on a dual-tower structure as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Multi-modal information retrieval method, device and equipment, readable storage medium and computer program product
CN117909555A