Multi-modal data cross-modal matching method, system, equipment and medium
By introducing a multimodal data cross-modal matching method in the cross-modal matching technology, the problem of difficulty in dealing with the information asymmetry between weakly aligned data and modals in the prior art is solved, and higher matching accuracy and system stability are achieved.
Patent Information
- Application Number
- CN202510623704.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing cross-modal matching technology is difficult to effectively deal with the asymmetry between information density and coverage between modes in weakly aligned data and complex scenarios, resulting in a decrease in the stability of the matching results and the retrieval robustness cannot be guaranteed.
A multimodal data cross-modal matching method is proposed. By obtaining multimodal data, it is input to a preset cross-modal matching model. The training process includes obtaining the model training set, calculating the training loss value, performing correlation label correction, and optimizing the model to obtain the trained cross-modal matching model. This method is compatible with weakly aligned data and dynamically adjusts the information weight between modes.
It improves the cross-modal matching accuracy and system stability in complex application environments, and ensures the retrieval robustness of multimodal data matching.
Smart Images

Figure CN120123788A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimodal machine learning technologies, and in particular, to a cross-modal matching method, system, device, and medium for multimodal data. Background Art
[0002] With the development of information technology, the requirements for managing multiple devices have gradually increased. Since different devices have different data modalities, it is necessary to establish semantic associations between different modality data to achieve cross-modal retrieval and interaction. Currently, cross-modal matching technologies mainly map heterogeneous modality data to a unified common subspace for similarity measurement. In the application scenario of image-text matching, this specifically includes a coarse-grained matching method that aligns modalities by extracting the overall features of images and texts, and a fine-grained matching method that improves the matching accuracy by establishing one-to-one correspondences between image regions and text words. Therefore, current cross-modal matching technologies rely on strict alignment of training data.
[0003] However, in real scenarios, multimodal data exhibits characteristics of weak alignment or partial misalignment due to high acquisition costs, annotation noise, etc., which easily leads to overfitting or semantic deviation. Moreover, existing technologies are difficult to effectively balance the asymmetry of information density and coverage between modalities in complex scenarios, resulting in a decrease in the stability of matching results, and thus unable to guarantee the retrieval robustness of multimodal data matching. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0005] The main objective of the embodiments of this disclosure is to propose a cross-modal matching method, system, device, and storage medium for multimodal data, which can be compatible with weakly aligned data and dynamically adjust the information weights between modalities to improve the cross-modal matching accuracy and system stability in complex application environments.
[0006] The first aspect of the embodiments of this application provides a cross-modal matching method for a central controller. The method includes: Obtain target data, where the target data is a cross-modal data pair containing multiple cross-modal data; Input the target data into a preset cross-modal matching model to obtain a matching result output by the cross-modal matching model; Wherein, the training process of the cross-modal matching model includes: Obtain a model training set, where the model training set contains multimodal matching sample pairs and relevance labels; Input the model training set into an initial cross-modal matching model for training to obtain an initial matching result; Calculate the training loss value of each multi-modal matching sample pair in the model training set according to the initial matching result; Based on each of the training loss values, correct the relevance labels for each multi-modal matching sample pair in the model training set to obtain positive samples and negative samples; Optimize the initial cross-modal matching model according to the positive samples and negative samples until a trained cross-modal matching model is obtained.
[0007] In some embodiments of the present application, the calculation formula for calculating the training loss value of each multi-modal matching sample pair in the model training set according to the initial matching result includes: ; ; where is the th image in the model training set, is the text corresponding to in the model training set, is 's feature vector, is 's feature vector, is a preset boundary value, is the th initial negative sample text except , is the th initial negative sample image except , is and 's similarity score, is and 's similarity score, is and 's similarity score.
[0008] In some embodiments of the present application, the method of correcting the relevance labels for each multi-modal matching sample pair in the model training set based on each of the training loss values to obtain positive samples and negative samples includes: Using a variational Bayesian Gaussian mixture model, calculate the posterior value of the clean probability of each multi-modal matching sample according to the training loss value; According to the posterior value of the clean probability, divide the model training set into a clean subset and a noise subset; Perform label correction on the clean subset and the noise subset respectively to obtain the positive samples and the negative samples.
[0009] In some embodiments of the present application, optimizing the initial cross-modal matching model according to the positive samples and negative samples until a trained cross-modal matching model is obtained includes: Optimizing the positive samples and negative samples respectively through a preset initial decoupling loss function to obtain optimized positive samples and optimized negative samples; Calculating the similarity of each sample according to the optimized positive samples and optimized negative samples; Adjusting the weight factor of the preset initial decoupling loss function according to the similarity of each sample to obtain a decoupling loss function; Updating the initial cross-modal matching model based on the decoupling loss function according to the model training set to obtain the cross-modal matching model.
[0010] In some embodiments of the present application, respectively correcting the labels of the clean subset and the noise subset to obtain the positive samples and the negative samples includes: Correcting the labels of the clean subset according to the relevance label and the model prediction label of the clean subset to obtain first positive samples and first negative samples, where the model prediction label is obtained by matching the clean subset using the initial cross-modal matching model; Inputting the noise subset into a preset learning model to obtain corresponding prediction results; Correcting the noise subset according to the prediction results to obtain second positive samples and second negative samples; Taking the first positive samples and the second positive samples as the positive samples; Taking the first negative samples and the second negative samples as the negative samples.
[0011] In some embodiments of the present application, the calculation formula of the preset initial decoupling loss function includes: ; ; ; Wherein, and are respectively the weight factor of the th positive sample pair and the weight factor of the th negative sample pair, is a preset factor, and are respectively the boundary values of the positive sample pair and the negative sample pair, is the similarity score of the image and the matching text , is the similarity score of the image and the negative sample text The similarity score, is a preset sample factor, is the ideal optimal boundary of the positive sample pair, is the ideal optimal boundary of the negative sample pair, is the total number of image-text pairs in the model dataset.
[0012] In some embodiments of the present application, the multimodal matching sample pairs include image-text sample pairs. Based on the decoupled loss function, the initial cross-modal matching model is updated according to the model training set to obtain the cross-modal matching model, including: Input the image in the image-text sample pair into the initial cross-modal matching model to obtain a corresponding first matching result; Input the text in the image-text sample pair into the initial cross-modal matching model to obtain a corresponding second matching result; Use the decoupled loss function to optimize the initial cross-modal matching model according to the image-text sample pair, the first matching result, and the second matching result to obtain the cross-modal matching model.
[0013] To achieve the above object, a second aspect of the embodiments of the present invention provides a multimodal data cross-modal matching system, the system includes: An acquisition module, configured to acquire target data, where the target data is a cross-modal data pair including multiple types of cross-modal data; A matching module, configured to input the target data into a preset cross-modal matching model to obtain a matching result output by the cross-modal matching model; Wherein, the training process of the cross-modal matching model includes: Obtain a model training set, where the model training set includes multimodal matching sample pairs and relevance labels; Input the model training set into the initial cross-modal matching model for training to obtain an initial matching result; Calculate the training loss value of each multimodal matching sample pair in the model training set according to the initial matching result; Based on each of the training loss values, correct the relevance labels of each multimodal matching sample pair in the model training set to obtain positive samples and negative samples; Optimize the initial cross-modal matching model according to the positive samples and negative samples until a trained cross-modal matching model is obtained.
[0014] To achieve the above object, a third aspect of the embodiments of the present invention provides an electronic device, including: at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the above multi-modal data cross-modal matching method.
[0015] To achieve the above object, a fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute the above multi-modal data cross-modal matching method.
[0016] The embodiments of the present application provide a multi-modal data cross-modal matching method, which includes obtaining multi-modal data; inputting the multi-modal data into a preset cross-modal matching model to obtain a matching result output by the cross-modal matching model; the training process of the cross-modal matching model includes: obtaining a model training set; inputting the model training set into an initial cross-modal matching model for training to obtain an initial matching result; calculating a training loss value for each multi-modal matching sample pair in the model training set according to the initial matching result; based on each training loss value, correcting the relevance labels of each multi-modal matching sample pair in the model training set to obtain positive samples and negative samples; optimizing the initial cross-modal matching model according to the positive samples and negative samples until a trained cross-modal matching model is obtained, which can be compatible with weakly aligned data and dynamically adjust the information weights between modalities to improve the cross-modal matching accuracy and system stability in complex application environments.
[0017] It can be understood that the beneficial effects of the above second to fourth aspects compared with the related art are the same as those of the above first aspect compared with the related art. For details, please refer to the relevant descriptions in the above first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, where: Figure 1 is a flowchart of a multi-modal data cross-modal matching method provided by an embodiment of the present application; Figure 2 is a structural diagram of a multi-modal data cross-modal matching training model provided by an embodiment of the present application; Figure 3 is a structural diagram of a multi-modal data cross-modal matching system provided by an embodiment of the present application; Figure 4 is a hardware structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0019] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.
[0020] In the description of the present application, if the first, second, etc. are described, they are only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.
[0021] In the description of the present application, it should be understood that for the orientation description, such as up, down, etc., the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.
[0022] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.
[0023] With the development of information technology, the requirements for managing multiple devices are gradually increasing. Since different devices have different data modalities, it is necessary to establish semantic associations between different modality data to achieve cross-modal retrieval and interaction. As an important technology in the field of multi-modal learning, the core of cross-modal matching technology lies in mapping data from different modalities to the same common subspace, so as to be able to retrieve data of other modalities related to it by using the input of one modality. The current cross-modal matching technology mainly maps heterogeneous modality data to a unified common subspace for similarity measurement. In the application scenario of image-text matching, the existing image-text matching methods can be roughly divided into two categories: one is coarse-grained matching, that is, matching is achieved by calculating the overall features of images and texts; the other is fine-grained matching, which focuses on the fine-grained correspondence between local regions of images and words in texts.
[0024] However, in the actual application scenarios of cross-modal matching technology, researchers still face some technical problems. First of all, most cross-modal matching methods are based on the assumption of highly aligned training data. However, in reality, due to the high cost of data collection and dataset annotation, it is often impossible to obtain a dataset that meets expectations for training. In addition, data differences and information imbalance between modalities in complex scenarios will also affect the matching results, which will ultimately lead to unstable system performance.
[0025] Therefore, in real scenarios, multi-modal data exhibits weak alignment or partial misalignment characteristics due to problems such as high acquisition costs and annotation noise, which is prone to overfitting or semantic deviation. Moreover, existing technologies are difficult to effectively balance the asymmetry of information density and coverage between modalities in complex scenarios, resulting in a decline in the stability of matching results, and thus unable to guarantee the retrieval robustness of multi-modal data for matching.
[0026] Based on this, the embodiments of this application provide a cross-modal matching method, system, electronic device and medium for multi-modal data, aiming to be compatible with weakly aligned data and dynamically adjust the information weights between modalities to improve the matching accuracy and system stability in complex application environments.
[0027] The cross-modal matching method, system, electronic device and medium provided by the embodiments of this application are specifically described through the following embodiments. First, the cross-modal matching method for multi-modal data in the embodiments of this application is described.
[0028] The embodiments of this application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0029] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0030] The cross-modal matching method for multi-modal data provided by the embodiments of this application relates to the field of multi-modal machine learning technology. The cross-modal matching method for multi-modal data provided by the embodiments of this application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the cross-modal matching method for multi-modal data, etc., but is not limited to the above forms.
[0031] This application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0032] It should be noted that in each specific embodiment of this application, when it comes to relevant processing that needs to be carried out based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.
[0033] For this reason, refer to Figure 1, embodiments of the present application provide a multimodal data cross-modal matching method. This method is applied to a central controller, which can be a server, an electronic device, a mobile terminal, etc., and is not specifically limited here. The method includes the following steps S110 to step S120.
[0034] Step S110: Obtain target data, where the target data is a cross-modal data pair containing multiple types of cross-modal data.
[0035] In this step, the target data can be real-time collected cross-modal data (such as images captured by a camera and text input by voice), or historical data called from a database, such as image data or text data, and then match the target data according to user needs.
[0036] Specifically, it refers to the cross-modal data pair that needs to be processed in actual applications, such as the image and text to be matched, video and audio, etc. In the actual deployment process of the cross-modal matching system, it is necessary to perform standardization processing on the input target data (image modal data / text modal data) to ensure that different modal inputs have consistent encoding specifications and feature space expression capabilities, so as to ensure the reliability and stability of the matching effect.
[0037] In some embodiments, for image data, preprocessing is performed by scaling to a unified resolution (such as 224×224) to eliminate the influence of different source image size differences on feature extraction. Then, a pre-trained Faster-RCNN (Faster Region-based Convolutional Neural Network) network is used to perform region target detection on the image, extract the target region features in the image, and map all region features to a feature vector representation of a unified dimension through a fully connected layer, which is convenient for subsequent cross-modal similarity calculation. For text data, standardization preprocessing such as word segmentation, stop word removal, and unified lowercase is performed on the text samples to avoid text noise interference. Then, a deep learning model based on gated recurrent units (Bidirectional Gated Recurrent Unit, Bi-GRU) is used to perform bidirectional semantic encoding on the text, capture phrase-level and sentence-level context semantics, and encode the text into a feature vector representation of a unified dimension.
[0038] Furthermore, align the cross-modal features by projecting the extracted image features and text features into the same common feature space respectively to unify the similarity metric. Specifically, it includes calculating the similarity score between the image-text feature pairs using the cosine similarity function, providing a basis for subsequent retrieval, matching, and optimization.
[0039] Furthermore, the data collected automatically or accessed online may have problems such as incorrect labels or inaccurate modal correspondence.
[0040] Step S120: Input the target data into a preset cross-modal matching model to obtain the matching result output by the cross-modal matching model.
[0041] In this step, the target data is input into a preset cross-modal matching model to obtain the matching result output by the trained cross-modal matching model. The trained model is used to achieve the associated matching of cross-modal data and output the matching probability or similarity score.
[0042] Specifically, in practical applications, the obtained initial target data may contain noisy or incompletely annotated samples. Therefore, after the initial target data is preprocessed (such as image normalization, text tokenization and encoding), it needs to be input into the model to perform cross-modal matching on the characteristics of the target data through the model, so as to achieve efficient, robust and accurate cross-modal correlation analysis of multi-modal data.
[0043] In some embodiments of the present invention, as Figure 2 shown, first, an image-text matching model is constructed through multiple components. Among them, the visual branch uses Faster-RCNN to extract the target region features and maps them to the semantic space through a fully connected layer; the text branch constructs a hierarchical representation based on Bi-GRU to capture phrase-level and semantic-level semantics. At the same time, to enhance the cross-modal alignment effect, a graph reasoning module is designed to calculate the inter-modal correlation matrix, and the local fine-grained features and global semantic features are fused through an attention mechanism to establish a joint embedding space to obtain a similarity matrix. In addition, to address the noise problem, a variational Bayesian Gaussian mixture model is introduced in the model construction to dynamically evaluate the sample confidence and improve the performance of the model under noisy data.
[0044] Furthermore, based on a common cross-modal matching framework, first, a variational Bayesian Gaussian mixture model is used to accurately model the sample loss distribution, determine the clean subset and the noise subset, so as to obtain the probability density function for adaptively calculating the optimization boundary of the sample, that is, for soft label estimation, and then calculate the soft boundary weight factor; then, a decoupled optimization loss function is constructed based on asymmetric learning, and the learning processes of positive and negative samples are independently adjusted in combination with the adaptive boundary, accurately punishing potential noisy positive samples while ensuring that the optimization of negative samples is not disturbed; finally, a two-way matching strategy is used to train the model to obtain a noise-robust cross-modal matching model, improving the cross-modal matching performance in the case of data noise correlation, solving the problems of multi-modal data noise correlation and insufficient learning of positive and negative samples in cross-modal matching, and effectively establishing the association between different modalities through the adaptive boundary calculation based on the variational Bayesian Gaussian mixture model and the asymmetric learning of positive and negative samples, enhancing the performance of downstream tasks such as image-text and video-text cross-modal matching.
[0045] Specifically, as Figure 2As shown, in asymmetric learning, an image or text sample is selected as an anchor point, and then positive and negative sample pairs are constructed based on the anchor point. Among them, the positive sample pair is the given image-text pair in the dataset, and the negative sample pair is composed of the anchor point and a non-paired sample of another modality (for example, assuming the anchor point is an image sample, then a text sample that is not paired with the anchor point is selected to form a negative sample pair). True positive indicates the case where the positive sample pair is a semantically similar pair, and false positive indicates the case where the positive sample pair is a semantically dissimilar pair.
[0046] The following explains the specific training steps of the cross-modal matching model: Step S210: Obtain a model training set, which includes multi-modal matching sample pairs and relevance labels; Step S220: Input the model training set into the initial cross-modal matching model for training to obtain an initial matching result; Step S230: Calculate the training loss value of each multi-modal matching sample pair in the model training set according to the initial matching result; Step S240: Based on each training loss value, correct the relevance label of each multi-modal matching sample pair in the model training set to obtain positive samples and negative samples; Step S250: Optimize the initial cross-modal matching model according to the positive samples and negative samples until a trained cross-modal matching model is obtained.
[0047] In this step, a model training set for training the model is obtained. Among them, the model training set is specifically a dataset of multi-modal matching sample pairs and relevance labels.
[0048] Specifically, the multi-modal matching sample pair includes a heterogeneous data set containing various different types of information, usually composed of data from different modalities (such as images and texts). Common types include visual data (such as images, videos), auditory data (such as speech, audio), text data (such as natural language descriptions), sensor data (such as temperature, acceleration), etc. The relevance label indicates whether the current cross-modal matching sample pair matches (that is, whether they describe the same content), and is used to train and verify the matching performance of the model, providing a supervision signal for the model to learn the association relationship between modalities.
[0049] Furthermore, input the model training set into the initial cross-modal matching model for training to obtain an initial matching result. Among them, the initial cross-modal matching model is a pre-constructed matching model.
[0050] Specifically, each multimodal sample is used to train the initial cross-modal matching model, so that the initial cross-modal matching model learns to map information of different modalities into a common space, and then performs matching according to the similarity between modalities to obtain an initial matching result, thereby reflecting the understanding degree of the initial cross-modal matching model on the relationship between sample pairs in an unoptimized state through the initial matching result.
[0051] Furthermore, according to the initial matching result, the training loss value of each multimodal matching sample pair in the model training set is calculated. Among them, the loss value quantifies the gap between the model prediction and the actual label, and is an important indicator for evaluating the model performance, thus providing a basis for subsequent adjustment.
[0052] Specifically, the calculation formula for the training loss value of each multimodal matching sample pair in the model training set according to the initial matching result includes: ; ; where is the th image in the model training set, is the text corresponding to in the model training set, is 's feature vector, is 's feature vector, is a preset boundary value, is the th initial negative sample text except , is the th initial negative sample image except , is and 's similarity score, is and 's similarity score, is and 's similarity score.
[0053] Furthermore, the calculated training loss value is used to distinguish the quality of each sample pair, so as to more accurately define the positive and negative sample sets, laying the foundation for targeted model optimization. For example, a low loss value may mean a high-quality positive sample (a correctly matched sample pair), while a high loss value may be a noise or incorrect negative sample. Then, the model parameters are refined according to the identified positive and negative samples to reduce errors and enhance the generalization ability of the model. Specifically, different processing strategies for positive and negative samples (such as the design of asymmetric loss functions) are used to more effectively guide model learning, so that it can not only perform well on clean data, but also better cope with noisy data. Finally, after multiple iterative optimizations, a cross-modal matching model with excellent performance and strong robustness is finally obtained.
[0054] In some embodiments, the model training set is preferably a manually annotated dataset and an automatically collected dataset, wherein: the manually annotated dataset is preferably a graphic recognition dataset (flickr30k dataset) or a visual dataset (Microsoft Common Objects in Context, MS-COCO), which is used to build a controlled noise experimental environment; the automatically collected dataset is preferably CC152K, which is used to simulate real scenes.
[0055] In some embodiments, the cross-modal matching model is an image-text matching model, and its training process includes: first obtaining a Model training set of image-text pairs ,in, is an image sample, For text samples, ∈[0,1] represents the cross-modal matching relevance label, which is preferably a binary label in this embodiment. indicates that the image-text pair is matched, Indicates that the image-text pair is not matched.
[0056] Further, the image-text matching model is preliminarily trained, wherein the preliminarily training process preferably first calculates the image-text pair training loss value of each sample pair by triple loss , and use the memory effect of deep neural networks to capture the loss distribution characteristics of clean samples to obtain loss calculation results .
[0057] Furthermore, the image-text matching model is warmed up. During the warm-up training, for each image-text pair ,The image branch uses Faster-RCNN to extract the target area features, and the text branch uses Bi-GRU to construct a hierarchical representation, capture phrase-level and semantic-level semantics, and generate text feature representation.
[0058] Furthermore, project the features of the encoded image and text into a common representation space and calculate the image-text pair loss , where is the image encoder for extracting image features; is the text encoder for extracting text features, is the similarity function for calculating the similarity between image and text features, is an abbreviation of, representing the triple , is the total number of image-text pairs in the model training set.
[0059] Specifically, the specific formula for calculating the loss is: ; ; where is the feature vector of the image , is the feature vector of the text , the numerator is the dot product of the two feature vectors, and the denominator is the product of their magnitudes; is a hyperparameter greater than 0, called the margin value, which is used to control the similarity score gap between positive and negative sample pairs and is an important adjustment parameter for model training, represents the th image in the model training set; represents the text corresponding to ; represents any negative sample (mismatch) text other than ; represents any negative sample (mismatch) image other than ; calculates the similarity score between the image and the text , and are the similarity scores between the image and the negative sample text , the negative sample image and the text respectively. The in the formula is a truncation operation defined as , representing taking the larger value between and 0. Calculate the loss using the initial memory effect of the deep neural network, that is, the characteristic of tending to fit simple samples first and resulting in a low loss for clean samples, providing a basis for distinguishing the nature of samples in the subsequent stage.
[0060] Furthermore, a variational Bayesian Gaussian mixture model is used to model the sample loss distribution, and by introducing the prior distribution of parameters and Bayesian inference, the posterior probability of the loss is calculated As the clean probability of the
[0061] th sample, it lays a foundation for the subsequent accurate division of the dataset Preferably, a probability model containing Gaussian components is established, where = 0 corresponds to the clean sample distribution, and ≥ 1 corresponds to the noise sample distribution. Then, the posterior probability is calculated through variational Bayesian inference. When the posterior probability of the sample is higher than the threshold
[0062] , it is determined as a clean sample Specifically, it is preset that the parameters of the variational Bayesian Gaussian mixture model, that is, the mean of each Gaussian distribution component, follows a normal distribution, and the covariance follows an inverse Wishart distribution. Among them, the variational Bayesian Gaussian mixture model introduces the prior distribution of model parameters and Bayesian inference. The mean and covariance refer to the mean and variance obtained by fitting each Gaussian distribution component in the clustering modeling of all sample triple loss values
[0063] Furthermore, the posterior distribution is approximated by the following formula ; In the formula, is the set of loss values of all image - text pairs , is the variational Bayesian Gaussian mixture model, representing the set structure of the model is the hidden variable of the model , representing the soft probability of which Gaussian component each sample loss belongs to represents the set of hidden variables belonging to the k - th Gaussian component is the set of parameters of the variational Bayesian Gaussian mixture model, including the parameters of each Gaussian component is the specific parameter of the th Gaussian component, is the component weight, is the mean, is the covariance is the approximate posterior distribution of the model parameters is the solution of the variational Bayesian Gaussian mixture model. Through the expectation maximization algorithm, for example, variational Bayes expectation maximization (VBEM), based on the set of loss values optimize and solve.
[0064] Specifically, the variational Bayesian Gaussian mixture model uses the current parameters to calculate the soft probability that each sample belongs to each Gaussian component The variational Bayesian Gaussian mixture model uses the calculated soft probability to update the parameters of the Gaussian component and solve iteratively.
[0065] Furthermore, after being processed by the variational Bayesian Gaussian mixture model, the loss of each sample is fitted. The formula is: ; where is the loss of the th sample pair, =2 means that the samples are divided into two categories: clean samples and noise samples, is the probability density function of the th component (i.e., clean sample or noise sample), indicating the proportion of this probability component in the entire dataset.
[0066] Finally, based on the above calculation results, according to the following Bayes' formula, calculate the posterior probability and use it as the clean probability of the th sample: ; In the formula, , =1 means that the image-text pair is a clean sample, =0 means it is a noise sample, is the prior probability that the sample belongs to a certain category (clean or noise).
[0067] It is calculated through the following formula: ; In the formula, is the probability that the loss is when the sample belongs to a certain category.
[0068] It is calculated through the following formula: ; In the formula, is the loss of The probability is calculated through the total probability formula. The posterior probability calculated therefrom represents the probability that the sample belongs to the clean sample component under the condition that the loss value of the known sample is . The higher the probability value, the greater the probability that the sample belongs to the clean sample, quantifying the probability that each sample is a clean sample and laying a foundation for accurately dividing the clean / noise data set.
[0069] Furthermore, according to the calculated posterior probability and the set threshold (usually taken as 0.5), the training data set is divided into a clean subset and a noise subset. Different adaptive prediction functions are used for label correction of the two subsets respectively, and at the same time, a co-training strategy is adopted to alleviate the accumulation of training errors caused by the model using its own prediction results.
[0070] Specifically, according to the calculated posterior probability and the set threshold (usually taken as 0.5), the data set is divided into a clean subset and a noise subset , specifically: ; ; For the clean subset and the noise subset, different adaptive prediction functions are used to reassign the sample labels according to the subset characteristics, aiming to convert the original discrete binary labels {0, 1} into continuous soft labels in the interval [0, 1]. Among them, the binary labels are the existing relevance annotation information in the model training set .
[0071] Preferably, for the clean subset, the original label of the sample and the prediction result of the model are used to correct together, and the corresponding adaptive prediction function is: ; where represents the soft label of the sample in the clean subset, is the confidence weight of the current sample, used to balance the original label and the prediction result of the model, is the original binary label taken from the model training set, which can take values of 0 or 1, It is the predicted probability of Model A or B, representing the predicted probability that the current sample pair is a matching pair, depending on whether it is Network A or Network B. Here, Network A or B are two A / B networks with the same structure but different parameter initialization methods (i.e., two independent asymmetric similarity learning models), and then they are co-trained through the A / B networks, supervised with each other, and predicted respectively. and .
[0072] For the noise subset, discard the original labels and correct them according to the prediction results of the A / B networks. The specific formula is: ; In the formula, is the soft label of the sample in the noise subset, is the predicted probability of Network A for the current sample, is the predicted probability of Network B for the current sample, and take the average of the two as the soft label of the probability that the sample is a matching sample.
[0073] Furthermore, based on the image-text pairs in the model training set and the prior distribution knowledge of the preset model parameters, design an asymmetric decoupled optimization loss , impose a dynamic boundary and a penalty factor on the positive sample pairs, set an independent boundary and a weight factor on the negative sample pairs, and then calculate the gradients of the loss function for the positive sample features and the negative sample features respectively, establishing an asymmetric optimization mechanism. When calculating the positive sample gradient, introduce an adaptive upper boundary to effectively optimize the positive samples.
[0074] Specifically, based on the image-text pairs in the model training set and the prior distribution knowledge of the model parameters, construct a decoupled optimization loss function , and the formula is: ; In the formula, and are the weight factors of the th positive sample pair and the th negative sample pair respectively, is the preset sample factor, is the ideal optimal boundary of the positive sample pair, is the ideal optimal boundary of the negative sample pair, is the total number of image-text pairs in the model dataset. is a scaling factor used to control the scale of the loss function, affecting the convergence speed and performance of model training. and represent the boundary values of positive and negative sample pairs respectively. In this embodiment, is the mean of the clean sample distribution , which can be adaptively adjusted according to the clean probability of the samples to adjust to meet . is the mean of the clean sample distribution . and are non - negative weight factors used to adjust the importance of positive and negative samples in loss calculation. Among them, is positively correlated with the noise intensity of the samples, finally making the noise samples obtain higher penalty weights. is the image and the matching text 's similarity score. is the image and the negative sample text 's similarity score.
[0075] In some embodiments, positive and negative samples are processed separately, and an independent optimization goal is set for each. Among them, a boundary constraint term is introduced into the positive sample gradient calculation to suppress the gradient propagation of potential noisy positive samples; a normalization weighting mechanism is adopted for the negative sample gradient calculation to retain the supervision signals of effective negative samples. In addition, by dynamically adjusting the boundary parameters , in the loss function, it is adapted to the current training - stage noise distribution, and the regularization coefficient is adaptively adjusted according to the validation - set performance to balance the model's fault - tolerance ability and fitting ability.
[0076] Furthermore, by decoupling and optimizing the loss function the non - negative weight factors are calculated to achieve precise optimization of positive and negative sample pairs: ; In the formula, and respectively represent the weight factor of the th positive sample pair and the soft - boundary weight factor of the th negative sample pair. is a coefficient determined by the soft label and the relaxation factor. The relaxation factor here is used to relax the upper - bound limit of the positive - sample similarity, preventing high - confidence samples from being misjudged as noise due to overly strict boundaries. is the ideal optimal boundary of the positive sample pair. is the ideal optimal boundary for negative sample pairs, is the image and its corresponding text between the similarity scores, is the image and the negative sample text between the similarity scores.
[0077] Furthermore, to optimize the model and enhance the utilization of negative sample supervision, a decoupled optimization loss function is used to measure the difference between the model prediction and the real situation, and a weight factor is combined to reflect the importance of different samples. The sensitivity of the decoupled optimization loss function to the changes of positive and negative samples is calculated respectively.
[0078] Specifically, first calculate the partial derivatives of the decoupled optimization loss function with respect to the similarity of positive and negative samples, which are used to represent the sensitivity of positive and negative samples respectively; then compare the sensitivities. The larger the sensitivity value, the greater the impact of the sample on the current loss, and it needs to be given priority attention; finally, the image-text matching model dynamically adjusts the weight allocation according to the sensitivity, so that the matching degree of high-sensitivity positive samples increases and the matching degree of negative samples decreases, and an initially optimized image-text matching model is obtained, so as to achieve the purpose of clarifying the impact of positive and negative sample changes on the model loss by calculating the sensitivity, and then adjusting the model parameters according to these changes, so that the matching model can learn the negative sample supervision information more fully and improve the performance and generalization ability.
[0079] Furthermore, perform bidirectional matching training from image to text and from text to image, synchronously optimize the model parameters, and finally dynamically adjust the model hyperparameters according to the evaluation results to improve the cross-modal matching robustness in a noisy environment.
[0080] Specifically, start training with the initially optimized image-text matching model. Here, the model parameters mainly refer to the weights and biases of each layer of the neural network. Among them, the weights determine the strength of the connection between neurons and affect the transmission and transformation of information in the network; while the biases provide additional learnable parameters for the activation of neurons, which helps the model better fit the data. At the same time, to ensure the consistent performance of the image and text modalities, a bidirectional matching strategy is adopted, and text→image and image→text retrievals are performed simultaneously, and the asymmetric dynamic matching losses in both directions are calculated respectively to jointly optimize the initially optimized image-text matching model.
[0081] In some embodiments, in the NVIDIA RTX 3090 + Pytorch 2.0.1 environment, the Flickr30K dataset is warmed up for 5 epochs, trained for a total of 40 epochs, the MS-COCO and CC152K are each warmed up for 10 epochs, and trained for a total of 30 and 40 epochs respectively. The batch size is 180, the initial learning rate is 0.0002, and the Adam optimizer is used to adjust the model parameters to improve the cross-modal matching effect. Furthermore, the recall rate R@K (K = 1, 5, 10) and the total recall rate rSum are used to evaluate the image-text matching performance of the image-text and text-image retrieval models. Among them, the recall rate R@K refers to the proportion of correct matches in the top K results retrieved given a query, and the calculation formula is:
[0082] And rSum is the sum of the three recall rates of R@K (K = 1, 5, 10), that is, rSum = R@1 + R@5 + R@10, which further provides a quantitative standard for measuring the cross-modal matching ability of the model after dealing with the problem of noisy correspondences.
[0083] Furthermore, according to the evaluation results, the hyperparameters of the image-text matching model are adjusted (such as etc.), to improve the performance of the image-text matching model in a noisy environment, ensure stable and accurate retrieval of cross-modal information, effectively handle the problem of noisy correspondences, and enhance the robustness and matching accuracy of the model.
[0084] In some embodiments, in step S240, based on each training loss value, the relevance label of each multi-modal matching sample pair in the model training set is corrected to obtain positive and negative samples, including the following steps S310 to step S330: Step S310: Use the variational Bayesian Gaussian mixture model to calculate the clean probability posterior value of each multi-modal matching sample according to the training loss value; Step S320: According to the clean probability posterior value, divide the model training set into a clean subset and a noise subset; Step S330: Correct the labels of the clean subset and the noise subset respectively to obtain positive and negative samples.
[0085] In this embodiment, first, based on the preliminarily trained cross-modal matching model, using the memory effect of the deep neural network, clean samples are preferentially fitted, and the loss value of each multi-modal matching sample pair (such as an image-text pair) is calculated, providing a basis for subsequent identification of noise samples. Then, the variational Bayesian Gaussian mixture model is used to model these loss values. Among them, the variational Bayesian Gaussian mixture model estimates the probability that each sample belongs to different Gaussian components by introducing a prior distribution of parameters and Bayesian inference to obtain the posterior probability distribution of each sample loss value.
[0086] Furthermore, based on the results of the variational Bayesian Gaussian mixture model, the posterior probability of each sample as a clean sample is calculated to divide the model set according to the posterior value of the clean probability. Preferably, the division is made by comparing the clean probability with a preset threshold. In this embodiment, 0.5 is selected as the threshold. If the posterior value of the clean probability of a sample is greater than or equal to 0.5, it is classified into the clean subset; otherwise, it is classified into the noise subset, thus achieving the purpose of dividing the entire training set into a clean subset and a noise subset.
[0087] Furthermore, for the samples in the clean subset, an adaptive prediction function is used to combine the original binary label and the model prediction result for label correction to correct potential mislabeling and enhance the model's learning of this part of high-quality data; for the samples in the noise subset, since there may be large errors in the original labels, the prediction results of the model A / B network are directly used for label correction. Preferably, a co-training strategy is adopted, that is, two networks (A and B) with the same structure but different initializations are trained simultaneously to reduce the error accumulation caused by the self-reinforcement of a single model.
[0088] Furthermore, for each multi-modal matching sample pair, a more accurate label (positive sample or negative sample) is obtained, which not only improves the data quality of model training but also provides a basis for subsequent construction of an asymmetric decoupled optimization loss function and optimization of positive and negative sample pairs, thereby enhancing the robustness and accuracy of the cross-modal matching model in a noisy environment and improving the overall model performance.
[0089] In some embodiments, in step S330, label correction is performed on the clean subset and the noise subset respectively to obtain positive samples and negative samples, including the following steps S410 to S450: Step S410: Perform label correction on the clean subset according to the relevance label and the model prediction label of the clean subset to obtain the first positive samples and the first negative samples. The model prediction label is obtained by matching the clean subset using the initial cross-modal matching model. Step S420: Input the noise subset into a preset learning model to obtain the corresponding prediction results. Step S430: Correct the noise subset according to the prediction results to obtain second positive samples and second negative samples; Step S440: Use the first positive samples and the second positive samples as positive samples; Step S440: Use the first negative samples and the second negative samples as negative samples.
[0090] In this embodiment, for each sample in the clean subset, the pre - trained cross - modal matching model is used to calculate the predicted label. The predicted label of the model is combined with the original relevance label in the clean subset, and the label correction is performed through an adaptive prediction function to obtain the first positive samples and the first negative samples. Among them, the predicted label is obtained based on the initial cross - modal matching model according to the clean subset, reflecting the matching performance of the initial cross - modal matching model on whether the current sample is a matching pair (i.e., positive sample or negative sample).
[0091] Furthermore, for the noise subset, the noise subset is input into a preset learning model, and the prediction result of each sample is obtained to correct the label of the noise subset through the prediction result to obtain second positive samples and second negative samples. Among them, the preset learning model is two network models with the same structure but different initializations.
[0092] Furthermore, the first positive samples obtained from the clean subset and the second positive samples obtained from the noise subset are combined to form the final positive sample set; similarly, the first negative samples and the second negative samples are combined to form the final negative sample set for the subsequent model optimization process, realizing the effective use of reliable information in the clean subset to guide model learning, and at the same time minimizing the negative impact of the noise subset on the model performance, thereby improving the robustness and accuracy of the entire cross - modal matching system and enhancing the applicability of the model in complex environments.
[0093] In some embodiments, in step S250, the initial cross - modal matching model is optimized according to the positive samples and negative samples until the trained cross - modal matching model is obtained. Obtaining the positive samples and negative samples includes the following steps S510 to S540: Step S510: Optimize the positive samples and negative samples respectively through a preset initial decoupling loss function to obtain optimized positive samples and optimized negative samples; Step S520: Calculate the similarity of each sample according to the optimized positive samples and optimized negative samples; Step S530: Adjust the weight factor of the preset initial decoupling loss function according to the similarity of each sample to obtain the decoupling loss function; Step S540: Based on the decoupling loss function, update the initial cross - modal matching model according to the model training set to obtain the cross - modal matching model.
[0094] In this embodiment, a decoupled loss function is first defined to independently process positive samples (matching pairs) and negative samples (non-matching pairs), so as to solve the problem of unbalanced learning of positive and negative samples that may exist in traditional methods. Furthermore, the decoupled loss function is used to optimize the positive and negative samples obtained in the first stage respectively.
[0095] Specifically, for positive samples, maximize their similarity scores, that is, enhance the model's ability to recognize correct matching pairs; for negative samples, ensure that their similarity scores are low enough to avoid incorrect matching.
[0096] Furthermore, after the preliminary optimization is completed, according to the state of the current matching model, calculate the similarity between all samples (including the optimized positive and negative samples). Specifically, it is preferably to use cosine similarity to measure the similarity between the image feature vector and the text feature vector.
[0097] Furthermore, adjust the weight factor in the decoupled loss function according to the calculated similarity value. By adjusting the weight factor, the weight factor of the target sample can be increased, making the model pay more attention to the learning of this part of the data. After the adjustment of the weight factor, the final decoupled loss function is formed, which not only considers the original difference between positive and negative samples, but also incorporates the consideration of the importance of different samples, thus guiding the model training more effectively.
[0098] Furthermore, use the final decoupled loss function to update the initial cross-modal matching model, so that it can better adapt to the noise situation in the dataset and improve its performance in the cross-modal matching task. It can not only effectively distinguish and optimize positive and negative samples, but also dynamically adjust the learning focus of the model, enhancing its robustness and accuracy in the face of noisy data.
[0099] In some embodiments, in step S540, based on the decoupled loss function, update the initial cross-modal matching model according to the model training set to obtain the cross-modal matching model, including the following steps S610 to S630: Step S610: Input the image in the image-text sample pair into the initial cross-modal matching model to obtain the corresponding first matching result; Step S620: Input the text in the image-text sample pair into the initial cross-modal matching model to obtain the corresponding second matching result; Step S630: Use the decoupled loss function to optimize the initial cross-modal matching model according to the image-text sample pair, the first matching result, and the second matching result to obtain the cross-modal matching model.
[0100] In this embodiment, the text part and the image part in the same image-text sample pair are respectively input into the initial cross-modal matching model to obtain the first matching result and the second matching result output by the initial cross-modal matching model, that is, to obtain the current cross-modal retrieval performance of the initial cross-modal matching model.
[0101] Furthermore, based on the image-text sample pair, the first matching result, and the second matching result, the initial cross-modal matching model is optimized. The decoupled loss function is used to adjust and optimize the parameters of the initial cross-modal matching model until the model performance reaches a predetermined standard or no longer improves significantly. Finally, a trained cross-modal matching model is obtained, which can establish a more accurate matching relationship between images and texts, especially showing more robustness in a dataset containing noise.
[0102] In addition, to further verify the effectiveness of the model, the recall rate (R@K, K = 1, 5, 10) of the model can also be evaluated through datasets such as Flickr30K and MS-COCO, and the model hyperparameters can be adjusted according to the evaluation results to improve the overall performance, so as to improve the accuracy of the cross-modal matching task and enhance the adaptability and stability of the model when facing noisy data.
[0103] In some embodiments, based on the decoupled loss function, the initial cross-modal matching model is updated according to the model training set to obtain the cross-modal matching model. Preferably, the update of the initial cross-modal matching model is achieved by using bidirectional matching training, which specifically includes: (1) Image-to-text retrieval: Extract image and text features respectively, and then use the similarity metric function to calculate the matching degree between the query image and all candidate texts and sort them; (2) Text-to-image retrieval: Similarly, extract image and text features respectively, and then use the similarity metric function (such as cosine similarity) to calculate the matching degree between the query text and all candidate images and sort them; (3) Jointly optimize a specific objective function, aiming to make the bidirectional cross-modal similarity of image-to-text retrieval and text-to-image retrieval as close as possible the distance between the anchor sample and the positive sample while pulling away the distance from the negative sample. In addition, the constructed adaptive boundary is used to penalize potential noisy positive samples to reduce their negative impact on the model.
[0104] In some of these embodiments, a variational Bayesian Gaussian mixture model is used to model the sample loss distribution, automatically identify potential noise samples, and utilize a soft label strategy and an asymmetric optimization mechanism to reduce the negative impact of noise samples on model training, ensure the robustness of system operation, overcome the problem of noise correspondence in multi-modal datasets in cross-modal matching tasks, and improve the performance of existing models in a noisy environment. Moreover, during the model training process, according to a set similarity threshold, it is determined whether an image-text pair matches, and at the same time, according to metrics such as R@1, R@5, and R@10, the matching situation in the top K retrieved results is evaluated to provide quantitative matching performance feedback, so as to achieve deployable, controllable, and scalable cross-modal matching applications in a real and complex multi-modal environment.
[0105] In some embodiments, first, a dataset containing image-text pairs and binary labels is given. After warm-up training to calculate the loss, the variational Bayesian Gaussian mixture model divides the clean subset and the noise subset according to the characteristics of the loss distribution. Then, a decoupled loss function is constructed, setting unique positive and negative sample pair boundaries and weight factors to achieve decoupled optimization of positive and negative sample pairs, accurately punishing potential noise positive samples, and reducing the interference with the optimization of negative samples. Finally, a two-way matching strategy is used to train the model, and the optimization is evaluated according to the recall rate metric and verified on datasets such as Flickr30K, MS-COCO, and CC152K.
[0106] Specifically, taking the Flickr30K dataset as an example, which contains 31,000 images collected from the Flickr website and corresponding 5 description texts. According to the principle of the present invention, first, the image-text pairs are input into the model. After warm-up training to calculate the loss, the variational Bayesian Gaussian mixture model divides the subsets according to the loss distribution, constructs a decoupled optimization loss function to optimize the sample pairs, and the two-way matching training is combined with the recall rate metric for evaluation and optimization.
[0107] As Figure 3 shown, some embodiments of the present application provide a multi-modal data cross-modal matching system, which includes an acquisition module 310 and a matching module 320. Specifically: The acquisition module 310 is used to acquire target data, and the target data is a cross-modal data pair containing various cross-modal data; The matching module 320 is used to input the target data into a preset cross-modal matching model to obtain a matching result output by the cross-modal matching model; Among them, the training process of the cross-modal matching model includes: Obtain a model training set, which contains multi-modal matching sample pairs and relevance labels; Input the model training set into an initial cross-modal matching model for training to obtain an initial matching result; Calculate the training loss value of each multi-modal matching sample pair in the model training set according to the initial matching result; Based on each training loss value, correct the relevance label of each multi-modal matching sample pair in the model training set to obtain positive samples and negative samples; Optimize the initial cross-modal matching model according to the positive samples and negative samples until the trained cross-modal matching model is obtained.
[0108] In some embodiments, the matching module 320 may include: ; ; Wherein, is the th image in the model training set, is the text corresponding to in the model training set, is 's feature vector, is 's feature vector, is the preset boundary value, is the th initial negative sample text except , is the th initial negative sample image except , is and 's similarity score, is and 's similarity score, is and 's similarity score.
[0109] In some embodiments, the matching module 320 may include: using the variational Bayesian Gaussian mixture model to calculate the clean probability posterior value of each multi-modal matching sample according to the training loss value.
[0110] In some embodiments, the matching module 320 may include: dividing the model training set into a clean subset and a noise subset according to the clean probability posterior value.
[0111] In some embodiments, the matching module 320 may include: respectively correcting the labels of the clean subset and the noise subset to obtain positive samples and negative samples.
[0112] In some embodiments, the matching module 320 may include: optimizing the positive samples and negative samples respectively through a preset initial decoupling loss function to obtain the optimized positive samples and the optimized negative samples.
[0113] In some embodiments, the matching module 320 may include: calculating the similarity of each sample according to the optimized positive samples and the optimized negative samples.
[0114] In some embodiments, the matching module 320 may include: adjusting the weight factor of the preset initial decoupling loss function according to the similarity of each sample to obtain the decoupling loss function.
[0115] In some embodiments, the matching module 320 may include: updating the initial cross-modal matching model based on the decoupling loss function according to the model training set to obtain the cross-modal matching model.
[0116] In some embodiments, the matching module 320 may include: correcting the labels of the clean subset according to the relevance labels and model prediction labels of the clean subset to obtain the first positive samples and the first negative samples, and the model prediction labels are obtained by matching the clean subset using the initial cross-modal matching model.
[0117] In some embodiments, the matching module 320 may include: inputting the noise subset into a preset learning model to obtain the corresponding prediction results.
[0118] In some embodiments, the matching module 320 may include: correcting the noise subset according to the prediction results to obtain the second positive samples and the second negative samples.
[0119] In some embodiments, the matching module 320 may include: using the first positive samples and the second positive samples as positive samples; In some embodiments, the matching module 320 may include: using the first negative samples and the second negative samples as negative samples.
[0120] In some embodiments, the matching module 320 may include: ; ; ; Wherein, and are respectively the weight factor of the th positive sample pair and the weight factor of the th negative sample pair, is a preset factor, and are respectively the boundary values of the positive sample pair and the negative sample pair, is the image and the matching text similarity score of, is the image and the negative sample text similarity score of, is the preset sample factor, is the ideal optimal boundary of the positive sample pair, is the ideal optimal boundary of the negative sample pair, is the total number of image - text pairs in the model dataset.
[0121] In some embodiments, the matching module 320 may include: inputting the image in the image - text sample pair into the initial cross - modal matching model to obtain the corresponding first matching result.
[0122] In some embodiments, the matching module 320 may include: inputting the text in the image - text sample pair into the initial cross - modal matching model to obtain the corresponding second matching result.
[0123] In some embodiments, the matching module 320 may include: using the decoupled loss function to optimize the initial cross - modal matching model according to the image - text sample pair, the first matching result, and the second matching result to obtain the cross - modal matching model.
[0124] It should be noted that the multi - modal data cross - modal matching system provided in this embodiment and the above - mentioned multi - modal data cross - modal matching method are based on the same inventive concept. Therefore, the relevant content of the above - mentioned multi - modal data cross - modal matching method also applies to the content of the multi - modal data cross - modal matching system. Therefore, it will not be elaborated here.
[0125] For, the system obtains the target data; inputs the target data into the preset cross - modal matching model to obtain the matching result output by the cross - modal matching model; the training process of the cross - modal matching model includes: obtaining the model training set; inputting the model training set into the initial cross - modal matching model for training to obtain the initial matching result; calculating the training loss value of each multi - modal matching sample pair in the model training set according to the initial matching result; based on each training loss value, correcting the relevance label of each multi - modal matching sample pair in the model training set to obtain positive samples and negative samples; optimizing the initial cross - modal matching model according to the positive samples and negative samples until the trained cross - modal matching model is obtained. In this way, it can achieve compatibility with weakly aligned data and dynamically adjust the information weights between modalities to improve the cross - modal matching accuracy and system stability in complex application environments.
[0126] This application embodiment also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above - mentioned multi - modal data cross - modal matching method.
[0127] As shown in Figure 4 , Figure 4 which is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application. The electronic device includes: At least one battery; At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the multi-modal data cross-modal matching method described above in the present disclosure.
[0128] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0129] The following will introduce the electronic device of the embodiment of the present application in detail.
[0130] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure; The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700, and are called by the processor 1600 to execute a multi-modal data cross-modal matching method of the present disclosure.
[0131] The input / output interface 1800 is used to implement information input and output; The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.); The bus 2000 transmits information between the various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.
[0132] The embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above multi-modal data cross-modal matching method.
[0133] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0134] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0135] Those skilled in the art can understand that the technical solutions shown in the figures do not limit the embodiments of the present disclosure, and may include more or fewer steps than shown in the figures, or combine some steps, or different steps.
[0136] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0137] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0138] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0139] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0140] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0141] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0142] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0143] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0144] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above-mentioned implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the embodiments of the present application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the embodiments of the present application.
[0145] The above has described the embodiments of the present application in detail with reference to the drawings, but the present application is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can also be made without departing from the purpose of the present application.
Claims
1. A cross-modal matching method for multimodal data, characterized in that: The method comprises: Acquire target data, where the target data is a cross-modal data pair including a plurality of cross-modal data; Inputting the target data into a preset cross-modal matching model to obtain a matching result output by the cross-modal matching model; The training process of the cross-modal matching model includes: Obtaining a model training set, wherein the model training set includes multimodal matching sample pairs and relevance labels; Inputting the model training set into an initial cross-modal matching model for training to obtain an initial matching result; Calculate the training loss value of each multimodal matching sample pair in the model training set according to the initial matching result; Based on each of the training loss values, performing correlation label correction on each multimodal matching sample pair in the model training set to obtain positive samples and negative samples; The initial cross-modal matching model is optimized according to the positive samples and the negative samples until a trained cross-modal matching model is obtained.
2. The multimodal data cross-modal matching method according to claim 1, characterized in that: The calculation formula for calculating the training loss value of each multimodal matching sample pair in the model training set according to the initial matching result includes: ; ; in, The first images, For model training, focus on The corresponding text, yes The characteristic vector of yes The characteristic vector of is the preset boundary value, For Other than Initial negative sample texts, For Other than Initial negative sample images, for and The similarity score between yes and The similarity score between yes and The similarity score between .
3. The multimodal data cross-modal matching method according to claim 1, characterized in that: Based on each of the training loss values, performing correlation label correction on each multimodal matching sample pair in the model training set to obtain positive samples and negative samples includes: Utilizing a variational Bayesian Gaussian mixture model, calculating a clean probability posterior value of each multimodal matching sample according to the training loss value; According to the clean probability posterior value, the model training set is divided into a clean subset and a noise subset; Label correction is performed on the clean subset and the noise subset respectively to obtain the positive sample and the negative sample.
4. The multimodal data cross-modal matching method according to claim 3, characterized in that: Optimizing the initial cross-modal matching model according to the positive samples and the negative samples until a trained cross-modal matching model is obtained includes: By presetting the initial decoupling loss function, the positive samples and negative samples are optimized respectively to obtain the optimized positive samples and the optimized negative samples; Calculating the similarity of each of the samples according to the optimized positive samples and the optimized negative samples; Adjusting the weight factor of the preset initial decoupling loss function according to the similarity of each of the samples to obtain a decoupling loss function; Based on the decoupling loss function, the initial cross-modal matching model is updated according to the model training set to obtain the cross-modal matching model.
5. The multimodal data cross-modal matching method according to claim 3, characterized in that: The step of respectively performing label correction on the clean subset and the noise subset to obtain the positive sample and the negative sample includes: Performing label correction on the clean subset according to the relevance label and the model prediction label of the clean subset to obtain a first positive sample and a first negative sample, wherein the model prediction label is obtained by matching the clean subset using the initial cross-modal matching model; Inputting the noise subset into a preset learning model to obtain a corresponding prediction result; Correcting the noise subset according to the prediction result to obtain a second positive sample and a second negative sample; Taking the first positive sample and the second positive sample as the positive samples; The first negative sample and the second negative sample are used as the negative samples.
6. The multimodal data cross-modal matching method according to claim 4, characterized in that: The calculation formula of the preset initial decoupling loss function includes: ; ; ; in, and Respectively The weight factor of the positive sample pair and the The weight factor of negative sample pairs, is the preset factor, and are the boundary values of positive sample pairs and negative sample pairs respectively, is an image and matching text The similarity score of is an image And negative sample text The similarity score of is the preset sample factor, is the ideal optimal boundary for positive sample pairs, is the ideal optimal boundary of negative sample pairs, is the total number of image-text pairs in the model dataset.
7. The multimodal data cross-modal matching method according to claim 4, characterized in that: The multimodal matching sample pairs include image-text sample pairs, and the initial cross-modal matching model is updated based on the decoupling loss function and the model training set to obtain the cross-modal matching model, including: Inputting the image in the image-text sample pair into the initial cross-modal matching model to obtain a corresponding first matching result; Inputting the text in the image-text sample pair into the initial cross-modal matching model to obtain a corresponding second matching result; The decoupling loss function is used to optimize the initial cross-modal matching model according to the image-text sample pair, the first matching result, and the second matching result to obtain a cross-modal matching model.
8. A multimodal data cross-modal matching system, characterized in that: The system comprises: An acquisition module is used to acquire target data; A matching module, used to input the target data into a preset cross-modal matching model to obtain a matching result output by the cross-modal matching model; The training process of the cross-modal matching model includes: Obtaining a model training set, wherein the model training set includes multimodal matching sample pairs and relevance labels; Inputting the model training set into an initial cross-modal matching model for training to obtain an initial matching result; Calculate the training loss value of each multimodal matching sample pair in the model training set according to the initial matching result; Based on each of the training loss values, performing correlation label correction on each multimodal matching sample pair in the model training set to obtain positive samples and negative samples; The initial cross-modal matching model is optimized according to the positive samples and the negative samples until a trained cross-modal matching model is obtained.
9. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute a multimodal data cross-modal matching method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute a multimodal data cross-modal matching method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal named entity recognition method based on uncertainty perception
CN114386412A
Noise data set training method based on sample screening and label correction
CN115641480A
Cited By
Multimodal emotion recognition method, device and system based on direction-amplitude decomposition and readable storage medium
CN121786762A