Cross-modal feature extraction model training method, retrieval method, system and product

CN120832506BActive Publication Date: 2026-09-08SUZHOU KEDA TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510935273.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2026-09-08
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

TIPR方法面临的核心挑战是图像和文本之间的显著模态异构性,即两者之间的表示差异,如何实现对图像和文本的跨模态特征提取以提高基于图像特征和文本特征进行匹配的准确性是该方法应用中需要解决的问题

Benefits of technology

通过采用该跨模态特征提取模型训练方法,首先采集训练数据并获得掩码图像和掩码文本,基于特征提取模型提取得到对应的图像特征和文本特征,以用于后续的建模步骤,在跨语言蒸馏掩码图像建模中,基于图像信息更少的掩码图像来实现挖掘第一语言文本和第二语言文本之间的隐层关系,更关注于第一语言表达和第二语言表达之间的联系,在跨语言掩码语言建模中,基于第二语言的掩码文本来减少因为翻译不准确导致的特征提取和图文匹配错误情况,深入挖掘第一语言表达和第二语言表达之间的隐层关系,并且通过图文结合进行建模和挖掘,解决图像和文本之间的显著模态异构性问题,提高文本特征提取的准确性,从而有利于提高图文匹配检索的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832506B_ABST
    Figure CN120832506B_ABST
Patent Text Reader

Abstract

The application provides a cross-modal feature extraction model training method, a retrieval method, a system and a product. The method comprises the following steps: collecting training data, the training data comprising multiple images and corresponding description text groups; obtaining a mask image and a mask text, obtaining training image features, training text features, mask image features and mask text features based on a feature extraction model; modeling based on the mask image features and the training text features in a first language and a second language, constructing a first loss function of a cross-language distillation mask image model; modeling based on the training image features and the mask text features in the first language and the second language, constructing a second loss function of a cross-language mask language model; and optimizing training based on the first loss function and the second loss function to obtain an optimized feature extraction model. The application is beneficial to improving the accuracy of cross-modal cross-language feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a cross-modal feature extraction model training method, retrieval method, system, and product. Background Technology

[0002] Text-to-image person retrieval (TIPR) aims to retrieve target pedestrians using descriptive text, finding the most relevant pedestrian images from a pedestrian image database. This task utilizes a free and flexible text query approach, which is more universal than the structured image query approach used in person re-identification (Re-ID) tasks, and has significant application potential in the field of public safety. A core challenge of TIPR methods is the significant modal heterogeneity between images and text, i.e., the representational differences between them. How to achieve cross-modal feature extraction from images and text to improve the accuracy of matching based on image and text features is a problem that needs to be solved in the application of this method. Furthermore, existing TIPR methods are centered on single-language (e.g., English) descriptive text, limiting their application in multilingual environments. Summary of the Invention

[0003] To address the problems in the existing technology, the purpose of this application is to provide a cross-modal feature extraction model training method, retrieval method, system, and product, which is beneficial to improving the accuracy of cross-modal and cross-linguistic feature extraction.

[0004] The first aspect of this application provides a method for training a cross-modal feature extraction model, comprising the following steps: Training data is collected, which includes multiple images and their corresponding descriptive text groups. Each descriptive text group includes descriptive text in the first language and the second language with the same meaning. Based on the feature extraction model, feature extraction is performed on the training data to obtain training image features and training text features. Based on the training data, mask images and mask text are obtained. Feature extraction is performed on the mask images to obtain mask image features, and feature extraction is performed on the mask text to obtain mask text features. Modeling is performed based on mask image features and training text features in the first and second languages, and a first loss function is constructed for the cross-language distilled mask image model. A second loss function for cross-lingual masked language models is constructed based on training image features and masked text features of the first and second languages. The optimized feature extraction model is obtained by optimizing the training based on the first and second loss functions.

[0005] In some embodiments, modeling is performed based on mask image features and training text features of the first and second languages ​​to construct a first loss function for the cross-language distilled mask image model, including the following steps: A feature fusion model is used to obtain the first fused feature of training image features and first language training text features, and a second fused feature of mask image features and second language training text features is obtained. Modeling is performed based on the first and second fusion features to construct the first loss function for the cross-language distillation mask image model.

[0006] In some embodiments, modeling is performed based on the first fusion feature and the second fusion feature to construct a first loss function for the cross-language distillation mask image model, including the following steps: The first fusion feature is used as the supervision signal, and the second fusion feature is used as the input to the cross-language distillation mask image model for modeling. The first loss function is constructed based on the similarity between the output of the cross-language distillation mask image model and the supervision signal.

[0007] In some embodiments, a second loss function for a cross-lingual masked language model is constructed based on training image features and masked text features of the first and second languages, including the following steps: A feature fusion model is used to obtain the third fusion feature of training image features and masked text features of the first language, and a fourth fusion feature of training image features and masked text features of the second language is obtained. The third and fourth fusion features are used as inputs to the cross-lingual masked language model, and the training text features of the first language are used as labels. A second loss function is constructed based on the output and labels of a cross-language masked language model.

[0008] In some embodiments, the following steps are also included: The third loss function of the cross-language image-text comparison learning model is constructed by modeling based on training image features and training text features of the first and second languages. The optimization training includes optimization training based on the first loss function, the second loss function, and the third loss function.

[0009] In some embodiments, the following steps are also included: The fourth loss function of the cross-language image-text matching model is constructed based on training image features, mask image features, and training text features of the first and second languages. The optimization training includes optimization training based on the first loss function, the second loss function, and the fourth loss function.

[0010] In some embodiments, a fourth loss function is constructed based on training image features, masked image features, and training text features of the first and second languages, including the following steps: A feature fusion model is used to obtain the fifth fusion feature of training image features and first language training text features, and a sixth fusion feature of mask image features and second language training text features is obtained. Based on the fifth fusion feature as the first input to the cross-language image-text matching model, the first part of the fourth loss function is constructed. Based on the sixth fusion feature as the second input to the cross-language image-text matching model, the second part of the fourth loss function is constructed. The complete fourth loss function is obtained from the first and second parts of the fourth loss function.

[0011] A second aspect of this application also provides a text and image retrieval method, comprising the following steps: Retrieve the description text to be searched; The descriptive text to be retrieved is input into the feature extraction model obtained by the cross-modal feature extraction model training method based on the first aspect, and the features of the descriptive text to be retrieved are obtained. The image features that match the descriptive text features to be retrieved are determined by matching them with the candidate image features.

[0012] A third aspect of this application also provides an image and text retrieval system, including: The text acquisition module is used to acquire the descriptive text to be retrieved; The feature extraction module is used to input the descriptive text to be retrieved into the feature extraction model obtained based on the cross-modal feature extraction model training method of the first aspect, and to obtain the features of the descriptive text to be retrieved. The feature matching module is used to match the features of the descriptive text to be retrieved with the features of the candidate images, and to determine the image features that match the features of the descriptive text to be retrieved.

[0013] The fourth aspect of this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the image and text retrieval method of the second aspect.

[0014] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application.

[0015] The cross-modal feature extraction model training method, retrieval method, system, and product of this application have the following beneficial effects: By employing this cross-modal feature extraction model training method, training data is first collected to obtain mask images and mask text. Based on the feature extraction model, corresponding image features and text features are extracted for subsequent modeling steps. In cross-lingual distillation mask image modeling, the hidden relationships between first-language and second-language texts are mined based on mask images with less image information, focusing more on the connection between first-language and second-language expressions. In cross-lingual mask language modeling, the second-language mask text is used to reduce feature extraction and image-text matching errors caused by inaccurate translation, deeply mining the hidden relationships between first-language and second-language expressions. Furthermore, by combining images and text for modeling and mining, the significant modal heterogeneity problem between images and text is solved, improving the accuracy of text feature extraction, thereby enhancing the accuracy of image-text matching retrieval. Attached Figure Description

[0016] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings.

[0017] Figure 1 This is a flowchart of a cross-modal feature extraction model training method according to an embodiment of this application; Figure 2 This is a flowchart illustrating the specific implementation of a cross-modal feature extraction model training method according to an embodiment of this application; Figure 3 This is a schematic diagram of various models and various modeling processes according to an embodiment of this application; Figure 4 This is a schematic diagram of a cross-language distillation mask image modeling process according to an embodiment of this application; Figure 5 This is a schematic diagram of a cross-language masked language modeling process according to an embodiment of this application; Figure 6 This is a schematic diagram of a cross-language text-image contrast learning modeling process according to an embodiment of this application; Figure 7 This is a schematic diagram of cross-language image-text matching modeling according to an embodiment of this application; Figure 8 This is a flowchart of an embodiment of the image and text retrieval method of this application; Figure 9 This is a schematic diagram of the structure of a text and image retrieval system according to an embodiment of this application. Detailed Implementation

[0018] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0019] Furthermore, the accompanying drawings are merely illustrative of this application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices. Although the terms "first" or "second," etc., are used in this specification to denote certain features, these are only for indicating function and not as a limitation on the number or importance of specific features.

[0020] The flowchart shown in the attached diagram is merely an illustrative example and does not necessarily include all steps. For example, some steps may be broken down, while others may be combined or partially combined. Therefore, the actual execution order may change depending on the specific circumstances.

[0021] like Figure 1 As shown in the figure, this application provides a method for training a cross-modal feature extraction model, including the following steps: S110: Collect training data. The training data includes multiple images and their corresponding descriptive text groups. Each descriptive text group includes descriptive text in the first language and the second language with the same meaning. In this embodiment, each image corresponds to a set of matching descriptive text. The feature extraction model obtained through this training method is used for cross-modal extraction of images and descriptive text during image-text retrieval to solve the problem of significant modal heterogeneity between images and text, reduce the representation difference between the two, and improve the accuracy of image-text retrieval. At the same time, through cross-language learning training in multiple languages ​​during training, it is possible to achieve accurate retrieval based on non-English descriptive text without restricting the language of the input descriptive text during the inference stage when applied to image-text retrieval scenarios. Based on this, the first language is the source language, such as English, which is the language type supported by conventional image and text retrieval methods. The second language is the target language that may actually be used, such as Chinese, French, German, etc. There can be one second language, in which case each description text group includes two description texts, forming a description text pair. There can also be multiple second languages. For example, if the second language includes Chinese and French, each description text group includes three description texts, namely English description text, Chinese description text, and French description text. The following will use the example of the second language including one language to illustrate the implementation of this method, and will provide supplementary explanations when using multiple second languages ​​with different implementation methods. S120: Based on the feature extraction model, feature extraction is performed on the training data to obtain training image features and training text features. Based on the training data, a mask image and mask text are obtained. Feature extraction is performed on the mask image to obtain mask image features, and feature extraction is performed on the mask text to obtain mask text features. In this embodiment, the images in the training data are masked to obtain masked images, and each descriptive text in each descriptive text group is masked to obtain masked text. Based on the feature extraction model, features are extracted from the images in the training data to obtain training image features, features are extracted from the descriptive text in the first language to obtain training text features in the first language, features are extracted from the descriptive text in the second language to obtain training text features in the second language, features are extracted from the masked images to obtain masked image features, features are extracted from the masked text features in the first language to obtain masked text features in the first language, and features are extracted from the masked text features in the second language to obtain masked text features in the second language. In this embodiment, the feature extraction model includes a text encoder and an image encoder, wherein the text encoder is used to extract text features and the image encoder is used to extract image features; S130: Modeling is performed based on mask image features and training text features of the first and second languages ​​to construct the first loss function of the cross-language distilled mask image model; wherein, the fusion features of the training image features and the training text features of the first language are used as the supervision signal, and the mask image features and the training text features of the second language are used as the input signals to construct the first loss function; This step is the cross-lingual distillation mask image modeling step. Through this modeling step, the hidden relationships between the training text features of the first language and the training text features of the second language can be mined. Because the image is masked, the image information is reduced compared to before masking. To improve the accuracy of reconstruction during modeling, the cross-lingual distillation mask image model focuses more on the intrinsic connections between the two languages, thereby improving the fine-grainedness and accuracy of mining the hidden relationships between the two languages. The modeling and training of the cross-lingual distillation mask image model is used to predict the fused features based on the mask image features and the training text features of the second language, gradually making it closer to the fusion of the training image features and the training text features of the first language. The feature extraction method delves deeper into the semantic relationships between the training text features of the second language and the training text features of the first language by gradually narrowing the gap between the two fused features. This ensures that the features obtained by encoding the second language descriptive text through the text encoder are as close as possible to the features of the first language descriptive text with the same semantics, making the semantic expression of the second language text features more accurate. Furthermore, this is achieved during training by combining the fusion of mask image features and text features, enabling the fusion analysis of features from different modalities. The semantic relationships between text features and image features are also deeply explored, making the feature expressions of text features and image features more suitable for subsequent cross-modal matching. S140: Modeling is performed based on training image features and masked text features of the first and second languages ​​to construct the second loss function of the cross-language masked language model; wherein, the training text features based on the first language are used as labels, the fusion features of training image features and masked text features of the first language are used as input data to construct the first part of the second loss function, the fusion features of training image features and masked text features of the second language are used as input data to construct the second part of the second loss function, and the second loss function is obtained based on the first part and the second part; This step is the cross-language masked language modeling step, which can uncover the hidden relationships between the first language and the second language. In the modeling process of cross-language masked language model, the training text features of the first language are used as labels. The training objective is to predict the true and complete text features based on the training image features and the masked text features of the first language, so that the predicted true and complete text features are close to the labels. Similarly, the training image features and the masked text features of the second language are used to predict the true and complete text features, so that the predicted true and complete text features are close to the labels, thereby restoring the true and complete semantics of the masked text features of the first and second languages ​​to the greatest extent. In practical applications, the descriptive text input into the text encoder may be translated from other languages. There may be translation errors or omissions, resulting in incomplete or missing content in the input descriptive text. This can lead to semantic defects. For example, second-language text may be translated from first-language text, and inaccurate translations between multiple languages ​​can prevent a complete one-to-one correspondence between the second-language and first-language texts. By masking the second-language text, parts of its content can be hidden, simulating local translation errors and allowing the text encoder to restore the original text as much as possible even with incomplete second-language information. Complete semantics helps improve the accuracy of second-language text encoding when the translation is inaccurate, thereby reducing errors in feature extraction and image-text matching caused by inaccurate translation. Furthermore, joint modeling and mining of text and images from multiple languages ​​allows for deeper exploration of the relationships between text and images, improving the accuracy of hidden relationship mining between text and images from different languages. Similarly, masking the first-language text can optimize the text encoder, enabling it to restore the true and complete semantics of the descriptive text to the greatest extent possible when the first-language descriptive text has defects and incomplete semantics. During training, the fusion of text features and image features, along with the analysis and mining of the fused features, further aligns the expression of image features and text features, further improving the accuracy of cross-modal feature extraction. S150: Optimization training is performed based on the first and second loss functions to obtain an optimized feature extraction model. Thus, through optimized training, better model parameters can be obtained for the feature extraction model, enabling the text encoder of the feature extraction model to extract text features more accurately. This allows the extracted text features to not only address the difference between text and image features, thus achieving cross-modal feature extraction, but also cross-lingual feature extraction. Regardless of whether the input descriptive text is in its first or second language, text features that can accurately match the image can be extracted. Furthermore, the image encoder of the feature extraction model can more accurately extract image features, which is beneficial for obtaining image features that more accurately match the semantics expressed by the text features, thereby improving the accuracy of text and image feature matching.

[0022] This application employs multiple modeling and training methods to obtain a final optimized feature extraction model. The feature extraction model extracts text and image features, and its output is processed (e.g., fused with images) as input to a cross-lingual masking language model and a cross-lingual masking language model, respectively. This results in a unified training model comprising a feature extraction model and a prediction part. The prediction part includes two independent branches of the cross-lingual masking language model. Optimization training based on a first and a second loss function involves determining a total loss function based on both loss functions, and then jointly optimizing each branch of the feature extraction model and the prediction part within the unified training model based on the total loss function. Specifically, in each batch, the images and text of that batch are input into the feature extraction model for feature extraction, and then input into the prediction part to obtain the output data of each branch. A loss function is constructed for each branch based on its label and output data. A total loss function is constructed based on the loss functions of all branches. The model parameters of the feature extraction model and the prediction part are then optimized in reverse before proceeding to the next batch. Therefore, through iterative optimization based on the total loss function, not only can the model parameters of each branch of the prediction part be optimized, but the model parameters of the feature extraction model can also be jointly optimized and trained, thereby improving the accuracy of the feature extraction model in representing features.

[0023] By employing this cross-modal feature extraction model training method, training data is first collected and mask images and mask text are obtained through steps S110 and S120. Then, in step S130, corresponding image and text features are extracted based on the feature extraction model for use in the modeling steps S140 and S150. In cross-lingual distillation mask image modeling, the hidden relationships between the first and second language texts are mined based on the mask image, which contains less image information, focusing more on the connection between the first and second language expressions. In cross-lingual mask language modeling, the second language mask text is used to reduce feature extraction and image-text matching errors caused by inaccurate translation, deeply mining the hidden relationships between the first and second language expressions. Furthermore, by combining images and text for modeling and mining, the significant modal heterogeneity between images and text is addressed, improving the accuracy of text feature extraction and thus enhancing the accuracy of image-text matching retrieval.

[0024] Traditional methods focus only on single-modal feature extraction, and cross-modal alignment often relies on prior rules. This application uses a total loss function to perform end-to-end joint optimization of multiple loss functions on the overall training model. The loss function of the prediction part is used to inversely optimize the model parameters of the feature extraction part, enabling bidirectional language reasoning between images and text, and between texts in different languages. This requires no prior knowledge and can adaptively learn implicit associations between different modalities, effectively solving the modal heterogeneity problem and improving cross-modal matching accuracy. Specifically, in the cross-language distillation mask image model, the mask image and language text are fused and then learned, narrowing the gap between different fused features, promoting semantic alignment between different language texts and images, and optimizing the feature representation of image and text features. In the cross-language mask language model, image features are also fused with mask text features to predict complete text features. Image semantics assists in text semantic reconstruction, which helps enhance the semantic constraint effect of images on text and improves the alignment effect of language text and images in feature representation.

[0025] Figure 1 The execution order of the steps in the cross-modal feature extraction model training method shown is merely an example and is not intended to limit the scope of protection of this application. In other embodiments, the execution order of the steps can be adjusted as needed without exceeding the scope of protection of this application. For example, the cross-lingual distillation mask image modeling step in step S130 and the cross-lingual distillation mask image modeling step in step S140 can be performed in parallel, or step S130 can be performed first and then step S140, or step S140 can be performed first and then step S130. The step of obtaining the mask text and mask text features can be performed after the cross-lingual distillation mask image modeling step in step S130, as long as the step of obtaining the mask text and mask text features is performed before the cross-lingual mask language modeling step in step S140.

[0026] Existing solutions to the heterogeneity problem in image-text retrieval mainly include representation learning and cross-modal alignment. Cross-modal alignment is further divided into two categories: global matching methods and local matching methods. Global matching methods focus on coarse-grained alignment by designing a cross-modal matching loss function to align the global representations of text and images. Local matching methods focus on fine-grained alignment by establishing the correspondence between text entities and body parts in images.

[0027] The following is a brief introduction to representation learning and cross-modal alignment.

[0028] 1. Representation Learning Methods: These methods focus on developing robust representation learning networks to extract discriminative feature representations. Early approaches used Convolutional Neural Networks (CNNs) as image encoders and Long Short-Term Memory Networks (LSTMs) as text encoders. Later improvements used Visual Transformer (VIT) and BERT as image and text encoders, respectively. Recent advancements utilize Visual Language Pre-trained Models (VLPs) as representation learning neural networks, enabling the extraction of more discriminative feature representations. However, representation learning methods have limited generalization ability and low robustness.

[0029] 2. Cross-modal alignment methods: These methods aim to design effective alignment strategies to achieve good matching between image and text modalities. Global matching methods directly align global text and visual representations by designing a reasonable cross-modal matching loss function. However, this method ignores fine-grained information, affecting matching accuracy. Local matching methods align fine-grained image and text information, enhancing cross-modal alignment. Local matching restricts cross-modal alignment to defined rule boundaries and requires prior information, leading to additional resource requirements and increased costs.

[0030] By employing the cross-modal feature extraction model training method of this application, bidirectional implicit relational reasoning between images and text, and between text in the first language and text in the second language, is achieved. This enhances robustness to noise by reducing feature extraction problems caused by translation issues. Cross-linguistic masking language modeling improves fine-grained alignment in multilingual environments, and cross-linguistic distillation masking image modeling enables bidirectional implicit reasoning between text and image content. Therefore, the model training method of this application fully considers fine-grained information, improves matching accuracy, and requires no prior information or additional resources, thus reducing costs.

[0031] In this embodiment, the cross-modal feature extraction model training method further includes the following steps: modeling based on training image features and training text features of the first and second languages, and constructing a third loss function for the cross-language image-text comparison learning model. Step S150, optimization training, includes optimization training based on the first loss function, the second loss function, and the third loss function.

[0032] In this embodiment, a fourth loss function for the cross-language image-text matching model is constructed based on training image features, mask image features, and training text features of the first and second languages. Step S150 involves optimization training based on the first, second, and fourth loss functions.

[0033] In one implementation, both the third and fourth loss functions are used as the basis for joint training. For example... Figure 2As shown, the training method for this cross-modal feature extraction model also includes the following steps: S141: Modeling is performed based on training image features and training text features of the first and second languages ​​to construct the third loss function of the cross-language image-text comparison learning model; wherein, after calculating the similarity between the training text features and image features based on the first and second languages ​​respectively, the calculated similarity is used as the model output, and the third loss function is constructed based on the true matching labels of the training image features and training text features and the model output. Training based on the third loss function aims to optimize the expression of text and image features by the feature extraction model, so that the expressions of text and image features with the same semantics are as close as possible, while the expressions of text and image features with different semantics are as far apart as possible. This improves the semantic accuracy of text and image features and enables the text features extracted by the feature extraction model to be adapted to matching with image features, and the extracted image features to be adapted to matching with text features, thereby achieving cross-modal feature alignment and accurate matching. S142: A fourth loss function for a cross-language image-text matching model is constructed based on training image features, masked image features, and training text features of the first and second languages. Specifically, the fused features of the training image features and the training text features of the first language are used as model input. The first part of the fourth loss function is constructed based on the model output and the true matching labels of the training image features and the training text features of the first language. The fused features of the masked image features and the training text features of the second language are used as model input. The second part of the fourth loss function is constructed based on the model output and the true matching labels of the masked image features and the training text features of the second language. The complete fourth loss function can be obtained based on the first and second parts of the fourth loss function. Training based on the fourth loss function aims to optimize the expression of text and image features by optimizing the feature extraction model, so that the expressions of text and image features with the same semantics are as close as possible, while the expressions of text and image features with different semantics are as far apart as possible, thereby improving the accuracy of semantic expression of text and image features. Cross-lingual image-text contrast learning and cross-lingual image-text matching models employ two different approaches to compare and match image and text features. These models can uncover implicit relationships between image and text features from different perspectives. The cross-lingual image-text contrast learning model primarily achieves global feature distribution alignment, while the cross-lingual image-text matching model achieves fine-grained semantic consistency judgment. These approaches respectively address the problems of cross-modal feature space fragmentation and cross-lingual semantic bias, ensuring global semantic consistency between multilingual text and images and enhancing accurate matching capabilities in complex scenarios. Furthermore, the cross-lingual image-text matching model fuses training text in the second language with masked image features. By introducing constraints related to missing image information, it forces the model to strengthen the mining of implicit associations between cross-lingual text semantics and image content. This also enhances the model's robustness to noise and information-deficient scenarios. Finally, by learning the complementary descriptive ability of cross-lingual text to image content, it addresses the modal heterogeneity problem between images and multilingual text.

[0034] Step S150 includes S151: performing optimization training based on the first loss function, the second loss function, the third loss function, and the fourth loss function.

[0035] Figure 2 The execution order of the steps in the cross-modal feature extraction model training method shown is merely an example and is not intended to limit the scope of protection of this application. In other embodiments, the execution order of the steps can be adjusted as needed without exceeding the scope of protection of this application. For example, step S151 can be executed before acquiring the mask image, mask text, mask image features, and mask text features, and step S152 can be executed before acquiring the mask text and mask text features. Steps S140, S150, S151, and S152 can be executed in parallel or sequentially, and the execution order is not limited.

[0036] When applied to a two-language environment, cross-language masked language modeling is equivalent to bilingual masked language modeling, cross-language graph-text comparison learning modeling is equivalent to bilingual graph-text comparison learning modeling, and cross-language graph-text matching modeling is equivalent to bilingual graph-text matching modeling. When there are multiple second languages, modeling can be performed by combining each language with the first language separately. For example, when the first language includes English and the second language includes Chinese and French, in cross-lingual masked language modeling, modeling can be performed based on the masked text features of English and Chinese, and then modeling can be performed based on the masked text features of English and French. In cross-lingual image-text contrastive learning modeling, similarity scores can be calculated between the text features of English and Chinese and the image features respectively to obtain a third loss function for contrastive learning modeling, and then similarity scores can be calculated between the text features of English and French and the image features to obtain a third loss function for contrastive learning modeling. In cross-lingual image-text matching modeling, matching can be performed between the text features of English and Chinese and the image features to obtain a fourth loss function for contrastive learning modeling, and then matching can be performed between the text features of English and French and the image features to obtain a fourth loss function for contrastive learning modeling.

[0037] The following example uses first-language descriptive text as the source text and second-language descriptive text as the target text, combined with... Figures 3-7 This document details the specific implementation steps of the training method for the cross-modal feature extraction model.

[0038] In this embodiment, step S110, collecting training data, includes the following steps: A dataset consisting of images and multilingual descriptive text is selected, containing both training and test data. The training data includes images and their corresponding source and target texts, where the source text is in English and the target text is translated from the source text. The test data contains query text (multilingual) and a database of images to be retrieved. Each data set in the dataset consists of one image and one descriptive text set, with each set formatted as a triple (...). , , ),in It is an image. Is with images Matched source text, Is with the source file Target texts that have the same meaning but are in different languages.

[0039] For example, when applying this method to pedestrian image-text retrieval, the dataset includes 40,206 pedestrian images and 80,440 text descriptions (including English, Chinese, French, and German), involving 13,003 identities. The dataset is divided into training and testing data. Preprocessed images are resized to 224×224 pixels and divided into M=196 non-overlapping image blocks (each block is 16×16 pixels). Text segmentation uses L=77 tokens, and CLS tokens are added to extract global representations.

[0040] In this embodiment, the feature extraction model includes a text encoder and an image encoder. The text encoder and image encoder are, for example, each composed of 12 Transformer layers. In step S120, feature extraction is performed on the training data based on the feature extraction model to obtain training image features and training text features. A mask image and mask text are obtained based on the training data. Feature extraction is performed on the mask image to obtain mask image features, and feature extraction is performed on the mask text to obtain mask text features. This includes the following steps: Each image The image is divided into M non-overlapping image patches, and the training image features are obtained through an image encoder. ,in It is a global representation feature of the image. This represents the feature of the i-th image patch. (Source text) and target text Each word is segmented into L text tags, and the training text features in the first language are obtained through a text encoder. (exist Figure 3 (represented as source text features) and training text features of the second language. (exist Figure 3 (The text is represented by the target text feature). and It is a global representation feature. and It is the feature representation corresponding to the i-th label.

[0041] For images Obtain the mask image by performing a masking operation. For example, masking a random image patch (setting it to a fixed pixel value), and then using the masked image... The input image encoder obtains the mask image features. Mask source text features and target text features Replace the random markers with mask markers. The masked text of the first language and the masked text of the second language are respectively input into the text encoder to obtain the masked text features of the first language. (exist Figure 3 (represented as masked source text features) and masked text features of the second language. (exist Figure 3 (This is represented as a mask target text feature).

[0042] like Figure 3 and Figure 4 As shown, in this embodiment, step S130 involves modeling based on the mask image features and the training text features of the first and second languages ​​to construct the first loss function of the cross-language distilled mask image model, including the following steps: A feature fusion model is used to obtain the first fused feature of training image features and first language training text features, and a second fused feature of mask image features and second language training text features is obtained. In this embodiment, the feature fusion model is implemented using a multimodal interactive encoder; the training image features are... and source text features Input the multimodal interactive encoder to obtain the first fused feature. , mask image features and target text features Inputting the multimodal interactive encoder yields the second fused feature. ; Modeling is performed based on the first and second fusion features to construct the first loss function for the cross-language distilled mask image model; Modeling is performed based on the first and second fusion features, constructing the first loss function for the cross-language distilled mask image model, including the following steps: Using the first fusion feature as the supervision signal Second fusion feature Modeled as input to a cross-language distillation mask image model; In this embodiment, the cross-language distillation mask image model is implemented using a first multilayer perceptron. The cross-language distillation mask image model is used to predict the fusion features of the corresponding image and source language text based on the input features. The multilayer perceptron (MLP) is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset. The first loss function is constructed based on the similarity between the output of the cross-language distillation mask image model and the supervision signal; In this embodiment, the modeling is performed using the first fusion feature. As a supervisory signal, the second fusion feature The first multilayer perceptron is used to reconstruct features; the first loss function is constructed as follows:

[0043] in, Expressing expectations, This indicates the label for the trained model. This represents the output features of the model after the second fusion feature is input into the cross-lingual distillation mask image model. This represents the cosine similarity. Here, corresponds to the first fusion feature of the supervision signal for each input. The corresponding first language training text features and the input second fusion features The corresponding training text features of the second language belong to the same text pair, meaning that the two have the same semantics.

[0044] like Figure 3 and Figure 5 As shown, in this embodiment, step S140 involves modeling based on training image features and masked text features of the first and second languages ​​to construct a second loss function for a cross-language masked language model, including the following steps: A feature fusion model is used to obtain the third fusion feature of training image features and masked text features of the first language, and a fourth fusion feature of training image features and masked text features of the second language is obtained. In this embodiment, the training image features and mask source text features The third fused feature is obtained by inputting a multimodal interactive encoder. , train image features and mask target text features The fourth fused feature is obtained by inputting the multimodal interactive encoder. ; The third and fourth fusion features are used as inputs to the cross-lingual masked language model, and the training text features of the first language are used as labels. A second loss function is constructed based on the output and labels of a cross-language masked language model; In this embodiment, the cross-language masking language model is implemented using a second multilayer perceptron, and incorporates the third fusion feature. and the fourth fusion feature The input is a second-level perceptron, which is used to predict mask labels by incorporating source text features. As a label; The first part of the constructed second loss function is as follows:

[0045] The second part of the constructed second loss function is as follows:

[0046] The second loss function is further constructed as follows:

[0047] Where H represents the cross-entropy loss, This represents the true label, i.e., the source text features corresponding to the third / fourth fusion feature. This represents the output obtained after inputting the third fused feature into the second multilayer perceptron. This represents the output obtained after inputting the fourth fused feature into the second multilayer perceptron.

[0048] like Figure 3 and Figure 6 As shown, in this embodiment, in step S141, a third loss function for the cross-language image-text comparison learning model is constructed based on the training image features and the training text features of the first and second languages. This loss function includes: for image-text pairs... and ), using contrastive learning for optimization, calculating separately and The similarity scores are used to construct a third loss function based on these two similarity scores, and then comparative learning is performed to optimize the loss function.

[0049] In this embodiment, the cross-language image-text comparison learning model is implemented using a similarity calculation module to calculate the similarity between image features and text features, such as using cosine similarity. A first image pair is constructed based on the trained image features and the trained text features of the first language. A second image pair is constructed based on the features of the training images and the training text features of the second language. Contrastive learning is used to model separately, and calculations are performed on the first image pairs respectively. The similarity score is calculated based on the second image pair. The similarity scores are used to optimize the image-text contrast loss function based on these two similarities. The first part of the third loss function, constructed based on the similarity scores calculated for the first image pair, is as follows:

[0050] The second part of the third loss function, constructed based on the similarity scores obtained from the second image pair, is as follows:

[0051] The third loss function is further constructed as follows:

[0052] in, and This represents the true label for regularization. A value of 1 indicates that the image features and text features in the corresponding image pair match, while a value of 0 indicates that the image features and text features in the corresponding image pair do not match. The superscript i2t indicates whether the image matches the text, and t2i indicates whether the text matches the image. and The similarity score is represented by the superscript i2t, which corresponds to the similarity between the image and the text, and t2i, which represents the similarity between the text and the image.

[0053] like Figure 3 and Figure 7 As shown, in step S142, a fourth loss function for the cross-language image-text matching model is constructed based on the training image features, mask image features, and training text features of the first and second languages. This includes the following steps: A feature fusion model is used to obtain the fifth fusion feature of training image features and first language training text features, and a sixth fusion feature of mask image features and second language training text features is obtained. In this embodiment, the training image features and first language training text features The input feature fusion model yields the fifth fused feature. When steps S130 and S152 are executed sequentially, since the fifth fusion feature and the first fusion feature are the same, only the training image features need to be used. and first language training text features The fused features obtained by the input feature fusion model after one fusion are respectively used as the first fused feature and the fifth fused feature; In this embodiment, the mask image features Training text features of a second language Inputting the multimodal interactive encoder yields the sixth fused feature. When steps S130 and S152 are executed sequentially, since the sixth fusion feature and the second fusion feature are the same, only the mask image features need to be processed. Training text features of a second language The fused features obtained by the input feature fusion model after one fusion are respectively used as the second fused feature and the sixth fused feature; Based on the fifth fusion feature as the first input to the cross-language image-text matching model, the first part of the fourth loss function is constructed. Based on the sixth fusion feature as the second input to the cross-language image-text matching model, the second part of the fourth loss function is constructed. Obtain the complete fourth loss function based on the first and second parts of the fourth loss function; In this embodiment, based on the fifth fusion feature Obtain the first global fusion representation features (The fusion part of the global representation features in the fifth fusion feature), based on the sixth fusion feature Obtain the second global fusion representation features (The fusion part of the global representation features in the sixth fusion feature) The cross-language image-text matching model is implemented using a third multilayer perceptron. The first global fusion representation and the second global fusion feature are input into the third multilayer perceptron, and the loss function is constructed using ITM (Image-Text Matching Loss).

[0054] The first part of constructing the fourth loss function is as follows:

[0055] The second part of constructing the fourth loss function is as follows:

[0056] The fourth loss function is constructed as follows:

[0057] in, This represents the true label. A value of 1 indicates a match between the image and text in the global fusion feature set, while a value of 0 indicates a mismatch. and These represent the first global fusion representation features, respectively. Second global fusion representation features The output features are obtained by inputting into the third multilayer perceptron.

[0058] In step S151, optimization training is performed based on the first loss function, the second loss function, the third loss function, and the fourth loss function, including: constructing the overall loss function based on the first loss function, the second loss function, the third loss function, and the fourth loss function as follows:

[0059] in, and These are hyperparameters, representing the weights of the corresponding loss function, which can be set or adjusted as needed. The overall training is optimized based on the total loss function. During the optimization training process, the parameters of the text encoder and image encoder in the feature extraction model are optimized, while the model parameters of the multimodal interactive encoder, the cross-lingual distillation mask image model, the cross-lingual mask language model, and the cross-lingual image-text matching model are also optimized. Therefore, in this embodiment, the overall training model includes a feature extraction model, a feature fusion model, and a prediction part. The prediction part includes four independent branches: the cross-lingual distillation mask image model, the cross-lingual mask language model, the cross-lingual image-text comparison learning model, and the cross-lingual image-text matching model. A total loss function is constructed using the four loss functions of the four branches of the prediction part. The overall training model is trained in reverse to achieve end-to-end model training. In each batch of training, the model parameters of the feature extraction model, feature fusion model, and each branch of the prediction part are optimized. The cross-language image-text comparison learning model is directly built based on similarity scores and does not involve internal model parameters. It only provides a third loss function as part of the total loss function and does not involve model parameter optimization training. Therefore, in each batch of training, the images and mask images, the first language description text and mask text, and the second language description text and mask text of that batch are input into the overall training model. First, the feature extraction model extracts features from the images and text to obtain the required image features and text features. Then, the feature fusion model obtains various fused features. The image features, text features, and fused features are input into the corresponding branches of the prediction part. The loss function of each branch is constructed based on the output value and label value of each branch. Then, the total loss function of the overall training model is obtained. After optimizing the model parameters of each part in the overall training model based on the total loss function, the training of the next batch continues. By employing this end-to-end multi-loss function joint optimization, bidirectional semantic reasoning between images and text, and between texts in different languages, is achieved. No prior knowledge is required, and implicit associations between different modalities and languages ​​can be adaptively learned, thereby effectively solving the modal heterogeneity problem and improving the accuracy of cross-modal and cross-language matching.

[0060] After model training, the testing process using test data includes: To verify the effectiveness of the above model training method, test data is used for testing. The test data includes a query text set Q (in different languages) and an image set gallery (a collection of multiple images). The goal of image retrieval is to find the pedestrian image in the gallery that best matches the semantics of a given query description text q. The testing method involves inputting q into a text encoder to obtain text features, and inputting all images in the image set into an image encoder to obtain image features. By calculating the similarity between the text features and image features (e.g., using cosine similarity), the matching degree between the text features of q and all images in the gallery is calculated. The pedestrian image with the highest matching degree is the result. By comparing the matching results with the true labels, the accuracy of the trained model can be evaluated by comparing the number of accurate matching results with the total number of queries.

[0061] like Figure 8 As shown in the embodiments of this application, an image and text retrieval method is also provided, including the following steps: S210: Obtain the description text to be retrieved; In this embodiment, the descriptive text to be retrieved is the given text that needs to be searched; S220: Input the descriptive text to be retrieved into the feature extraction model obtained based on the cross-modal feature extraction model training method to obtain the features of the descriptive text to be retrieved; Since the feature extraction model obtained by the above training method is applicable to image and text retrieval of descriptive text in various languages, after the descriptive text in various languages ​​is input into the feature extraction model, text features that can accurately express the descriptive text can be obtained. S230: Based on the features of the descriptive text to be retrieved, match the features of the candidate images to determine the image features that match the features of the descriptive text to be retrieved; Image features can be obtained by pre-inputting each candidate image from the candidate image set into a feature extraction model trained using a cross-modal feature extraction model. In practical image-text retrieval applications, it is necessary to obtain the image features of each candidate image beforehand. Specifically, the image encoder in the feature extraction model extracts image features from each candidate image. By optimizing the image encoder using the aforementioned cross-modal feature extraction model training method, the accuracy of image feature extraction is improved, making image features more suitable for matching with text features. This achieves cross-modal alignment of image and text features, improving the accuracy of image and text feature comparison. In other alternative implementations, image features can also be extracted using other feature extraction models.

[0062] By adopting the image-text retrieval method of this application, it can be applied to cross-modal image-text retrieval in multiple language environments, improve image-text retrieval efficiency, and effectively reduce matching errors caused by translation issues.

[0063] like Figure 9 As shown in the embodiments of this application, an image and text retrieval system is also provided, including: The text acquisition module M100 is used to acquire the descriptive text to be retrieved; The feature extraction module M200 is used to input the descriptive text to be retrieved into the feature extraction model obtained based on the cross-modal feature extraction model training method, and obtain the features of the descriptive text to be retrieved. The feature matching module M300 is used to match the features of the descriptive text to be retrieved with the features of the candidate images to determine the image features that match the features of the descriptive text to be retrieved.

[0064] In the image and text retrieval system of this application, the functions of each module can be implemented using the specific implementation method of the cross-modal feature extraction model training method described above, which will not be elaborated here.

[0065] By adopting the image and text retrieval system of this application, it can be applied to cross-modal image and text retrieval in multiple language environments, improve image and text retrieval efficiency, and effectively reduce matching errors caused by translation issues.

[0066] An exemplary embodiment of this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the cross-modal feature extraction model training method described above.

[0067] In one embodiment, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium can be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid-state drive (SSD), etc. Exemplarily, the computer program product can be implemented as a non-volatile storage medium storing a computer program, such as read-only memory, NAND flash memory, etc.

[0068] In one implementation, the computer program product can be an intangible product containing a computer program. For example, the computer program product can be implemented as a virtual digital product, such as an executable file, installation package, or other digital file storing the computer program.

[0069] Computer program code can be written in one or more programming languages. Examples of programming languages ​​include C, Java, C++, and Python. Program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, such as a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via an internet connection provided by a mobile network operator).

[0070] Computer programs can be carried or transmitted via signals such as electricity, magnetism, light, electromagnetic fields, and infrared radiation. Electronic devices can convert the signals carrying computer programs into digital signals, thereby running the computer programs. When a computer program runs on an electronic device, its code is used to cause the electronic device to execute (more specifically, the processor of the electronic device to execute) the method steps of various exemplary embodiments of this application, such as the steps of the cross-modal feature extraction model training method described above.

[0071] When the computer program is executed by the processor, it implements the steps of the cross-modal feature extraction model training method described above. Therefore, the computer program product can also obtain the technical effects of the cross-modal feature extraction model training method described above.

[0072] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of this application, and all such modifications or substitutions should be considered within the scope of protection of this application.

Claims

1. A method for training a cross-modal feature extraction model, characterized in that, Includes the following steps: Collect training data, which includes multiple images and their corresponding descriptive text groups, each of which includes descriptive text in a first language and a second language with the same meaning; Based on the feature extraction model, feature extraction is performed on the training data to obtain training image features and training text features. Based on the training data, a mask image and mask text are obtained. Feature extraction is performed on the mask image to obtain mask image features, and feature extraction is performed on the mask text to obtain mask text features. A first loss function for a cross-language distilled mask image model is constructed based on the mask image features and the training text features of the first and second languages. Specifically, a first fusion feature of the training image features and the training text features of the first language is obtained, and a second fusion feature of the mask image features and the training text features of the second language is obtained. The first fusion feature is used as a supervision signal, and the second fusion feature is used as the input signal for the cross-language distilled mask image model to construct the first loss function. A second loss function for a cross-language masked language model is constructed based on the training image features and the masked text features of the first and second languages. The optimized feature extraction model is obtained by performing optimization training based on the first loss function and the second loss function.

2. The cross-modal feature extraction model training method according to claim 1, characterized in that, A first fused feature is obtained by using a feature fusion model to obtain the training image features and the training text features of the first language, and a second fused feature is obtained by using the feature fusion model to obtain the mask image features and the training text features of the second language.

3. The cross-modal feature extraction model training method according to claim 2, characterized in that, Constructing the first loss function includes the following steps: A first loss function is constructed based on the similarity between the output of the cross-language distillation mask image model and the supervision signal.

4. The cross-modal feature extraction model training method according to claim 1, characterized in that, Modeling is performed based on the training image features and the masked text features of the first and second languages ​​to construct a second loss function for the cross-language masked language model, including the following steps: A feature fusion model is used to obtain a third fusion feature of the training image features and the masked text features of the first language, and a fourth fusion feature of the training image features and the masked text features of the second language is obtained. The third and fourth fusion features are used as inputs to the cross-language masked language model, and the training text features of the first language are used as labels. A second loss function is constructed based on the output of the cross-language masked language model and the labels.

5. The cross-modal feature extraction model training method according to claim 1, characterized in that, It also includes the following steps: The third loss function of the cross-language image-text comparison learning model is constructed based on the training image features and the training text features of the first and second languages. The optimization training includes optimizing the training based on the first loss function, the second loss function, and the third loss function.

6. The cross-modal feature extraction model training method according to claim 1, characterized in that, It also includes the following steps: A fourth loss function for the cross-language image-text matching model is constructed based on the training image features, the mask image features, and the training text features of the first and second languages. The optimization training includes optimizing the training based on the first loss function, the second loss function, and the fourth loss function.

7. The cross-modal feature extraction model training method according to claim 6, characterized in that, A fourth loss function is constructed based on the training image features, the mask image features, and the training text features of the first and second languages, including the following steps: A feature fusion model is used to obtain a fifth fusion feature of the training image features and the training text features of the first language, and a sixth fusion feature of the mask image features and the training text features of the second language is obtained. Based on the fifth fusion feature as the first input of the cross-language image-text matching model, the first part of the fourth loss function is constructed. Based on the sixth fusion feature as the second input of the cross-language image-text matching model, the second part of the fourth loss function is constructed. The complete fourth loss function is obtained based on the first and second parts of the fourth loss function.

8. A method for image and text retrieval, characterized in that, Includes the following steps: Retrieve the description text to be searched; The descriptive text to be retrieved is input into a feature extraction model obtained based on the cross-modal feature extraction model training method according to any one of claims 1 to 7, to obtain the features of the descriptive text to be retrieved; Based on the matching of the descriptive text features to be retrieved with the candidate image features, the image features that match the descriptive text features to be retrieved are determined.

9. A text and image retrieval system, characterized in that, include: The text acquisition module is used to acquire the descriptive text to be retrieved; The feature extraction module is used to input the descriptive text to be retrieved into the feature extraction model obtained based on the cross-modal feature extraction model training method according to any one of claims 1 to 7, and obtain the features of the descriptive text to be retrieved; The feature matching module is used to match the features of the descriptive text to be retrieved with the features of the candidate images to determine the image features that match the features of the descriptive text to be retrieved.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the image and text retrieval method according to claim 8.