Computer-implemented method, system for computer learning, and medium

CN115374798BActive Publication Date: 2026-09-11BAIDU USA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210554578.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-07
Filing Date
2022-05-19
Publication Date
2026-09-11
Estimated Expiration
2042-05-19

Smart Images

  • Figure CN115374798B_ABST
    Figure CN115374798B_ABST
Patent Text Reader

Abstract

The present disclosure relates to computer-implemented methods, systems for computer learning, and media. Current pre-trained vision-language models for English cross-modal retrieval tasks depend on the availability of many annotated image-caption datasets for pre-training with English text. However, the text is not necessarily in English. While machine translation (MT) tools can be used to translate the text into English, the performance is largely dependent on the quality of the MT and can be impacted by high latency issues in real-world applications. Embodiments herein address these issues by learning cross-language cross-modal representations for matching images and their associated captions in multiple languages. Embodiments seamlessly combine cross-language pre-training objectives and cross-modal pre-training objectives in a unified framework to learn images and text in a joint embedding space from available English image-caption data, monolingual corpora, and parallel corpora.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This patent application relates to jointly pending and jointly owned U.S. Patent Application No. 63 / 190,667 (Case No. 28888-2494P), filed May 19, 2021, entitled “SYSTEMS AND METHODS FOR CROSS-LINGUAL CROSS-MODAL TRAINING FOR MULTIMODALRETRIEVAL”, and claims priority to it under 35 USC §119(e), the patent document of which is incorporated herein by reference in its entirety and for all purposes. Technical Field

[0003] This disclosure generally relates to systems and methods for computer learning that can provide improved computer performance, features, and uses. More specifically, this disclosure relates to systems and methods for cross-language, cross-modal training for multimodal retrieval and for deploying trained multimodal retrieval models. Background Technology

[0004] Recently, converter-based pre-trained visual-language models have achieved outstanding performance in cross-modal retrieval, image captioning, and visual question answering (VQA) tasks in English. For example, in recent VQA competitions, most leading competitors relied on converter-based pre-trained visual-language models. However, their success largely depends on the availability of large amounts of annotated image-captioning pre-training datasets (e.g., concept captions). In reality, such data for other languages ​​is limited.

[0005] When extending to cross-language and cross-modal applications, a straightforward approach is to use machine translation (MT) tools to translate non-English text into English and reuse pre-trained English models. However, performance is highly dependent on the capabilities of the MT tools and is hampered by high latency issues in real-world applications.

[0006] To learn multilingual and multimodal representations, recent researchers have used multilingual datasets to model images and text captions in a joint embedding space. Based on a shared feature space learning approach, two types exist: word-level alignment and sentence-level alignment. These models can capture a certain degree of semantic similarity between language and images. However, they only model the relevance of global features between text and images. This limitation may prevent these models from effectively detecting relevance locally. Meanwhile, cross-lingual language models such as multilingual BERT and XLM, as well as pre-trained vision-language models, are already common in bridging different languages ​​and modalities. These models use a transformer architecture trained simultaneously from multilingual or image-caption pairs to build the encoder and fine-tune it for downstream task-specific objectives. This process achieves full cross-lingual and modal interaction. However, current cross-lingual and cross-modal models are trained on multilingual corpora and English caption data, respectively. Therefore, the resulting pre-trained models are not directly applicable to downstream cross-modal tasks involving non-English languages.

[0007] Therefore, what is needed is a system and method that provides cross-language, cross-modal pre-training framework implementations to learn language-invariant representations across image and text modalities. Attached Figure Description

[0008] Reference will be made to embodiments of the present disclosure, and examples of embodiments may be illustrated in the accompanying drawings. These drawings are intended to be illustrative and not restrictive. Although the present disclosure has been generally described in the context of these embodiments, it should be understood that this is not intended to limit the scope of the disclosure to these particular embodiments. Items in the drawings may not be drawn to scale.

[0009] Figure 1 The illustrations depict cross-language and cross-modal relationships between data according to embodiments of the present disclosure.

[0010] Figure 2 A pre-trained model according to an embodiment of the present disclosure is illustrated.

[0011] Figure 3 The flowchart of a cross-modal text recovery system and method according to embodiments of the present disclosure is illustrated.

[0012] Figure 4 A method for pre-training a cross-language, cross-modal model according to embodiments of the present disclosure is described.

[0013] Figure 5 An architecture for fine-tuning a pre-trained cross-language cross-modal (CLCM) network according to an embodiment of the present disclosure is illustrated.

[0014] Figure 6 A method for fine-tuning according to embodiments of the present disclosure is described.

[0015] Figure 7 A method for finding a set of one or more related images using a cross-language cross-modal (CLCM) system and query text, according to embodiments of the present disclosure, is described.

[0016] Figure 8 A method for finding one or more related texts using a CLCM system and an input image, according to embodiments of the present disclosure, is described.

[0017] Figure 9 Table 2 is included, which depicts the cross-modal retrieval results (expressed as a percentage %) in English according to embodiments of this disclosure.

[0018] Figure 10 Table 3 is included, which depicts the cross-modal retrieval results for Japanese (dataset 2) and German (dataset 1) according to embodiments of this disclosure.

[0019] Figure 11 A simplified block diagram of a computing device / information processing system according to embodiments of the present disclosure is depicted. Detailed Implementation

[0020] In the following description, specific details are set forth for purposes of explanation in order to provide an understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these details. Furthermore, those skilled in the art will recognize that the embodiments of this disclosure described below can be implemented in various ways, such as as a process, apparatus, system, device, or method on a tangible computer-readable medium.

[0021] The components or modules shown in the accompanying drawings are illustrative of exemplary embodiments of this disclosure and are intended to avoid obscuring this disclosure. It should also be understood that throughout the discussion, a component can be described as a separate functional unit that may include subunits; however, those skilled in the art will recognize that various components or portions thereof may be divided into separate components or may be integrated together, including, for example, within a single system or component. It should be noted that the functions or operations discussed herein can be implemented as components. Components can be implemented in software, hardware, or a combination thereof.

[0022] Furthermore, the connections between components or systems shown in the accompanying drawings are not intended to be limited to direct connections. Rather, data between these components can be modified, reformatted, or otherwise altered via intermediate components. Additionally, additional or fewer connections may be used. It should also be noted that the terms “coupled,” “connection,” “communicationally coupled,” “interface connection,” “interface,” or any derivative thereof should be understood to include direct connections, indirect connections via one or more intermediate devices, and wireless connections. It should also be noted that any communication, such as signals, responses, replies, acknowledgments, messages, queries, etc., may involve one or more exchanges of information.

[0023] The use of terms such as "one or more embodiments," "preferred embodiments," "embodiments," and "multiple embodiments" in the specification means that a particular feature, structure, characteristic, or function described in connection with an embodiment is included in at least one embodiment of this disclosure and may be included in more than one embodiment. Furthermore, the above phrases appearing in various places in the specification do not necessarily refer to the same one or more embodiments.

[0024] The use of certain terms throughout this specification is illustrative and should not be construed as limiting. Services, functions, or resources are not limited to a single service, function, or resource; the use of these terms may refer to a group of related services, functions, or resources that may be distributed or aggregated. The terms “include,” “includes,” “contains,” and “comprises” should be understood as open-ended terms, and any subsequent listings are illustrative and do not imply limitation. A “layer” may include one or more operations. The terms “best,” “optimal,” “optimized,” etc., refer to improvements in results or processes and do not require that a particular result or process has reached an “optimal” or peak state. The use of terms such as memory, database, repository, data storage, table, hardware, cache, etc., may be used herein to refer to one or more system components in which information can be input or otherwise recorded.

[0025] In one or more embodiments, the stopping condition may include: (1) a certain number of iterations have been performed; (2) a processing time has been reached; (3) convergence (e.g., the difference between successive iterations is less than a first threshold); (4) divergence (e.g., performance degradation); (5) an acceptable result has been achieved; and (6) all data in the data has been processed.

[0026] Those skilled in the art will recognize that: (1) certain steps may be performed optionally; (2) the steps may not be limited to the specific order set forth herein; (3) certain steps may be performed in different orders; and (4) certain steps may be performed simultaneously.

[0027] Any headings used herein are for organizational purposes only and should not be used to limit the scope of the specification or claims. Every reference / document mentioned in this patent document is incorporated herein by reference in its entirety.

[0028] It should be noted that any experiments and results provided herein are provided by way of illustration and performed under specific conditions using one or more specific embodiments; therefore, these experiments and their results should not be used to limit the scope of disclosure of this patent document.

[0029] A. Overview

[0030] Pre-trained visual-language models have achieved excellent performance in multimodal applications involving the English language. However, these successes largely depend on the availability of large-scale image-text data for pre-training. A key issue is the lack of large-scale datasets in other languages. To address this lack of large-scale datasets in other languages, implementations seek to transfer knowledge between non-English languages ​​and visual modalities via English as a bridge, such as... Figure 1 This is illustrated in diagram form. For example... Figure 1 As shown, one or more large datasets of source language text 115 and visual data (e.g., images and videos) 120 form source language-image data. Therefore, source language data 115 can serve as a bridge between image data 120 and one or more target languages ​​(e.g., target languages ​​105, 110).

[0031] This paper presents a system and method for a cross-linguistic, cross-modal pre-training framework to learn language-invariant representations across image and text modalities. Introducing pre-training objectives related to other languages ​​and modeling interactions between English and other languages ​​yields better representations that generalize well to downstream tasks. The embodiment incorporates monolingual and parallel corpora related to other languages ​​to further refine the shared latent space, thereby extending visual-linguistic pre-training efforts based on English image captioning data to adjust parameters.

[0032] Figure 2 A schematic view of an embodiment of a pre-training framework according to embodiments of the present disclosure is depicted in a graphical manner. In one or more embodiments, Figure 2 The framework embodiments depicted are built upon a vision-language natural language processing (e.g., VL-BERT) model 220, with more data sources and more pre-trained targets covering different languages ​​and modalities. In one or more embodiments, the backbone network is a single-stream multimodal BERT variant with cross-attention between text and image bounding box features.

[0033] In the depicted embodiment, the input data includes three sources: source language captions (e.g., English captions) and corresponding visual bounding box features 215; parallel sentences 210 involving the source language (e.g., English) and other languages ​​for establishing connections between the other languages ​​and the source language, as well as connections from the source language to the visual domain; and a monolingual text corpus 205. There is a correspondence between the data sources and pre-trained tasks encoded with different line types. Therefore, the MLM (Masked Language Modeling) task, the MRC (Masked Region Classification) task, and the CMTR (Cross-Modal Text Retrieval) task 230 are associated with the source language caption input 215, which includes classification tokens (CLS), source language captions (i.e., w1, w2, ...), separator tokens (SEP), and corresponding image features (i.e., B1, ..., B...). N The TLM (Translation Language Modeling) task and CLTR (Cross-Language Text Recovery) task 235 are associated with a parallel corpus 210, which consists of source language captions and corresponding (or parallel) target language captions separated by one or more separator tokens. Finally, the MLM task 240 is associated with a monolingual text 205, which can be text in any language.

[0034] The acronyms for the pre-training tasks are summarized in Table 1 below:

[0035] Table 1—Commonly Used Acronyms

[0036]

[0037] For the language component, in one or more embodiments, masked language modeling (MLM) is used on monolingual text corpora (Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, 2019, BERT: Pre-Training Of Deep Bidirectional Transformers For Language Understanding, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pp. 4171-4186, Minneapolis, MN (hereinafter referred to as "Devlin et al., 2019")), and translated language modeling (TLM) derived from XLM is used on parallel text corpora (Alexis Conneau and Guillaume). Lample, 2019, Cross-lingual Language Model Pretraining, Advances in Neural Information Processing Systems (NeurIPS), Vancouver, Canada, 7057-7069 (hereinafter referred to as "Conneau and Lample, 2019"), the cited literature is incorporated herein by reference in its entirety. In one or more embodiments, the visual-language portion follows a standard visual-language pretrained model and MLM is used for text captioning and masked region classification (MRC).Cross-lingual text recovery (CLTR) tasks have been used, for example in Unicoder (Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang and Ming Zhou, 2019, Unicoder: A Universal Language Encoder by PretrainingWith Multiple Cross-Lingual Tasks, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 2485-2494, Hong Kong, China (hereinafter referred to as "Huang et al., 2019"), which is incorporated herein by reference in its entirety. A related but distinct task (Cross-modal Text Recovery (CMTR)) was developed, and an implementation of CMTR is presented in this paper. Like CLTR, CMTR utilizes an attention matrix between image-caption pairs to learn the alignment between words and regions of interest in an image.

[0038] In one or more embodiments, text-to-image and image-to-text retrieval tasks are performed on two multimodal, multilingual image captioning benchmarks: Dataset 1 (German and English) captions and Dataset 2 (English and Japanese). State-of-the-art (SOTA) results are achieved for retrieval tasks involving Japanese and German, compared to machine translation baselines and other recently published results.

[0039] B. Some related work

[0040] 1. Visual-Language Pre-trained Model

[0041] Recently, BERT-based vision-language pre-trained models have emerged. In these models, pre-training typically includes three types of tasks: 1) masked language modeling, 2) masked region modeling, and 3) text-image matching. By leveraging cross-modal attention and pre-training on large-scale datasets, cross-modal BERT methods have achieved state-of-the-art performance on many text-visual understanding tasks. However, all of the aforementioned models only handle single-language English and image or video domains.

[0042] 2. Cross-language pre-trained models

[0043] Cross-lingual pre-trained language models can encode text from multiple languages ​​simultaneously. Most notably, the multilingual BERT (Devlin et al., 2019, cited in its entirety) employs the same model structure and training objectives as BERT, but is pre-trained on Wikipedia for more than 100 languages. XLM models are pre-trained using MLM and TLM to leverage parallel sentence resources (if available). Evaluations on a range of cross-lingual transfer tasks have demonstrated the significant utility of these cross-lingual LMs for knowledge transfer between languages. For example, U.S. Patent Application 17 / 027,560 (Case No. 28888-2427), filed March 19, 2021, by inventors Fei Hongliang and Li Ping, entitled "CROSS-LINGUAL UNSUPERVISED SENTIMENT CLASSIFICATION WITH MULTI-VIEW TRANSFER LEARNING," evaluates a series of cross-lingual transfer tasks. The aforementioned U.S. Patent Application pursuant to 35 USC §119(e), filed June 16, 2020, by inventors Fei Hongliang and Li Ping, entitled "CROSS-LINGUAL UNSUPERVISED SENTIMENT CLASSIFICATION WITH MULTI-VIEW TRANSFER," also addresses this issue. The priority interest in co-pending and co-owned U.S. Patent Application 63 / 039,967 (Case No. 28888-2427P) entitled “LEARNING [Cross-linguistic Unsupervised Emotion Classification with Multi-view Transmission Learning]” (each patent document is incorporated herein by reference in its entirety and for all purposes).

[0044] The embodiments in this paper present the integration of cross-lingual pre-training tasks with visual language pre-training to obtain general multilingual multimodal representations.

[0045] C. Method Implementation Examples

[0046] The framework implementation follows the network structure of VL-BERT presented in the following literature: Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai, “VL-BERT: Pretraining of Generic Visual-Linguistic Representations”, Proceedings of the 8th International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia (2020) (hereinafter referred to as “Su et al., 2020”), which is incorporated herein by reference in its entirety. VL-BERT is a single-stream cross-modal model that concatenates word features from text with bounding box features from images and feeds the concatenated sequence into a series of transformer blocks. While the embodiments described herein may use or be adapted to VL-BERT as part of the framework, it should be noted that other models and network structures may be used.

[0047] 1. Example of a pre-training task

[0048] The model implementations used a vision-based masked language modeling (MLM) task and a text-based masked region classification (MRC) task for image-caption data. In one or more implementations, masked language modeling based on other language texts was also used due to the introduction of an auxiliary multilingual text corpus.

[0049] The pre-trained model can be further improved by involving more tasks; therefore, in one or more embodiments, two additional cross-language pre-training tasks and one cross-modal task are introduced and employed to improve performance.

[0050] a) Masked Language Modeling (MLM) Examples

[0051] Masked Language Modeling (MLM) is a language modeling in which a portion of the input is masked and the model learns to predict the missing tokens. An example training object is discussed in Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu, UNITER: Universal Image-Text Representation Learning (2020) (hereinafter referred to as "Chen et al., 2020"), which is incorporated herein by reference in its entirety. In one or more embodiments, an image region can be represented as r = {r1, ..., r...} K The input words can be represented as w = {w1, ..., w}. T}, and the mask index is represented as in, `m` is a natural number, `M` is the number of mask tokens, and `m` is the mask index set. In MLM, the input word is masked with a given probability (e.g., 15%). The masked word `w` m Replaced with a special token (e.g., [MASK]). The goal is based on the words surrounding these masked words. \m And for all observations of image region v, these masked words are predicted by minimizing the following negative log-likelihood equation:

[0052]

[0053] Here, θ represents the trainable parameters. Each pair (w,v) can be sampled from the training set D.

[0054] b) Mask Region Classification (MRC) Example

[0055] Similar to MLM, image regions can be sampled, and their visual features are masked with a certain probability (e.g., 15%). Given a remaining region v... \m Given all words w, the model can be trained to reconstruct the masked region v. m The visual features of the masked region can be replaced by zero. Because visual features are high-dimensional and continuous, they cannot be supervised via similar methods. The basic objective can be defined as:

[0056]

[0057] As discussed in Chen et al., 2020, MRC learns to predict the semantic category of the object for each masked region. The masked regions can be first... The transformer output is fed into a fully connected layer to predict scores for K object categories, which are further transformed into a normalized distribution by a softmax function. Note that there may be no real labels. Therefore, object detection outputs from object detection models (such as Faster R-CNN (Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018, Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering), Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6077-6086, Salt Lake City, Utah (hereinafter referred to as "Anderson et al., 2018"), which are incorporated herein by reference in their entirety) can be used, and the detected object category with the highest confidence score can be used as the label of a mask region, which is then converted into a one-hot vector. The ultimate goal can be defined as minimizing the cross-entropy (CE) loss, as shown below:

[0058]

[0059] c) Cross-modal text recovery (CMTR) example

[0060] Figure 3 A cross-modal text recovery system and method flow 300 according to embodiments of the present disclosure is illustrated. As shown, the system includes a set of embedding layers or inputs—token embedding 305, position embedding 310, segment embedding 315, and visual feature embedding 320. The token embedding represents the embedding of the input word. The visual feature embedding is a vector representation of different bounding boxes extracted from the input image using an object detection model such as Faster R-CNN (Anderson et al., 2018). The position embedding represents the position of the input and can be performed sequentially; note that in this embodiment, the visual feature embeddings of different bounding boxes from the input image all have the same position embedding value. Finally, the segment embedding represents this type of input. Figure 3 In the depicted embodiments, "T" indicates that the input is text, and "I" indicates that the input is obtained from an image. System 300 also includes an attention matrix 325 that measures the similarity or relevance between the input text token and input visual features. The output of the attention matrix is ​​used to compute an attention representation of the input text token and bounding box features. In one or more embodiments, the attention representation is then fed into a transformer layer 330, which may be a VL-BERT implementation (e.g., backbone network 220), and a recovery loss is evaluated. Descriptions of embodiments of the method are described below.

[0061] In one or more embodiments, the CMTR system 300 directly learns the potential alignment between words in an image and regions of interest, and generates attentional input to stacked transformer layers to recover all input words. Note that the attention matrix 325 is transposed.

[0062] like Figure 3 As shown, the CMTR embodiment uses image-caption pairs as input, but does not use the original caption words. Instead, this CMTR embodiment computes the alignment between word features extracted by a tool (e.g., Faster-RCNN) and bounding box features, and uses attention features to simultaneously recover all input words. Specifically, let (B, E) be the image-caption input pair, where B = (b1, b2, ..., b...). n ) is the bounding box feature embedding and E = (e1, e2, ..., e m () is the word embedding. In one or more embodiments, the CMTR embodiment first calculates the attention representation of the caption word and the bounding box features as follows: in, And h represents the embedding dimension. It is calculated using bilinear attention. The attention matrix, where W represents the trainable parameters. Finally, in one or more embodiments, the attention matrix is... As input, predict the words in the original caption. In one or more embodiments, the objective function is:

[0063]

[0064] Where Δ(.,.) is the sum of token-level cross-entropy losses, and e(.) is the encoder component that includes the input layer, attention layer, and converter layer. d(.) is the decoder applied to the output of the converter, which can be a linear projection layer shared with other MLM and CLTR tasks described below.

[0065] d) Cross-Language Text Recovery (CLTR) Example

[0066] This task (CLTR) can be considered an adaptation of Unicoder (Huang et al., 2019), which employs parallel sentence pairs (X,Y) and enables a pre-trained model to learn latent word alignments between the two languages. In one or more embodiments, the model structure of the CLTR task is similar to... Figure 3 The model structure of the CMTR shown is the same. Therefore, similar to the CMTR implementation, a bilinear attention mechanism (e.g., attention matrix 325) is also used to compute the attention representations of one sentence X and another sentence Y in the source language. And then try to use attention input To recover the input X. In the CLTR task, the same objective function in equation (1) can be optimized. Note that in one or more embodiments, CLTR and CMTR do not share attention parameters because a large modal gap still exists between the text and the image before cross-attention is applied.

[0067] Therefore, given a pair of two statements (X,Y), where X=(x1,x2,…,x…), m Y is a sentence with m words from source language s, where Y = (y1, y2, ..., ym). n Given a sentence with n words from the target language t, the task is to first embed each x into all the words of Y. i Represented as

[0068]

[0069] in, and They represent x respectively i and y j Let h represent the word embedding dimension, and Ω∈R. m×n The attention matrix is ​​calculated using the following formula:

[0070]

[0071] Where, V∈R m*n These are trainable weights. Then, the system will... As input, it attempts to predict the original word sequence X.

[0072] e) Examples of Translation Language Models

[0073] This task (TLM) can be considered an adaptation of XLM (Conneau and Lample, 2019) which takes parallel sentence pairs with random mask tokens in different languages ​​as input. The model is trained to predict mask tokens by paying attention to both the local context and the remote context of the other language.

[0074] It should be noted that TLM essentially shares the same objective function as MLM, but takes parallel sentences as input. Therefore, instead of considering monolingual text streams, it concatenates parallel sentences, such as... Figure 2 As shown in item 210, words in both the source and target sentences can be randomly masked. To predict the masked words in the first sentence of the first language or the source language, the model can focus on surrounding words in the first or source language, or on the translation of the second or target language. Therefore, the model is encouraged to align the source and target language representations. To facilitate alignment, the position of the target sentence can also be reset.

[0075] f) Examples of pre-training methods

[0076] Figure 4 An example overview of a pre-training method according to embodiments of the present disclosure is described. In one or more embodiments, given a first batch of training data including source language captions and corresponding visual features, the first batch of training data is used as input (405) to train a cross-language, cross-modal network. Losses for MLM, MRC, and CMTR tasks are computed (410) based on the first batch of training data. Note that for the CMTR task, such as... Figure 3 As shown, unlike the MLM and MRC tasks, the embeddings of the first batch of training data are directly input into the transformer layer, but instead are input into an attention matrix that provides cross-attention between the text input and the visual features of the image bounding box.

[0077] Given a second batch of training data consisting of a set of source language texts and a corresponding set of target language texts, the second batch of training data is used as input (415) to train the cross-language, cross-modal network. The losses for the (420) TLM and CLTR tasks are calculated based on the second batch of training data. Note that, as with the CMTR task, for the CLTR task, the embeddings of the second batch of training data are not directly input into the transformer layer, but rather into the attention matrix.

[0078] Given a third training data batch containing monolingual text, the third training data batch of monolingual text is used as input (425) to train a cross-language cross-modal network, and the loss of the (430) MLM task is calculated based on the third training data batch.

[0079] The cross-lingual, cross-modal network is updated via backpropagation using aggregated losses from various tasks, including MLM, MRC, CMTR, TLM, CLTR, and monolingual MLM. The aggregated losses can be combined uniformly or in a weighted manner, where the weights can be training parameters or hyperparameters. In one or more embodiments, the computation graph can be used to track the loss and relevant parameters associated with specific loss components for updating; that is, the attention matrices for the CMTR and CLTR tasks can be updated appropriately.

[0080] If the stopping condition has not been met, select (445) the next first training data batch, second training data batch and third training data batch (which can be collectively referred to as super batch), and repeat the process by returning to step 405.

[0081] If the stopping condition has been met, the pre-trained cross-lingual cross-modal model is output. In one or more embodiments, the pre-trained cross-lingual cross-modal model can then be fine-tuned, which will be discussed in the next section.

[0082] It should be noted that for pre-training, the MLM, TLM, CLTR, and CMTR tasks can share a linear projection layer (which can have a size of hidden_dimension * vocabulary size) 335 at each output token. Furthermore, the MRC task can have its own linear projection layer (which can have a size of hidden_dimension * object_type_size).

[0083] It should also be noted that, in one or more embodiments, the source language in the first and second training datasets is a bridging language used to help support cross-language development—for example, the source language is English. The language of the single sentences in the third batch of data can be any language.

[0084] 2. Fine-tuning of the cross-modal retrieval embodiment

[0085] To help improve performance, especially for non-source languages, fine-tuning of the pre-trained network can be performed. One of the benefits of fine-tuning is that the one or more non-source languages ​​used in fine-tuning do not require as large datasets as would typically be needed to achieve the same or similar levels of performance.

[0086] Figure 5 An architecture for fine-tuning a pre-trained cross-language cross-modal (CLCM) network according to embodiments of the present disclosure is illustrated. In one or more embodiments, the CLCM network or system 510 includes a pre-trained backbone network 520 fed into a model head 525 (e.g., Figure 2 backbone network 220 / Figure 3 The converter layer 330), the model head can be a feedforward network that generates output 535 or another type of network.

[0087] For fine-tuning, in one or more embodiments, the triple ranking loss can be minimized to fine-tune the retrieval model embodiment (e.g., CLCM system 510). To improve performance, hard negative mining can be used.

[0088] Figure 6 A method for fine-tuning according to embodiments of the present disclosure is described. For each text query, in one or more embodiments, there is one positive (i.e., relevant) image sample, while the remainder are negative (irrelevant) image samples (605). Accordingly, for each image, there is one positive (i.e., relevant) text, while the remainder are negative (irrelevant) text (605). For each text, a loss can be determined or obtained based on a comparison between the relevance output of the CLCM system for the text given a positive image and the relevance output of the CLCM system for the text given a negative image (e.g., the worst negative image). For example, by... This represents a mini-batch of training samples, where query q i With Image I i Relatedly, in one or more embodiments, only the most difficult negative image in that mini-batch is penalized by the following formula:

[0089]

[0090] Where m is a margin set to 0.2 by default (although other values ​​can be used), and [x] + =max(0, x) is the clip function. R(q, I) is a function used to evaluate the similarity between query q and image I, parameterized by u and b:

[0091]

[0092] Where u represents a linear layer attached to the pooled VL-BERT representation (e.g., the backbone network) and receives classified [CLS] tokens.

[0093] For each image, the loss can be determined or obtained based on a comparison between the CLCM system’s relevance output for the image given positive text and the CLCM system’s relevance output for the image given negative text (e.g., the worst negative text).

[0094] For each image, in one or more embodiments, only the hardest negative query in a mini-batch is penalized by the following formula:

[0095]

[0096] Considering the entire mini-batch of images and text, the final loss function can be calculated using the following formula:

[0097]

[0098] The loss can be used to update (620) the CLCM system 510 to fine-tune it. After fine-tuning, the output (625) is the fine-tuned CLCM system.

[0099] It should be noted that in one or more embodiments, more than one negative sample can be used to obtain any of the losses described above.

[0100] 3. Use / Deployment of CLCM Network Implementation Examples

[0101] Figure 7 A method for finding a set of one or more relevant images using a cross-language cross-modal (CLCM) system and query text, according to embodiments of the present disclosure, is described. In one or more embodiments, query text in a certain language is input (705) into the CLCM system. Given the query text and a set of images, the CLCM system is used (710) to obtain relevance scores for at least some images in the set of images. Note that in one or more embodiments, the language of the text query is one of the languages ​​used in fine-tuning. Based on the output relevance values ​​from the CLCM system, a set of the top K images relevant to the query text can be output (715), where K can be one or more.

[0102] Figure 8 A method for finding a set of one or more relevant texts using a CLCM system and an input image, according to embodiments of the present disclosure, is described. In one or more embodiments, a query image is input (805) into the CLCM system. Given a query image and a set of texts in one or more languages, the CLCM system is used (810) to obtain relevance scores for at least some of the texts in the set. Note that in one or more embodiments, the language of the text query is one of the languages ​​used in fine-tuning. Based on the output relevance values ​​from the CLCM system, a set of the top K texts relevant to the query image can be output (815), where K can be one or more.

[0103] D. Experiment

[0104] It should be noted that these experiments and results are provided by way of illustration and performed under specific conditions using one or more specific embodiments; therefore, these experiments and their results should not be used to limit the scope of disclosure of this patent document.

[0105] For pre-training, two English image-caption datasets were used: Dataset 3 and Dataset 4. A total of approximately 3.7 million text-image pairs were collected. For monolingual (en, de, ja) text and parallel corpora (en-de), 20 million sentences were randomly sampled from Wikipedia text and 9 million parallel sentences were randomly sampled from the MultiUN corpus. An additional 2.8 million en-ja parallel sentences were also collected.

[0106] For fine-tuning, two multilingual, multimodal datasets were used for retrieval: Dataset 1 and Dataset 2. Dataset 2 contains approximately 120,000 images, with five captions for each image. The English data was split into approximately 113,000 training samples, 5,000 validation samples, and 5,000 test samples.

[0107] A subset of approximately 33,700 images has Japanese subtitles generated for it. Within this subset, approximately 23,700 samples are used for training, 5,000 for validation, and 5,000 for testing. Dataset 1 contains approximately 32,000 images, each with five subtitles. This dataset is split into approximately 30,000 training samples, 1,000 validation samples, and 1,000 test samples.

[0108] R@K (K = 1, 5, 10) is used as the evaluation metric. R@K is the percentage of true matches that appear in the top K results.

[0109] 1. Experimental setup

[0110] A multilingual, case-insensitive version of BERT (Devlin et al., 2019) was used to initialize a test model instance with 12 layers of transformer blocks. Each block had 768 hidden units, 12 self-attention heads, and a vocabulary of 105,879. The maximum sequence length was set to 64. One hundred bounding boxes were detected for each image using Faster-RCNN (Anderson et al., 2018), pre-trained on an image dataset annotated with region descriptions, objects, attributes, and relationships. Pre-training was performed on 16 NVIDIA V100 GPUs (16GB of memory), and fine-tuning was performed on 8 NVIDIA V100 GPUs. FP16 (16-bit floating-point) was used to accelerate training and reduce memory usage. The Adam optimizer was used, and the batch size per GPU was set to 16. The initial learning rate was 1e-5. The model was pre-trained for 50 epochs and fine-tuned based on the average R@{1,5,10} on the validation set. The experiment was repeated five times and the average metric of the test set was recorded.

[0111] 2. Baseline

[0112] The model implementation is compared with several recent competing methods. VL-BERT (Su et al., 2020) and Unicoder-VL (Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang. 2020, Unicoder-VL: Universal Encoder for Vision and Language by Cross-Modal Pretraining, Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI), pp. 11336–11344, New York, NY (hereinafter referred to as "Li et al., 2020"), which is incorporated herein by reference in its entirety) are two well-known cross-modal BERT-based models. For VL-BERT, English results are reproduced by fine-tuning its official pre-trained model, and non-English results are generated from its published code following the same configuration as the model implementation. For Unicoder-VL, the English results reported in the paper are used.In addition to pre-trained models, several other methods were compared, including the cross-attention-based model SCAN (Lee et al., 2018), the multilingual word embedding alignment-based model AME (Alireza Mohammadshahi, Rémi Lebret and Karl Aberer. 2019, Aligning MultilingualWord Embeddings for Cross-Modal Retrieval Task, Proceedings of the Beyond Vision and Language: integrating Real-world knowledge (LANTERN@EMNLPIJCNLP), pp. 11-17, Hong Kong, China (hereinafter referred to as "Mohammadshahi et al., 2019"), which is incorporated herein by reference in its entirety, and the multilingual sentence alignment-based model LIME (Jonatas Wehrmann, Maurício Armani Lopes, Douglas...). M. Souza and Rodrigo C. Barros. 2019, Language-Agnostic Visual-Semantic Embeddings, Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 5803-5812, Seoul, South Korea (hereinafter referred to as "Wehrmann et al., 2019", the cited literature is incorporated herein by reference in its entirety). Performance was performed using SCAN, AME, and LIME as reported in their paper. Finally, a comparison was made with a machine translation baseline: "Translation-Test", which used Google Translate to translate test data from Japanese and German into English and was then evaluated on a finely tuned English VL-BERT retrieval model.

[0113] 3. Experimental Results

[0114] Included Figure 9Table 2 presents the results for English captions. Compared to Unicoder-VL (Li et al., 2020), the tested model instance performed slightly worse, but still achieved better results than VL-BERT. A possible reason is that Unicoder-VL is initialized using an English BERT specifically optimized for English. For cross-modal retrieval tasks involving non-English languages, Table 3 (which includes...) shows the results... Figure 10 The benefits of the tested model implementation are demonstrated in the (Chinese) section. The machine translation baseline “translation-test” was observed to achieve better results than VL-BERT, which is pre-trained on a multilingual corpus using an MLM objective and fine-tuned in the target language, thus demonstrating the importance of aligning different languages.

[0115] Furthermore, the average recall rate of the “translation test” was approximately 1%–2% lower than that of the tested method implementation. This result suggests that, for both benchmarks, pre-training with an additional cross-linguistic objective is more effective than translating the target language into English. While combining a more powerful machine translation tool with better fine-tuning of the English retrieval model can yield slightly better performance, the tested method implementation learns a general representation without relying on external machine translation tools specific to language pairs, making it more suitable for real-world applications.

[0116] Finally, the implementation with additional cross-lingual pre-training tasks resulted in performance improvements compared to VL-BERT (Su et al., 2020), which is pre-trained on multilingual corpora using only MLM tasks.

[0117] 4. Ablation Research

[0118] To understand the impact of different components, an ablation study was performed on the test set, and the average Recall@1 is reported in Table 4 below. Although the cross-lingual pre-training tasks (TLM and CLTR) did not significantly improve English-related retrieval tasks, they contributed more than 1% to the improvement in Japanese and German. This was expected, as these tasks effectively connect non-English languages ​​with the visual domain using English as a bridge. Among all components, CMTR consistently contributed approximately 1 point of improvement.

[0119] Table 4: Ablation studies of the mean R@1. Statistically significant results are marked in bold.

[0120]

[0121] 5. Some observations

[0122] This patent document presents embodiments of a multilingual corpus and three pre-trained objectives to improve a converter-based visual-language model for retrieval tasks. Extensive experiments demonstrate the effectiveness of the embodiments for cross-modal retrieval tasks. Detailed ablation studies prove the rationality of the modeling choices in the embodiments. Those skilled in the art will recognize that the embodiments can be extended for zero-shot transfers.

[0123] E. Computing System Implementation

[0124] In one or more embodiments, various aspects of this patent document may be directed to, may include, or may be implemented on one or more information processing systems (or computing systems). An information processing system / computing system may include any tool or collection of tools operable for calculating, estimating, determining, classifying, processing, transmitting, receiving, retrieving, initiating, routing, switching, storing, displaying, communicating, manifesting, detecting, recording, reproducing, disposing of, or utilizing information, intelligence, or data of any form. For example, a computing system may be or may include a personal computer (e.g., a laptop computer), a tablet computer, a mobile device (e.g., a personal digital assistant (PDA), a smartphone, a tablet phone, a tablet computer, etc.), a smartwatch, a server (e.g., a blade server or a rack server), a network storage device, a camera, or any other suitable device, and may vary in size, shape, performance, functionality, and price. A computing system may include random access memory (RAM), one or more processing resources (such as a central processing unit (CPU) or hardware or software control logic), read-only memory (ROM), and / or other types of memory. Additional components of a computing system may include one or more drives (e.g., hard disk drives, solid-state drives, or both), one or more network ports for communicating with external devices, and various input and output (I / O) devices (such as keyboards, mice, styluses, touchscreens, and / or video displays). The computing system may also include one or more buses operable for transmitting communication between various hardware components.

[0125] Figure 11 A simplified block diagram of an information processing system (or computing system) according to embodiments of the present disclosure is depicted. It will be understood that the functionality shown for system 1100 can operate to support various embodiments of the computing system—although it should be understood that the computing system can be configured differently and include different components, including those with fewer or more components, such as… Figure 11 The description.

[0126] like Figure 11As shown, the computing system 1100 includes one or more central processing units (CPUs) 1101 that provide computing resources and control the computer. The CPU 1101 may be implemented using a microprocessor or the like, and may also include one or more graphics processing units (GPUs) 1102 and / or floating-point coprocessors for mathematical calculations. In one or more embodiments, the one or more GPUs 1102 may be incorporated into a display controller 1109, such as part of one or more graphics cards. The system 1100 may also include system memory 1119, which may include RAM, ROM, or both.

[0127] Multiple controllers and peripherals can also be provided, such as Figure 11 As shown. Input controller 1103 represents an interface to various one or more input devices 1104 (such as a keyboard, mouse, touchscreen, and / or stylus). Computing system 1100 may also include a storage controller 1107 for interface connection to one or more storage devices 1108, each of which includes a storage medium (such as magnetic tape or disk, or an optical medium that can be used to record instruction programs for operating systems, utilities, and applications (which may include embodiments of programs implementing various aspects of this disclosure)). According to this disclosure, one or more storage devices 1108 may also be used to store processed data or data to be processed. System 1100 may also include a display controller 1109 for providing an interface to a display device 1111, which may be a cathode ray tube (CRT) display, a thin-film transistor (TFT) display, an organic light-emitting diode (OLED) display, an electroluminescent panel, a plasma panel, or any other type of display. Computing system 1100 may also include one or more peripheral controllers or interfaces 1105 for one or more peripheral devices 1106. Examples of peripheral devices may include one or more printers, scanners, input devices, output devices, sensors, etc. The communication controller 1114 can interface with one or more communication devices 1115, enabling the system 1100 to connect to remote devices via any of a variety of networks, including the Internet, cloud resources (e.g., Ethernet cloud, Ethernet Fibre Channel (FCoE) / Data Center Bridge (DCB) cloud, etc.), local area network (LAN), wide area network (WAN), storage area network (SAN), or via any suitable electromagnetic carrier signal including infrared signals. As depicted in the embodiments, the computing system 1100 includes one or more fans or fan disks 1118 and one or more cooling subsystem controllers 1117 that monitor one or more thermal temperatures of the system 1100 (or its components) and operate the fans / fan disks 1118 to help regulate the temperature.

[0128] In the illustrated system, all major system components can be connected to bus 1116, which can represent more than one physical bus. However, the various system components may be physically close to each other or physically distant. For example, input data and / or output data can be remotely transmitted from one physical location to another. Additionally, programs implementing various aspects of this disclosure can be accessed from a remote location (e.g., a server) via a network. Such data and / or programs can be transmitted via any of a variety of machine-readable media, including, for example: magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as optical discs (CDs) and holographic devices; magneto-optical media; and hardware devices specifically configured to store or store and execute program code, such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), flash memory devices, other non-volatile memory (NVM) devices (such as 3D XPoint-based devices), and ROM and RAM devices.

[0129] Various aspects of this disclosure may be encoded on one or more non-transitory computer-readable media as instructions for causing one or more processors or processing units to perform steps. It should be noted that the one or more non-transitory computer-readable media should include volatile and / or non-volatile memory. It should be noted that alternative implementations are possible, including hardware implementations or software / hardware implementations. The functionality of the hardware implementation can be implemented using one or more ASICs, programmable arrays, digital signal processing circuits, etc. Therefore, the term "module" in any claim is intended to cover both software and hardware implementations. Similarly, the term "one or more computer-readable media" as used herein includes software and / or hardware, or a combination thereof, having a program of instructions embodied thereon. In consideration of these alternative implementations, it should be understood that the accompanying drawings and description provide functional information required by those skilled in the art to write program code (i.e., software) and / or manufacture circuitry (i.e., hardware) to perform the desired processing.

[0130] It should be noted that embodiments of this disclosure may further relate to computer products having a non-transitory tangible computer-readable medium having computer code on it for performing various computer-implemented operations. The medium and computer code may be those specifically designed and constructed for the purposes of this disclosure, or they may be of types known or available to those skilled in the art. Examples of tangible computer-readable media include, for example: magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CDs and holographic devices; magneto-optical media; and hardware devices specifically configured to store or store and execute program code, such as ASICs, programmable logic devices (PLDs), flash memory devices, other non-volatile memory (NVM) devices (such as 3D XPoint-based devices), and ROM and RAM devices. Examples of computer code include machine code generated by a compiler and files containing higher-level code executed by a computer using an interpreter. Embodiments of this disclosure may be implemented, in whole or in part, as machine-executable instructions that may reside in program modules executed by a processing device. Examples of program modules include libraries, programs, routines, objects, components, and data structures. In a distributed computing environment, program modules can physically reside locally, remotely, or in a combination of both.

[0131] Those skilled in the art will recognize that no computing system or programming language is essential to the practice of this disclosure. They will also recognize that the multiple elements described above may be physically and / or functionally divided into modules and / or submodules or combined together.

[0132] Those skilled in the art will understand that the foregoing examples and embodiments are exemplary and do not limit the scope of this disclosure. All substitutions, enhancements, equivalents, combinations, and modifications that will be apparent to those skilled in the art upon reading the specification and studying the accompanying drawings are included within the true spirit and scope of this disclosure. It should also be noted that the elements of any claim may be arranged differently, including having multiple dependencies, configurations, and combinations.

Claims

1. A computer-implemented method, comprising: Given a first batch of training data, the first batch of training data is used as input to train a cross-language, cross-modal network, which includes subtitles in the source language and visual features of the corresponding images. The loss of the Masked Language Modeling (MLM) task, the Masked Region Classification (MRC) task, and the Cross-Modal Text Recovery (CMTR) task are calculated based on the first batch of training data. Given a second training data batch, the cross-language cross-modal network is trained using the second training data batch as input. The second training data batch includes a set of text in the source language and a set of corresponding text in the target language. The loss of the Translation Language Modeling (TLM) task and the Cross-Language Text Recovery (CLTR) task is calculated based on the second batch of training data. Given a third batch of training data, the cross-lingual, cross-modal network is trained using the third batch of training data as input, the third batch of training data comprising monolingual text. The loss for the monolingual MLM task is calculated based on the third batch of training data. The cross-lingual cross-modal network is updated using the losses from the MLM task, the MRC task, the CMTR task, the TLM task, the CLTR task, and the monolingual MLM task. In response to the failure to meet the stopping condition, the above steps are repeated using the next first batch of training data, the second batch of training data, and the third batch of training data. as well as In response to reaching the stopping condition, output the pre-trained cross-language cross-modal CLCM network.

2. The computer-implemented method as described in claim 1, wherein: For the CMTR task, the CLCM network includes an attention layer for learning the alignment between text features and visual features from the first training data batch.

3. The computer-implemented method as described in claim 1, wherein: For the CLTR task, the CLCM network includes an attention mechanism for computing attention representations of the input text in the source language and the corresponding text in the target language.

4. The computer-implemented method as described in claim 1, wherein: Fine-tuned data is used as input to the CLCM system, which includes the pre-trained CLCM network, wherein for each text, there exists a positive image associated with the text and the remaining images in the image are not associated with the text, and correspondingly, for each image, there exists a positive text associated with the image and the remaining text in the text is not associated with the image; For each text in a set of texts from the fine-tuning data, determining the loss involves comparing the relevance output of the CLCM system for the text given a corresponding positive image of the text with the relevance output of the CLCM system for the text given an unrelated image; For each image in a set of images from the fine-tuning data, determining the loss involves a comparison between the relevance output of the CLCM system for the image given the positive text of the image and the relevance output of the CLCM system for the image given the irrelevant text. The CLCM system is updated using a final loss based on the combination of said losses; and Output a finely tuned CLCM system.

5. The computer-implemented method as described in claim 4, wherein: Given the text, the irrelevant image produces the worst relevance output; and Given the image, the irrelevant text produces the worst relevance output.

6. The computer-implemented method as described in claim 4, wherein: The text in the fine-tuning data includes one or more non-source languages.

7. The computer-implemented method as described in claim 4, wherein: Receive query text in a non-source language as input to the fine-tuned CLCM system; Given the query text and a set of images, the CLCM system is used to obtain relevance scores of at least some images in the set of images relative to the query text; and The query text is then used to output a set of the top k images based on the relevance score.

8. The computer-implemented method as described in claim 4, wherein: The query image is received as input to the finely tuned CLCM system; Given a query image and a set of texts in one or more non-source languages, the CLCM system is used to obtain relevance scores of at least some texts in the set of texts relative to the query image; and The first k texts are output for the query image based on the relevance score.

9. A computer-implemented method, comprising: Receive query text or query images in a non-source language as input to the cross-language, cross-modal CLCM system; In response to the input being the query image, perform the following steps: Given the query image and a set of texts in one or more non-source languages, the CLCM system is used to obtain relevance scores of at least some of the texts in the set of texts relative to the query image. as well as Based on the relevance score, output a set of the top k texts for the query image; In response to the input being the query text, perform the following steps: Given the query text and a set of images, the CLCM system is used to obtain relevance scores of at least some of the images in the set of images relative to the query text; as well as Based on relevance scores, a set of the top k images is output for the query text; and The CLCM system is trained by performing the following steps: Given a first batch of training data, the first batch of training data is used as input to train a cross-language, cross-modal network, which includes subtitles in the source language and visual features of the corresponding images. The loss of the Masked Language Modeling (MLM) task, the Masked Region Classification (MRC) task, and the Cross-Modal Text Recovery (CMTR) task are calculated based on the first batch of training data. Given a second training data batch, the cross-language cross-modal network is trained using the second training data batch as input. The second training data batch includes a set of text in the source language and a set of corresponding text in the target language. The loss of the Translation Language Modeling (TLM) task and the Cross-Language Text Recovery (CLTR) task is calculated based on the second batch of training data. Given a third batch of training data, the cross-lingual, cross-modal network is trained using the third batch of training data as input, the third batch of training data comprising monolingual text. The loss for the monolingual MLM task is calculated based on the third batch of training data. The cross-lingual cross-modal network is updated using the losses from the MLM task, the MRC task, the CMTR task, the TLM task, the CLTR task, and the monolingual MLM task. In response to the failure to meet the stopping condition, the above steps are repeated using the next batch of first, second, and third training data; and In response to the termination condition being met, the cross-language, cross-modal CLCM network of the CLCM system is output.

10. The computer-implemented method of claim 9, wherein: For the CMTR task, the CLCM system includes an attention layer for learning the alignment between text features and visual features from the first training data batch.

11. The computer-implemented method of claim 9, wherein: For the CLTR task, the CLCM system includes an attention mechanism for calculating the attention representation of the input text in the source language and its corresponding text in the target language.

12. The computer-implemented method as described in claim 9, wherein, The CLCM system is further trained by performing the following steps: Fine-tuning data is used as input to the CLCM system, which includes the CLCM network, wherein for each text, there exists a positive image associated with the text and the remaining images in the image are not associated with the text, and correspondingly, for each image, there exists a positive text associated with the image and the remaining text in the text is not associated with the image; For each text in a set of texts from the fine-tuning data, determining the loss involves comparing the relevance output of the CLCM system for the text given a corresponding positive image of the text with the relevance output of the CLCM system for the text given an unrelated image; For each image in the set of images from the fine-tuning data, determining the loss involves comparing the relevance output of the CLCM system for the image given positive text with the relevance output of the CLCM system for the image given irrelevant text; and The CLCM system is updated using a final loss based on the combination of said losses to obtain the CLCM system.

13. The computer-implemented method of claim 12, wherein: Given the text, the irrelevant image produces the worst relevance output; and Given the image, the irrelevant text produces the worst relevance output.

14. The computer-implemented method of claim 12, wherein: The text in the fine-tuning data includes one or more non-source languages.

15. A system for computer learning, comprising: One or more processors; as well as One or more non-transitory computer-readable media, the one or more non-transitory computer-readable media comprising one or more instruction sets, the one or more instruction sets, when executed by at least one of the one or more processors, causing the following steps to be performed: Receive query text or query images in a non-source language as input to the cross-language, cross-modal CLCM system; In response to the input being the query image, perform the following steps: Given the query image and a set of texts in one or more non-source languages, the CLCM system is used to obtain relevance scores of at least some of the texts in the set of texts relative to the query image. as well as Based on the relevance score, output a set of the top k texts for the query image; In response to the input being the query text, perform the following steps: Given the query text and a set of images, the CLCM system is used to obtain relevance scores of at least some of the images in the set of images relative to the query text; as well as Based on relevance scores, a set of the top k images is output for the query text; and The CLCM system is trained by performing the following steps: Given a first batch of training data, the first batch of training data is used as input to train a cross-language, cross-modal network, which includes subtitles in the source language and visual features of the corresponding images. The loss of the Masked Language Modeling (MLM) task, the Masked Region Classification (MRC) task, and the Cross-Modal Text Recovery (CMTR) task are calculated based on the first batch of training data. Given a second training data batch, the cross-language cross-modal network is trained using the second training data batch as input. The second training data batch includes a set of text in the source language and a set of corresponding text in the target language. The loss of the Translation Language Modeling (TLM) task and the Cross-Language Text Recovery (CLTR) task is calculated based on the second batch of training data. Given a third batch of training data, the cross-lingual, cross-modal network is trained using the third batch of training data as input, the third batch of training data comprising monolingual text. The loss for the monolingual MLM task is calculated based on the third batch of training data. The cross-lingual cross-modal network is updated using the losses from the MLM task, the MRC task, the CMTR task, the TLM task, the CLTR task, and the monolingual MLM task. In response to the failure to meet the stopping condition, the above steps are repeated using the next batch of first, second, and third training data; and In response to the termination condition being met, the cross-language, cross-modal CLCM network of the CLCM system is output.

16. The system of claim 15, wherein: For the CMTR task, the CLCM system includes an attention layer for learning the alignment between text features and visual features from the first training data batch.

17. The system of claim 15, wherein: For the CLTR task, the CLCM system includes an attention mechanism for calculating the attention representation of the input text in the source language and its corresponding text in the target language.

18. The system of claim 15, wherein, The CLCM system is further trained by performing the following steps: Fine-tuning data is used as input to the CLCM system, which includes the CLCM network, wherein for each text, there exists a positive image associated with the text and the remaining images in the image are not associated with the text, and correspondingly, for each image, there exists a positive text associated with the image and the remaining text in the text is not associated with the image; For each text in a set of texts from the fine-tuning data, determining the loss involves comparing the relevance output of the CLCM system for the text given a corresponding positive image of the text with the relevance output of the CLCM system for the text given an unrelated image; For each image in the set of images from the fine-tuning data, determining the loss involves comparing the relevance output of the CLCM system for the image given positive text with the relevance output of the CLCM system for the image given irrelevant text; and The CLCM system is updated using a final loss based on the combination of said losses to obtain the CLCM system.

19. The system of claim 18, wherein: Given the text, the irrelevant image produces the worst relevance output; and Given the image, the irrelevant text produces the worst relevance output.

20. The system of claim 18, wherein: The text in the fine-tuning data includes one or more non-source languages.

21. A system for computer learning, comprising: One or more processors; as well as One or more non-transitory computer-readable media, the one or more non-transitory computer-readable media comprising one or more instruction sets, the one or more instruction sets causing, when executed by at least one of the one or more processors, to perform the method as described in any one of claims 1-8.

22. A non-transitory computer-readable medium comprising one or more instruction sets, which, when executed by one or more processors, cause the method of any one of claims 1-14 to be performed.

23. A computer program product comprising a computer program, wherein, The computer program, when executed by one or more processors, implements the method of any one of claims 1-14.

Citation Information

Patent Citations

  • Closet lighting system

    US11284495B2