A method and device for constructing a picture-text semantic alignment model

By constructing a semantic alignment model of graphic and text, the problem of inefficient matching of massive images and text information is solved, and efficient and diverse graphic and text matching is achieved, which is suitable for processing massive graphic and text information.

CN115455225BActive Publication Date: 2025-05-23GUANGZHOU YOUMI INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211108881.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2025-05-23
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently match massive images and text information, resulting in inefficient matching of graphic and text, which cannot meet the needs of processing massive graphic and text information.

Method used

A semantic alignment model of graphic text is constructed. By semantic alignment analysis is performed on the pre-determined graphic text input model, the semantic alignment results of each graphic text pair are obtained, and the model parameters are corrected according to the results until the convergence conditions are met.

Benefits of technology

It improves the efficiency and diversity of picture and text matching, can effectively process massive picture and text information, and realizes the matching degree prediction of any text, image or both.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115455225B_ABST
    Figure CN115455225B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for constructing a semantic alignment model of a picture and text, including: inputting a plurality of picture and text pairs into the semantic alignment model, so that the semantic alignment model analyzes each picture and text pair, and obtains the semantic alignment result of each picture and text pair, and the semantic alignment result is used to represent the matching degree of the sample image and the sample text in the corresponding picture and text pair; judging whether the semantic alignment model meets the convergence condition according to the semantic alignment results and the actual matching results of all picture and text pairs; if not, correcting the model parameters until a picture and text semantic alignment model that can be used to predict the matching degree between the image corresponding to the text, the text corresponding to the image, and the image and the text that meets the convergence condition is obtained. It can be seen that the implementation of the present invention trains the semantic alignment model through a plurality of picture and text pairs, and obtains a picture and text semantic alignment model that can be used for a variety of picture and text matching scenarios, which can improve the efficiency of picture and text matching and the diversity of picture and text matching methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification, and in particular to a method and device for constructing a picture-text semantic alignment model. Background Art

[0002] With the development of the digital age, there is a huge amount of image and text information on the Internet. People often have to process image and text information at work, for example, matching multiple images with multiple texts. When the number of images and texts is small, people can manually match images and texts. However, when the number of images and texts is large, the efficiency of manually matching images and texts is low and cannot meet people's needs for processing massive image and text information. It can be seen that how to build a semantic alignment model for images and texts to improve the efficiency of image and text matching is particularly important. Summary of the invention

[0003] The technical problem to be solved by the present invention is to provide a method and device for constructing a picture-text semantic alignment model, which can not only improve the efficiency of picture-text matching, but also increase the diversity of picture-text matching methods.

[0004] In order to solve the above technical problems, the first aspect of the present invention discloses a method for constructing a picture-text semantic alignment model, the method comprising:

[0005] Inputting a plurality of predetermined image-text pairs into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair, each image-text pair including a sample image and a sample text, and the semantic alignment result is used to represent the matching degree of the sample image and the sample text in the corresponding image-text pair;

[0006] According to the semantic alignment results of all the image-text pairs and the actual matching results of all the image-text pairs marked in advance, judging whether the semantic alignment model meets the convergence condition;

[0007] When the judgment result is no, the model parameters of the semantic alignment model are corrected, and the operation of inputting a plurality of predetermined image-text pairs into the semantic alignment model to be trained is re-executed so that the semantic alignment model analyzes each of the image-text pairs to obtain the semantic alignment result of each of the image-text pairs, and the operation of judging whether the semantic alignment model meets the convergence condition based on the semantic alignment results of all the image-text pairs and the actual matching results of all the image-text pairs marked in advance is performed, until a image-text semantic alignment model that meets the convergence condition is obtained, and the image-text semantic alignment model is used to predict one or more of the image corresponding to any text, the text corresponding to any image, and the matching degree between any image and any text.

[0008] As an optional implementation, in the first aspect of the present invention, the semantic alignment model includes an image processing structure, a text processing structure and an alignment structure;

[0009] The semantic alignment model analyzes each of the image-text pairs to obtain a semantic alignment result of each of the image-text pairs, including:

[0010] The image processing structure performs a feature extraction operation on the sample image of each of the image-text pairs to obtain the image features of each of the image-text pairs, and the text processing structure performs a feature extraction operation on the sample text of each of the image-text pairs to obtain the text features of each of the image-text pairs;

[0011] The alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment result of each image-text pair.

[0012] As an optional implementation, in the first aspect of the present invention, the semantic alignment model further includes one or more feature conversion structures, each of which includes at least a fully connected layer;

[0013] After the image processing structure performs a feature extraction operation on the sample image of each of the image-text pairs to obtain the image features of each of the image-text pairs, and before the alignment structure performs an analysis on the image-text splicing features obtained by splicing the image features and text features of each of the image-text pairs to obtain the semantic alignment results of each of the image-text pairs, the method further includes:

[0014] The fully connected layer performs feature conversion processing on the image features of each of the image-text pairs to update the image features of the image-text pair, wherein the feature conversion processing is used to match the feature attributes corresponding to the image features of each of the image-text pairs with the feature attributes corresponding to the text features of the image-text pair, wherein the feature attributes include feature dimensions and / or feature spaces;

[0015] The output result of each preceding feature conversion structure is the input content of its succeeding adjacent feature conversion structure.

[0016] As an optional implementation, in the first aspect of the present invention, each of the characteristic conversion structures further includes a nonlinear processing layer;

[0017] After the fully connected layer performs feature conversion processing on the image features of each of the image-text pairs to update the image features of the image-text pairs, the method further includes:

[0018] The nonlinear processing layer performs nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair;

[0019] The nonlinear processing layer performs nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair, including:

[0020] The nonlinear processing layer performs activation function operation processing on the image features of each image-text pair processed by the fully connected layer based on a preset activation function;

[0021] The nonlinear processing layer randomly hides the values ​​of one or more output neurons in the neural network layer corresponding to each of the image-text pairs based on a preset random hiding method and random hiding probability to update the image features of the image-text pair. The neural network corresponding to each of the image-text pairs includes a neural network layer corresponding to the image features obtained after the image-text pair is processed by the activation function.

[0022] As an optional implementation, in the first aspect of the present invention, the alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment result of each image-text pair, including:

[0023] The vector processing structure of the alignment structure performs vector conversion processing on the image-text splicing features obtained by splicing the image features and text features of each image-text pair, so as to obtain a target matrix corresponding to each image-text pair;

[0024] The fully connected layer of the alignment structure processes the target matrix corresponding to each of the image-text pairs to obtain the confidence level of the semantics of the sample image and the semantics of the sample text of the image-text pair matching as the semantic alignment result of the image-text pair.

[0025] As an optional implementation, in the first aspect of the present invention, the semantic alignment result of each of the image-text pairs includes a confidence level of matching between the semantics of the sample image and the semantics of the sample text of the image-text pair;

[0026] The step of judging whether the semantic alignment model satisfies a convergence condition according to the semantic alignment results of all the image-text pairs and the actual matching results of all the pre-annotated image-text pairs includes:

[0027] Calculating the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each of the image-text pairs and the target confidence corresponding to the actual matching result of the pre-annotated image-text pair;

[0028] Determining whether the predicted loss value is less than a preset loss value threshold;

[0029] When the judgment result is yes, it is determined that the semantic alignment model meets the convergence condition, and when the judgment result is no, it is determined that the semantic alignment model does not meet the convergence condition.

[0030] As an optional implementation, in the first aspect of the present invention, all of the image-text pairs include at least one positive image-text pair and / or at least one negative image-text pair, the actual matching result of the positive image-text pair is a first matching result in which the sample image and the sample text match, and the actual matching result of the negative image-text pair is a second matching result in which the sample image and the sample text do not match;

[0031] Before calculating the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each of the image-text pairs and the target confidence corresponding to the actual matching result of the pre-labeled image-text pair, the method further includes:

[0032] According to a preset label smoothing coefficient, the initial confidence corresponding to each actual matching result is updated to obtain a target confidence corresponding to each actual matching result;

[0033] The target confidence corresponding to the first matching result and the target confidence corresponding to the second matching result are respectively:

[0034] P1=1-ε,

[0035] P2=ε / (N-1),

[0036] Among them, P1 is used to represent the target confidence corresponding to the first matching result, P2 is used to represent the target confidence corresponding to the second matching result, ε is used to represent the label smoothing coefficient, and N is used to represent the number of all the negative example image-text pairs.

[0037] As an optional implementation, in the first aspect of the present invention, before determining whether the predicted loss value is less than a preset loss value threshold, the method further includes:

[0038] Determine the similarity between the target image feature and the target text feature of each of the image-text pairs determined based on the semantic alignment model as the similarity corresponding to the image-text pair;

[0039] updating the predicted loss value according to the similarities corresponding to all the image-text pairs and the actual matching results of all the image-text pairs;

[0040] Furthermore, before determining the similarity between the target image feature and the target text feature of each image-text pair, the method further comprises:

[0041] For each of the image-text pairs, according to the input feature dimension corresponding to the input content of the vector processing structure of the semantic alignment model in the process of the semantic alignment model analyzing the image-text pair, the target matrix corresponding to the image-text pair output by the vector processing structure is divided into target image features and target text features, wherein the input content includes the image-text splicing features determined by the semantic alignment model based on the sample images and sample texts of each of the image-text pairs.

[0042] A second aspect of the present invention discloses a device for constructing a picture-text semantic alignment model, the device comprising:

[0043] An input module, used for inputting a plurality of predetermined image-text pairs into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair, each image-text pair including a sample image and a sample text, and the semantic alignment result is used to represent the matching degree of the sample image and the sample text in the corresponding image-text pair;

[0044] A judgment module, used to judge whether the semantic alignment model meets the convergence condition according to the semantic alignment results of all the image-text pairs and the actual matching results of all the image-text pairs marked in advance;

[0045] A correction module is used to correct the model parameters of the semantic alignment model when the judgment module determines that the semantic alignment model does not meet the convergence condition, and trigger the input module to re-execute the operation of inputting a plurality of predetermined image-text pairs into the semantic alignment model to be trained so that the semantic alignment model analyzes each of the image-text pairs to obtain the semantic alignment result of each of the image-text pairs, and trigger the judgment module to execute the operation of judging whether the semantic alignment model meets the convergence condition based on the semantic alignment results of all the image-text pairs and the actual matching results of all the image-text pairs marked in advance, until a image-text semantic alignment model that meets the convergence condition is obtained, and the image-text semantic alignment model is used to predict one or more of the image corresponding to any text, the text corresponding to any image, and the matching degree between any image and any text.

[0046] As an optional implementation, in the second aspect of the present invention, the semantic alignment model includes an image processing structure, a text processing structure and an alignment structure;

[0047] The semantic alignment model analyzes each of the image-text pairs to obtain the semantic alignment result of each of the image-text pairs in a specific manner including:

[0048] The image processing structure performs a feature extraction operation on the sample image of each of the image-text pairs to obtain the image features of each of the image-text pairs, and the text processing structure performs a feature extraction operation on the sample text of each of the image-text pairs to obtain the text features of each of the image-text pairs;

[0049] The alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment result of each image-text pair.

[0050] As an optional implementation, in the second aspect of the present invention, the semantic alignment model further includes one or more feature conversion structures, each of which includes at least a fully connected layer;

[0051] The fully connected layer is used for performing a feature extraction operation on the sample image of each of the image-text pairs in the image processing structure to obtain the image features of each of the image-text pairs, and then the alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each of the image-text pairs, and before obtaining the semantic alignment results of each of the image-text pairs, performing a feature conversion process on the image features of each of the image-text pairs to update the image features of the image-text pairs, wherein the feature conversion process is used to match the feature attributes corresponding to the image features of each of the image-text pairs with the feature attributes corresponding to the text features of the image-text pairs, and the feature attributes include feature dimensions and / or feature spaces;

[0052] The output result of each preceding feature conversion structure is the input content of its succeeding adjacent feature conversion structure.

[0053] As an optional implementation, in the second aspect of the present invention, each of the feature conversion structures further includes a nonlinear processing layer;

[0054] The nonlinear processing layer is used for performing nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair after the fully connected layer performs feature conversion processing on the image features of each image-text pair to update the image features of the image-text pair;

[0055] The nonlinear processing layer performs nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair in a specific manner including:

[0056] The nonlinear processing layer performs activation function operation processing on the image features of each image-text pair processed by the fully connected layer based on a preset activation function;

[0057] The nonlinear processing layer randomly hides the values ​​of one or more output neurons in the neural network layer corresponding to each of the image-text pairs based on a preset random hiding method and random hiding probability to update the image features of the image-text pair. The neural network corresponding to each of the image-text pairs includes a neural network layer corresponding to the image features obtained after the image-text pair is processed by the activation function.

[0058] As an optional implementation, in the second aspect of the present invention, the alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair, and the specific method of obtaining the semantic alignment result of each image-text pair includes:

[0059] The vector processing structure of the alignment structure performs vector conversion processing on the image-text splicing features obtained by splicing the image features and text features of each image-text pair, so as to obtain a target matrix corresponding to each image-text pair;

[0060] The fully connected layer of the alignment structure processes the target matrix corresponding to each of the image-text pairs to obtain the confidence level of the semantics of the sample image and the semantics of the sample text of the image-text pair matching as the semantic alignment result of the image-text pair.

[0061] As an optional implementation, in the second aspect of the present invention, the semantic alignment result of each of the image-text pairs includes a confidence level of matching between the semantics of the sample image and the semantics of the sample text of the image-text pair;

[0062] The specific manner in which the judging module judges whether the semantic alignment model meets the convergence condition according to the semantic alignment results of all the image-text pairs and the actual matching results of all the pre-annotated image-text pairs includes:

[0063] Calculating the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each of the image-text pairs and the target confidence corresponding to the actual matching result of the pre-annotated image-text pair;

[0064] Determining whether the predicted loss value is less than a preset loss value threshold;

[0065] When the judgment result is yes, it is determined that the semantic alignment model meets the convergence condition, and when the judgment result is no, it is determined that the semantic alignment model does not meet the convergence condition.

[0066] As an optional implementation, in the second aspect of the present invention, all of the image-text pairs include at least one positive image-text pair and / or at least one negative image-text pair, the actual matching result of the positive image-text pair is a first matching result in which the sample image and the sample text match, and the actual matching result of the negative image-text pair is a second matching result in which the sample image and the sample text do not match;

[0067] The device also includes:

[0068] A first updating module is used to update the initial confidence corresponding to each actual matching result according to a preset label smoothing coefficient before the judgment module calculates the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each image-text pair and the target confidence corresponding to the actual matching result of the image-text pair marked in advance, so as to obtain the target confidence corresponding to each actual matching result;

[0069] The target confidence corresponding to the first matching result and the target confidence corresponding to the second matching result are respectively:

[0070] P1=1-ε,

[0071] P2=ε / (N-1),

[0072] Among them, P1 is used to represent the target confidence corresponding to the first matching result, P2 is used to represent the target confidence corresponding to the second matching result, ε is used to represent the label smoothing coefficient, and N is used to represent the number of all the negative example image-text pairs.

[0073] As an optional implementation, in the second aspect of the present invention, the device further includes:

[0074] A determination module, used for determining the similarity between the target image feature and the target text feature of each of the image-text pairs determined based on the semantic alignment model before the judgment module determines whether the predicted loss value is less than a preset loss value threshold, as the similarity corresponding to the image-text pair;

[0075] A second updating module, used for updating the predicted loss value according to the similarities corresponding to all the image-text pairs and the actual matching results of all the image-text pairs;

[0076] And, the device also includes:

[0077] A feature segmentation module is used to segment the target matrix corresponding to the image-text pair output by the vector processing structure into target image features and target text features according to the input feature dimensions corresponding to the input content of the vector processing structure of the semantic alignment model in the process of the semantic alignment model analyzing the image-text pair for each of the image-text pairs, wherein the input content includes the image-text splicing features determined by the semantic alignment model based on the sample images and sample texts of each of the image-text pairs.

[0078] The third aspect of the present invention discloses another device for constructing a picture-text semantic alignment model, the device comprising:

[0079] A memory storing executable program code;

[0080] a processor coupled to the memory;

[0081] The processor calls the executable program code stored in the memory to execute the method for constructing a graphic-text semantic alignment model disclosed in the first aspect of the present invention.

[0082] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the method for constructing a graphic-text semantic alignment model disclosed in the first aspect of the present invention.

[0083] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0084] In an embodiment of the present invention, a plurality of predetermined image-text pairs are input into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair, each image-text pair includes a sample image and a sample text, and the semantic alignment result is used to represent the matching degree of the sample image and the sample text in the corresponding image-text pair; according to the semantic alignment results of all image-text pairs and the actual matching results of all pre-annotated image-text pairs, it is judged whether the semantic alignment model meets the convergence condition; when the judgment result is no, the model parameters of the semantic alignment model are corrected until a image-text semantic alignment model that meets the convergence condition is obtained, and the image-text semantic alignment model is used to predict one or more of the matching degrees between an image corresponding to any text, a text corresponding to any image, and any image and any text. It can be seen that the implementation of the present invention can train a semantic alignment model through a plurality of image-text pairs to obtain an image-text semantic alignment model that can be used to predict the matching degree between an image corresponding to any text, a text corresponding to any image, and any image and text, which can not only improve the efficiency of image-text matching, but also improve the diversity of image-text matching methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0086] Figure 1 It is a flow chart of a method for constructing a graphic-text semantic alignment model disclosed in an embodiment of the present invention;

[0087] Figure 2 is a structural schematic diagram of a semantic alignment model disclosed in an embodiment of the present invention;

[0088] Figure 3 It is a flowchart of another method for constructing a graphic-text semantic alignment model disclosed in an embodiment of the present invention;

[0089] Figure 4 It is a structural schematic diagram of a device for constructing a picture-text semantic alignment model disclosed in an embodiment of the present invention;

[0090] Figure 5 It is a structural schematic diagram of another device for constructing a picture-text semantic alignment model disclosed in an embodiment of the present invention;

[0091] Figure 6 It is a structural schematic diagram of another device for constructing a graphic-text semantic alignment model disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0092] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0093] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or end including a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or ends.

[0094] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0095] The present invention discloses a method and device for constructing a semantic alignment model of a picture and text, which can train a semantic alignment model through a plurality of picture and text pairs, and obtain a picture and text semantic alignment model that can be used to predict the image corresponding to any text, the text corresponding to any image, and the matching degree between any image and text, which can not only improve the efficiency of picture and text matching, but also improve the diversity of picture and text matching methods. The following are detailed descriptions.

[0096] Embodiment 1

[0097] See also Figure 1 , Figure 1 : is a flow chart of a method for constructing a text-image semantic alignment model disclosed in an embodiment of the present invention. Figure 1 The method for constructing a graph-text semantic alignment model described above can be applied to the construction process of a graph-text semantic alignment model based on any architecture, and the embodiment of the present invention does not limit this. Figure 1 As shown, the method for constructing the image-text semantic alignment model may include the following operations:

[0098] 101. Input a plurality of predetermined image-text pairs into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair.

[0099] In an embodiment of the present invention, optionally, each image-text pair may include a sample image and a sample text, and the semantic alignment result is used to represent the degree of match between the sample image and the sample text in the corresponding image-text pair. Further optionally, all image-text pairs include at least one positive image-text pair and / or at least one negative image-text pair, and the positive image-text pair is used to represent an image-text pair with matching images and texts, such as the sample text is "raccoon cat" and the sample image is a raccoon cat image. The negative image-text pair is used to represent an image-text pair with mismatched images and texts, such as the sample text is "golden retriever" and the sample image is "Alaskan dog". This can reduce the occurrence of overfitting in the training of the semantic alignment model.

[0100] In the embodiment of the present invention, optionally, the semantic alignment result of each image-text pair may include a confidence level of matching the semantics of the sample image and the semantics of the sample text of the image-text pair.

[0101] In an embodiment of the present invention, optionally, a set of image-text pairs used to train a semantic alignment model may include multiple image-text pairs that are classified based on arbitrary granularity, which is not limited in the embodiment of the present invention. Further optionally, the arbitrary granularity may include a granularity based on a basic category (such as birds, dogs, cats, etc.) and / or a granularity based on multiple subclasses of a basic category (such as cuckoos, woodpeckers, swallows, etc.), which is not limited in the embodiment of the present invention. Preferably, the set of image-text pairs used to train a semantic alignment model includes image-text pairs corresponding to multiple subclasses of a basic category. In this case, the text label corresponding to the subclass is the sample text. This can improve the accuracy of the trained image-text semantic alignment model in image classification.

[0102] In the embodiment of the present invention, optionally, Figure 2 As shown, the semantic alignment model may include an image processing structure, a text processing structure, and an alignment structure. Further, optionally, the alignment structure may include a vector processing structure and a fully connected layer. Optionally, the image processing structure may include an image encoder, the text processing structure may include a text encoder, the vector processing structure is used to perform semantic analysis on image features and text features, and the vector processing structure may include a vector conversion structure based on a self-attention mechanism; preferably, the image encoder may be a CNN encoder, the text encoder may be a BERT encoder, and the vector conversion structure may be a Transformer structure. In this way, the matching degree between the image encoding result and the text encoding result and the image and text can be improved, and the correlation and globality between the internal information of the image features and the text features themselves can be improved by adopting a vector conversion structure based on a self-attention mechanism.

[0103] In the embodiment of the present invention, further optional, such as Figure 2 As shown, the semantic alignment model may further include one or more feature conversion structures, each feature conversion structure includes at least a fully connected layer, and further optionally, each feature conversion structure may also include a nonlinear processing layer.

[0104] As an optional implementation, Figure 2 As shown, the semantic alignment model analyzes each image-text pair to obtain the semantic alignment results of each image-text pair, which may include:

[0105] The image processing structure performs a feature extraction operation on the sample image of each image-text pair to obtain the image feature of each image-text pair, and the text processing structure performs a feature extraction operation on the sample text of each image-text pair to obtain the text feature of each image-text pair;

[0106] The alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment results of each image-text pair.

[0107] It can be seen that implementing this optional implementation method can respectively extract the image features of the sample image and the text features of the sample text in the image-text pair, and analyze the splicing results obtained after splicing the image-text features and the text features to obtain the semantic alignment results, thereby increasing the dimension of the image-text features and improving the neural network complexity of the semantic alignment model, which is beneficial to improving the accuracy and reliability of training the image-text semantic alignment model.

[0108] In this optional embodiment, optionally, as Figure 2 As shown, the alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment results of each image-text pair, which may include:

[0109] The vector processing structure of the alignment structure performs vector conversion processing on the image-text splicing features obtained by splicing the image features and text features of each image-text pair, so as to obtain the target matrix corresponding to each image-text pair;

[0110] The target matrix corresponding to each image-text pair is processed by the fully connected layer of the alignment structure to obtain the confidence that the semantics of the sample image of the image-text pair and the semantics of the sample text match each other, which is used as the semantic alignment result of the image-text pair.

[0111] It can be seen that implementing this optional implementation can also utilize the vector processing structure to understand the semantics of image features and text features, and determine the confidence of the image-text matching of the image-text pair through the fully connected layer, thereby improving the efficiency and accuracy of determining the semantic alignment results of the image-text pair through the semantic alignment model.

[0112] In this optional implementation, further optionally, the fully connected layer of the alignment structure processes the target matrix corresponding to each image-text pair to obtain the confidence that the semantics of the sample image of the image-text pair and the semantics of the sample text match each other, as the semantic alignment result of the image-text pair, which may include:

[0113] The fully connected layer of the alignment model processes the target matrix corresponding to each image-text pair to obtain one or more classification results and the confidence corresponding to each classification result, and determines the highest target confidence among the confidences corresponding to all classification results as the confidence that the semantics of the sample image of the image-text pair and the semantics of the sample text match, as the semantic alignment result of the image-text pair.

[0114] It can be seen that implementing this optional implementation scheme can also perform linear processing on the target matrix output by the vector processing structure in the fully connected layer and obtain the highest target confidence among the confidences corresponding to multiple classification results as the confidence of the semantic match between the image and text, thereby matching the downstream tasks of the model training samples with the processing method of the fully connected layer, thereby improving the accuracy and reliability of the semantic alignment model in obtaining the confidence of the semantic match between the image and text.

[0115] 102. According to the semantic alignment results of all image-text pairs and the actual matching results of all pre-annotated image-text pairs, it is determined whether the semantic alignment model meets the convergence condition.

[0116] As an optional implementation, judging whether the semantic alignment model meets the convergence condition according to the semantic alignment results of all image-text pairs and the actual matching results of all pre-annotated image-text pairs may include:

[0117] Calculate the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each image-text pair and the target confidence corresponding to the actual matching result of the pre-annotated image-text pair;

[0118] Determine whether the predicted loss value is less than the preset loss value threshold;

[0119] When the judgment result is yes, it is determined that the semantic alignment model meets the convergence condition, and when the judgment result is no, it is determined that the semantic alignment model does not meet the convergence condition.

[0120] It can be seen that implementing this optional implementation method can calculate the loss value of the semantic alignment model according to the difference between the confidence of the semantic alignment of the image and text and the pre-set target confidence, so as to judge whether the semantic alignment model meets the convergence conditions, thereby improving the accuracy and reliability of judging whether the semantic alignment model meets the convergence conditions, and thus improving the matching degree between the training results of the semantic alignment model and the training objectives.

[0121] 103. When the judgment result of step 102 is no, the model parameters of the semantic alignment model are corrected, and steps 101 and 102 are re-executed.

[0122] 104. When the judgment result of step 102 is yes, the current process ends and a graphic-text semantic alignment model that meets the convergence condition is obtained.

[0123] In the embodiment of the present invention, optionally, the image-text semantic alignment model can be used to predict one or more of the image corresponding to any text, the text corresponding to any image, and the matching degree between any image and any text.

[0124] It can be seen that the implementation of the embodiment of the present invention can train a semantic alignment model through several image-text pairs to obtain an image-text semantic alignment model that can be used to predict the image corresponding to any text, the text corresponding to any image, and the matching degree between any image and text. This can not only improve the efficiency of image-text matching, but also improve the diversity of image-text matching methods.

[0125] In an optional embodiment, if Figure 2As shown, after the image processing structure performs a feature extraction operation on the sample image of each image-text pair to obtain the image features of each image-text pair, the alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment results of each image-text pair. The method may also include:

[0126] The fully connected layer performs feature conversion processing on the image features of each image-text pair to update the image features of the image-text pair, and the feature conversion processing is used to match the feature attributes corresponding to the image features of each image-text pair with the feature attributes corresponding to the text features of the image-text pair, and the feature attributes include feature dimensions and / or feature spaces;

[0127] The output result of each preceding feature conversion structure is the input content of its succeeding adjacent feature conversion structure.

[0128] In this optional embodiment, preferably, the semantic alignment model may include two feature conversion structures, the fully connected layer in the first feature conversion structure is used to match the feature dimensions corresponding to the image features of each image-text pair with the feature dimensions corresponding to the text features of the image-text pair, and the fully connected layer in the second feature conversion structure is used to match the feature space corresponding to the image features of each image-text pair with the feature space corresponding to the text features of the image-text pair.

[0129] It can be seen that implementing this optional embodiment can match the feature attributes corresponding to the image features of the image-text pair with the feature attributes corresponding to the text features through the fully connected layer, thereby reducing the distribution differences between the image features and the text features, and increasing the possibility of successful splicing of the image features and the text features. By comparing the image features and the text features under the premise of the same feature attributes, it is helpful to further improve the accuracy and reliability of the semantic alignment results of the image-text pairs.

[0130] In this optional embodiment, as an optional implementation, as Figure 2 As shown, after the fully connected layer performs feature conversion processing on the image features of each image-text pair to update the image features of the image-text pair, the method may further include:

[0131] The nonlinear processing layer performs nonlinear processing on the image features of each image-text pair after being processed by the fully connected layer to update the image features of the image-text pair.

[0132] In this optional implementation, optionally, the nonlinear processing layer performs nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair, which may include:

[0133] The nonlinear processing layer performs activation function operation on the image features of each image-text pair after being processed by the fully connected layer based on a preset activation function;

[0134] The nonlinear processing layer randomly hides the values ​​of one or more output neurons in the neural network layer corresponding to each image-text pair based on a preset random hiding method and random hiding probability to update the image features of the image-text pair. The neural network corresponding to each image-text pair includes a neural network layer corresponding to the image features obtained after the image-text pair is processed by the activation function.

[0135] In this optional implementation, further optionally, the nonlinear processing layer randomly hides the values ​​of one or more output neurons in the neural network layer corresponding to each image-text pair based on a preset random hiding method and random hiding probability to update the image features of the image-text pair, which may include:

[0136] The nonlinear processing layer randomly transforms the values ​​of one or more output neurons in the neural network layer corresponding to each image-text pair into 0 based on a preset random hiding method and random hiding probability, so as to update the image features of the image-text pair.

[0137] In this optional implementation, preferably, the activation function may be a GELU activation function, the random hiding method may be a dropout hiding method, and the random hiding probability may be 0.2.

[0138] It can be seen that the implementation of this optional implementation method performs activation function operation processing on the image features after feature conversion processing through the activation function, thereby being able to introduce nonlinear factors in the image features, which is beneficial for the semantic alignment model to have the ability to solve nonlinear classification, and further improve the image-text matching ability of the semantic alignment model. In addition, by randomly hiding the values ​​of neurons in the neural network layer corresponding to the image features after the activation function operation processing, it is possible to reduce the dependency between fixed neuron combinations, reduce the occurrence of overfitting in the semantic alignment model training, and improve the generalization ability of the semantic alignment model.

[0139] In yet another optional embodiment, Figure 2 As shown, the semantic alignment model can also include a splicing structure;

[0140] Furthermore, before the vector processing structure of the alignment structure performs vector conversion processing on the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the target matrix corresponding to each image-text pair, the method may further include:

[0141] The splicing structure splices the image features and text features of each image-text pair according to the feature dimensions of the image features and the feature dimensions of the text features, so as to obtain the image-text splicing features corresponding to each image-text pair.

[0142] For example, the feature dimension of the image feature of a certain image-text pair is [64,768], and the feature dimension of the text feature is [64,768]. Then the feature dimension of the image-text splicing feature obtained after splicing is [64,1536].

[0143] It can be seen that implementing this optional embodiment can splice image features and text features based on feature dimensions, so that each feature dimension of the image features and each feature dimension of the text features in the image-text feature splicing process are spliced ​​one-to-one with each feature dimension of the text features, thereby improving the accuracy and reliability of the image-text feature splicing.

[0144] Embodiment 2

[0145] See also Figure 3 , Figure 3 is a flow chart of another method for constructing a text-image semantic alignment model disclosed in an embodiment of the present invention. Figure 3 The described method for constructing a graph-text semantic alignment model can be applied to a construction process of a graph-text semantic alignment model based on any architecture, and the embodiment of the present invention does not limit this.

[0146] like Figure 3 As shown, the method for constructing the image-text semantic alignment model may include the following operations:

[0147] 201. Input a plurality of predetermined image-text pairs into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair.

[0148] 202. Calculate the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each image-text pair and the target confidence corresponding to the actual matching result of the image-text pair marked in advance.

[0149] In the embodiment of the present invention, optionally, the actual matching result of the positive example image-text pair is a first matching result in which the sample image and the sample text match, and the actual matching result of the negative example image-text pair is a second matching result in which the sample image and the sample text do not match.

[0150] In the embodiment of the present invention, optionally, for a positive example image-text pair, when the confidence of the image-text semantic matching of the image-text pair is greater than or equal to the target confidence corresponding to the first matching result, the difference between the target confidence corresponding to the semantic alignment result of the image-text pair and the actual matching result of the image-text pair is 0, and for a negative example image-text pair, when the confidence of the image-text semantic matching of the image-text pair is less than or equal to the target confidence corresponding to the second matching result, the difference between the target confidence corresponding to the semantic alignment result of the image-text pair and the actual matching result of the image-text pair is 0. For example, if the target confidence corresponding to the first matching result is 0.8, if the confidence of the image-text semantic matching of a positive example image-text pair is 0.9, indicating that the semantic alignment model accurately predicts the semantic alignment result of the positive example image-text pair, then the difference between the target confidence of the semantic alignment result of the positive example image-text pair and the actual matching result of the positive example image-text pair is 0. This can reduce the occurrence of a situation where the confidence of the image-text pair matching determined by the semantic alignment model deviates from the actual confidence due to the use of the label smoothing method.

[0151] As an optional implementation, calculating the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each image-text pair and the target confidence corresponding to the actual matching result of the pre-annotated image-text pair may include:

[0152] The prediction loss value of the semantic alignment model is calculated based on the binary cross entropy loss function and the difference between the semantic alignment result of each image-text pair and the target confidence corresponding to the actual matching result of the pre-annotated image-text pair.

[0153] It can be seen that implementing this optional implementation method can use the binary cross entropy loss function to calculate the prediction loss value of the semantic alignment model, thereby treating the semantic alignment model as a classification model based on two categories to calculate the model loss value, reducing the difficulty of calculating the prediction loss value of the semantic alignment model and improving the accuracy of the loss calculation.

[0154] 203. Determine whether the predicted loss value is less than a preset loss value threshold.

[0155] 204. When the judgment result of step 203 is no, the model parameters of the semantic alignment model are corrected, and steps 201, 202 and 203 are re-executed.

[0156] 205. When the judgment result of step 202 is yes, the current process ends and a graphic-text semantic alignment model that meets the convergence condition is obtained.

[0157] In the embodiment of the present invention, for other descriptions of step 201, step 204, and step 205, please refer to the detailed description of step 101, step 103, and step 104 in embodiment 1, and the embodiment of the present invention will not be repeated here.

[0158] It can be seen that the implementation of the embodiment of the present invention can train the semantic alignment model through several image-text pairs to obtain an image-text semantic alignment model that can be used to predict the image corresponding to any text, the text corresponding to any image, and the matching degree between any image and text. This can not only improve the efficiency of image-text matching, but also improve the diversity of image-text matching methods. In addition, by calculating the loss value of the semantic alignment model according to the difference between the confidence of the semantic alignment of the image-text pairs and the pre-set target confidence, it is determined whether the semantic alignment model meets the convergence conditions. This improves the accuracy and reliability of determining whether the semantic alignment model meets the convergence conditions, thereby improving the matching degree between the training results of the semantic alignment model and the training objectives.

[0159] In an optional embodiment, before calculating the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each image-text pair and the target confidence corresponding to the actual matching result of the pre-annotated image-text pair, the method may further include:

[0160] According to the preset label smoothing coefficient, the initial confidence corresponding to each actual matching result is updated to obtain the target confidence corresponding to each actual matching result;

[0161] In this optional embodiment, optionally, the target confidence corresponding to the first matching result and the target confidence corresponding to the second matching result are respectively:

[0162] P1=1-ε,

[0163] P2=ε / (N-1),

[0164] Among them, P1 is used to represent the target confidence corresponding to the first matching result, P2 is used to represent the target confidence corresponding to the second matching result, ε is used to represent the label smoothing coefficient, and N is used to represent the number of all negative example image-text pairs.

[0165] For example, the initial confidence corresponding to the first matching result is 1, and the initial confidence corresponding to the second matching result is 2. If ε=0.2 is set, the target confidence corresponding to the first matching result is P1=0.8, and the target confidence corresponding to the second matching result is P2=0.2 / (N-1).

[0166] In this optional embodiment, the target confidence corresponding to each actual matching result is the similarity label corresponding to the actual matching result.

[0167] It can be seen that the implementation of this optional embodiment reduces the overfitting of the semantic alignment model training by utilizing the label smoothing system to perform label smoothing on the required target confidence, i.e., similarity labels, so that image-text pairs with incomplete semantic matching and image-text pairs with similarities between different subclasses can be used as training samples during the model training process, thereby improving the robustness of the semantic alignment of the semantic alignment model.

[0168] In another optional embodiment, before determining whether the predicted loss value is less than a preset loss value threshold, the method may further include:

[0169] Determine the similarity between the target image feature and the target text feature of each image-text pair determined based on the semantic alignment model as the similarity corresponding to the image-text pair;

[0170] Update the predicted loss value based on the similarities corresponding to all image-text pairs and the actual matching results of all image-text pairs.

[0171] It can be seen that the implementation of this optional embodiment improves the accuracy and comprehensiveness of calculating the model loss by taking the similarity between the target image features and the target text features of each image-text pair as a factor in calculating the loss value of the semantic alignment model, thereby improving the semantic alignment accuracy of the semantic alignment model.

[0172] In this optional embodiment, as an optional implementation, updating the predicted loss value according to the similarities corresponding to all image-text pairs and the actual matching results of all image-text pairs may include:

[0173] Calculate the cosine loss value of the semantic alignment model based on the cosine loss function, the similarities corresponding to all image-text pairs, and the actual matching results of all image-text pairs;

[0174] According to the cosine loss value, update the predicted loss value.

[0175] It can be seen that implementing this optional implementation method can use the cosine loss function to calculate the cosine loss value of the semantic alignment model of the image-text pair, which can improve the accuracy and reliability of calculating the prediction loss value of the semantic alignment model.

[0176] In this optional embodiment, as an optional implementation, before determining the similarity between the target image feature and the target text feature of each image-text pair, the method may further include:

[0177] For each image-text pair, according to the input feature dimension corresponding to the input content of the vector processing structure of the semantic alignment model in the process of the semantic alignment model analyzing the image-text pair, the target matrix corresponding to the image-text pair output by the vector processing structure is divided into target image features and target text features, wherein the input content includes the image-text splicing features determined by the semantic alignment model based on the sample images and sample texts of each image-text pair.

[0178] For example, for a certain image-text pair, the feature dimension corresponding to the image feature determined by the semantic alignment model based on the sample image of the image-text pair is [64,768], and the feature dimension corresponding to the text feature determined based on the sample text of the image-text pair is [64,768]. The feature dimension corresponding to the image-text splicing feature determined after splicing the two is [64,1536], which is the input feature dimension corresponding to the input content of the input vector processing structure. Therefore, the feature dimensions corresponding to the target image features and target text features obtained by segmenting the target matrix corresponding to the image-text pair input by the vector processing structure are also [64,768].

[0179] It can be seen that the implementation of this optional implementation method can obtain target image features and target text features by splitting the target matrix input by the vector processing structure, so that the trained image-text semantic alignment model can directly apply the similarity of the image features and text features of the image-text pairs to be predicted to perform semantic alignment, thereby reducing unnecessary splicing operations in the actual application of the image-text semantic alignment model and improving the analysis efficiency of the image-text semantic alignment model.

[0180] In another optional embodiment, a plurality of predetermined image-text pairs are input into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair. The method may further include:

[0181] Combining a sample image of any positive example image-text pair among a plurality of positive example image-text pairs prepared in advance with a sample text of any other positive example image-text pairs to obtain a plurality of negative example image-text pairs;

[0182] One or more positive example image-text pairs and one or more negative example image-text pairs are determined as image-text pairs for training the semantic alignment model to be trained.

[0183] It can be seen that the implementation of this optional embodiment can shuffle and reorganize the sample images and sample texts of multiple positive example image-text pairs to obtain negative example image-text pairs, thereby improving the efficiency and quantity of obtaining negative example image-text pairs.

[0184] Embodiment 3

[0185] See also Figure 4 , Figure 4is a schematic diagram of the structure of another apparatus for constructing a semantic alignment model of images and texts disclosed in an embodiment of the present invention. Figure 4 The described apparatus for constructing a graph-text semantic alignment model may be applied to a process for constructing a graph-text semantic alignment model based on any architecture, and the embodiment of the present invention does not limit this.

[0186] like Figure 4 As shown, the device for constructing the image-text semantic alignment model may include:

[0187] An input module 301 is used to input a plurality of predetermined image-text pairs into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair, each image-text pair includes a sample image and a sample text, and the semantic alignment result is used to indicate the matching degree of the sample image and the sample text in the corresponding image-text pair;

[0188] A judgment module 302 is used to judge whether the semantic alignment model meets the convergence condition according to the semantic alignment results of all image-text pairs and the actual matching results of all pre-annotated image-text pairs;

[0189] The correction module 303 is used to correct the model parameters of the semantic alignment model when the judgment module 302 determines that the semantic alignment model does not meet the convergence conditions, and trigger the input module 301 to re-execute the above-mentioned operation of inputting a plurality of predetermined image-text pairs into the semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain the semantic alignment result of each image-text pair, and trigger the judgment module 302 to execute the above-mentioned operation of judging whether the semantic alignment model meets the convergence conditions based on the semantic alignment results of all image-text pairs and the actual matching results of all pre-annotated image-text pairs, until a image-text semantic alignment model that meets the convergence conditions is obtained. The image-text semantic alignment model is used to predict one or more of the matching degrees between an image corresponding to any text, a text corresponding to any image, and any image and any text.

[0190] It can be seen that implementation Figure 4 The described device can train a semantic alignment model through a number of image-text pairs to obtain an image-text semantic alignment model that can be used to predict the image corresponding to any text, the text corresponding to any image, and the matching degree between any image and text. This can not only improve the efficiency of image-text matching, but also increase the diversity of image-text matching methods.

[0191] In an optional embodiment, if Figure 4 As shown, the semantic alignment model includes an image processing structure, a text processing structure, and an alignment structure;

[0192] The semantic alignment model analyzes each image-text pair, and the specific method of obtaining the semantic alignment result of each image-text pair may include:

[0193] The image processing structure performs a feature extraction operation on the sample image of each image-text pair to obtain the image feature of each image-text pair, and the text processing structure performs a feature extraction operation on the sample text of each image-text pair to obtain the text feature of each image-text pair;

[0194] The alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment results of each image-text pair.

[0195] It can be seen that implementation Figure 4 The described device can also extract the image features of the sample image and the text features of the sample text in the image-text pair respectively, and analyze the splicing results obtained after splicing the image-text features and the text features to obtain the semantic alignment results, thereby increasing the dimension of the image-text features and improving the neural network complexity of the semantic alignment model, which is beneficial to improving the accuracy and reliability of training the image-text semantic alignment model.

[0196] In another optional embodiment, Figure 4 As shown, the semantic alignment model also includes one or more feature conversion structures, each of which includes at least a fully connected layer;

[0197] A fully connected layer, which is used to perform feature extraction operations on the sample images of each image-text pair in the image processing structure, and after obtaining the image features of each image-text pair, the alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair, and before obtaining the semantic alignment results of each image-text pair, perform feature conversion processing on the image features of each image-text pair to update the image features of the image-text pair, and the feature conversion processing is used to match the feature attributes corresponding to the image features of each image-text pair with the feature attributes corresponding to the text features of the image-text pair, and the feature attributes include feature dimensions and / or feature spaces;

[0198] The output result of each preceding feature conversion structure is the input content of its succeeding adjacent feature conversion structure.

[0199] It can be seen that implementation Figure 4 The described device can also match the feature attributes corresponding to the image features of the image-text pair with the feature attributes corresponding to the text features through a fully connected layer, thereby reducing the distribution differences between the image features and the text features, and increasing the possibility of successful splicing of the image features and the text features. By comparing the image features and the text features under the premise of the same feature attributes, it is helpful to further improve the accuracy and reliability of the semantic alignment results of the image-text pair.

[0200] In yet another optional embodiment, Figure 4 As shown, each feature conversion structure also includes a nonlinear processing layer;

[0201] A nonlinear processing layer, configured to perform feature conversion processing on the image features of each image-text pair in the fully connected layer to update the image features of the image-text pair, and then perform nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair;

[0202] The nonlinear processing layer performs nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair. The specific method includes:

[0203] The nonlinear processing layer performs activation function operation on the image features of each image-text pair processed by the fully connected layer based on a preset activation function;

[0204] The nonlinear processing layer randomly hides the values ​​of one or more output neurons in the neural network layer corresponding to each image-text pair based on a preset random hiding method and random hiding probability to update the image features of the image-text pair. The neural network corresponding to each image-text pair includes a neural network layer corresponding to the image features obtained after the image-text pair is processed by the activation function.

[0205] It can be seen that implementation Figure 4 The described device can also use an activation function to perform activation function operation processing on the image features after feature conversion processing, so as to introduce nonlinear factors in the image features, which is beneficial for the semantic alignment model to have the ability to solve nonlinear classification and further improve the image-text matching ability of the semantic alignment model. In addition, by randomly hiding the values ​​of neurons in the neural network layer corresponding to the image features after the activation function operation processing, the dependency between fixed neuron combinations can be reduced, the occurrence of overfitting in the semantic alignment model training can be reduced, and the generalization ability of the semantic alignment model can be improved.

[0206] In yet another optional embodiment, Figure 4 As shown, the alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair, and the specific method of obtaining the semantic alignment result of each image-text pair may include:

[0207] The vector processing structure of the alignment structure performs vector conversion processing on the image-text splicing features obtained by splicing the image features and text features of each image-text pair, so as to obtain the target matrix corresponding to each image-text pair;

[0208] The target matrix corresponding to each image-text pair is processed by the fully connected layer of the alignment structure to obtain the confidence that the semantics of the sample image of the image-text pair and the semantics of the sample text match each other, which is used as the semantic alignment result of the image-text pair.

[0209] It can be seen that implementation Figure 4 The described device can also use the vector processing structure to understand the semantics of image features and text features, and determine the confidence of image-text matching of image-text pairs through the fully connected layer, thereby improving the efficiency and accuracy of determining the semantic alignment results of image-text pairs through the semantic alignment model.

[0210] In yet another optional embodiment, Figure 4 As shown, the semantic alignment result of each image-text pair includes the confidence level of the semantics of the sample image and the semantics of the sample text of the image-text pair matching;

[0211] The specific manner in which the judging module 302 judges whether the semantic alignment model meets the convergence condition based on the semantic alignment results of all image-text pairs and the actual matching results of all pre-annotated image-text pairs may include:

[0212] Calculate the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each image-text pair and the target confidence corresponding to the actual matching result of the pre-annotated image-text pair;

[0213] Determine whether the predicted loss value is less than the preset loss value threshold;

[0214] When the judgment result is yes, it is determined that the semantic alignment model meets the convergence condition, and when the judgment result is no, it is determined that the semantic alignment model does not meet the convergence condition.

[0215] It can be seen that the implementation of this optional embodiment can calculate the loss value of the semantic alignment model according to the difference between the confidence of the semantic alignment of the image and text and the pre-set target confidence, so as to judge whether the semantic alignment model meets the convergence conditions, thereby improving the accuracy and reliability of judging whether the semantic alignment model meets the convergence conditions, and further improving the matching degree between the training results of the semantic alignment model and the training objectives.

[0216] In yet another optional embodiment, Figure 5 As shown, all image-text pairs include at least one positive image-text pair and / or at least one negative image-text pair, the actual matching result of the positive image-text pair is a first matching result in which the sample image and the sample text match, and the actual matching result of the negative image-text pair is a second matching result in which the sample image and the sample text do not match;

[0217] The device may also include:

[0218] The first updating module 304 is used to update the initial confidence corresponding to each actual matching result according to a preset label smoothing coefficient before the judgment module 302 calculates the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each image-text pair and the target confidence corresponding to the actual matching result of the image-text pair marked in advance, so as to obtain the target confidence corresponding to each actual matching result;

[0219] Among them, the target confidence corresponding to the first matching result and the target confidence corresponding to the second matching result are respectively:

[0220] P1=1-ε,

[0221] P2=ε / (N-1),

[0222] Among them, P1 is used to represent the target confidence corresponding to the first matching result, P2 is used to represent the target confidence corresponding to the second matching result, ε is used to represent the label smoothing coefficient, and N is used to represent the number of all negative example image-text pairs.

[0223] It can be seen that the implementation Figure 5 The described device can utilize a label smoothing system to perform label smoothing on the required target confidence, i.e., similarity label, thereby reducing the occurrence of overfitting in semantic alignment model training, and enabling image-text pairs with incomplete semantic matching and image-text pairs with similarity between different subclasses to be used as training samples during model training, thereby improving the robustness of the semantic alignment of the semantic alignment model.

[0224] In yet another optional embodiment, Figure 5 As shown, the device may also include:

[0225] The determination module 305 is used to determine the similarity between the target image feature and the target text feature of each image-text pair determined based on the semantic alignment model before the judgment module 302 determines whether the predicted loss value is less than a preset loss value threshold, as the similarity corresponding to the image-text pair;

[0226] A second updating module 306 is used to update the predicted loss value according to the similarities corresponding to all image-text pairs and the actual matching results of all image-text pairs;

[0227] And, the device may also include:

[0228] The feature segmentation module 307 is used to segment the target matrix corresponding to the image-text pair output by the vector processing structure into target image features and target text features according to the input feature dimensions corresponding to the input content of the vector processing structure of the semantic alignment model in the process of the semantic alignment model analyzing the image-text pair, wherein the input content includes the image-text splicing features determined by the semantic alignment model based on the sample image and sample text of each image-text pair.

[0229] It can be seen that implementation Figure 5 The described device can also use the similarity between the target image features and the target text features obtained by segmenting the target matrix input by the vector processing structure as a factor for calculating the loss value of the semantic alignment model, thereby improving the accuracy and comprehensiveness of the calculation model loss, thereby improving the semantic alignment accuracy of the semantic alignment model, and can enable the trained image-text semantic alignment model to directly apply the similarity between the image features and text features of the image-text pair for semantic alignment, reducing unnecessary splicing operations in the actual application process of the image-text semantic alignment model, and improving the analysis efficiency of the image-text semantic alignment model.

[0230] Embodiment 4

[0231] See also Figure 6 , Figure 6 is a schematic diagram of the structure of another apparatus for constructing a text-image semantic alignment model disclosed in an embodiment of the present invention. Figure 6 As shown, the device for constructing the image-text semantic alignment model may include:

[0232] A memory 401 storing executable program codes;

[0233] a processor 402 coupled to the memory 401;

[0234] The processor 402 calls the executable program code stored in the memory 401 to execute the steps in the method for constructing the image-text semantic alignment model described in the first embodiment of the present invention or the second embodiment of the present invention.

[0235] Embodiment 5

[0236] An embodiment of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the steps in the method for constructing a graphic-text semantic alignment model described in Embodiment 1 or Embodiment 2 of the present invention.

[0237] Embodiment 6

[0238] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps in the method for constructing a graphic-text semantic alignment model described in Embodiment 1 or Embodiment 2.

[0239] The device embodiments described above are only illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, i.e., they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.

[0240] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution can be essentially or partly contributed to the prior art in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0241] Finally, it should be noted that the method and device for constructing a graphic-text semantic alignment model disclosed in the embodiment of the present invention disclose only the preferred embodiments of the present invention, which are only used to illustrate the technical scheme of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical schemes described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical schemes from the spirit and scope of the technical schemes of the embodiments of the present invention.

Claims

1. A method for constructing a semantic alignment model of images and texts. It is characterized in that The method comprises: Inputting a plurality of predetermined image-text pairs into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair, each image-text pair includes a sample image and a sample text, and the semantic alignment result is used to indicate the matching degree of the sample image and the sample text in the corresponding image-text pair; the semantic alignment result of each image-text pair includes the confidence degree of matching the semantics of the sample image and the semantics of the sample text of the image-text pair; Calculating the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each of the image-text pairs and the target confidence corresponding to the actual matching result of the pre-annotated image-text pair; Determining whether the predicted loss value is less than a preset loss value threshold; When the judgment result is yes, determining that the semantic alignment model meets the convergence condition; When the judgment result is no, it is determined that the semantic alignment model does not meet the convergence condition, the model parameters of the semantic alignment model are corrected, and the operation of inputting a plurality of predetermined image-text pairs into the semantic alignment model to be trained is re-executed so that the semantic alignment model analyzes each of the image-text pairs to obtain the semantic alignment result of each of the image-text pairs, and the operation of calculating the predicted loss value of the semantic alignment model according to the difference between the semantic alignment result of each of the image-text pairs and the target confidence corresponding to the actual matching result of the image-text pair marked in advance is performed; the operation of judging whether the predicted loss value is less than a preset loss value threshold is performed until a image-text semantic alignment model that meets the convergence condition is obtained, and the image-text semantic alignment model is used to predict one or more of the matching degrees between an image corresponding to any text, a text corresponding to any image, and any image and any text; All of the image-text pairs include at least one positive image-text pair and / or at least one negative image-text pair, the actual matching result of the positive image-text pair is a first matching result in which the sample image and the sample text match, and the actual matching result of the negative image-text pair is a second matching result in which the sample image and the sample text do not match; Before calculating the prediction loss value of the semantic alignment model according to the difference between the semantic alignment result of each of the image-text pairs and the target confidence corresponding to the actual matching result of the pre-labeled image-text pair, the method further includes: According to a preset label smoothing coefficient, the initial confidence corresponding to each actual matching result is updated to obtain a target confidence corresponding to each actual matching result; The target confidence corresponding to the first matching result and the target confidence corresponding to the second matching result are respectively: P1=1-ε, P2=ε / (N-1), Among them, P1 is used to represent the target confidence corresponding to the first matching result, P2 is used to represent the target confidence corresponding to the second matching result, ε is used to represent the label smoothing coefficient, and N is used to represent the number of all the negative example image-text pairs.

2. The method for constructing a graph-text semantic alignment model according to claim 1, It is characterized in that The semantic alignment model includes an image processing structure, a text processing structure and an alignment structure; The semantic alignment model analyzes each of the image-text pairs to obtain a semantic alignment result of each of the image-text pairs, including: The image processing structure performs a feature extraction operation on the sample image of each of the image-text pairs to obtain the image features of each of the image-text pairs, and the text processing structure performs a feature extraction operation on the sample text of each of the image-text pairs to obtain the text features of each of the image-text pairs; The alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment result of each image-text pair.

3. The method for constructing a graph-text semantic alignment model according to claim 2, It is characterized in that The semantic alignment model further includes one or more feature conversion structures, each of which includes at least a fully connected layer; After the image processing structure performs a feature extraction operation on the sample image of each of the image-text pairs to obtain the image features of each of the image-text pairs, and before the alignment structure performs an analysis on the image-text splicing features obtained by splicing the image features and text features of each of the image-text pairs to obtain the semantic alignment results of each of the image-text pairs, the method further includes: The fully connected layer performs feature conversion processing on the image features of each of the image-text pairs to update the image features of the image-text pair, wherein the feature conversion processing is used to match the feature attributes corresponding to the image features of each of the image-text pairs with the feature attributes corresponding to the text features of the image-text pair, wherein the feature attributes include feature dimensions and / or feature spaces; The output result of each preceding feature conversion structure is the input content of its succeeding adjacent feature conversion structure.

4. The method for constructing a picture-text semantic alignment model according to claim 3, It is characterized in that Each of the feature conversion structures also includes a nonlinear processing layer; Furthermore, after the fully connected layer performs feature conversion processing on the image features of each of the image-text pairs to update the image features of the image-text pairs, the method further includes: The nonlinear processing layer performs nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair; The nonlinear processing layer performs nonlinear processing on the image features of each image-text pair processed by the fully connected layer to update the image features of the image-text pair, including: The nonlinear processing layer performs activation function operation processing on the image features of each image-text pair processed by the fully connected layer based on a preset activation function; The nonlinear processing layer randomly hides the values ​​of one or more output neurons in the neural network layer corresponding to each of the image-text pairs based on a preset random hiding method and random hiding probability to update the image features of the image-text pair. The neural network corresponding to each of the image-text pairs includes a neural network layer corresponding to the image features obtained after the image-text pair is processed by the activation function.

5. The method for constructing a picture-text semantic alignment model according to any one of claims 2 to 4, It is characterized in that The alignment structure analyzes the image-text splicing features obtained by splicing the image features and text features of each image-text pair to obtain the semantic alignment result of each image-text pair, including: The vector processing structure of the alignment structure performs vector conversion processing on the image-text splicing features obtained by splicing the image features and text features of each image-text pair, so as to obtain a target matrix corresponding to each image-text pair; The fully connected layer of the alignment structure processes the target matrix corresponding to each of the image-text pairs to obtain the confidence level of the semantics of the sample image and the semantics of the sample text of the image-text pair matching as the semantic alignment result of the image-text pair.

6. The method for constructing a picture-text semantic alignment model according to claim 1, It is characterized in that Before determining whether the predicted loss value is less than a preset loss value threshold, the method further includes: Determine the similarity between the target image feature and the target text feature of each of the image-text pairs determined based on the semantic alignment model as the similarity corresponding to the image-text pair; updating the predicted loss value according to the similarities corresponding to all the image-text pairs and the actual matching results of all the image-text pairs; Furthermore, before determining the similarity between the target image feature and the target text feature of each image-text pair, the method further comprises: For each of the image-text pairs, according to the input feature dimension corresponding to the input content of the vector processing structure of the semantic alignment model in the process of the semantic alignment model analyzing the image-text pair, the target matrix corresponding to the image-text pair output by the vector processing structure is divided into target image features and target text features, wherein the input content includes the image-text splicing features determined by the semantic alignment model based on the sample images and sample texts of each of the image-text pairs.

7. A device for constructing a picture-text semantic alignment model, It is characterized in that The device is used to execute the method for constructing a graphic-text semantic alignment model according to any one of claims 1 to 6, and the device includes: An input module, used for inputting a plurality of predetermined image-text pairs into a semantic alignment model to be trained, so that the semantic alignment model analyzes each image-text pair to obtain a semantic alignment result of each image-text pair, each image-text pair including a sample image and a sample text, and the semantic alignment result is used to represent the matching degree of the sample image and the sample text in the corresponding image-text pair; A judgment module, used to judge whether the semantic alignment model meets the convergence condition according to the semantic alignment results of all the image-text pairs and the actual matching results of all the image-text pairs marked in advance; A correction module is used to correct the model parameters of the semantic alignment model when the judgment module determines that the semantic alignment model does not meet the convergence condition, and trigger the input module to re-execute the operation of inputting a plurality of predetermined image-text pairs into the semantic alignment model to be trained so that the semantic alignment model analyzes each of the image-text pairs to obtain the semantic alignment result of each of the image-text pairs, and trigger the judgment module to execute the operation of judging whether the semantic alignment model meets the convergence condition based on the semantic alignment results of all the image-text pairs and the actual matching results of all the image-text pairs marked in advance, until a image-text semantic alignment model that meets the convergence condition is obtained, and the image-text semantic alignment model is used to predict one or more of the image corresponding to any text, the text corresponding to any image, and the matching degree between any image and any text.

8. A device for constructing a picture-text semantic alignment model, It is characterized in that The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the method for constructing a graphic-text semantic alignment model as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Entity alignment method and device, electronic equipment and storage medium

    CN111563192A

  • Assigning labels to images

    US8873867B1