A matching degree determination method and device, electronic equipment, storage medium and product

By using a multimedia model to process the text and images under test and obtain overall and element-level matching degrees, the problem of low accuracy in image generation model text-image matching degree evaluation is solved, and higher accuracy and interpretability of matching degree determination are achieved.

CN122116380APending Publication Date: 2026-05-29BEIJING ZITIAO NETWORK TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-28
Publication Date
2026-05-29

Smart Images

  • Figure CN122116380A_ABST
    Figure CN122116380A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a matching degree determination method and device, electronic equipment, storage medium and product. The matching degree determination method comprises inputting a to-be-tested text and an image generated based on the to-be-tested text into a processing model to obtain a first relationship set and a second relationship set, the first relationship set comprising data indicating a relationship between the to-be-tested text as a whole and the image, and the second relationship set comprising data indicating a relationship between elements included in the to-be-tested text and the image; determining an overall matching degree between the to-be-tested text as a whole and the image based on the first relationship set; and determining an element-level matching degree between the to-be-tested text element dimension and the image based on the second relationship set. The matching degree determination method realizes processing of the to-be-tested text and the image by the processing model, can determine the matching degree of the image and the text from the overall dimension and the element dimension of the to-be-tested text, and improves the accuracy of the image and text matching degree determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to computer technology, and more particularly to a matching degree determination method, apparatus, electronic device, storage medium, and product. Background Technology

[0002] With the development of deep learning, text-to-image tasks have become a popular research area. Text-to-image tasks refer to the process of generating a corresponding image based on a given text description using computer techniques, such as image generation models.

[0003] When generating images based on image generation models, image-text matching accuracy is a metric for evaluating the model's performance. However, current methods for assessing the image-text matching accuracy of image generation models have low accuracy. Summary of the Invention

[0004] This disclosure provides a matching degree determination method, apparatus, electronic device, storage medium, and product to improve the accuracy of the processing model in determining the matching degree of images and text.

[0005] Firstly, a method for determining the matching degree includes:

[0006] The text to be tested and the image generated based on the text to be tested are input into the processing model to obtain a first relation set and a second relation set. The first relation set includes data indicating the relationship between the text to be tested as a whole and the image, and the second relation set includes data indicating the relationship between the elements included in the text to be tested and the image.

[0007] Based on the first set of relationships, the overall matching degree between the text to be tested and the image is determined, wherein the overall matching degree indicates the degree of matching between the text to be tested and the image.

[0008] Based on the second set of relations, the element-level matching degree between the element dimension of the text to be tested and the image is determined, and the element-level matching degree indicates the degree of matching between the text to be tested and the image in the element dimension.

[0009] Secondly, embodiments of this disclosure also provide a matching degree determination device, comprising:

[0010] An input module is used to input the text to be tested and an image generated based on the text to be tested into a processing model to obtain a first relation set and a second relation set. The first relation set includes data indicating the relationship between the text to be tested as a whole and the image, and the second relation set includes data indicating the relationship between the elements included in the text to be tested and the image.

[0011] The first determining module is used to determine the overall matching degree between the text to be tested and the image based on the first relationship set, wherein the overall matching degree indicates the matching degree between the text to be tested and the image.

[0012] The second determining module is used to determine the element-level matching degree between the element dimension of the text to be tested and the image based on the second relationship set, wherein the element-level matching degree indicates the matching degree between the text to be tested and the image in the element dimension.

[0013] Thirdly, embodiments of this application provide an electronic device, characterized in that the electronic device comprises:

[0014] One or more processing devices;

[0015] Storage device for storing one or more programs.

[0016] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the matching degree determination method as provided in the embodiments of this disclosure.

[0017] Fourthly, embodiments of this application provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the matching degree determination method as provided in embodiments of this disclosure.

[0018] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements a matching degree determination method according to embodiments of this disclosure.

[0019] In this embodiment of the disclosure, the matching degree determination method enables the processing of the text and image to be tested by a processing model, which can determine the matching degree of the text and image from the overall dimension and element dimension of the text to be tested, thereby improving the accuracy of the matching degree determination. Attached Figure Description

[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0021] Figure 1 This is a flowchart illustrating a matching degree determination method provided in an embodiment of this disclosure;

[0022] Figure 2 This is a flowchart illustrating a matching degree determination method provided in an embodiment of this disclosure;

[0023] Figure 3 This is a flowchart illustrating another matching degree determination method provided in this embodiment of the present disclosure;

[0024] Figure 4 This is a schematic diagram of the structure of a processing model provided in an embodiment of this disclosure;

[0025] Figure 5 This is a schematic diagram of the structure of a matching degree determination device provided in an embodiment of this disclosure;

[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0033] Image generation models can be considered as models that generate images based on text. The accuracy of the images generated by these models can be used to measure their performance. To address the low accuracy and poor usability of image-text matching determination methods in image generation models, this disclosure provides a matching degree determination method. The processing model used can be considered a multimedia model with fine-grained perception capabilities, enabling the determination of image-text matching degree. Because the fine-grained matching is applied down to the element dimension of the text, it not only determines the matching degree but also identifies the reasons for poor matching based on the element dimension, thus improving interpretability. The image-text matching degree refers to the degree of fit between the image generated from a given text description and the text in terms of content, semantics, etc. Therefore, the image-text matching degree can be considered an indicator of the strength of the correlation between the image generated by the image generation model based on the text and the given text document.

[0034] Figure 1 This is a flowchart illustrating a matching degree determination method provided in this embodiment. This embodiment is applicable to situations where the matching degree of an application processing model is determined. The method can be executed by a matching degree determination device, which can be implemented in software and / or hardware. Optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0035] like Figure 1 As shown, the method includes:

[0036] S110. Input the text to be tested and the image generated based on the text to be tested into the processing model to obtain the first relation set and the second relation set.

[0037] The text to be tested can be considered as the text for which the image-text matching degree needs to be determined during the application stage of the processing model. The method for generating the image based on the text to be tested is not limited here; for example, an image of the text to be tested can be generated using an image generation model.

[0038] The processing model can be considered as a model that determines the matching degree between images and text, such as a model that determines the degree of fit between an image based on the text to be tested and the text to be tested.

[0039] The first set of relationships includes data indicating the relationship between the text under test as a whole and the image, which can be used to determine the degree of matching between the text under test and the image from the perspective of the text under test as a whole.

[0040] The second set of relationships includes data indicating the relationship between the elements included in the text to be tested and the image, which can determine the degree of matching between the text to be tested and the image from the element dimension of the elements included in the text to be tested.

[0041] In this embodiment, the processing model can be determined using the matching degree determination method provided in this embodiment.

[0042] In this embodiment, after obtaining the text to be tested and the image generated based on the text to be tested, the text to be tested and the image are input into the processing model to obtain a first set of relations and a second set of relations, so as to determine the matching degree between the text to be tested and the image.

[0043] S120. Based on the first set of relationships, determine the overall matching degree between the text to be tested and the image.

[0044] The overall matching degree indicates the degree of matching between the text to be tested and the image as a whole.

[0045] The method for determining the overall matching degree can be the same as the method for determining the third matching degree during the model training and processing stage.

[0046] This operation can perform mathematical operations on the data within the first relation set to determine the overall matching degree. For example, the mean of the data in the first relation set can be used to determine the overall matching degree, thus characterizing the overall matching degree between the text to be tested and the image.

[0047] The third matching degree is the matching degree between the overall text information and the image, determined during the model training phase. The overall matching degree is the matching degree between the overall text information and the image, determined during the model application phase.

[0048] S130. Based on the second relationship set, determine the element-level matching degree between the dimension of the text element to be tested and the image.

[0049] Element-level matching degree indicates the degree to which the text under test matches the image in the element dimension.

[0050] The second relation set can be directly used as the element-level matching degree between the text element dimension and the image. Each data in the second relation set reflects the matching degree between the corresponding element and the image.

[0051] This embodiment can also retain the matching degree between the valid elements in the second relation set and the image, and determine the retained content as the element-level matching degree.

[0052] The determination of element validity is not limited here; elements that affect the image generation model are considered valid. These include elements that explicitly describe the things to be presented in the image, such as people, animals, objects, and scenes. Other examples include elements describing various characteristics of an object, including appearance attributes (color, shape, size), material attributes (wood, metal), and state attributes (worn, new). Still others include elements that reflect the object's state and actions. Invalid elements, on the other hand, are those that do not affect the image generation model's ability to generate an image. These include elements that serve as syntactic connectors.

[0053] In this embodiment, a processing model processes the text and image under test to determine the overall matching degree and element-level matching degree, thus realizing the determination of the text-image matching degree at both the overall and element levels. The element-level text-image matching degree can reflect which elements in the text do not match the image, making the output of the processing model interpretable and pinpointing the cause of the mismatch to specific elements.

[0054] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.

[0055] In one embodiment, the text to be tested and an image generated based on the text to be tested are input into a processing model to obtain an identifier set, the identifier set including validity identifier information indicating the validity of elements in the text to be tested.

[0056] Validity indicators can be considered as markers indicating whether elements in the text under test are valid. The means of marking are not limited; different numerical values ​​can represent whether elements in the text under test are valid.

[0057] In this embodiment, after processing the text and image to be tested, the processing model also outputs an identifier set. That is, the text and image to be tested are input into the processing model to obtain a first relation set, a second relation set, and an identifier set.

[0058] In one embodiment, the matching degree determination method further includes:

[0059] Multiply the data at the same position in the second relation set and the identifier set to obtain an intermediate data set;

[0060] The mean of the data in the intermediate dataset and the mean of the data in the first relation set are used to determine the target matching degree between the text to be tested and the image.

[0061] In this embodiment, the processing method for the second relation set and the identifier set is the same as the processing method for determining the third dataset and the second dataset in the processing model.

[0062] The intermediate dataset corresponds to the target vector, and the intermediate dataset is determined during the application phase of the processing model. The target vector is determined during the training phase of the processing model.

[0063] In this embodiment, the data at the same position in the second relationship set and the identifier set are multiplied to obtain an intermediate data set, so that the matching degree between valid elements and the image can be retained in the intermediate data set, and the matching degree between invalid elements and the image is set to 0.

[0064] The mean of the data in the intermediate dataset can be considered as the overall image-text matching degree of all elements, determined based on the matching degree of each element dimension.

[0065] The mean of the data in the first relation set can be considered as the image-text matching degree of the overall dimension of the text to be tested. The mean of the two means is determined as the target matching degree. The target matching degree can be considered as the matching degree determined by combining the overall dimension of the text to be tested and the dimension of the elements of the text to be tested. The target matching degree can be used to measure the matching degree between the text to be tested and the image.

[0066] Figure 2 This is a flowchart illustrating a matching degree determination method provided in this embodiment. This embodiment refines the method for determining the processing model. For details not covered in this embodiment, please refer to the above embodiments.

[0067] like Figure 2 As shown, this embodiment includes the following steps:

[0068] S210. Obtain a training sample set, the training sample set including text information, an image generated based on the text information, a first matching degree, a second matching degree, and a first identifier information of the validity of elements in the text information, wherein the first matching degree indicates the matching degree between the text information and the image determined by annotation, and the second matching degree indicates the matching degree between the elements included in the text information and the image determined by annotation.

[0069] Text information can be descriptive text used to generate images. For example, textual descriptions input into an image generation model convey the requirements for the desired image generation.

[0070] This embodiment does not describe how to generate images based on text information, but rather how to generate images corresponding to text information through a model. Images can be generated based on different image generation models.

[0071] The text information and the generated images can be used to train a processing model so that the trained model can determine the accuracy of the generated images.

[0072] In this embodiment, the text information included in the training sample set may include text information selected from the database. The selection method is not limited here; for example, it may be a hybrid integer linear programming approach to ensure the balance of the selected text information. Alternatively, it may be text information constructed from elements to achieve training on specific content words. This could be used to accurately detect the generation effects of nouns, verbs, and adjectives.

[0073] The first and second matching degrees can be considered as data evaluating the matching degree between text information and images from different dimensions of text information. The first matching degree, second matching degree, and first identifier information can be determined through annotation. The first matching degree determines the matching degree between text information and images from the overall dimension of text information. For example, the first matching degree can be determined by directly comparing the text information as a whole with the image.

[0074] The second matching degree can be determined from the element dimension of the text information to determine the matching degree between the text information and the image.

[0075] For example, text information is segmented using a model, with segmentation based on different types, such as objects, people / animals, attributes, actions, locations, colors, shapes, materials, food, quantities, spatial relationships, and / or others. Then, the segmented elements are labeled for matching or not. A first value indicates a mismatch, and a second value indicates a match. These first and second values ​​can have different values. The labeled content can be considered the second matching degree, which can be stored as a vector. The second matching degree can include the matching degree between elements in the text information and the image, i.e., fine-grained element dimensions, to determine the matching degree.

[0076] The first identifier can be considered as an identifier representing the validity or invalidity of elements in the text information. The means of identification are not limited; different numerical values ​​can be used to represent whether an element in the text information is valid.

[0077] The value of the first identifier information can be different based on whether it is valid or not. For example, a valid value can be a third value, and an invalid value can be a fourth value. The third value can be 1, and the fourth value can be 0.

[0078] This operation can obtain a pre-determined training sample set or directly generate a training sample set, such as filtering text information from a database and constructing text information through elements. Then, the text information is processed by different image generation models to generate corresponding images. Finally, the first matching degree, the second matching degree, and the first identifier information of the validity of elements in the text information are labeled.

[0079] S220. Train the processing model to be trained according to the training sample set.

[0080] After obtaining the training sample set, this operation can train the processing model to be trained based on the training sample set. For example, text information and images generated from that text information are input into the processing model to be trained to obtain the model's output. Then, based on the model's output and the first matching degree, second matching degree, and the first identifier information of the validity of elements in the text information, a target loss is determined to adjust the model parameters of the model to be trained. The target loss can be a loss function determined based on the model output and labels, used to adjust the model parameters of the processing model to be trained. The method for determining the target loss is not limited here.

[0081] In this embodiment, the processing model to be trained can be iteratively trained using a training sample set, and the termination condition for training is not limited here.

[0082] The model structure of the processing model to be trained is not limited here, as long as it can be determined from at least two dimensions: the overall text information and the element-level matching degree.

[0083] S230. Input the text to be tested and the image generated based on the text to be tested into the processing model to obtain the first relation set and the second relation set.

[0084] S240. Based on the first set of relationships, determine the overall matching degree between the text to be tested and the image.

[0085] S250. Based on the second set of relations, determine the element-level matching degree between the dimension of the text element to be tested and the image.

[0086] The technical solution of this disclosure improves the accuracy of image-text matching by training a processing model with a first matching degree determined based on the overall dimension of text information, a second matching degree based on the element dimension of text information, and a first identifier representing the validity of elements in the text information. This enables the trained processing to judge the matching degree of images and text from multiple dimensions, including the overall dimension and the element dimension of text information.

[0087] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.

[0088] In one embodiment, before training the processing model to be trained based on the training sample set, the method further includes:

[0089] Obtain a data augmentation dataset, which includes text to be augmented and images generated based on the text to be augmented;

[0090] Replace the elements in the text to be enhanced that match the image with mutually exclusive elements in the corresponding set of mutually exclusive elements to obtain the enhanced text;

[0091] Determine the fifth degree of matching between the text to be enhanced and the image generated from the text to be enhanced;

[0092] Determine the sixth degree of matching between the enhanced text and the image generated from the text to be enhanced;

[0093] The multimodal visual text language model in the processing model is trained based on the fifth and sixth matching degrees.

[0094] Data augmentation datasets can be considered as datasets to be augmented. Text to be augmented can be considered as the text to be augmented in the context of determining image-text matching. The text to be augmented can include text information from the training sample set, as well as text information other than that from the training sample set. Images generated from the text to be augmented can be considered as images generated from the text using text-to-image generation techniques, such as images generated from the text to be augmented using image generation models.

[0095] After obtaining the data augmentation dataset, this embodiment can replace the text to be augmented. The replacement content includes some or all elements in the text to be augmented that match the image generated from the text to be augmented.

[0096] A set of mutually exclusive elements can be considered as a collection of multiple mutually exclusive elements. Elements within a set of mutually exclusive elements have mutually exclusive meanings, such as [dog, person, table, sky]. Mutually exclusive elements can also be considered as elements with mutually exclusive meanings, such as those representing completely different types of things.

[0097] Enhanced text can be considered as text generated by replacing elements in the text to be enhanced. Replacing elements in the text to be enhanced that match the image with mutually exclusive elements from a set of mutually exclusive elements can be done by randomly selecting an element from any text to be enhanced; if that element matches the image generated from the text to be enhanced, then selecting another element from the set of mutually exclusive elements containing that element and replacing it. For example, replacing an image of a dog with an image of a table.

[0098] The fifth matching degree can be considered as the matching degree between the text to be enhanced and the image. The sixth matching degree can be considered as the matching degree between the enhanced text and the image. The fifth and sixth matching degrees can be matching degrees obtained through manual annotation.

[0099] After determining the fifth and sixth matching degrees, the fifth and sixth matching degrees can be compared and learned to train the multimodal visual text language model in the processing model.

[0100] A multimodal visual-text language model can be considered a multimodal model that processes the relationship between text and images, fusing visual and textual elements to perform multimodal tasks. A modality can be a different way information is present or presented, such as a visual modality and a textual modality.

[0101] Multimodal visual text language models can be pre-trained based on data augmentation datasets. This embodiment trains the multimodal visual text language model based on the fifth and sixth matching degrees, ensuring that the trained model prioritizes a fifth matching degree greater than the sixth. For example, cross-entropy loss is used to minimize the difference between the model's predicted ranking of the fifth and sixth matching degrees and the actual ranking.

[0102] Figure 3 This is a flowchart illustrating another matching degree determination method provided in this disclosure embodiment, which details the operation of training the processing model. For example... Figure 3 As shown, the matching degree determination method provided in this embodiment includes the following operations:

[0103] S310. Obtain the training sample set.

[0104] S320. Input the text information and the image into the processing model to be trained to obtain the first dataset, the second dataset, and the third dataset.

[0105] The first dataset includes data indicating the relationship between the overall text information and the image. The data included in the first dataset can indicate the relationship between the overall text information and the image, and this relationship can be used to determine the matching degree between the text information and the image from the overall dimension of the text information.

[0106] The second dataset includes data indicating the relationship between the elements included in the text information and the image. This relationship can be considered a relationship along the dimension of the text information elements, which can be the degree of matching between the elements and the image.

[0107] The third dataset includes second identification information indicating the validity of elements in the text information. This second identification information can be identification information regarding the validity of elements in the text information determined based on matching degree.

[0108] The first set of relations can correspond to the first dataset, which is the dataset processed during the model training phase. The first set of relations is the dataset processed during the model application phase.

[0109] The second set of relations can correspond to the second dataset. The second set of relations is the dataset output during the model application phase, and the second dataset is the dataset output during the model training phase.

[0110] The identifier set can correspond to the third dataset, which is the dataset output during the application phase of the processing model, while the third dataset is the dataset output during the training phase of the processing model.

[0111] Figure 4 This is a schematic diagram of the structure of a processing model provided in an embodiment of this disclosure. See also... Figure 4 The processing model to be trained takes text information and images processed by an image encoder as input and outputs three datasets: the first dataset, the second dataset, and the third dataset. Specifically, the cross-attention module in the processing model processes queries, such as the query itself, text information, and images (concatenating the query with the text and then performing cross-attention processing with the image), and then passes it through a multi-layer perceptron (MLP) to output the first and second datasets.

[0112] The query component can be a parameter used to train the model. It integrates information from images and text and can be used to generate the first dataset to determine the overall matching degree of the text.

[0113] Each position in the text information portion can be used to generate a second dataset, which is the alignment result along the element dimension. The second dataset can include the matching degree of each element in the text information.

[0114] The third dataset can be a dataset obtained by processing text information through self-attention and then using an MLP. The processing model maps the split elements to the text information, marking the corresponding positions as 1 and the non-corresponding parts as 0, thus forming the third dataset. The third dataset includes a second identifier for the validity of each element, such as a mask.

[0115] S330. Determine the third matching degree based on the first dataset.

[0116] The third matching degree indicates the degree of matching between the text information and the image determined by the processing model. The determination method is not limited here; for example, the third matching degree can be obtained by performing mathematical operations on the data in the first dataset.

[0117] In one embodiment, determining the third matching degree based on the first dataset includes:

[0118] The mean of the data in the first dataset is determined as the third matching degree.

[0119] In this embodiment, the mean of the data in the first dataset is calculated, and the determined mean is used as the third matching degree. For example... Figure 3 The data in the first dataset are averaged to determine the first loss. The first loss can be considered as the loss determined based on the third matching degree.

[0120] S340. Determine the fourth matching degree based on the second dataset and the third dataset.

[0121] The fourth matching degree indicates the degree of matching between the text information and the image in the element dimension, as determined by the processing model. The fourth matching degree characterizes the degree of matching between elements in the text information and the image. Since the fourth matching degree is determined in conjunction with the third dataset, it ensures that the determined fourth matching degree represents the effective matching degree between elements and the image.

[0122] This operation combines data from the third dataset and the second dataset, such as by multiplying corresponding positions, so that the matching degree represented by invalid elements in the generated fourth matching degree is 0, while the matching degree represented by valid elements retains its value. The fourth matching degree can include the matching degree between each element in the text information and the image.

[0123] In one embodiment, the first dataset and the second dataset are each dataset represented in vector form, and determining the fourth matching degree based on the second dataset and the third dataset includes:

[0124] Multiply the data at the same position in the third dataset and the second dataset to obtain the target vector;

[0125] The target vector is determined as the fourth matching degree.

[0126] The same position can correspond to the same element. In the second dataset, each data point represents the matching degree between the corresponding element and the image. In the third dataset, each data point represents whether the corresponding element is valid data. In the target vector obtained by multiplying the data at the same position in the third dataset and the second dataset, the value corresponding to invalid elements can be 0, and the value corresponding to valid elements can be any value found in the second dataset.

[0127] S350. The processing model is jointly trained using the first matching degree, the second matching degree, the third matching degree, the fourth matching degree, the first identification information, and the third dataset.

[0128] After determining the third matching degree, the fourth matching degree, and the third dataset, the processing model can be jointly trained by combining the first matching degree, the second matching degree, and the first identifier information from the training sample set. This aims to make the first and third matching degrees as close as possible, the second and fourth matching degrees as close as possible, and the first identifier information and the third dataset as close as possible.

[0129] S360. Input the text to be tested and the image generated based on the text to be tested into the processing model to obtain the first relation set and the second relation set.

[0130] S370. Based on the first set of relationships, determine the overall matching degree between the text to be tested and the image.

[0131] S380. Based on the second set of relations, determine the element-level matching degree between the dimension of the text element to be tested and the image.

[0132] In this embodiment, the processing model is jointly trained using the third matching degree, the fourth matching degree, the third dataset, and the first matching degree, the second matching degree, and the first identifier information in the training set. This enables the processing model to determine the image-text matching degree from both the overall dimension and the element dimension of the text information, thereby improving the accuracy of the determination.

[0133] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.

[0134] In one embodiment, the step of jointly training the processing model using the first matching degree, the second matching degree, the third matching degree, the fourth matching degree, the first identifier information, and the third dataset includes:

[0135] Determine the first loss between the first matching degree and the third matching degree;

[0136] Determine the second loss for the second matching degree and the fourth matching degree;

[0137] Determine the third loss of the third dataset and the first identification information;

[0138] The processing model is trained by jointly using the first loss, the second loss, and the third loss.

[0139] This embodiment does not limit the means of determining the first loss, the second loss, and the third loss; they can be determined using different loss functions.

[0140] The first loss can characterize the overall text-image matching degree of the text information output by the processing model, that is, the loss between the third matching degree and the first matching degree of manual annotation.

[0141] The second loss can characterize the image-text matching degree of the text information element dimension of the processing model output, that is, the loss of the fourth matching degree and the second matching degree of manual annotation.

[0142] The third loss can characterize the loss of the identification information of effective elements in the text information output by the processing model and the loss of the identification information of effective elements in the manually labeled text information.

[0143] The first loss can be used to optimize the model's performance on overall relationships. The second loss can be used to optimize the model's performance on element-level relationships. The third loss can be used to optimize the model's performance on element splitting.

[0144] In this embodiment, when training the processing model based on the first loss, the second loss, and the third loss, the three losses can be fused to obtain the target loss. The model parameters of the processing model can then be adjusted based on the target loss, so that the adjusted processing model ensures that the values ​​of the three losses are as small as possible.

[0145] like Figure 4 The third dataset and the second dataset are multiplied to obtain a fourth matching degree. A second loss is determined based on the fourth matching degree and the second matching degree. A third loss is determined based on the third dataset and the first identification information.

[0146] In one embodiment, jointly training the processing model using the first loss, the second loss, and the third loss includes:

[0147] The first loss, the second loss, and the third loss are each weighted and summed with a set weight coefficient to obtain the target loss;

[0148] The processing model is trained based on the target loss.

[0149] In this embodiment, when jointly training the processing model, the first loss, second loss, and third loss can be weighted and summed to obtain the target loss, which is then used to train the processing model. The weight coefficients of the first, second, and third losses are not limited here; different weight coefficients can be set for the three losses.

[0150] In one embodiment, the weighting coefficient includes target data of a set multiple, the target data being determined based on the variance of the matching degree between the text information and the generated different images.

[0151] In this embodiment, the weight coefficients of the first loss, the second loss, and the third loss are all set multiples of the target data. The set multiples corresponding to different losses can be different, and the set multiples can be preset values, which are not limited here.

[0152] The target data can be a method that matches a text message with the matching degree of multiple different images corresponding to that text message.

[0153] In image-text matching tasks, overly simple or complex text information can lead to variations in human annotation, resulting in similar matching scores for generated images and making it difficult to effectively differentiate the performance of different generation models. This embodiment quantifies the ability of text information to distinguish between different generated images using target data. The target data is determined based on the variance of the matching scores of the same text information and different images generated based on that text information.

[0154] If the variance of e is used as the target coefficient, the loss weights are adjusted based on the target data during training, improving the image-text matching effect and enhancing generalization. This enables the processing model to determine the alignment level between images and text.

[0155] In one embodiment, inputting the text information and the image into a processing model to be trained to obtain a first dataset, a second dataset, and a third dataset includes:

[0156] The text information and the image are input into the processing model to be trained;

[0157] The first dataset and the second dataset are determined through the image-text matching task in the processing model.

[0158] The third dataset is determined through the image-text comparison task of the processing model.

[0159] Image-text matching can be considered as determining whether text information and images match. Image-text comparison can be considered as teaching the model the alignment relationship between text and images, enabling the model to bring matching image and text representations closer together in the feature space, while pushing mismatched image and text representations further apart.

[0160] The first dataset output by the image-text matching task can measure the overall degree of matching between the image and the text information.

[0161] The second dataset output by the image-text matching task can measure the degree of matching between image and text information elements.

[0162] Image-text matching tasks determine whether an image matches text information. The model needs to take an input image and text, and output a result indicating whether they match.

[0163] This embodiment can directly fit a first matching degree based on the mean determined by the first dataset, and use a first loss for optimization. Fitting the first matching degree can be considered as making the mean as close as possible to or equal to the first matching degree. The goal of optimization can be to minimize the difference between the mean and the first matching degree.

[0164] For fine-grained matching, the second dataset is mapped to terms, such as tokens. This establishes a correspondence between the matching degree between elements and images and the corresponding terms.

[0165] When processing text, in order for computers to better understand and analyze the text content, the text is broken down into basic units, which are called tokens.

[0166] This embodiment fits the second dataset and the third dataset output by the image-text matching task. The fitting can be a second matching degree. The processing model breaks down the text information into elements, each element having a corresponding word unit, and the matching degree between the element and the image is also mapped to a word unit.

[0167] The third dataset output by the image-text comparison task can be the features output by the multimodal visual-text language model. These features are then processed by an MLP to predict which lexical units in the text information are relevant to the matching task. The matching task can be considered as determining whether image-text pairs match.

[0168] Image-text comparison tasks can learn the semantic alignment relationships between images and text. By comparing different image-text pairs, matching (positive examples) and non-matching (negative examples) pairs can be identified, enabling the model to accurately understand the semantic connections between images and text.

[0169] The processing model can output a first dataset determined by the multimodal visual-text language model, as well as valid lexical units in the text information and a second dataset. The first dataset is used to determine the overall matching degree, and the second dataset is used to determine the matching degree of the element set.

[0170] The text comparison task is not integrated with the query; instead, it extracts the features of the text information itself and processes them using an MLP.

[0171] Figure 5 This is a schematic diagram of a matching degree determination device provided in an embodiment of this disclosure, as shown below. Figure 5 As shown, the device includes:

[0172] The input module 510 is used to input the text to be tested and the image generated based on the text to be tested into the processing model to obtain a first relationship set and a second relationship set. The first relationship set includes data indicating the relationship between the text to be tested as a whole and the image, and the second relationship set includes data indicating the relationship between the elements included in the text to be tested and the image.

[0173] The first determining module 520 is used to determine the overall matching degree between the text to be tested and the image based on the first relationship set, wherein the overall matching degree indicates the matching degree between the text to be tested and the image.

[0174] The second determining module 530 is used to determine the element-level matching degree between the element dimension of the text to be tested and the image based on the second relationship set, wherein the element-level matching degree indicates the matching degree between the text to be tested and the image in the element dimension.

[0175] The technical solution provided in this disclosure uses a processing model to process the text and image under test, determine a first set of relations and a second set of relations, and thus determine the text-image matching degree of the overall text and its elements. The text-image matching degree at the element level can reflect which elements in the text under test do not match the image, making the output of the processing model interpretable and pinpointing the cause of the mismatch to a specific element.

[0176] In one embodiment, the input module 510 is specifically used for:

[0177] The test text and the image generated based on the test text are input into the processing model to obtain an identifier set, which includes validity identifier information indicating the validity of elements in the test text.

[0178] In one embodiment, the matching degree determination device further includes a product module, configured to:

[0179] Multiply the data at the same position in the second relation set and the identifier set to obtain an intermediate data set;

[0180] The mean of the data in the intermediate dataset and the mean of the data in the first relation set are used to determine the target matching degree between the text to be tested and the image.

[0181] In one embodiment, the matching degree determination method further includes a training module comprising:

[0182] A determining unit is configured to determine the acquisition of a training sample set, wherein the training sample set includes text information, an image generated based on the text information, a first matching degree, a second matching degree, and a first identifier information indicating the validity of elements in the text information. The first matching degree indicates the matching degree between the text information and the image as determined by annotation, and the second matching degree indicates the matching degree between the elements included in the text information and the image as determined by annotation.

[0183] The training unit is used to train the processing model to be trained based on the training sample set.

[0184] In one embodiment, the training unit includes:

[0185] An input subunit is used to input the text information and the image into the processing model to be trained, thereby obtaining a first dataset, a second dataset, and a third dataset. The first dataset includes data indicating the relationship between the text information as a whole and the image. The second dataset includes data indicating the relationship between the elements included in the text information and the image. The third dataset includes second identification information indicating the validity of the elements in the text information.

[0186] A first determining subunit is configured to determine a third matching degree based on the first dataset, wherein the third matching degree indicates the matching degree between the text information and the image as determined by the matching degree;

[0187] The second determining subunit is used to determine a fourth matching degree based on the second dataset and the third dataset, wherein the fourth matching degree indicates the degree of matching between the text information determined by the matching degree and the image in the element dimension;

[0188] The training subunit is used to jointly train the processing model using the first matching degree, the second matching degree, the third matching degree, the fourth matching degree, the first identification information, and the third dataset.

[0189] In one embodiment, the first dataset and the second dataset are respectively datasets represented in vector form, and the second determined sub-unit is specifically used for:

[0190] Multiply the data at the same position in the third dataset and the second dataset to obtain the target vector;

[0191] The target vector is determined as the fourth matching degree.

[0192] In one embodiment, the first determined subunit is specifically used for:

[0193] The mean of the data in the first dataset is determined as the third matching degree.

[0194] In one embodiment, the training subunit includes:

[0195] Determine the first loss between the first matching degree and the third matching degree;

[0196] Determine the second loss for the second matching degree and the fourth matching degree;

[0197] Determine the third loss of the third dataset and the first identification information;

[0198] The processing model is trained by jointly using the first loss, the second loss, and the third loss.

[0199] In one embodiment, training a subunit includes jointly training the processing model using the first loss, the second loss, and the third loss, including:

[0200] The first loss, the second loss, and the third loss are each weighted and summed with a set weight coefficient to obtain the target loss;

[0201] The processing model is trained based on the target loss.

[0202] In one embodiment, the weighting coefficient includes target data of a set multiple, the target data being determined based on the variance of the matching degree between the text information and the generated different images.

[0203] In one embodiment, the input subunit is specifically used for:

[0204] The text information and the image are input into the processing model to be trained;

[0205] The first dataset and the second dataset are determined through the image-text matching task in the processing model.

[0206] The third dataset is determined through the image-text comparison task of the processing model.

[0207] In one embodiment, the matching degree determination device further includes: an enhancement module, configured to: acquire a data augmentation dataset before training a processing model to be trained based on the training sample set, the data augmentation dataset including text to be augmented and an image generated based on the text to be augmented;

[0208] Replace the elements in the text to be enhanced that match the image with mutually exclusive elements in the corresponding set of mutually exclusive elements to obtain the enhanced text;

[0209] Determine the fifth degree of matching between the text to be enhanced and the image generated from the text to be enhanced;

[0210] Determine the sixth degree of matching between the enhanced text and the image generated from the text to be enhanced;

[0211] The multimodal visual text language model in the processing model is trained based on the fifth and sixth matching degrees.

[0212] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. The following refers to... Figure 6 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 6 A structural diagram of the terminal device or server in the 500.

[0213] Electronic equipment 500, including:

[0214] One or more processing devices 501;

[0215] Storage device 508, for storing one or more programs,

[0216] When the one or more programs are executed by the one or more processing devices 501, the one or more processing devices 501 implement any of the matching degree determination methods provided in this disclosure.

[0217] The terminal devices in this disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (e.g., vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0218] like Figure 6 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.

[0219] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0220] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0221] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0222] The electronic device provided in this embodiment and the matching degree determination method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0223] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the methods provided in the above embodiments.

[0224] It should be noted that the computer-readable medium described above in this disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination thereof.

[0225] The computer storage medium may be a storage medium for computer-executable instructions, which, when executed by a computer processor, are used to perform the methods provided in this disclosure.

[0226] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0227] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0228] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0229] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0230] The text to be tested and the image generated based on the text to be tested are input into the processing model to obtain a first relation set and a second relation set. The first relation set includes data indicating the relationship between the text to be tested as a whole and the image, and the second relation set includes data indicating the relationship between the elements included in the text to be tested and the image.

[0231] Based on the first set of relationships, the overall matching degree between the text to be tested and the image is determined, wherein the overall matching degree indicates the degree of matching between the text to be tested and the image.

[0232] Based on the second set of relations, the element-level matching degree between the element dimension of the text to be tested and the image is determined, and the element-level matching degree indicates the degree of matching between the text to be tested and the image in the element dimension.

[0233] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0234] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0235] The modules or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of modules or units do not necessarily limit the specific unit; for example, the first determining module can also be described as an "overall matching degree determining module".

[0236] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0237] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0238] A computer program product includes a computer program that, when executed by a processor, implements any of the methods provided according to embodiments of this disclosure. It possesses the corresponding functional modules and beneficial effects for implementing the methods.

[0239] According to one or more embodiments of this disclosure, [Example 1] provides a matching degree determination method, including:

[0240] The text to be tested and the image generated based on the text to be tested are input into the processing model to obtain a first relation set and a second relation set. The first relation set includes data indicating the relationship between the text to be tested as a whole and the image, and the second relation set includes data indicating the relationship between the elements included in the text to be tested and the image.

[0241] Based on the first set of relationships, the overall matching degree between the text to be tested and the image is determined, wherein the overall matching degree indicates the degree of matching between the text to be tested and the image.

[0242] Based on the second set of relations, the element-level matching degree between the element dimension of the text to be tested and the image is determined, and the element-level matching degree indicates the degree of matching between the text to be tested and the image in the element dimension.

[0243] According to one or more embodiments of this disclosure, Example 2 provides the method described in Example 1, further comprising:

[0244] The test text and the image generated based on the test text are input into the processing model to obtain an identifier set, which includes validity identifier information indicating the validity of elements in the test text.

[0245] According to one or more embodiments of this disclosure, Example 3 provides the method described in Example 1, further comprising:

[0246] Multiply the data at the same position in the second relation set and the identifier set to obtain an intermediate data set;

[0247] The mean of the data in the intermediate dataset and the mean of the data in the first relation set are used to determine the target matching degree between the text to be tested and the image.

[0248] According to one or more embodiments of this disclosure, [Example 4] provides the method of Example 1, wherein the processing model is trained in the following manner:

[0249] Obtain a training sample set, which includes text information, an image generated based on the text information, a first matching degree, a second matching degree, and a first identifier information indicating the validity of elements in the text information. The first matching degree indicates the matching degree between the text information and the image as determined by annotation, and the second matching degree indicates the matching degree between the elements included in the text information and the image as determined by annotation.

[0250] The processing model to be trained is trained based on the training sample set.

[0251] According to one or more embodiments of this disclosure, [Example 5] provides the method described in Example 4, wherein training the processing model to be trained based on the training sample set includes:

[0252] The text information and the image are input into the processing model to be trained to obtain a first dataset, a second dataset and a third dataset. The first dataset includes data indicating the relationship between the text information as a whole and the image. The second dataset includes data indicating the relationship between the elements included in the text information and the image. The third dataset includes second identification information indicating the validity of the elements in the text information.

[0253] A third matching degree is determined based on the first dataset, the third matching degree indicating the matching degree between the text information and the image as determined by the matching degree;

[0254] Based on the second dataset and the third dataset, a fourth matching degree is determined, the fourth matching degree indicating the degree of matching between the text information determined by the matching degree and the image in the element dimension;

[0255] The processing model is jointly trained using the first matching degree, the second matching degree, the third matching degree, the fourth matching degree, the first identification information, and the third dataset.

[0256] According to one or more embodiments of this disclosure, [Example 6] provides the method of Example 5, wherein the first dataset and the second dataset are datasets represented in vector form, and the step of determining a fourth matching degree based on the second dataset and the third dataset includes:

[0257] Multiply the data at the same position in the third dataset and the second dataset to obtain the target vector;

[0258] The target vector is determined as the fourth matching degree.

[0259] According to one or more embodiments of this disclosure, [Example 7] provides the method of Example 5, wherein determining the third matching degree based on the first dataset includes:

[0260] The mean of the data in the first dataset is determined as the third matching degree.

[0261] According to one or more embodiments of this disclosure, [Example 8] provides the method described in Example 5, wherein jointly training the processing model using the first matching degree, the second matching degree, the third matching degree, the fourth matching degree, the first identification information, and the third dataset includes:

[0262] Determine the first loss between the first matching degree and the third matching degree;

[0263] Determine the second loss for the second matching degree and the fourth matching degree;

[0264] Determine the third loss of the third dataset and the first identification information;

[0265] The processing model is trained by jointly using the first loss, the second loss, and the third loss.

[0266] According to one or more embodiments of this disclosure, [Example 9] provides the method of Example 8, wherein jointly training the processing model using the first loss, the second loss, and the third loss includes:

[0267] The first loss, the second loss, and the third loss are each weighted and summed with a set weight coefficient to obtain the target loss;

[0268] The processing model is trained based on the target loss.

[0269] According to one or more embodiments of this disclosure, [Example 10] provides the method of Example 9, wherein the weighting coefficient includes target data of a set multiple, the target data being determined based on the variance of the matching degree between the text information and different generated images.

[0270] According to one or more embodiments of this disclosure, [Example 11] provides the method of Example 5, wherein inputting the text information and the image into a processing model to be trained to obtain a first dataset, a second dataset, and a third dataset includes:

[0271] The text information and the image are input into the processing model to be trained;

[0272] The first dataset and the second dataset are determined through the image-text matching task in the processing model.

[0273] The third dataset is determined through the image-text comparison task of the processing model.

[0274] According to one or more embodiments of this disclosure, [Example 12] provides the method of Example 4, which further includes, before training the processing model to be trained based on the training sample set:

[0275] Obtain a data augmentation dataset, which includes text to be augmented and images generated based on the text to be augmented;

[0276] Replace the elements in the text to be enhanced that match the image with mutually exclusive elements in the corresponding set of mutually exclusive elements to obtain the enhanced text;

[0277] Determine the fifth degree of matching between the text to be enhanced and the image generated from the text to be enhanced;

[0278] Determine the sixth degree of matching between the enhanced text and the image generated from the text to be enhanced;

[0279] The multimodal visual text language model in the processing model is trained based on the fifth and sixth matching degrees.

[0280] According to one or more embodiments of this disclosure, [Example 13] provides a matching degree determination apparatus, comprising:

[0281] An input module is used to input the text to be tested and an image generated based on the text to be tested into a processing model to obtain a first relation set and a second relation set. The first relation set includes data indicating the relationship between the text to be tested as a whole and the image, and the second relation set includes data indicating the relationship between the elements included in the text to be tested and the image.

[0282] The first determining module is used to determine the overall matching degree between the text to be tested and the image based on the first relationship set, wherein the overall matching degree indicates the matching degree between the text to be tested and the image.

[0283] The second determining module is used to determine the element-level matching degree between the element dimension of the text to be tested and the image based on the second relationship set, wherein the element-level matching degree indicates the matching degree between the text to be tested and the image in the element dimension.

[0284] According to one or more embodiments of this disclosure, [Example 14] provides an electronic device, the electronic device comprising:

[0285] One or more processing devices;

[0286] Storage device for storing one or more programs.

[0287] When the one or more programs are executed by the one or more processing devices, the one or more processing devices perform the method as described in any of Examples 1-12.

[0288] According to one or more embodiments of this disclosure, [Example 15] provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform a matching degree determination method as described in any of Examples 1-12.

[0289] According to one or more embodiments of this disclosure, [Example 16] provides a computer program product including a computer program that, when executed by a processor, implements the matching degree determination method according to any one of Examples 1-12.

[0290] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0291] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0292] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for determining matching degree, characterized in that, include: The text to be tested and the image generated based on the text to be tested are input into the processing model to obtain a first relation set and a second relation set. The first relation set includes data indicating the relationship between the text to be tested as a whole and the image, and the second relation set includes data indicating the relationship between the elements included in the text to be tested and the image. Based on the first set of relationships, the overall matching degree between the text to be tested and the image is determined, wherein the overall matching degree indicates the degree of matching between the text to be tested and the image. Based on the second set of relations, the element-level matching degree between the element dimension of the text to be tested and the image is determined, and the element-level matching degree indicates the degree of matching between the text to be tested and the image in the element dimension.

2. The method according to claim 1, characterized in that, Also includes: The test text and the image generated based on the test text are input into the processing model to obtain an identifier set, which includes validity identifier information indicating the validity of elements in the test text.

3. The method according to claim 2, characterized in that, Also includes: Multiply the data at the same position in the second relation set and the identifier set to obtain an intermediate data set; The mean of the data in the intermediate dataset and the mean of the data in the first relation set are used to determine the target matching degree between the text to be tested and the image.

4. The method according to claim 1, characterized in that, The processing model was trained in the following way: Obtain a training sample set, which includes text information, an image generated based on the text information, a first matching degree, a second matching degree, and a first identifier information indicating the validity of elements in the text information. The first matching degree indicates the matching degree between the text information and the image as determined by annotation, and the second matching degree indicates the matching degree between the elements included in the text information and the image as determined by annotation. The processing model to be trained is trained based on the training sample set.

5. The method according to claim 4, characterized in that, The step of training the processing model to be trained based on the training sample set includes: The text information and the image are input into the processing model to be trained to obtain a first dataset, a second dataset and a third dataset. The first dataset includes data indicating the relationship between the text information as a whole and the image. The second dataset includes data indicating the relationship between the elements included in the text information and the image. The third dataset includes second identification information indicating the validity of the elements in the text information. A third matching degree is determined based on the first dataset, wherein the third matching degree indicates the matching degree between the text information and the image as determined by the processing model; Based on the second dataset and the third dataset, a fourth matching degree is determined, which indicates the degree of matching between the text information determined by the processing model and the image in the element dimension; The processing model is jointly trained using the first matching degree, the second matching degree, the third matching degree, the fourth matching degree, the first identification information, and the third dataset.

6. The method according to claim 5, characterized in that, The first dataset and the second dataset are datasets represented in vector form, and the step of determining the fourth matching degree based on the second dataset and the third dataset includes: Multiply the data at the same position in the third dataset and the second dataset to obtain the target vector; The target vector is determined as the fourth matching degree.

7. The method according to claim 5, characterized in that, Determining the third matching degree based on the first dataset includes: The mean of the data in the first dataset is determined as the third matching degree.

8. The method according to claim 5, characterized in that, The step of jointly training the processing model using the first matching degree, the second matching degree, the third matching degree, the fourth matching degree, the first identifier information, and the third dataset includes: Determine the first loss between the first matching degree and the third matching degree; Determine the second loss for the second matching degree and the fourth matching degree; Determine the third loss of the third dataset and the first identification information; The processing model is trained by jointly using the first loss, the second loss, and the third loss.

9. The method according to claim 8, characterized in that, The process of jointly training the processing model using the first loss, the second loss, and the third loss includes: The first loss, the second loss, and the third loss are each weighted and summed with a set weight coefficient to obtain the target loss; The processing model is trained based on the target loss.

10. The method according to claim 9, characterized in that, The weighting coefficients include target data with a set multiple, and the target data is determined based on the variance of the matching degree between the text information and the generated different images.

11. The method according to claim 5, characterized in that, The step of inputting the text information and the image into the processing model to be trained to obtain a first dataset, a second dataset, and a third dataset includes: The text information and the image are input into the processing model to be trained; The first dataset and the second dataset are determined through the image-text matching task in the processing model. The third dataset is determined through the image-text comparison task of the processing model.

12. The method according to claim 4, characterized in that, Before training the processing model to be trained based on the training sample set, the process also includes: Obtain a data augmentation dataset, which includes text to be augmented and images generated based on the text to be augmented; Replace the elements in the text to be enhanced that match the image with mutually exclusive elements in the corresponding set of mutually exclusive elements to obtain the enhanced text; Determine the fifth degree of matching between the text to be enhanced and the image generated from the text to be enhanced; Determine the sixth degree of matching between the enhanced text and the image generated from the text to be enhanced; The multimodal visual text language model in the processing model is trained based on the fifth and sixth matching degrees.

13. A matching degree determination device, characterized in that, include: An input module is used to input the text to be tested and an image generated based on the text to be tested into a processing model to obtain a first relation set and a second relation set. The first relation set includes data indicating the relationship between the text to be tested as a whole and the image, and the second relation set includes data indicating the relationship between the elements included in the text to be tested and the image. The first determining module is used to determine the overall matching degree between the text to be tested and the image based on the first relationship set, wherein the overall matching degree indicates the matching degree between the text to be tested and the image. The second determining module is used to determine the element-level matching degree between the element dimension of the text to be tested and the image based on the second relationship set, wherein the element-level matching degree indicates the matching degree between the text to be tested and the image in the element dimension.

14. An electronic device, characterized in that, The electronic device includes: One or more processing devices; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processing devices, the one or more processing devices perform the method as described in any one of claims 1-12.

15. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the matching degree determination method as described in any one of claims 1-12.

16. A computer program product comprising a computer program that, when executed by a processor, implements the matching degree determination method according to any one of claims 1-12.