Data processing method and device, computer readable medium and computer equipment
By splitting and adjusting the template words of the description text, generating new graphic and text samples, and training the graphic and text model, the shortcomings of the graphic and text model in fine-grained recognition and consistency discrimination are solved, and the recognition accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510406516.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The existing graphic and text models have shortcomings in fine-grained recognition and graphic and text consistency discrimination, and cannot accurately identify the details of images and text, resulting in poor accuracy of consistency discrimination.
By splitting the description text into descriptive words corresponding to the template words, adjusting the words and calculating the description score, generating new graphic and text samples, training the graphic and text model, enhancing the fine-grained recognition ability.
The recognition accuracy and robustness of the graphic and text model for descriptive text is improved, ensuring more accurate and stable performance in different types of graphic and text matching tasks.
Smart Images

Figure CN120336567A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer and communication technologies, and in particular, to a data processing method, apparatus, computer-readable medium, and computer device. Background Art
[0002] With the development of cross-modal image-text models, the effect of image-text consistency discrimination has been significantly improved and can be widely applied to scenarios such as image retrieval and text-to-image generation. However, in the actual application process, the image-text models proposed in the related technologies tend to be more focused on the general expression of things and cannot achieve fine-grained recognition of image-text details, thus resulting in poor accuracy of image-text consistency discrimination by the image-text models. Summary of the Invention
[0003] Embodiments of this application provide a data processing method, apparatus, computer-readable medium, and computer device, which can improve the fine-grained recognition ability of the image-text model and enhance the accuracy of image-text consistency discrimination.
[0004] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0005] According to one aspect of the embodiments of this application, a data processing method is provided, including: obtaining an image-text pair sample, where the image-text pair sample includes an image sample and a description text corresponding to the image sample; splitting the description text into description words corresponding to each of the template words according to the set template words to obtain the description words of the description text; adjusting the description words of the description text to obtain an adjusted text that is different from the description text, and calculating a description score of the adjusted text according to the change situation of the adjusted text compared with the description text, where the description score is used to represent the correct rate of the adjusted text in describing the image sample; generating a new image-text pair sample according to the adjusted text and the image sample, and training an image-text model according to the new image-text pair sample and the description score of the adjusted text.
[0006] According to one aspect of the embodiments of the present application, a data processing device is provided, including: an acquisition unit configured to acquire a text-image pair sample, where the text-image pair sample includes an image sample and a description text corresponding to the image sample; a splitting unit configured to split the description text into description words corresponding to each of the template words according to the set template words, to obtain the description words of the description text; a processing unit configured to adjust the description words of the description text to obtain an adjusted text that is different from the description text, and calculate a description score of the adjusted text according to the change situation of the adjusted text compared with the description text, where the description score is used to represent the correct rate of the adjusted text in describing the image sample; a training unit configured to generate a new text-image pair sample according to the adjusted text and the image sample, and train a text-image model according to the new text-image pair sample and the description score of the adjusted text.
[0007] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: replace at least one description word of the description text with other description words of the same type to obtain the adjusted text.
[0008] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: calculate the weight proportion of the description words that are not replaced in the description words of the description text according to the description words replaced in the description text and the weights of the respective description words of the description text; use the weight proportion as the description score of the adjusted text.
[0009] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: delete at least one description word of the description text to obtain the adjusted text.
[0010] In some embodiments of the present application, based on the foregoing solution, the set template words include a subject word; where deleting at least one description word of the description text includes: deleting at least one description word of the description text other than the description word corresponding to the subject word.
[0011] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: calculate the ratio between the weight of the description word deleted in the description text and a set scale factor; use the difference between a preset maximum description score and the ratio as the description score of the adjusted text.
[0012] In some embodiments of the present application, based on the foregoing solution, the processing unit is further configured to: calculate the maximum description score of the first adjusted text obtained by replacing the description words in the description text according to the weights of the respective description words in the description text, and the description score of the second adjusted text obtained by deleting all the description words in the description text except the specified description words; calculate the value of the scaling factor under the constraint that the description score of the second adjusted text is greater than the maximum description score of the first adjusted text.
[0013] In some embodiments of the present application, based on the foregoing solution, the obtaining unit is configured to: obtain a candidate text-image pair, where the candidate text-image pair includes a candidate image and a candidate description text; split the candidate description text into description words corresponding to the respective template words according to the template words to obtain the description words of the candidate description text; if the content described by all the description words of the candidate description text appears in the candidate image, use the candidate text-image pair as the text-image pair sample.
[0014] In some embodiments of the present application, based on the foregoing solution, the obtaining unit is further configured to: extract the image feature vector of the candidate image and the text feature vectors of the respective description words of the candidate description text; calculate the similarity between the image feature vector and the text feature vectors of the respective description words to determine whether the content described by the respective description words appears in the candidate image according to the similarity.
[0015] In some embodiments of the present application, based on the foregoing solution, the text-image model includes a text encoding module and an image encoding module, and the training unit is configured to: generate the input data of the text encoding module according to the template words and the word representations corresponding to the description words of the adjusted text; obtain the text feature vector output by the text encoding module according to the input data, and obtain the image feature vector output by the image encoding module for the image sample; calculate the similarity score between the text feature vector and the image feature vector, calculate the loss data according to the similarity score and the description score of the adjusted text; adjust the model parameters of the text-image model according to the loss data to obtain a trained text-image model, and the trained text-image model is used to evaluate the matching degree between text and image.
[0016] In some embodiments of the present application, based on the foregoing solution, the image coding network includes a feature extraction network for extracting visual features of an image, a first attention pooling module and a second attention pooling module respectively connected to the feature extraction network, and a feature fusion module for fusing the output data of the first attention pooling module and the output data of the second attention pooling module; wherein, the training unit is configured to: adjust the parameters of the second attention pooling module, the parameters of the feature fusion module, and the word representation parameters corresponding to the description words according to the loss data.
[0017] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: divide the training data into multiple batches, each batch containing at least one text-image pair sample; calculate the loss data of each batch according to the loss data corresponding to each text-image pair sample, and adjust the model parameters of the text-image model according to the loss data of each batch; if the model parameters of the text-image model are adjusted according to the loss data of the multiple batches, then complete one round of training process of the text-image model; if the training of the text-image model reaches the training end condition, then determine that the training of the text-image model is completed.
[0018] In some embodiments of the present application, based on the foregoing solution, the processing unit is further configured to: obtain the model service parameters of each service scenario obtained by training the text-image model using text-image pair samples in different service scenarios; associate and store the scenario identification information of each service scenario with the model service parameters, so as to select the associated model service parameters according to the scenario identification information carried in the service request, and respond to the service request through the text-image model.
[0019] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data processing method as described in the above embodiments.
[0020] According to one aspect of the embodiments of the present application, there is provided a computer device, including: one or more processors; a storage device for storing one or more computer programs, and when the one or more computer programs are executed by the one or more processors, the computer device implements the data processing method as described in the above embodiments.
[0021] According to one aspect of the embodiments of the present application, there is provided a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads and executes the computer program from the computer-readable storage medium, so that the computer device executes the data processing methods provided in the above various alternative embodiments.
[0022] In the technical solutions provided by some embodiments of the present application, after obtaining a text-image pair sample including an image sample and a description text corresponding to the image sample, the description text can be split into description words corresponding to each template word according to the set template words to obtain the description words of the description text. Then, the description words of the description text are adjusted to obtain an adjusted text that is different from the description text, and according to the change situation of the adjusted text compared with the description text, a description score of the adjusted text is calculated. This description score is used to represent the correct rate of the adjusted text in describing the image sample. Furthermore, a new text-image pair sample is generated based on the adjusted text and the image sample, and the text-image model is trained according to the new text-image pair sample and the description score of the adjusted text.
[0023] It can be seen that the technical solution of the embodiment of the present application splits the description text into multiple specific description words (such as categories, clothing, actions, etc.) by using the set template words, which enables the text-image model to more carefully understand and process the key information in the text. This fine-grained text parsing ability enhances the recognition accuracy of the text-image model for the description text. By adjusting the description text (such as modifying or deleting some description words), a large number of new texts with different degrees of differences can be generated. These new texts not only enrich the training data set but also simulate various error situations that may occur in actual application scenarios, thereby improving the robustness of the text-image model. By calculating the corresponding description score according to the change situation of the adjusted text relative to the original description text to quantify the correct rate of the adjusted text in describing the image sample, it is possible to consider both the completely correct description situation (i.e., the original description text) and the situation of partial information loss or error (i.e., the adjusted text), and further make the training data more comprehensive and reasonable, ensuring that the text-image model can perform more accurately and stably when facing different types of text-image matching tasks.
[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solution of the embodiment of the present application can be applied;
[0026] Figure 2 A schematic diagram of an image detected by applying the technical solution of the embodiment of the present application;
[0027] Figure 3 A schematic diagram showing the multi-service joint application by applying the technical solution of the embodiment of the present application;
[0028] Figure 4Shows a flowchart of a data processing method according to an embodiment of the present application;
[0029] Figure 5 Shows a schematic diagram of the training process of a graphic-text model according to an embodiment of the present application;
[0030] Figure 6 and Figure 7 Shows a schematic diagram of a new sample obtained by random modification according to an embodiment of the present application;
[0031] Figure 8 and Figure 9 Shows a schematic diagram of a new sample obtained by random erasing according to an embodiment of the present application;
[0032] Figure 10 Shows a schematic diagram of fine-grained modeling according to an embodiment of the present application;
[0033] Figure 11 Shows a schematic diagram of the structure of a graphic-text model according to an embodiment of the present application;
[0034] Figure 12 Shows a block diagram of a data processing device according to an embodiment of the present application;
[0035] Figure 13 Shows a schematic diagram of the structure of a computer system of a computer device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0036] Now, the exemplary embodiments will be described in a more comprehensive manner with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as being limited to these examples; on the contrary, these embodiments are provided so that the present application is more comprehensive and complete, and the concept of the exemplary embodiments is fully conveyed to those skilled in the art.
[0037] In addition, the features, structures, or characteristics described in the present application can be combined in any suitable manner in one or more embodiments. In the following description, there are many specific details so that the embodiments of the present application can be fully understood. However, those skilled in the art should be aware that when implementing the technical solutions of the present application, not all the detailed features in the embodiments are required, one or more specific details can be omitted, or other methods, elements, devices, steps, etc. can be adopted.
[0038] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.
[0039] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0040] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0041] It should be noted that: "a plurality of" mentioned in this article means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0042] It can be understood that before and during the process of collecting relevant data of the user (such as image data, text data, etc.) in the present application, a prompt interface or a pop-up window can be displayed. The prompt interface or the pop-up window is used to prompt the user that their relevant data is currently being collected, so that the present application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation issued by the user for the prompt interface or the pop-up window. Otherwise (that is, when the confirmation operation issued by the user for the prompt interface or the pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, all user data collected by the present application is collected with the consent and authorization of the user, and the collection, use, and processing of the relevant user data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0043] The technical solution of the embodiment of the present application relates to the related processing of the graphic-text model. Specifically, due to the development of the cross-modal graphic-text model, the effect of graphic-text consistency discrimination has been significantly improved, and it can be widely applied to scenarios such as image retrieval (such as retrieving matching images according to the description text input by the user), text-to-image generation, etc. For example, in the scenario of text-to-image generation, after generating an image according to the text, in order to evaluate whether the generated image matches the text, the graphic-text model can output the feature vector of the image and the feature vector of the text respectively, and then score the consistency of the image and the text according to the similarity between the feature vector of the image and the feature vector of the text.
[0044] However, different business scenarios (such as images in game scenarios, images in anime scenarios, pet images, human images, etc.) have different requirements for the text granularity that can be expressed by the graphic-text similarity, and the focuses are also different. Some business scenarios focus on foreground description, or both foreground and background, or focus on the consistency of multiple fine-grained information. For example, in the scenario of text-to-image selection for human-related images, the focus requires the consistency of the human appearance, clothing, and actions with the text description, and to a certain extent, the background should also match. However, the technical solutions proposed in the related art lack the modeling of the graphic-text relevance with different focuses, resulting in the inability to score the similarity according to different focuses, and the application effect is poor.
[0045] Specifically, in the scenario of text-to-image generation, in the related art, the CLIP (Contrastive Language-Image Pre-training) model is directly used to extract the feature vector of the image and the feature vector of the text, and then the similarity between the feature vector of the image and the feature vector of the text is calculated to obtain the score of the generated image, that is, the higher the similarity, the higher the score of the generated image, and then the image with the highest score can be returned. Although this solution performs well in general expressions, its ability to distinguish differences at a specific level is limited. For example, for unique individuals (such as a white and yellow-spotted pet dog and a white and black-spotted pet dog), CLIP cannot accurately distinguish these subtle differences. In addition, this method based on general features is difficult to adjust its judgment criteria according to the different contents of the image foreground and background, resulting in limited ability to correctly identify foreground objects in complex backgrounds. At the same time, due to the lack of specific optimization for different business scenarios, this method performs poorly when dealing with different categories of things (such as pet dogs in anime and pet dogs in game videos).
[0046] Based on the above technical problems, an embodiment of the present application proposes a new data processing solution, which can split the description text into multiple specific description words (such as categories, clothing, actions, etc.) by using the set template words, enabling the image-text model to more carefully understand and process the key information in the text. This fine-grained text parsing ability enhances the recognition accuracy of the image-text model for the description text. At the same time, by adjusting the description text (such as modifying or deleting certain description words), a large number of new texts with different degrees of differences can be generated, enriching the training data set, making the training data more comprehensive and reasonable, and ensuring that the image-text model can perform more accurately and stably when facing different types of image-text matching tasks.
[0047] Specifically, in one example, as Figure 1 shown is a schematic diagram of an exemplary system architecture to which the technical solution of the embodiment of the present application can be applied. The system architecture includes a model training device 101, a training sample device 102, and a model application device 103. The data processing solution provided by the present application can be executed by the model training device 101 or the model application device 103. Among them, the model training device 101 can be directly or indirectly connected to the training sample device 102 in a wired or wireless manner.
[0048] It should be noted that Figure 1 the number and form of the devices shown are for illustration purposes and do not constitute a limitation on the embodiments of the present application. Optionally, the model training device 101 and the model application device 103 can be the same electronic device or two different electronic devices. Optionally, the model training device 101 can also be the same device as the training sample device 102 or different electronic devices; or the model training device 101 can also be the same device as the training sample device 102 and the model application device 103 or different electronic devices. The present application does not make any limitations in this regard.
[0049] Among them, the model training device 101, the training sample device 102, and the model application device 103 may specifically be terminal devices or servers. The terminal devices may include, but are not limited to: smart phones (such as Android phones, IOS phones, etc.), tablet computers, portable personal computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, wearable devices, etc. The embodiments of the present application do not make limitations in this regard; the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The embodiments of the present application do not make limitations in this regard.
[0050] In an embodiment of the present application, the training sample device 102 may provide an image-text pair sample to the model training device 101. The image-text pair sample includes an image sample and the description text corresponding to the image sample. After obtaining the image-text pair sample, the model training device 101 may split the description text into description words corresponding to each template word according to the set template words, and obtain the description words of the description text. For example, in an example, if the description text is "a woman wearing a white T-shirt walks on the road and meets a giraffe", and the set template words are [class], [clothing], [doing], [others], then the description text can be split into the description word "a woman" corresponding to the template word [class], the description word "wearing a white T-shirt" corresponding to the template word [clothing], the description word "walks" corresponding to the template word [doing], and the description word "on the road and meets a giraffe" corresponding to the template word [others].
[0051] Then the model training device 101 may adjust the description words of the description text to obtain an adjusted text that is different from the description text. For example, for the above example, the description word "wearing a white T-shirt" corresponding to the template word [clothing] can be adjusted to "wearing a blue T-shirt", so that an adjusted text that is different from the original description text can be obtained.
[0052] After obtaining the adjusted text, the model training device 101 can calculate the description score of the adjusted text according to the change of the adjusted text compared with the description text. This description score is used to represent the correct rate of the adjusted text in describing the image sample. For example, if the weight values corresponding to the template words [class], [clothing], [doing], and [others] are 2, 2, 1, and 1 respectively, then for the above example, since the description word corresponding to the template word [clothing] is adjusted, the description score of the adjusted text can be 1 - 2 / (2 + 2 + 1 + 1) = 2 / 3. In other words, compared with the description score of the original description text being 1, the description score of the adjusted text is 2 / 3. Furthermore, a new image-text pair sample can be generated based on the adjusted text and the image sample, and the image-text model can be trained according to the new image-text pair sample and the description score of the adjusted text to obtain a trained image-text model, which can be used to evaluate the matching degree between the image and the text.
[0053] In some alternative embodiments, after obtaining the trained image-text model, the trained image-text model can be deployed in the model application device 103. Furthermore, the model application device 103 can use the trained image-text model to evaluate the matching degree between the image and the text. Specifically, for example, in an image retrieval scenario, when the user inputs a description text, the model application device 103 can extract a text feature vector according to the description text input by the user, and then match the text feature vector with the image feature vector of the candidate image, and return several images to the user in the order of similarity from high to low.
[0054] The following are the application scenarios of the technical solutions of the embodiments of the present application:
[0055] In an application scenario of this application, the technical solution of the embodiments of this application can be applied to the scenario of custom text retrieval and recall of images. Specifically, an instance database of the user can be established first, such as establishing an instance image database according to the instance images in the user's electronic photo album, and then relevant images can be searched according to the user's search intention. It should be noted that in the embodiments of this application, an instance refers to a specific object, person, animal, etc. For example, if the description text is "A little dog wearing red overalls is playing in the park", then the "little dog" is an instance. For example, if the user wants to retrieve the pet dog in the electronic photo album, then the technical solution of the embodiments of this application can be adopted to quickly fine-tune the cross-modal model using the text-image pairs of the pet dog as samples, and then extract the fine-tuned image representations for the images in the electronic photo album. After the user inputs the desired instance description text (such as the pet dog lying on the sofa), the text representation can be extracted for the instance description text input by the user, and then the similarity between the text representation and the image representations of each image in the electronic photo album can be calculated, and several images can be returned in the order of similarity from high to low.
[0056] In an application scenario of this application, the technical solution of the embodiments of this application can be applied to the scenario of general text retrieval and recall of images. Specifically, if it is necessary to retrieve images in an anime scene that are consistent with the user's description, video anime images can be collected first, and image descriptions can be generated by means of image-to-text conversion to obtain text-image pair samples. According to the text-image pair samples, the technical solution of the embodiments of this application can be used to train the text-image model, and the model service can be deployed after the text-image model is trained. When applying, the user inputs the description text corresponding to the image to be obtained, and then the anime image corresponding to the description text can be retrieved. For example, if the description text input by the user is "Rural natural style, two little birds flying in the sky, surrounded by green trees and white clouds", then the retrieved anime image can be as Figure 2 shown. In this example, a style word (i.e., "Rural natural style") is added to the description text. In this case, when training the text-image model, the style word can be processed as a set template word.
[0057] In an application scenario of this application, the technical solution of the embodiment of this application can be applied to the scenario of graphic-text consistency scoring. Specifically, for a certain text-to-image generation service, graphic-text pair samples for a certain instance or graphic-text pair samples for multiple specific instances in the service scenario can be collected, and then template words for the text can be designed based on this (for example, if the instance in the service scenario is a cute pet, the template words can be: style word + subject + adverb + predicate, etc.). After that, the text-to-image generation model can be fine-tuned, and the graphic-text model for evaluating similarity can be trained by using the technical solution of the embodiment of this application. Then, the text-to-image generation model and the graphic-text model for evaluating similarity can be deployed. When applying, the description text input by the user is collected, multiple images are generated by using the text-to-image generation model, the similarity between the description text and the generated images is calculated by using the graphic-text model for evaluating similarity, and then the generated images with similarity scores greater than or equal to the threshold (for example, the threshold is 0.3. In this case, the instances in the selected images may be correct, but there may be inconsistencies in other aspects such as clothing and actions) are selected and returned to the user in descending order of similarity scores.
[0058] Similar to the above description, the instance in the embodiment of this application refers to a specific object, person, animal, etc. If the instance is a pet, then for pet image generation, the appearance and clothing of the pet need to be consistent (for a pet dog, the clothing can be considered as the characteristics on the pet dog. For example, if the pet dog has spots on its body, then the spots on the pet dog can be considered as clothing) to confirm the pet's identity. Then, images with similarity greater than the threshold but inconsistent actions can be returned. When there are no such images, the system can prompt the user "Insufficient samples, please regenerate the image"; when multiple generations do not meet the image requirements, the system can prompt the user "Add more pet images with different actions and different lighting conditions". If the instance is a person, then for person image generation, the appearance of the person needs to be consistent, and the clothing and actions can vary (since the recognition rate of the human face is higher, even if the clothing is different, it can still be considered the same person based on the facial appearance). Then, images with similarity greater than the threshold but inconsistent clothing and actions can be returned.
[0059] In an application scenario of the present application, the technical solution of the embodiments of the present application can be used for multi-service joint applications. Specifically, for each business scenario (such as business scenarios of images in gaming scenarios, images in animation scenarios, pet images, human images, etc.), the text-image model in the embodiments of the present application can be trained to obtain model business parameters corresponding to each business scenario. For example, model parameter 1 is obtained by training with data of business 1, and model parameter 2 is obtained by training with data of business 2. At the same time, the basic parameters for model inference need to be obtained. The basic parameters for model inference are the native parameters of the text-image model. For example, if the embodiments of the present application are obtained by fine-tuning on the basis of the CLIP model, then the native parameters of the CLIP model are the basic parameters for model inference. When a model inference is requested in a specific business scenario, as Figure 3 shown, the text-image consistency service is the service obtained by deploying the text-image model trained in the embodiments of the present application. When a user sends a business request through interaction, the business id can be added to the business request. Then, in the embodiments of the present application, the corresponding model business parameters can be determined according to the business id (such as Figure 3 the text operator, image operator, fusion operator, etc. shown), and then the basic parameters for model inference are combined with the model business parameters for inference to obtain the final output. For example, since model parameter 1 is obtained by training with data of business 1, then model parameter 1 is used as the model business parameter when performing model inference for business 1.
[0060] The implementation details of the technical solution of the embodiments of the present application are elaborated in detail below:
[0061] Figure 4 FIG. shows a flowchart of a data processing method according to an embodiment of the present application. The data processing method can be executed by a computer device, which can be a server, a terminal device, or other devices with computing and processing capabilities, etc. Referring to Figure 4 shown, the data processing method includes at least S410 to S440, which are introduced in detail as follows:
[0062] In S410, a text-image pair sample is obtained, and the text-image pair sample includes an image sample and a description text corresponding to the image sample.
[0063] In an embodiment of the present application, the "image-text pair sample" refers to a set of data containing an image and its corresponding descriptive text. The image-text sample pair is the basis for training and evaluating the image-text model and is used for the image-text model to perform the image-text consistency discrimination task, that is, to determine whether the image accurately reflects the content described in the text. Optionally, the image sample can be a specific image, which can depict any object or scene, such as a person, an animal, a landscape, etc. The image sample can be an image in any business scenario, such as an image obtained by shooting with a camera, an image intercepted from a video (such as an anime video, a game video, a TV drama, etc.), or an image generated by artificial intelligence content generation (AIGC) technology, etc.
[0064] Optionally, the descriptive text corresponding to the image sample is a literal description of the content in the image sample, and usually can include a detailed description of the main objects (such as people, animals, landscapes, etc.) in the image sample, such as the category, clothing, actions, and other relevant information of the described object. For example, the descriptive text can be "A girl in a blue dress is feeding pigeons in the park".
[0065] In some alternative embodiments, in order to obtain accurate image-text pair samples (accurate image-text pair samples refer to those where the descriptive text can completely and accurately describe the content in the image sample), one or more candidate image-text pairs can be obtained, and then accurate image-text pair samples can be selected from the candidate image-text pairs.
[0066] Optionally, the candidate image-text pair can include a candidate image and a candidate descriptive text. When selecting an accurate image-text pair sample from the candidate image-text pairs, the candidate descriptive text can be split into descriptive words corresponding to each template word according to the template word to obtain the descriptive words of the candidate descriptive text. If all the content described by the descriptive words of the candidate descriptive text appears in a certain candidate image, then the candidate image-text pair is used as the image-text pair sample.
[0067] In some alternative embodiments, splitting the candidate description text into description words corresponding to each template word according to the template words may be as follows: first determine the description words corresponding to each template word in the candidate description text, and then split out the description words corresponding to each template word. In a specific example, if the candidate description text is "A woman wearing a white T-shirt walks on the road and meets a giraffe", and the set template words are [class], [clothing], [doing], [others], then the description words corresponding to the template word [class] in the candidate description text are "A woman", the description words corresponding to the template word [clothing] are "wearing a white T-shirt", the description words corresponding to the template word [doing] are "walks", and the description words corresponding to the template word [others] are "on the road and meets a giraffe". Furthermore, the candidate description text can be split into the description words "A woman" corresponding to the template word [class], the description words "wearing a white T-shirt" corresponding to the template word [clothing], the description words "walks" corresponding to the template word [doing], and the description words "on the road and meets a giraffe" corresponding to the template word [others].
[0068] In some alternative embodiments, when determining whether the content described by the description words in the candidate description text appears in a certain candidate image, the image feature vector of the candidate image and the text feature vectors of the respective description words in the candidate description text may be extracted, and then the similarity between the image feature vector and the text feature vectors of the respective description words is calculated. Furthermore, based on the similarity, it is determined whether the content described by each description word appears in the candidate image. For example, a similarity threshold can be set. If the similarity between the image feature vector of the candidate image and the text feature vector of a certain description word is greater than or equal to the similarity threshold, then it can be determined that the content described by the description word appears in the candidate image.
[0069] Optionally, different similarity thresholds can be set for different template words because the importance of different template words may be different. For example, the template word [class] describes the main object in the image, and the template word [doing] describes the action of the object in the image. Then the importance of the template word [class] is generally higher than that of the template word [doing]. Therefore, the similarity threshold set for the template word [class] can be greater than or equal to the similarity threshold set for the template word [doing].
[0070] In S420, the description text is split into description words corresponding to each template word according to the set template words, obtaining the description words of the description text.
[0071] In some alternative embodiments, splitting the description text into description words corresponding to each template word according to the template words is similar to the process of splitting the candidate description text into each description word in the above embodiments, that is, the description words corresponding to each template word in the description text can be determined, and then the description words corresponding to each template word are split out from the description text.
[0072] In S430, the description words of the description text are adjusted to obtain an adjusted text that is different from the description text, and according to the change situation of the adjusted text compared with the description text, the description score of the adjusted text is calculated. The description score is used to represent the correct rate of the adjusted text for describing the image sample.
[0073] In the embodiments of the present application, adjusting the description words of the description text to obtain an adjusted text that is different from the description text is mainly for data augmentation, that is, increasing the number of training samples, and at the same time, various error situations that may occur in the actual application scenario can be simulated, thereby improving the robustness of the image-text model.
[0074] In some alternative embodiments, the process of adjusting the description words of the description text to obtain an adjusted text that is different from the description text may be: replacing at least one description word of the description text with other description words of the same type to obtain the adjusted text. In this embodiment, the reason for replacing at least one description word of the description text with other description words of the same type is mainly to reduce the adjustment range of the description text and avoid a large adjustment range of the description text resulting in a large difference between the adjusted text and the image sample. Of course, in other embodiments of the present application, the description words of the description text may also be replaced with other description words of different types.
[0075] In an example, if the description text is "A woman wearing a white T-shirt walks on the road and meets a giraffe", then the description words obtained according to the technical solution in the above embodiments are: the description word "A woman" corresponding to the template word [class], the description word "wearing a white T-shirt" corresponding to the template word [clothing], the description word "walks" corresponding to the template word [doing], and the description word "on the road and meets a giraffe" corresponding to the template word [others]. In this case, when replacing the description words of the same type, for example, the description word "wearing a white T-shirt" corresponding to the template word [clothing] can be replaced with "wearing a blue T-shirt"; the description word "A woman" corresponding to the template word [class] can be replaced with "A man"; the description word "walks" corresponding to the template word [doing] can be replaced with "sits", etc.
[0076] In some alternative embodiments, if the adjusted text is obtained by replacing descriptive words, then when calculating the descriptive score of the adjusted text, the weight proportion of the descriptive words that are not replaced in the descriptive words of the descriptive text can be calculated based on the descriptive words replaced in the descriptive text and the weights of the respective descriptive words in the descriptive text, and then this weight proportion can be used as the descriptive score of the adjusted text.
[0077] For example, if the weight values corresponding to the template words [class], [clothing], [doing], and [others] are 2, 2, 1, and 1 respectively (it should be noted that the weight value of the descriptive word corresponding to the template word is the same as the weight value corresponding to the template word), then the descriptive score of the original descriptive text can be 1. If the descriptive word corresponding to the template word [clothing] is replaced, then the weight of the descriptive words that are not replaced in the descriptive text is 2 + 1 + 1 = 4, and the weight proportion of the descriptive words that are not replaced in the descriptive words of the descriptive text is 4 / (2 + 2 + 1 + 1) = 2 / 3. If the descriptive words corresponding to the template words [clothing] and [doing] are replaced respectively, then the weight of the descriptive words that are not replaced in the descriptive text is 2 + 1 = 3, and the weight proportion of the descriptive words that are not replaced in the descriptive words of the descriptive text is 3 / (2 + 2 + 1 + 1) = 1 / 2.
[0078] In some alternative embodiments, the process of adjusting the descriptive words of the descriptive text to obtain an adjusted text that is different from the descriptive text can be: deleting at least one descriptive word of the descriptive text to obtain the adjusted text. In one example, if the descriptive text is "a woman wearing a white T-shirt walks on the road and meets a giraffe", then the descriptive words obtained according to the technical solution in the above embodiment are: the descriptive word "a woman" corresponding to the template word [class], the descriptive word "wearing a white T-shirt" corresponding to the template word [clothing], the descriptive word "walks" corresponding to the template word [doing], and the descriptive word "on the road and meets a giraffe" corresponding to the template word [others]. In this case, when deleting descriptive words, for example, the descriptive word corresponding to the template word [clothing] can be deleted; or the descriptive word "walks" corresponding to the template word [doing] can be deleted, etc.
[0079] In some alternative embodiments, the set template words can include subject words. In this case, deleting at least one descriptive word of the descriptive text can be: deleting at least one descriptive word other than the descriptive word corresponding to the subject word in the descriptive text. The technical solution of this embodiment can avoid the description text changing too much due to deleting the subject word and affecting the final model training effect.
[0080] In some alternative embodiments, if the adjusted text is obtained by deleting descriptive words, then when calculating the description score of the adjusted text, the ratio between the weight of the deleted descriptive words in the descriptive text and a set scale factor can be calculated, and then the difference between the preset maximum description score and this ratio is used as the description score of the adjusted text. Optionally, since in the embodiments of the present application, there may be both ways of replacing descriptive words and deleting descriptive words, the description scores of the adjusted texts obtained by these two ways can be adjusted by the set scale factor, so as to avoid the description scores obtained by the way of replacing descriptive words and the way of deleting descriptive words differing too much and affecting the training effect of the image-text model. In some alternative embodiments, the preset maximum description score can be 1 (or it can also be other values). If the preset maximum description score is 1 and the ratio between the weight of the deleted descriptive words and the scale factor is 1 / 3, then the description score of the obtained adjusted text is 1 - 1 / 3 = 2 / 3.
[0081] In some alternative embodiments, since the adjusted text after deleting some descriptive words still corresponds to the image sample, except that some descriptive words are reduced, while the adjusted text obtained by replacing some descriptive words has inconsistent and incorrect content with the image sample. Therefore, even if all descriptive words in the descriptive text except the specified descriptive words (such as the subject) are deleted, the description score of the obtained adjusted text should still be greater than the description score obtained by replacing descriptive words. Based on this, the embodiments of the present application propose a technical solution for calculating the above-mentioned scale factor, which is specifically as follows:
[0082] In some alternative embodiments, the maximum description score of the first adjusted text obtained by replacing the description words in the description text can be calculated according to the weights of the respective description words in the description text, as well as the description score of the second adjusted text obtained by deleting all description words in the description text except the specified description words. Then, with the constraint that the description score of the second adjusted text is greater than the maximum description score of the first adjusted text, the value of the scaling factor can be calculated. For example, if the weight values corresponding to the template words [class], [clothing], [doing], and [others] are 2, 2, 1, and 1 respectively, that is, the weight values of the description words corresponding to the template words [class], [clothing], [doing], and [others] are 2, 2, 1, and 1 respectively, then the maximum description score of the first adjusted text obtained by replacing the description words in the description text (such as replacing the description word corresponding to the template word [doing] or the description word corresponding to the template word [others]) is 1 - 1 / (2 + 2 + 1 + 1) = 5 / 6. The description score of the second adjusted text obtained by deleting all description words in the description text except the specified description word (such as the description word corresponding to [class]) can be 1 - (2 + 1 + 1) / M, where M represents the value of the scaling factor set. Therefore, when 1 - (2 + 1 + 1) / M > 5 / 6 is satisfied, M > 24 can be calculated, and then the value of the scaling factor can be selected, such as 25, 16, etc.
[0083] Continue to refer to Figure 4 As shown, in S440, new text-image pair samples are generated based on the adjusted text and the image samples, and the text-image model is trained according to the new text-image pair samples and the description scores of the adjusted text.
[0084] In some alternative embodiments, in addition to training the text-image model based on the description scores of the new text-image pairs for the samples and the adjusted text, the text-image model can also be trained based on the description scores of the original text-image pairs for the samples and the original description text (the description score of the original description text can be the maximum description score, for example, it can be 1), and the trained text-image model can be used to evaluate the matching degree between the text and the image. For example, if a user inputs a piece of text and an image, the trained text-image model can evaluate the matching degree between this piece of text and this image, that is, evaluate the description score of this piece of text for this image; another example is that the text-to-image model generates an image based on a piece of text, then the trained text-image model of the present application can be used to evaluate the matching degree between this piece of text and the generated image, and further evaluate the performance of the text-to-image model; yet another example is that a user inputs a piece of text and wants to retrieve a matching image from a database, then the trained text-image model of the present application can evaluate the matching degree between each image in the database and the text input by the user, and further select one or more images according to the matching degree and return them to the user.
[0085] In some alternative embodiments, the text-image model may include a text encoding module and an image encoding module. As the name implies, the text encoding module is used to perform encoding processing on the text content to obtain a text feature vector, and the image encoding module is used to perform encoding processing on the image content (such as an image sample) to obtain an image vector. In this case, the process of training the text-image model can be as follows: according to the template words and the word representations corresponding to the description words of the adjusted text, generate the input data of the text encoding module, and then the text feature vector output by the text encoding module according to the input data can be obtained, and at the same time, the image feature vector output by the image encoding module for the image sample can also be obtained; then the similarity score between the text feature vector and the image feature vector can be calculated, the loss data can be calculated according to this similarity score and the description score of the adjusted text, and the model parameters of the text-image model can be adjusted according to this loss data to obtain the trained text-image model.
[0086] In the above embodiments, by generating the input data of the text encoding module according to the template words and the word representations corresponding to the description words of the adjusted text, it is possible to perform separate modeling on each template word with fine-grained division, and further enable the text-image model to more carefully understand and process the key information in the text, so as to enhance the recognition accuracy of the text-image model for the description text. When calculating the loss data according to the similarity score and the description score of the adjusted text, the mean square error between the similarity score and the description score of the adjusted text can be calculated, and this mean square error can be used as the calculated loss data.
[0087] In some alternative embodiments, the image encoding network includes a feature extraction network for extracting visual features of an image, a first attention pooling module and a second attention pooling module respectively connected to the feature extraction network, and a feature fusion module for fusing the output data of the first attention pooling module and the output data of the second attention pooling module. On this basis, the process of adjusting the model parameters of the text-image model according to the loss data may be: adjusting the parameters of the second attention pooling module, the parameters of the feature fusion module, and the word representation parameters corresponding to the description words. In this embodiment, the first attention pooling module and the feature extraction network in the image encoding network may be network modules in a pre-trained model. When the technical solution of the embodiment of the present application is used for training, the model parameters of these two network modules may not be adjusted. Instead, a new attention pooling module and a feature fusion module are introduced into the image encoding network in the embodiment of the present application. In this way, the attention of the text-image model to different regions in the image can be enhanced through the newly introduced attention pooling module in the image encoding network, so as to more accurately locate and describe the image elements associated with the text. For example, when processing the description text "a woman in a white T-shirt meets a giraffe on the road", the model can pay more attention to the regions related to the woman, the white T-shirt, and the giraffe in the image. It can be seen that the technical solution of the embodiment of the present application can enhance the ability of the text-image model to capture fine-grained features. By setting learnable word representation parameters for the description text, the text-image model can better capture and understand the fine-grained information in the text description by fine-tuning the word representation parameters, and thus can improve the accuracy and effect of the text-image model in text-image consistency discrimination.
[0088] In some alternative embodiments, when adjusting the model parameters of the text-image model according to the loss data, the training data may be divided into multiple batches, each batch containing at least one text-image pair sample. Then, according to the loss data corresponding to each text-image pair sample, the loss data of each batch is calculated, and the model parameters of the text-image model are adjusted according to the loss data of each batch. If the model parameters of the text-image model are adjusted according to the loss data of multiple batches, a round of training process for the text-image model is completed. If the training of the text-image model reaches the training end condition, it is determined that the training of the text-image model is completed. Optionally, when calculating the loss data of each batch, the sum of the loss data corresponding to the text-image pair samples included in a batch may be calculated, or the loss data corresponding to the text-image pair samples included in a batch may be averaged to obtain the loss data of each batch. Then, the Stochastic Gradient Descent (SGD) method may be used to backpropagate the loss data of each batch back into the text-image model to obtain the gradient of the model parameters and update the model parameters.
[0089] Optionally, the training of the image-text model reaches the training end condition, which can be, for example, meeting one or more of the following conditions: completing a set number of rounds of training for the image-text model, the training duration of the image-text model reaching a set duration, the change in the loss function of the image-text model being less than a set threshold, the accuracy rate of the image-text model reaching a set threshold, and so on. It should be noted that if the change in the loss function of the image-text model is less than a set threshold, further training may not bring significant performance improvement, so it can be determined that the training of the image-text model is completed.
[0090] In some alternative embodiments, image-text pairs in different business scenarios can be used to train the image-text model respectively to obtain the model business parameters for each business scenario. For example, image-text pairs in business scenarios such as game scenario image-text pairs, anime scenario image-text pairs, pet-related image-text pairs, and person-related image-text pairs can be used to train the image-text model in the embodiments of the present application to obtain the model business parameters corresponding to each business scenario. For example, the image-text pairs of business 1 (such as game scenario image-text pairs) are used to train model parameter 1, and the image-text pairs of business 2 (such as anime scenario image-text pairs) are used to train model parameter 2.
[0091] In the case of the above embodiments, the scenario identification information of each business scenario can be associated and stored with the model business parameters. Then, in the process of applying the image-text model, the associated model business parameters can be selected according to the scenario identification information carried in the business request, and the business request can be responded to through the image-text model. For example, in the game scenario, if image-text consistency evaluation is required, the identification information of the game scenario can be carried in the business request. Then, the model business parameters corresponding to the game scenario can be obtained according to this identification information, and then the basic parameters of model reasoning are combined with this model business parameter for reasoning to obtain the final output. The technical solution of this embodiment enables the joint processing of multiple business scenarios and improves the flexibility and adaptability of the image-text model.
[0092] It should be noted that the image-text model in the embodiments of the present application can be applied to many application scenarios, such as the scenarios of customized text retrieval and recall of images, general text retrieval and recall of images, image-text consistency scoring, etc. mentioned in the foregoing embodiments. Taking image-text consistency scoring as an example, in combination with Figures 5 to 11 , the implementation details of the technical solution of the embodiments of the present application are described in detail again:
[0093] Refer to Figure 5The following shows the training process of the text-image model in an embodiment of the present application. For each description text, four representations to be fine-tuned can be established according to preset template words (such as [class], [clothing], [doing], [others]), which respectively represent the category [class], clothing [clothing], action [doing], and others [others] (or it can also be the background, etc.) of an instance (an instance refers to a specific object, person, animal, etc.). The established representations to be fine-tuned and the template words are passed through a text encoder to generate a text description representation (i.e., a text feature vector). At the same time, the image sample passes through an image encoder to generate an image representation (i.e., an image feature vector). Then, the similarity between the text feature vector and the image feature vector is calculated. For text-image sample pairs with consistent text and image, the learnable representation parameters at the description text end and the parameters in the image encoder can be adjusted with the goal of maximizing the similarity score to obtain the trained text-image model. Optionally, during model training, multiple enhanced samples can be generated by using enhancement methods such as different degrees of text deletion or incorrect text. On the one hand, it enriches the training data, and on the other hand, it can generate training data that meets text-image consistency to a certain extent, meeting the requirements of the training goal. The following is a detailed description:
[0094] In an embodiment of the present application, the training data includes image samples and the description texts of the image samples. A quick method is to directly use the text-image pairs for training the text-to-image model as the training samples in the embodiment of the present application to avoid secondary data collection. If there is no ready-made text-image pair data, the corresponding text-image data can be collected according to the specific business scenario. For example, if text-image retrieval or text-image consistency evaluation of text-to-image in movies and TV shows is required, data can be collected by matching movie and TV show images with description texts. For example, images of the same instance (such as characters, roles, heroes, etc.) with different clothing, different actions, and in different environments can be collected through movies and TV shows. Such rich images can train a detailed representation model. After obtaining the movie and TV show images, the corresponding description texts can be generated through a pre-trained model, such as the Bootstrapping Language-Image Pre-training (BLIP) model.
[0095] In an embodiment of the present application, the obtained description text can be aligned according to a predetermined format. For example, if the predetermined format is [class]+[clothing]+[doing]+[others], then it is necessary to determine the nouns (corresponding to [class]), clothing descriptors (corresponding to [clothing]), actions (corresponding to [doing]), and other words (corresponding to [others], and other words can include the environment, object, atmosphere words, and other related descriptions) in the description text. When processing, the words in the description text can be split and then mapped separately, that is, split into subject nouns, descriptive adjectives, actions, objects, adverbs, etc. The subject (or the first noun) is used as the word represented by [class], the subject descriptor is used as the clothing description [clothing], the verb is used as [doing], and the remaining words can be classified as [others] (i.e., the words with less impact in the model application scenario), so as to generate the description words of the description text. For example, for the description text "A woman wearing a white T-shirt walks on the road and meets a giraffe", after splitting, it can be obtained: [class] (A woman) + [clothing] (wearing a white T-shirt) + [doing] (walks) + [others] (meets a giraffe on the road). Optionally, when splitting the words in the description text, a tokenization tool such as the Natural Language Toolkit (NLTK) can be used for splitting processing.
[0096] In an embodiment of the present application, in order to obtain sufficient training data, sample data associated with [class], [clothing], [doing], etc. can also be collected in advance. For example, collect the types that may be encountered in the business scenario. For example, [class] includes: various people (men, women, the elderly, children, students, etc.), animals, pets, etc.; [clothing] includes clothes of various colors, various styles (windbreakers, work vests, jumpsuits, dresses, etc.), various accessories (schoolbags, hats, shoes, etc.); [doing] includes various actions that [class] can have (such as running, jumping, walking, lying prone, etc.).
[0097] An embodiment of the present application also proposes a data augmentation scheme. Specifically, after obtaining the description text, the description is first templatized, that is, the descriptive words in the description text are matched according to the set template word pairs. For example, for description texts in the form of subject-predicate-object, subject-predicate, etc., the sentence patterns can include: subject-predicate-object; subject-predicate; subject adjective + subject + predicate + object; subject adjective + subject + adverb + predicate + object adjective + object, etc. Optionally, for images of different styles, style words can also be added as prefixes to fine-tune the style information into the text-image model. For example, a style word can be added or subtracted based on the template word.
[0098] The following takes the template with the sentence pattern of subject + subject adjective + predicate + other words as an example for introduction.
[0099] In some alternative embodiments, the descriptive words in the description text can be randomly modified to obtain augmented text. Specifically, for the description text in the text-image pair sample, after aligning the description format, that is, after determining the descriptive words corresponding to [class], [clothing], [doing], and [others] respectively, random sampling can be performed on the descriptive words corresponding to [class], [clothing], [doing], and [others] respectively, so that they are inconsistent with the original description, and the inconsistent parts are recorded, with a total of k1 samplings. For example, if the description text is: [class] (a woman) + [clothing] (wearing a white T-shirt) + [doing] (walking) + [others] (meeting a giraffe on the road), then it can be sampled to get: [class] (a woman) + [clothing] (wearing blue clothes) + [doing] (walking) + [others] (meeting a giraffe on the road), and the inconsistent one is the descriptive word corresponding to [clothing]. In this way, k1 new texts can be obtained. These k1 new texts and the original image samples together obtain k1 text-image data, which can be recorded as: k1 (image, inconsistent text, inconsistent part, 1), where the "1" indicates the error text generated by the random modification method.
[0100] In one embodiment of the present application, after obtaining new graphic and text data through random modification, each new graphic and text data can be scored (i.e., the description score of the new graphic and text data is determined). Specifically, weights can be set for [class], [clothing], [doing], and [others] respectively: 2, 2, 1, 1. The higher the weight, the higher the importance. At this time, N = the sum of the weights, which is 6 (it should be noted that the specific values of the weights can be set according to actual needs. For example, the sum of the weights can also be set to 1). Then, the score after a description error can be determined according to the weights of different description words: The full score for scoring is 1 point, that is, the description words corresponding to [class], [clothing], [doing], and [others] are all correct. If the description word corresponding to [clothing] is incorrect, then the score is 1 - 2 / N; if the description word corresponding to [doing] is incorrect, then the score is 1 - 1 / N; if the description words corresponding to both [clothing] and [doing] are incorrect, then the score is 1 - 3 / N; if the description word corresponding to [class] is incorrect, then the score is 1 - 2 / N; if the description word corresponding to [others] is incorrect, then the score is 1 - 1 / N.
[0101] In one example of the present application, as Figure 6 shown, k1 new samples (i.e., k1 new texts) are obtained by randomly modifying the description words corresponding to [class] and [doing] in the original sample (i.e., the original description text). In this case, the similarity score between the text feature vector of the original sample and the image feature vector corresponding to the image sample is the maximum (i.e., it can be a full score, such as 1 point); while the similarity scores between the text feature vectors of the k1 new samples and the image feature vector corresponding to the image sample are partial scores (the partial scores are less than the maximum score, such as less than 1 point), which is caused by introducing error information by modifying the original description text.
[0102] In another example of the present application, as Figure 7 shown, k1 new samples (i.e., k1 new texts) are obtained by randomly modifying the description words corresponding to [class] and [clothing] in the original sample (i.e., the original description text). In this case, the similarity score between the text feature vector of the original sample and the image feature vector corresponding to the image sample is the maximum (i.e., it can be a full score, such as 1 point); while the similarity scores between the text feature vectors of the k1 new samples and the image feature vector corresponding to the image sample are partial scores (the partial scores are less than the maximum score, such as less than 1 point), which is caused by introducing error information by modifying the original description text.
[0103] In some alternative embodiments, the descriptive words in the descriptive text can be randomly erased (i.e., masked) to obtain enhanced text. Specifically, for the descriptive text in the image-text pair sample, after aligning the descriptive format, that is, after determining the descriptive words corresponding to [class], [clothing], [doing], and [others] respectively, the descriptive words corresponding to [clothing] and [doing] can be randomly erased. For example, the descriptive text is: [class] (a woman) + [clothing] (wearing a white T-shirt) + [doing] (walking) + [others] (meeting a giraffe on the road). Then, by randomly erasing the descriptive word corresponding to [clothing], we can get: [class] (a woman) + [doing] (walking) + [others] (meeting a giraffe on the road). If k2 random erasures are performed, then k2 new texts can be obtained. These k2 new texts together with the original image samples result in k2 image-text data, which can be denoted as, for example: k2 (image, partially erased text, erased part, 2), where the "2" indicates the text generated by the random erasure method.
[0104] In an embodiment of the present application, after obtaining new image-text data by the random erasure method, each new image-text data can be scored (i.e., determining the descriptive score of the new image-text data). Specifically, weights can be set for [class], [clothing], [doing], and [others] respectively: 2, 2, 1, 1. The higher the weight, the higher the importance (it should be noted that the specific values of the weights can be set according to actual needs). Then, the scoring after erasure can be determined according to the weights of different descriptive words: the full score is 1 point, that is, the descriptive words corresponding to [class], [clothing], [doing], and [others] are not erased; if the descriptive word corresponding to [clothing] is erased, then the score is 1 - 2 / M; if the descriptive word corresponding to [doing] is erased, then the score is 1 - 1 / M; if the descriptive words corresponding to [clothing] and [doing] are both erased, then the score is 1 - 3 / M; if the descriptive word corresponding to [others] is erased, then the score is 1 - 1 / M. Optionally, it can be set that the descriptive word corresponding to [class] is not erased, because after erasing the descriptive word corresponding to [class], the change in the descriptive text is relatively large, which may affect the effect of model training.
[0105] In the above example, the value of M is involved in the ratio adjustment in the scoring calculation. Especially when dealing with the scoring problem of randomly modifying and randomly erasing text, it is necessary to ensure that the text with randomly erased words has a higher score than the text with randomly modified words. This is because the text after erasing some descriptive words still corresponds to the image sample, except that some descriptive words are reduced, while the text obtained by modifying some descriptive words contains content that is inconsistent with and incorrect for the image sample. Therefore, even if all descriptive words except the specified descriptive words (such as the descriptive words corresponding to [class]) in the descriptive text are deleted, the score of the resulting text should be greater than the score of the text obtained by modifying the descriptive words. On this basis, if all descriptive words are erased and only the descriptive words corresponding to [class] have a score of 1 - 4 / M, and the minimum score of the text obtained by modifying the descriptive words (such as only the descriptive words corresponding to [doing] are incorrect) is 1 - 1 / N, so we can make 1 - 4 / M > 1 - 1 / N, and get that M needs to be greater than 4×N.
[0106] In an example of the present application, as Figure 8 shown, by erasing the descriptive words corresponding to [clothing] in the original sample (i.e., the original descriptive text), k2 new samples (i.e., k2 new texts) are obtained. In this case, the similarity score between the text feature vector of the original sample and the image feature vector corresponding to the image sample is the largest (i.e., it can be a full score, such as 1 point); while the similarity score between the text feature vectors of the k2 new samples and the image feature vector corresponding to the image sample is the occlusion (mask) score (this occlusion score is less than the maximum score, such as less than 1 point), which is caused by the information loss introduced by erasing the original descriptive text.
[0107] In another example of the present application, as Figure 9 shown, by erasing the descriptive words corresponding to [doing] in the original sample (i.e., the original descriptive text), k2 new samples (i.e., k2 new texts) are obtained. In this case, the similarity score between the text feature vector of the original sample and the image feature vector corresponding to the image sample is the largest (i.e., it can be a full score, such as 1 point); while the similarity score between the text feature vectors of the k2 new samples and the image feature vector corresponding to the image sample is the occlusion (mask) score (this occlusion score is less than the maximum score, such as less than 1 point), which is caused by the information loss introduced by erasing the original descriptive text.
[0108] In some alternative embodiments, due to the existence of the scaling factor M, the above scoring results may lead to too low scores, and thus the score differentiation and feature differentiation of different types may vary greatly. Therefore, the scoring results can be stretched to generate larger scores. For example, all the scoring results obtained in the above embodiments can be multiplied by Q, where Q is a constant and Q can be equal to M.
[0109] In some alternative embodiments, before performing the above data augmentation processing, the collected image-text pairs can also be processed first to select the image-text pairs with consistent images and texts. Specifically, the description text can be tokenized first using the tokenization method in the above embodiments to obtain the description words corresponding to the 4 template words (i.e., [class], [clothing], [doing], [others]). Then, each time only one description word (such as the description word corresponding to [class]) is input to generate a text feature vector, and then the similarity between the text feature vector and the image feature vector is calculated. Whether the content described by the description word appears correctly in the image is judged according to whether the similarity exceeds the similarity threshold, so as to obtain the result of whether all 4 description words appear correctly. Then, the image-text pairs with consistent images and texts can be obtained through manual inspection and correction, and then the data augmentation processing in the above embodiments can be performed based on the obtained image-text pairs with consistent images and texts.
[0110] In some alternative embodiments, the structure of the image-text model adopted in this application can be adjusted based on the CLIP multi-modal model. Specifically, considering that the objects in the image have their own unique features (such as game characters, pet dogs, etc.), the direct CLIP representation does not distinguish such unique features, making it difficult to measure the fine-grained image-text pairing effect. Therefore, in the embodiments of this application, instances in different business scenarios (such as game scenarios, anime scenarios, pet image scenarios, etc.) can be modeled separately. At the same time, as Figure 10 shown, the instances can be modeled separately from multiple perspectives such as the category, clothing attribute, and behavior attribute of the instance. Finally, the consistency between the image and the text is scored from multiple perspectives such as instance compliance, clothing compliance, and behavior compliance. As Figure 11 shown, a tunable attention pooling layer (i.e., attention pooling layer 2) and a tunable fusion module can be added based on the CLIP multi-modal model. Among them, the original structure of the image encoder in the CLIP multi-modal model only includes a feature extraction network for extracting image visual features and an attention pooling layer (i.e., attention pooling layer 1).
[0111] Among them, the tunable attention pooling layer (i.e., attention pooling layer 2) allows the model to dynamically adjust its attention to different regions of the image according to the information provided by the input text. For example, when processing the descriptive text "A woman in a white T-shirt meets a giraffe on the road", the model can pay more attention to the regions in the image related to the woman, the white T-shirt, and the giraffe. At the same time, since it is tunable, it means that the parameters in this attention pooling layer can be adjusted through the training process to adapt to specific business scenarios or specific types of image-text pairs. In this way, the image-text model can learn how to better focus its attention on the image elements most relevant to the given text description.
[0112] In some alternative embodiments, the fusion module can be a Multilayer Perceptron (MLP) fusion module. The role of the fusion module is to further fuse and process the features processed by the attention pooling layer. Optionally, a structure of multiple Fully Connected Layers (FC) + Rectified Linear Unit (ReLU) activation functions can be added before the MLP fusion module here to achieve non-linear feature transformation. Such processing can not only enable the image-text model to learn more complex feature representations but also flexibly adjust the feature dimensions to meet the requirements of subsequent tasks.
[0113] In some alternative embodiments, when training the model in the embodiments of the present application, as shown in Table 1 below, tunable vectors for category [class], clothing [clothing], and action [doing] can be established for each descriptive text respectively. These three types of features can each contain 5, 3, and 3 vectors (the number of vectors can be adjusted), and each vector can be the CLIP text dictionary feature dimension, such as 1×768. Since the category needs to record more information for differentiating instances, more vectors are required. The vectors of these three types of features can be initialized by adding random perturbations (adding 0.01×random Gaussian sampling to each position of the embedding vector) to template words (such as girl, in white, walking, etc.).
[0114] Adaptation parameter Parameter to be trained Parameter type Text template word 1 - [class] All parameters Weight Text template word 2 - [clothing] All parameters Weight Text template word 3 - [doing] All parameters Weight
[0115] Table 1
[0116] Optionally, for other or background in the text template words, a learnable feature vector can be added to represent it. For example, when the image is in a non-general background such as a comic (i.e., not present in the training data of the original CLIP model), a background representation learnable vector can be added.
[0117] In some alternative embodiments, as shown in Table 2, the parameters to be learned in the image encoder are the attention pooling layer and the subsequent feature fusion module designed for images of different service types to adapt to different text emphases.
[0118] Network layer Parameter to be trained Parameter type Attention pooling layer 2 All parameters Attention pooling Fusion module All parameters Weight
[0119] Table 2
[0120] In some alternative embodiments, when training the model, P text-image pair samples can be iterated for L rounds (such as 10 rounds). During each round of iteration, the full set of text-image pair samples are divided into batches of bs samples each, and the model parameters are updated once for each batch. When all batches (P / bs) have been trained once, that is, when all text-image pair samples have been trained in the model once, it is called one round of iteration. Optionally, the training process for each batch is as follows:
[0121] Parameter initialization before training the first batch in the first round: For all parameters to be learned, they are initialized using a normal distribution, with a learning rate of 0.0004. After every 5 rounds of learning, the learning rate becomes 0.1 times the original, and a total of 10 rounds of training are performed (it should be noted that the values in this example are only for illustration and can be adjusted according to actual needs in practical applications); bs text-image pair samples are extracted and input into the model to calculate the text-image similarity; the loss is calculated based on the scores of each text-image pair sample and the similarity output by the model, and the total loss of this batch of samples is statistically calculated. Then, using the method of stochastic gradient descent, the loss is backpropagated into the model to obtain the gradients of the model parameters and update the model parameters.
[0122] It should be noted that in other embodiments of the present application, if the training of the text-image model reaches other training end conditions, such as the training duration of the text-image model reaches a set duration, the change in the loss function of the text-image model is less than a set threshold, the accuracy of the text-image model reaches a set threshold, etc., then further training may not bring significant performance improvement, so it can also be determined that the training of the text-image model is completed.
[0123] In some alternative embodiments, when calculating the text-image similarity, the cosine similarity between the image feature vector and the text feature vector can be calculated. When calculating the similarity, the text feature vector and the image feature vector can also be normalized respectively. For example, in the above embodiments, the scores of all samples are usually restricted within a fixed range (such as between 0 and 1). However, in some cases, in order to better reflect the subtle differences between different samples, it is necessary to expand this score range. For example, the total score is set to 10 instead of 1. This can increase the score difference and make it easier for the model to identify which samples have a higher or lower matching degree. In this case, if the similarity needs to be calculated, the image feature vector can be normalized first, then the text feature vector can be normalized, then the text feature vector is multiplied by 10, and then the cosine similarity is calculated with the image feature vector, so as to allow the score to be greater than 1 or less than 1 due to text differences. This method allows highlighting the importance of text features and their impact on the final score through scaling.
[0124] In some alternative embodiments, since the technical solution of the present application is used for fine-grained text-image similarity evaluation and adopts similarity representation, the more consistent the text and image are, the higher the similarity. In the case of text-image consistency, the richer the text description information is, the higher the similarity is. Therefore, the mean-square error (MSE) can be used to calculate the loss data, which is specifically as follows:
[0125]
[0126] In the above formula, y i represents the score obtained by scoring the i-th text-image pair sample; y i p represents the predicted score of the i-th text-image pair sample output by the model; n represents the number of text-image pair samples in a batch during the training process.
[0127] It should be noted that in the above embodiments of the present application, mainly four template words are preset as an example for description. In other embodiments of the present application, other numbers of template words can also be used, and the type of template words can be selected according to actual needs. The embodiments of the present application do not limit this.
[0128] The advantage of the technical solution of the above embodiment of the present application is fine-grained text-level modeling, which significantly improves the discrimination of general graphic-text consistency. At the same time, it supports fine-grained graphic-text modeling in the business scenario of customized instance fine-tuning. When it is necessary to perform graphic-text alignment and graphic-text retrieval for this customized instance, better results can be obtained. At the same time, learnable representation parameters are added to the text end and a small number of new parameters are added to the image end. Since the added number of parameters is very small compared to the original CLIP parameters, adjustments can be made separately at the graphic and text ends to support fast business-targeted training, which brings an improvement in training efficiency and can improve the fine-grained recognition ability of the graphic-text model, and enhances the accuracy of graphic-text consistency discrimination.
[0129] The following introduces the device embodiments of the present application, which can be used to execute the data processing method in the above embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the embodiments of the above data processing method of the present application.
[0130] Figure 12 The block diagram of a data processing device according to an embodiment of the present application is shown. The data processing device can be applied to a computer device, which can be a terminal device, a server, or other devices.
[0131] Refer to Figure 12 As shown, a data processing device 1200 according to an embodiment of the present application includes: an acquisition unit 1202, a splitting unit 1204, a processing unit 1206, and a training unit 1208.
[0132] Among them, the acquisition unit 1202 is configured to acquire a graphic-text pair sample, and the graphic-text pair sample includes an image sample and a description text corresponding to the image sample; the splitting unit 1204 is configured to split the description text into description words corresponding to each of the template words according to the set template words, to obtain the description words of the description text; the processing unit 1206 is configured to adjust the description words of the description text to obtain an adjusted text that is different from the description text, and calculate a description score of the adjusted text according to the change situation of the adjusted text compared with the description text, and the description score is used to represent the correct rate of the adjusted text in describing the image sample; the training unit 1208 is configured to generate a new graphic-text pair sample according to the adjusted text and the image sample, and train a graphic-text model according to the new graphic-text pair sample and the description score of the adjusted text.
[0133] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: replace at least one description word of the description text with other description words of the same type to obtain the adjusted text.
[0134] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: calculate the weight proportion of the description words that are not replaced in the description words of the description text according to the description words replaced in the description text and the weights of the respective description words of the description text; and use the weight proportion as the description score of the adjusted text.
[0135] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: delete at least one description word of the description text to obtain the adjusted text.
[0136] In some embodiments of the present application, based on the foregoing solution, the set template words include subject words; wherein, deleting at least one description word of the description text includes: deleting at least one description word other than the description word corresponding to the subject word in the description text.
[0137] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: calculate the ratio between the weight of the deleted description word and the set scale factor according to the weight of the deleted description word in the description text and the set scale factor; and use the difference between the preset maximum description score and the ratio as the description score of the adjusted text.
[0138] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is further configured to: calculate the maximum description score of the first adjusted text obtained by replacing the description words in the description text and the description score of the second adjusted text obtained by deleting all the description words other than the specified description words according to the weights of the respective description words in the description text; and calculate the value of the scale factor under the constraint that the description score of the second adjusted text is greater than the maximum description score of the first adjusted text.
[0139] In some embodiments of the present application, based on the foregoing solution, the obtaining unit 1202 is configured to: obtain a candidate image-text pair, the candidate image-text pair including a candidate image and a candidate description text; split the candidate description text into description words corresponding to the respective template words according to the template words to obtain the description words of the candidate description text; and use the candidate image-text pair as the image-text pair sample if the content described by all the description words of the candidate description text appears in the candidate image.
[0140] In some embodiments of the present application, based on the foregoing solution, the obtaining unit 1202 is further configured to: extract the image feature vector of the candidate image and the text feature vectors of the respective description words of the candidate description text; calculate the similarity between the image feature vector and the text feature vectors of the respective description words, so as to determine whether the content described by the respective description words appears in the candidate image according to the similarity.
[0141] In some embodiments of the present application, based on the foregoing solution, the text-image model includes a text encoding module and an image encoding module, and the training unit 1208 is configured to: generate the input data of the text encoding module according to the template word and the word representations corresponding to the description words of the adjusted text; obtain the text feature vector output by the text encoding module according to the input data, and obtain the image feature vector output by the image encoding module for the image sample; calculate the similarity score between the text feature vector and the image feature vector, calculate the loss data according to the similarity score and the description score of the adjusted text; adjust the model parameters of the text-image model according to the loss data to obtain a trained text-image model, and the trained text-image model is used to evaluate the matching degree between the text and the image.
[0142] In some embodiments of the present application, based on the foregoing solution, the image encoding network includes a feature extraction network for extracting image visual features, a first attention pooling module and a second attention pooling module respectively connected to the feature extraction network, and a feature fusion module for fusing the output data of the first attention pooling module and the output data of the second attention pooling module; wherein, the training unit 1208 is configured to: adjust the parameters of the second attention pooling module, the parameters of the feature fusion module, and the word representation parameters corresponding to the description words according to the loss data.
[0143] In some embodiments of the present application, based on the foregoing solution, the training unit 1208 is configured to: divide the training data into multiple batches, and each batch contains at least one text-image pair sample; calculate the loss data of each batch according to the loss data corresponding to each text-image pair sample, and adjust the model parameters of the text-image model according to the loss data of each batch; if the model parameters of the text-image model are adjusted according to the loss data of the multiple batches, then complete one round of training process of the text-image model; if the training of the text-image model reaches the training end condition, then determine that the training of the text-image model is completed.
[0144] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is further configured to: obtain the model service parameters of each service scenario obtained by training the graphic-text model with graphic-text pair samples in different service scenarios; associate and store the scenario identification information of each service scenario with the model service parameters, so as to select the associated model service parameters according to the scenario identification information carried in the service request, and respond to the service request through the graphic-text model.
[0145] Figure 13 FIG. shows a schematic structural diagram of a computer system of a computer device suitable for implementing the embodiments of the present application. The computer device may be the computer device used to execute the data processing method in the foregoing embodiments.
[0146] It should be noted that Figure 13 The computer system 1300 of the computer device shown is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0147] As Figure 13 shown, the computer system 1300 may include a central processing unit (CPU) 1301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1302 or the program loaded from the storage part 1308 into the random access memory (RAM) 1303, such as executing the method described in the foregoing embodiments. In the RAM 1303, various programs and data required for system operation are also stored. The CPU 1301, ROM 1302, and RAM 1303 are connected to each other through a bus 1304. The input / output (I / O) interface 1305 is also connected to the bus 1304.
[0148] The following components can be connected to the I / O interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed so that a computer program read therefrom can be installed into the storage section 1308 as needed.
[0149] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, and the computer program is used to execute the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1309, and / or installed from the removable medium 1311. When the computer program is executed by a central processing unit (CPU) 1301, various functions defined in the system of the present application are executed.
[0150] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a computer program, and this computer program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and a computer program.
[0152] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation on the units themselves in some cases.
[0153] As another aspect, this application also provides a computer-readable medium, which may be included in the computer device described in the above embodiments; or it may exist separately without being assembled into the computer device. The above computer-readable medium carries one or more computer programs, and when the above one or more computer programs are executed by a computer device, the computer device implements the method described in the above embodiments.
[0154] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0155] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented in software or in the form of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computer device to execute the method according to the embodiments of this application.
[0156] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of this application. This application is intended to cover any variations, uses, or adaptations of this application, which follow the general principles of this application and include the known common knowledge or conventional technical means in the technical field not disclosed in this application.
[0157] It should be understood that this application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.
Claims
1. A data processing method, characterized in that Including: Obtaining a text-image pair sample, where the text-image pair sample includes an image sample and a description text corresponding to the image sample; Splitting the description text into description words corresponding to each of the template words according to the set template words to obtain the description words of the description text; Adjusting the description words of the description text to obtain an adjusted text that is different from the description text, and calculating a description score of the adjusted text according to the change situation of the adjusted text compared with the description text, where the description score is used to represent the correct rate of the adjusted text in describing the image sample; Generating a new text-image pair sample according to the adjusted text and the image sample, and training a text-image model according to the new text-image pair sample and the description score of the adjusted text.
2. The data processing method according to claim 1, wherein Adjusting the description words of the description text to obtain an adjusted text that is different from the description text, including: Replacing at least one description word of the description text with other description words of the same type to obtain the adjusted text.
3. The data processing method according to claim 2, wherein Calculating a description score of the adjusted text according to the change situation of the adjusted text compared with the description text, including: Calculating the weight proportion of the description words that are not replaced in the description words of the description text according to the description words replaced in the description text and the weights of the respective description words of the description text; Taking the weight proportion as the description score of the adjusted text.
4. The data processing method according to claim 1, wherein Adjusting the description words of the description text to obtain an adjusted text that is different from the description text, including: Deleting at least one description word of the description text to obtain the adjusted text.
5. The data processing method according to claim 4, wherein The set template words include subject words; Among them, deleting at least one description word of the description text includes: deleting at least one description word of the description text other than the description word corresponding to the subject word.
6. The data processing method according to claim 4, wherein Calculating a description score of the adjusted text according to the change situation of the adjusted text compared with the description text, including: Calculating the ratio between the weight of the description word deleted in the description text and a set scale factor; Taking the difference between the preset maximum description score and the ratio as the description score of the adjusted text.
7. The data processing method according to claim 6, wherein The data processing method further includes: Calculating the maximum description score of a first adjusted text obtained by replacing description words of the description text according to the weights of the respective description words of the description text, and the description score of a second adjusted text obtained by deleting all description words of the description text except specified description words; Calculating the value of the scale factor with the constraint that the description score of the second adjusted text is greater than the maximum description score of the first adjusted text.
8. The data processing method according to claim 1, wherein Obtaining a text-image pair sample, including: Obtaining candidate text-image pairs, where the candidate text-image pairs include candidate images and candidate description texts; Splitting the candidate description text into description words corresponding to each of the template words according to the template words to obtain the description words of the candidate description text. If all the content described by the description words of the candidate description text appears in the candidate image, then the candidate text-image pair is used as the text-image pair sample.
9. The data processing method according to claim 8, wherein The data processing method further includes: extracting an image feature vector of the candidate image and text feature vectors of the respective description words of the candidate description text; calculating a similarity between the image feature vector and the text feature vectors of the respective description words to determine whether the content described by the respective description words appears in the candidate image according to the similarity.
10. The data processing method according to claim 1, wherein The text-image model includes a text encoding module and an image encoding module. Training the text-image model according to the new text-image pair sample and the description score of the adjusted text includes: generating input data for the text encoding module according to the template word and the word representations corresponding to the description words of the adjusted text; obtaining a text feature vector output by the text encoding module according to the input data, and obtaining an image feature vector output by the image encoding module for the image sample; calculating a similarity score between the text feature vector and the image feature vector, and calculating loss data according to the similarity score and the description score of the adjusted text; adjusting model parameters of the text-image model according to the loss data to obtain a trained text-image model, and the trained text-image model is used to evaluate a matching degree between text and an image.
11. The data processing method according to claim 10, wherein The image encoding network includes a feature extraction network for extracting image visual features, a first attention pooling module and a second attention pooling module respectively connected to the feature extraction network, and a feature fusion module for fusing output data of the first attention pooling module and output data of the second attention pooling module; wherein, adjusting the model parameters of the text-image model according to the loss data includes: adjusting parameters of the second attention pooling module, parameters of the feature fusion module, and word representation parameters corresponding to the description words.
12. The data processing method according to claim 10, wherein Adjusting the model parameters of the text-image model according to the loss data includes: dividing training data into multiple batches, and each batch includes at least one text-image pair sample; calculating loss data of each batch according to the loss data corresponding to each text-image pair sample, and adjusting the model parameters of the text-image model according to the loss data of each batch; if the model parameters of the text-image model are adjusted according to the loss data of the multiple batches, then one round of training process of the text-image model is completed; if the training of the text-image model reaches a training end condition, it is determined that the training of the text-image model is completed.
13. The data processing method according to any one of claims 1 to 12, characterized in that The data processing method further includes: obtaining model service parameters of each service scenario obtained by training the text-image model respectively using text-image pair samples in different service scenarios; associatively storing the scenario identification information of each service scenario with the model service parameters, so as to select associated model service parameters according to the scenario identification information carried in a service request, and responding to the service request through the text-image model.
14. A data processing device, characterized in that, including: An acquisition unit configured to acquire text-image pair samples, where each text-image pair sample includes an image sample and a description text corresponding to the image sample; A splitting unit configured to split the description text into description words corresponding to respective template words according to the set template words, to obtain the description words of the description text; A processing unit configured to adjust the description words of the description text to obtain an adjusted text that is different from the description text, and calculate a description score of the adjusted text according to the change situation of the adjusted text compared with the description text, where the description score is used to represent the correct rate of the adjusted text in describing the image sample; A training unit configured to generate new text-image pair samples according to the adjusted text and the image sample, and train a text-image model according to the new text-image pair samples and the description score of the adjusted text.
15. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data processing method according to any one of claims 1 to 13.
16. A computer device, characterized in that, Comprising: One or more processors; A memory for storing one or more computer programs, and when the one or more computer programs are executed by the one or more processors, the computer device implements the data processing method according to any one of claims 1 to 13.
17. A computer program product, characterized in that, The computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of the computer device reads and executes the computer program from the computer-readable storage medium, so that the computer device executes the data processing method according to any one of claims 1 to 13.