Integrity identification model training method and device, electronic equipment and storage medium

By training the completeness recognition model, using the contrast loss function and the regression loss function, the problem of inaccurate assessment of subject integrity in the image is solved, and the fine portrayal and accurate evaluation of image integrity is achieved, and the quality of image synthesis and advertising is improved.

CN120031104APending Publication Date: 2025-05-23RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510087337.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to accurately evaluate and identify the integrity of the subject in the image, causing the image selection results to deviate from expectations, affecting the quality of image synthesis and advertising attractiveness.

Method used

By initializing the integrity recognition model, obtaining sample image data, and using the comparison loss function and regression loss function to train the model to generate the trained integrity recognition model. This model is able to refine the integrity of the subject in the image.

Benefits of technology

It realizes fine portrayal and accurate evaluation of image integrity, improves the accuracy of image feature extraction and the accuracy of subject integrity scores, and thus improves the quality of material screening and creative advertising synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031104A_ABST
    Figure CN120031104A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an integrity recognition model training method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining sample image data through initializing a first initial structure parameter in an integrity recognition model, inputting the sample image data into the integrity recognition model, and obtaining a first initial structure parameter in the integrity recognition model; determining a comparison loss value corresponding to the sample text feature and the sample image feature by using a comparison loss function so as to obtain a sample text feature, a sample image feature and a first sample main body integrity score, and determining a comparison loss value corresponding to the sample text feature and the sample image feature by using a comparison loss function; determining a first regression loss value corresponding to the first sample main body integrity score and the sample actual main body integrity score by adopting a first regression loss function, and adjusting the first initial structure parameter based on the comparison loss value and the first regression loss value to obtain a first target structure parameter, and generating a trained integrity identification model based on the first target structure parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for training an integrity recognition model. Background Art

[0002] The image resources on the Internet are rich and varied, but their quality varies, which has become a key factor affecting visual effects and application value. Among them, the core evaluation criterion of subject integrity is particularly prominent. However, the subject integrity problem of online images is the most significant. Subject integrity is not only related to the visual coherence and information transmission efficiency of the image, but also directly affects the accuracy of material screening and the synthesis quality of creative advertisements. Once the subject integrity is misjudged, it may cause the image selection results to deviate from expectations, and the image synthesis based on the image will have poor presentation effect, which will weaken the attractiveness of the advertisement and even affect the brand image. Summary of the invention

[0003] The main purpose of this application is to provide a completeness recognition model training method, device, electronic device and storage medium, aiming to provide a completeness recognition model training method that can finely identify the integrity of the subject in the image, and achieve fine characterization and accurate evaluation of the image integrity. The technical solution is as follows:

[0004] In a first aspect, an embodiment of the present application provides a completeness recognition model training method, comprising:

[0005] Initializing the first initial structural parameters in the integrity identification model;

[0006] Acquire sample image data; the sample image data includes a sample image, sample subject integrity description data and a sample actual subject integrity score;

[0007] Inputting the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score; the first sample body integrity score is confirmed based on the sample image features;

[0008] Using a contrast loss function to determine the contrast loss values ​​corresponding to the sample text features and the sample image features;

[0009] Determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score using a first regression loss function;

[0010] The first initial structural parameters are adjusted based on the contrast loss value and the first regression loss value to obtain first target structural parameters, and a trained integrity recognition model is generated based on the first target structural parameters.

[0011] In a second aspect, an embodiment of the present application provides a completeness recognition model training device, comprising:

[0012] An initialization unit, used for initializing a first initial structural parameter in the integrity recognition model;

[0013] An acquisition unit, configured to acquire sample image data; the sample image data includes a sample image, sample subject integrity description data, and a sample actual subject integrity score;

[0014] A prediction unit, configured to input the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score; the first sample body integrity score is obtained based on the sample image features;

[0015] A contrast loss determination unit, used to determine the contrast loss values ​​corresponding to the sample text features and the sample image features using a contrast loss function;

[0016] A first regression loss determining unit, configured to determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score by using a first regression loss function;

[0017] A training unit is used to adjust the first initial structure parameter based on the contrast loss value and the first regression loss value to obtain a first target structure parameter, and generate a trained integrity recognition model based on the first target structure parameter.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the above method when executed by the processor.

[0019] In a fourth aspect, an embodiment of the present application provides a storage medium having a computer program stored thereon, and the computer program implements the steps of the above method when executed by a processor.

[0020] In the embodiment of the present application, by initializing the first initial structure parameters in the integrity recognition model, sample image data is obtained. The sample image data is input into the integrity recognition model to obtain sample text features, sample image features, and the first sample main body integrity score. The first sample main body integrity score is determined based on the sample image features. A contrast loss function is used to determine the contrast loss value corresponding to the sample text features and the sample image features, and a first regression loss function is used to determine the first regression loss value corresponding to the first sample main body integrity score and the actual sample main body integrity score. The first initial structure parameters are adjusted based on the contrast loss value and the first regression loss value to obtain the first target structure parameters, and a trained integrity recognition model is generated based on the first target structure parameters. The contrast loss function is used to constrain the prediction accuracy of the sample image features of the sample image, so that the extracted sample image features are more consistent with the description of the sample main body integrity, thereby improving the extraction accuracy of the sample image features. Since the main body integrity score is inferred based on the sample image features, the accuracy of the main body integrity score can also be improved. Further, the first regression loss function is used to constrain the prediction accuracy of the main body integrity score, so that the trained integrity recognition model has better prediction accuracy of the main body integrity score, thereby realizing the fine description and accurate evaluation of the image integrity. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is an application schematic diagram of a method for training an integrity recognition model provided by an embodiment of the present application; Figure 2 It is a flowchart of a method for training an integrity recognition model provided by an embodiment of the present application; Figure 3 It is a flowchart of a method for training an integrity recognition model provided by an embodiment of the present application; Figure 4 It is an example schematic diagram of a method for training an integrity recognition model provided by an embodiment of the present application; Figure 5 It is a flowchart of a method for training an integrity recognition model provided by an embodiment of the present application; Figure 6 It is an example schematic diagram of a method for training an integrity recognition model provided by an embodiment of the present application; Figure 7It is a flowchart of a completeness recognition model training method provided in an embodiment of the present application; Figure 8 This is an example schematic diagram of a completeness recognition model training method provided in an embodiment of the present application; Fig. 9 It is a flowchart of a completeness recognition model training method provided in an embodiment of the present application; Fig.10 This is an example schematic diagram of a completeness recognition model training method provided in an embodiment of the present application; Fig.11 It is a structural schematic diagram of a completeness recognition model training device provided in an embodiment of the present application; Fig.12 It is a structural schematic diagram of a completeness recognition model training device provided in an embodiment of the present application; Fig.13 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0035] In the description of this specification, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise clearly specified and limited, "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood in specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are an "or" relationship.

[0036] Subject integrity usually refers to the degree of completeness of key objects in the image, such as food, retail products, etc. This indicator plays a vital role in areas such as material selection, creative synthesis and image review. High standards of subject integrity ensure that the objects displayed are presented intact, thus providing clear and attractive visual information to the audience.

[0037] In the related technology, it mainly relies on directly classifying images to judge their integrity, for example, dividing them into two categories: "complete" and "incomplete". However, such methods have significant limitations in defining the "incomplete" state, that is, their definition range is too broad and it is difficult to accurately capture the subtle differences in the degree of image integrity. This leads to ambiguity and uncertainty in the classification results to a certain extent, and cannot meet the high-precision requirements for image integrity assessment.

[0038] Based on the above problems, this specification provides a completeness recognition model training method. Figure 1, an application schematic diagram of a completeness recognition model training method is provided for an embodiment of the present application, by acquiring sample image data, inputting the sample image data into an initialized completeness recognition model, the sample image data including sample image and sample subject completeness description data, the completeness recognition model outputs sample text features, sample image features and a first sample subject completeness score predicted based on the sample image features according to the sample image data, the first sample subject completeness score is used to characterize the completeness of the subject in the sample image, a contrast loss function is used to determine a contrast loss value between the sample text features and the sample image features, a first regression loss function is used to determine a first regression loss value between the first sample subject completeness score and the sample actual subject completeness score, wherein the sample actual subject completeness score is the actual completeness score annotated for the sample image in the sample image data, the first regression loss value and the contrast loss value can be used to constrain the feature extraction capability and prediction capability of the completeness recognition model, so that the completeness recognition model can be trained to obtain a trained completeness recognition model.

[0039] The integrity recognition model training device in the embodiment of the present specification can be a terminal device such as a mobile phone, a computer, a tablet computer or a vehicle-mounted device, or it can be a module in the terminal device for implementing the integrity recognition model training method. The integrity recognition model training device can initialize the first initial structural parameters in the integrity recognition model, obtain sample image data, and input the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score. The first sample body integrity score is confirmed based on the sample image features, and a contrast loss function is used to determine the contrast loss values ​​corresponding to the sample text features and the sample image features. A first regression loss function is used to determine the first sample body integrity score and the first regression loss value corresponding to the sample actual body integrity score. The first initial structural parameters are adjusted based on the contrast loss value and the first regression loss value to obtain the first target structural parameters, and a trained integrity recognition model is generated based on the first target structural parameters.

[0040] The completeness recognition model training method provided in this specification is described in detail below in conjunction with specific embodiments.

[0041] See also Figure 2 , which is a flow chart of a completeness recognition model training method provided in the present application embodiment. Figure 2 As shown, the method in the embodiment of the present application may include the following steps S101 to S106.

[0042] S101, initializing the first initial structural parameter in the integrity recognition model;

[0043] In one embodiment, the integrity recognition model is used to recognize the integrity of the subject in an image. Before model training, the first initial structural parameters of the integrity recognition model are set to obtain an initialized integrity recognition model. Exemplarily, the first initial structural parameters include the network structure parameters of the model and the loss function, etc. The network structure of the model can refer to the existing convolutional neural network (CNN), which mainly consists of a convolutional layer, a pooling layer, and a fully connected layer. The loss function can be set according to actual needs, and this specification does not make specific limitations on it.

[0044] S102, Obtain sample image data;

[0045] In one embodiment, the sample image data refers to the sample data used for model training, usually including a batch of training sample data. The sample image data includes a sample image, sample subject integrity description data, and a sample actual subject integrity score. Among them, the sample image is preferably an image including a subject. For each sample image, the corresponding sample subject integrity description data and the subject integrity score can be marked. The sample subject integrity description data is text describing the integrity of the subject in the sample image. For example, it can be "image complete" / "image incomplete". The sample actual subject integrity score is the true subject integrity score of the sample image, and this score is a verified score, and the score range can be between 0 and 1. For example, it is "1 (complete), 0.5, 0.3, 0 (incomplete)".

[0046] S103, Input the sample image data into the integrity recognition model to obtain sample text features, sample image features, and a first sample subject integrity score;

[0047] In one embodiment, the sample image data is input into the initialized integrity recognition model, and the sample text features, sample image features, and a first sample subject integrity score are output through the integrity recognition model. The sample text features are the features extracted from the sample subject integrity description data in the sample image data. The sample image features are the features extracted from the sample image in the sample image data. The first sample subject integrity score is confirmed by the integrity recognition model based on the sample image features. The first sample subject integrity score is used to characterize the integrity degree of the subject in the sample image, and more specifically, the integrity degree is represented by a score value, so as to realize the fine description of the integrity degree of the subject.

[0048] S104, Use a contrast loss function to determine the contrast loss value corresponding to the sample text features and the sample image features;

[0049] In one embodiment, a contrast loss function is used to constrain the prediction accuracy of sample image features of sample images, so that the extracted sample image features are more consistent with the sample text features corresponding to the completeness description of the sample body, thereby improving the extraction accuracy of the sample image features, that is, enabling the model to output correct description information based on the input image. It should be noted that in contrastive learning, the model is required to map relevant text descriptions and image content to adjacent positions in space, while irrelevant ones are mapped to distant positions. In this way, the model learns how to distinguish between relevant and irrelevant text-image pairs. Exemplarily, the contrastive loss function can use InfoNCE commonly used in contrastive learning.

[0050] S105, using a first regression loss function to determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score;

[0051] In one embodiment, the first regression loss function is used to determine the difference between the first sample subject integrity score and the actual subject integrity score of the sample to obtain a first regression loss value. The score prediction accuracy is evaluated and trained by calculating the first regression loss value, so that the trained integrity recognition model has a better subject integrity score prediction accuracy, thereby achieving a detailed characterization and accurate evaluation of the image integrity. Exemplarily, the first regression loss function can use mean square error loss (MSE Loss).

[0052] S106: Adjust the first initial structural parameters based on the contrast loss value and the first regression loss value to obtain first target structural parameters, and generate a trained integrity recognition model based on the first target structural parameters.

[0053] In one embodiment, the first initial structural parameter of the initial subject integrity recognition model is adjusted according to the determined contrast loss value and the first regression loss value to obtain the first target structural parameter, which is the parameter after adjustment when the model performance meets the requirements. Exemplarily, the sum of the losses of the contrast loss value and the first regression loss value can be used as the optimization target of the integrity recognition model. If the sum does not converge, the step of inputting the sample image data into the integrity recognition model is continued to be executed to obtain the sample text features, the sample image features and the first sample subject integrity score, and the model is adjusted until the sum of the losses converges to obtain the final first target structural parameter, and the first target structural parameter is used as the parameter of the trained integrity recognition model.

[0054] In an embodiment of the present application, by initializing the first initial structural parameter in the integrity recognition model, obtaining sample image data, inputting the sample image data into the integrity recognition model, so as to obtain sample text features, sample image features and the first sample subject integrity score, the first sample subject integrity score is obtained based on the sample image feature confirmation, using the contrast loss function to determine the contrast loss value corresponding to the sample text feature and the sample image feature, using the first regression loss function to determine the first sample subject integrity score and the first regression loss value corresponding to the sample actual subject integrity score, adjusting the first initial structural parameter based on the contrast loss value and the first regression loss value to obtain the first target structural parameter, and generating a trained integrity recognition model based on the first target structural parameter. The integrity recognition model trained in this way can make a refined and continuous prediction of image integrity, and the trained subject integrity model is helpful to achieve accurate image selection and improve image quality in the process of material optimization and creative advertising synthesis, and ultimately serve higher levels of visual expression and communication needs.

[0055] See also Figure 3 , is a flow chart of a completeness recognition model training method provided in an embodiment of the present application. Figure 3 As shown, the method of the embodiment of the present application may include the following steps S201-S208.

[0056] S201, obtaining sample image data;

[0057] For details, please refer to step S102 in the above-mentioned embodiment of the specification, which will not be described in detail here.

[0058] S202, inputting the sample body completeness description data into the text encoder to obtain sample text features;

[0059] In one embodiment, the integrity recognition model includes a cutout module, a text encoder and an image encoder. Since the sample body integrity description data is text information, the sample body integrity description data is input into the text encoder, and the text encoder extracts the corresponding sample text features.

[0060] S203, inputting the sample image into the cutout module to obtain a sample main body mask image;

[0061] It is understandable that in traditional processing methods, the original image is often directly used as input, and a neural network is used for classification or regression tasks to solve the problem of subject integrity assessment. However, this method has a significant drawback: the background information is too rich and complex, which can easily interfere with the model's judgment, resulting in a significant reduction in the accuracy of the prediction results. In order to effectively eliminate the impact of background noise and increase the model's attention to the subject object, a cutout module is introduced into the integrity recognition model to accurately obtain the subject mask information of the image first. This step helps the subsequent neural network to focus more on the main part of the image, significantly reducing the difficulty of model prediction.

[0062] Specifically, the cutout module uses the MODNet model, which is composed of three key parts: semantic estimation, detail prediction, and semantic-detail fusion, which work together to achieve high-quality cutout effects. Among them, the Semantic Estimation module plays the role of Encoder, and its core function is to extract high-level semantic features from the complex original image (sample image). The output of this module is a semantic map with a dimension of [1, h, w]. The Detail Prediction module consists of a lightweight Encoder-Decoder structure, which focuses on capturing and reconstructing the foreground texture and detail features of each pixel in the image. The Semantic-Detail Fusion module fuses the outputs of the first two stages to finally obtain the sample main body mask image.

[0063] S204, inputting the sample image and the sample body mask image into the image encoder to obtain sample image features and a first sample body integrity score;

[0064] In one embodiment, after obtaining the sample subject mask image, the sample subject mask image and the sample image are input to the image encoder together, and the image encoder extracts features and processes the features of the body sample image and the first sample subject integrity score. By inputting the sample subject mask image to the image encoder together, the image encoder can better extract the subject features in the image according to the subject information provided by the sample subject mask image, so that the predicted first sample subject integrity score can correspond to the subject part in the image, thereby improving the prediction targeting.

[0065] S205, using a contrast loss function to determine contrast loss values ​​corresponding to the sample text feature and the sample image feature;

[0066] Specifically, after embedding text and images into a common semantic space through the text encoder and the image encoder, the purpose of the model is to make the representations of related text descriptions and image content in this space close to each other, while the representations of irrelevant ones are far away. For example, the sample text feature is the sample text feature vector, and the sample image feature is the sample image feature vector. The cosine similarity between the sample image feature vector and the sample text feature vector can be calculated, and then the contrast loss value can be calculated. The goal of the contrast loss function is to make the similarity of positive sample pairs higher and the similarity of negative sample pairs lower.

[0067] S206, using a first regression loss function to determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score;

[0068] For details, please refer to step S105 of the above-mentioned embodiment of the specification, which will not be elaborated here.

[0069] S207, adjusting the initial parameters of the image encoder based on the contrast loss value and the first regression loss value to obtain target parameters of the image encoder;

[0070] In one embodiment, the first initial structural parameters include initial parameters of the image encoder, and only the image encoder in the integrity recognition model is adjusted based on the contrast loss value and the first regression loss value, so that the image encoder achieves the required performance. The initial parameters of the image encoder are target parameters of the image encoder after the adjustment.

[0071] S208: Using the image encoder target parameters as the first target structure parameters, and generating a trained integrity recognition model based on the first target structure parameters.

[0072] In one embodiment, during the training process, the initial parameters of the cutout module and the initial parameters of the text encoder in the image integrity recognition model are fixed, the target parameters of the image encoder obtained after training are determined as the first target structure parameters, and the trained integrity recognition model is generated. The trained integrity recognition model includes the cutout module and the image encoder, and the text encoder is used to assist the image encoder training and is not needed in the actual reasoning process.

[0073] See also Figure 4 , Figure 4An example schematic diagram of a completeness recognition model training method provided in an embodiment of the present application, after acquiring sample image data, the sample image is input into a cutout module to obtain a sample body mask image, the sample image and the sample body mask image are input into an image encoder together to obtain sample image features, and a first sample body completeness score determined based on the sample image features, a first regression loss is calculated for the first sample body completeness score, the sample body completeness description data is input into a text encoder to obtain sample text features, and the contrast loss between the sample text features and the sample image features is calculated.

[0074] Optionally, in one embodiment, the method of the embodiment of the present application may include the following steps S2081-S2082:

[0075] S2081, acquiring a target image;

[0076] Specifically, after the integrity recognition model is trained, a target image can be acquired, where the target image is an image to be subjected to image quality detection.

[0077] S2082, inputting the target image into the trained integrity recognition model to obtain a target subject integrity score of the target image.

[0078] Specifically, the target image is input into a trained integrity recognition model, and the integrity recognition model can output a target subject integrity score of the target image. In a feasible implementation, the trained integrity recognition model includes a cutout module and an image encoder, and the target image is first input into the cutout module to obtain a target subject mask image, and the target image and the target subject mask image are input into the image encoder to obtain a target subject integrity score.

[0079] In an embodiment of the present application, by acquiring sample image data, inputting sample subject integrity description data into a text encoder to obtain sample text features, inputting sample images into a cutout module to obtain sample subject mask images, inputting sample images and sample subject mask images into an image encoder to obtain sample image features and a first sample subject integrity score, using a contrast loss function to determine the contrast loss value corresponding to the sample text features and the sample image features, using a first regression loss function to determine the first sample subject integrity score and the first regression loss value corresponding to the sample actual subject integrity score, adjusting the image encoder initial parameters based on the contrast loss value and the first regression loss value to obtain the image encoder target parameters, using the image encoder target parameters as the first target structural parameters, and generating a trained integrity recognition model based on the first target structural parameters. Adding a cutout model to the integrity recognition model and accurately obtaining the subject mask information of the image first helps the subsequent neural network to focus more on the subject part of the image, significantly reducing the difficulty of model prediction.

[0080] See also Figure 5 , is a flow chart of a completeness recognition model training method provided in an embodiment of the present application. Figure 5 As shown, the method of the embodiment of the present application may include the following steps S301-S308.

[0081] S301, obtaining sample image data;

[0082] In one embodiment, the completeness recognition model includes a cutout module, a text encoder and an image encoder. The image encoder includes an original image convolution layer, a subject image convolution layer, a cross attention module, an encoding layer and a first fully connected layer. A sample image, sample subject completeness description data of the sample image and a sample actual subject completeness score are obtained as sample image data.

[0083] S302, inputting the sample body completeness description data into the text encoder to obtain sample text features;

[0084] S303, inputting the sample image into the cutout module to obtain a sample main body mask image;

[0085] For details, please refer to steps S202-S203 of the above-mentioned embodiment of the specification, which will not be elaborated here.

[0086] S304, inputting the sample image into the original image convolution layer to obtain sample original image features;

[0087] Specifically, the image encoder includes an original image convolution layer for processing a sample image. The original image convolution layer is used to extract features from an unprocessed original sample image to obtain features of the sample original image. The original image convolution layer may include multiple convolution layers.

[0088] S305, inputting the sample subject mask image into the subject image convolution layer to obtain sample subject image features;

[0089] Specifically, the image encoder further includes a subject image convolution layer for processing the sample subject mask image, thereby extracting the sample subject image features. Correspondingly, the number of convolution layers in the subject image convolution layer can be the same as the number of original image convolution layers.

[0090] S306, inputting the sample original image features and the sample main image features into the cross attention module to obtain sample fusion image features;

[0091] Specifically, the image encoder includes a cross-attention module, which is used to fuse the sample original image features of the sample image and the sample subject image features of the sample subject mask image through a cross-attention mechanism to obtain fused sample fused image features.

[0092] S307, inputting the sample fusion image features into the encoding layer to obtain sample image features;

[0093] Specifically, the image encoder includes a coding layer. The coding layer is used to extract features of the sample fusion image features to obtain sample image features. Exemplarily, the coding layer can be selected from the coding layer in the transformer model, and there is no specific limitation. The sample image features are image features extracted by the model that may conform to the sample subject integrity description data, so that the subject integrity score can be better predicted based on the features.

[0094] S308: Input the sample image features into a fully connected layer to obtain a first sample body completeness score.

[0095] Specifically, the image encoder also includes a fully connected layer, which is used to map input features to output results, that is, output the first sample body integrity score according to the sample image features.

[0096] It is understandable that in order to effectively utilize the output result (subject mask) of the cutout module, so that the model can pay more attention to the subject area during the completeness prediction process, the Alpha-ClIP model is adopted, which includes a text encoder and an image encoder. In Alpha-CLIP, an additional AlphaConv layer is introduced in parallel with the RGB convolution layer, so that the CLIP image encoder can accept an additional alpha channel as input. The alpha channel input is set to a value in the range of [0,1], where 1 represents the foreground and 0 represents the background, which corresponds to the value of 0 / 1 at each pixel position in the sample subject mask image. When training the model, freeze the Text Encoder parameters in Alpha-ClIP.

[0097] See also Figure 6 , Figure 6 An example schematic diagram of a completeness recognition model training method provided in an embodiment of the present application, after obtaining the sample image data, the sample image is input into the matting module (Matte Module) to obtain the sample subject mask image, the sample image and the sample subject mask image are input into the image encoder (Image Encoder) together, the sample image features are extracted by the original image convolution layer (RGB conv) in the image encoder to obtain the sample original image features, the subject image convolution layer (Alpha conv) in the image encoder is used to obtain the sample subject image features, the sample original image features and the sample subject image features are input into the cross attention module (Attention Block) to obtain the sample fusion image features, the sample fusion image features are input into the encoding layer (Encoder), the sample image features are input into the fully connected layers (FC) to obtain the first sample subject completeness score, and the first regression loss function (MSE Loss) is calculated. The sample subject completeness description data is input into the text encoder (Text Encoder) to obtain the sample text features, and the contrast loss (Contrastive Loss) between the sample text features and the sample image features is calculated.

[0098] In an embodiment of the present application, by acquiring sample image data, inputting the sample subject completeness description data into a text encoder to obtain sample text features, inputting the sample image into a cutout module to obtain a sample subject mask image, inputting the sample image into an original image convolution layer to obtain sample original image features, inputting the sample subject mask image into a subject image convolution layer to obtain sample subject image features, inputting the sample original image features and the sample subject image features into a cross-attention module to obtain sample fusion image features, the model's understanding of the image can be enhanced, the sample fusion image features are input into an encoding layer to obtain sample image features, and the sample image features are input into a fully connected layer to obtain a first sample subject completeness score. Through this training method, the effect of processing and generating tasks can be improved.

[0099] See also Figure 7 , is a flow chart of a completeness recognition model training method provided in an embodiment of the present application. Figure 7 As shown, the method of the embodiment of the present application may include the following steps S401-S414.

[0100] S401, initializing the first initial structural parameter in the integrity recognition model;

[0101] S402, obtaining sample image data; the sample image data includes a sample image, sample subject integrity description data and a sample actual subject integrity score;

[0102] S403, inputting the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score; the first sample body integrity score is obtained based on the sample image features;

[0103] S404, using a contrast loss function to determine contrast loss values ​​corresponding to the sample text feature and the sample image feature;

[0104] S405, using a first regression loss function to determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score;

[0105] For details, please refer to steps S101-S105 in the above-mentioned embodiment of the specification, which will not be described in detail here.

[0106] S406, adjusting the initial parameters of the image encoder based on the contrast loss value and the first regression loss value to obtain target parameters of the image encoder;

[0107] In one embodiment, after calculating the loss function for the sample image features output by the image encoder in the integrity recognition model and the first sample subject integrity score using the subject integrity description data and the actual subject integrity score of the sample, the initial parameters of the image encoder are adjusted according to the loss function to obtain the target parameters of the image encoder.

[0108] S407, initializing the second initial structure parameters of the convolutional neural network;

[0109] In one embodiment, the trained image encoder needs to receive the subject mask information recognized by the cutout module each time to predict the subject integrity score, which increases the time-consuming and complexity of the overall process, and the cutout itself is a challenging task. Therefore, a lightweight convolutional neural network is used to distill the image encoder so that the model can directly infer the final integrity score from the native input image. Specifically, the structural parameters of the convolutional neural network are initialized to obtain the second initial structural parameters.

[0110] S408, generating a trained image encoder based on the image encoder target parameters;

[0111] In one embodiment, a trained image encoder is generated based on target parameters of an image encoder trained using sample image data, and the target parameters of the image encoder are not adjusted during the training process of the convolutional neural network.

[0112] S409, inputting the sample image into the convolutional neural network to obtain a second sample body integrity score;

[0113] In one embodiment, a sample image in the sample image data is input into a convolutional neural network to obtain a second sample subject integrity score predicted by the convolutional neural network. It is understandable that the sample image may be the same as a sample image previously used to train the image encoder, or may be an unused training sample, without any specific limitation.

[0114] S410, inputting the sample image into the cutout module to obtain a sample main body mask image;

[0115] In one embodiment, the same sample image is input into a cutout module to obtain a sample subject mask image.

[0116] S411, inputting the sample body encoded image and the sample image into the trained image encoder to obtain a third sample body integrity score;

[0117] In one embodiment, the sample body encoded image and the sample image are input into a trained image encoder to obtain a third sample body completeness score.

[0118] S412, using a second regression loss function to determine a second regression loss value corresponding to the second sample subject integrity score and the sample actual subject integrity score;

[0119] In one embodiment, a second regression loss value is calculated based on the second sample subject integrity score output by the convolutional neural network and the sample actual subject integrity score according to the second regression loss function. The second regression loss function can adopt a loss function such as L1Loss, and is not specifically limited. The gap between the output result of the convolutional neural network and the true label is evaluated according to the second regression loss value.

[0120] S413, using a distillation loss function to determine a distillation loss value corresponding to the second sample body integrity score and the third sample body integrity score;

[0121] In one embodiment, the trained image encoder can generate a prediction result that meets the required accuracy based on the sample image, and a learning goal of the convolutional neural network is to make its output close to the trained image encoder so that it can replace the image encoder. Specifically, the distillation loss value between the second sample body integrity score and the third sample body integrity score is determined by the distillation loss function.

[0122] S414: Adjust the second initial structure parameters based on the second regression loss value and the distillation loss value to obtain second target structure parameters, and generate a trained integrity recognition model based on the second target structure parameters.

[0123] Specifically, the second initial structural parameters of the convolutional neural network are adjusted according to the second regression loss value and the distillation loss value to obtain the adjusted second target structural parameters. It can be understood that whether the sum of the second regression loss value and the distillation loss value converges can be used as a basis for judging whether to stop training. If the sum of the second regression loss value and the distillation loss value does not converge, the convolutional neural network is adjusted, and the step of inputting the sample image into the convolutional neural network to obtain the second sample subject integrity score is executed until the sum of the second regression loss value and the distillation loss value converges to obtain the adjusted second target structural parameters, and the trained convolutional neural network is determined according to the second target structural parameters, and the trained convolutional neural network is used as the trained integrity recognition model. When the target image is obtained, the target image is input into the integrity recognition model, and the subject integrity score can be directly obtained without the need for cutout processing.

[0124] See also Figure 8, is an example schematic diagram of a completeness recognition model training method provided in an embodiment of the present application. After the image encoder is trained, a sample image is obtained, and the sample image is directly input into a convolutional neural network (CNN) to obtain a second sample image feature extracted by the convolutional neural network, and the second sample image feature is input into a fully connected layer to obtain a second sample subject completeness score; the sample image is input into a matting module (Matte Module) to obtain a sample subject mask image, and the sample image and the sample subject mask image are input into an image encoder (Image Encoder) together, and the sample image feature is extracted by the original image convolution layer (RGB conv) in the image encoder to obtain the sample original image feature, and the subject image convolution layer (Alpha conv) in the image encoder is used to obtain the sample subject image feature, and the sample original image feature and the sample subject image feature are input into a cross attention module (Attention Block) to obtain a sample fused image feature, and the sample fused image feature is input into a coding layer (Transfomer Encoder), and the sample image feature is input into a fully connected layer (Fully Connected Encoder). layers, FC) to obtain the third sample body integrity score, and calculate the distillation loss function (L1 Loss) between the second sample body integrity score and the third sample body integrity score. And calculate the second regression loss function between the second sample body integrity score and the sample actual body integrity score of the sample image. The convolutional neural network is trained by the second regression loss function and the distillation loss function.

[0125] In an embodiment of the present application, by initializing the first initial structural parameters in the integrity recognition model, sample image data is obtained, and the sample image data is input into the integrity recognition model to obtain sample text features, sample image features, and a first sample subject integrity score, a contrast loss function is used to determine the contrast loss values ​​corresponding to the sample text features and the sample image features, and a first regression loss function is used to determine the first sample subject integrity score and the first regression loss value corresponding to the sample actual subject integrity score. After adjusting the initial parameters of the image encoder based on the contrast loss value and the first regression loss value to obtain the target parameters of the image encoder, the second initial structural parameters of the convolutional neural network are initialized, and a trained image is generated based on the image encoder target parameters. encoder, input the sample image into the convolutional neural network to obtain the second sample subject integrity score, input the sample image into the cutout module to obtain the sample subject mask image, input the sample subject encoded image and the sample image into the trained image encoder to obtain the third sample subject integrity score, use the second regression loss function to determine the second regression loss value corresponding to the second sample subject integrity score and the sample actual subject integrity score, use the distillation loss function to determine the distillation loss value corresponding to the second sample subject integrity score and the third sample subject integrity score, adjust the second initial structure parameter based on the second regression loss value and the distillation loss value to obtain the second target structure parameter, and generate the trained integrity recognition model based on the second target structure parameter. The convolutional neural network is used to distill the trained image encoder, which reduces the model's dependence on the mask image during the reasoning process, significantly shortens the model's reasoning time, and thus greatly improves the performance.

[0126] See also Fig. 9 , is a flow chart of a completeness recognition model training method provided in an embodiment of the present application. Fig. 9 As shown, the method in the embodiment of the present application may include the following steps S501-S503.

[0127] S501, acquiring a first sample image including a complete subject and a first subject mask image corresponding to the first sample image;

[0128] In one embodiment, before training the completeness recognition model, training data needs to be prepared, specifically by obtaining a first sample image including a complete subject and a corresponding first subject mask image, and then performing image processing on the first sample image to obtain images with different subject completeness. The first subject mask image can be obtained by a cutout module.

[0129] S502, performing a first transformation process on the first sample image to obtain a second sample image including a complete subject;

[0130] Specifically, the first transformation process is an image processing method that does not destroy the integrity of the subject in the image, for example, random scaling, SD inpainting for background replacement, adding noise and other operations, each of which is controlled by a random factor to be implemented or not. By performing the first transformation process on the first sample image, a large number of second sample images can be obtained. The actual sample integrity score of the second sample image is the first value. The first value represents the integrity of the subject in the sample image.

[0131] S503: Perform a second transformation process on the first sample image and the first subject mask image to obtain a third sample image including the missing subject and a second subject mask image corresponding to the third sample image.

[0132] Specifically, the second transformation processing is an image processing method for changing the integrity of the subject in the image, and during the second transformation processing, the first sample image and the first subject mask image need to be processed together, for example, the first sample image and the first subject mask image are subjected to random cropping, template addition, radial transformation and other operations at the same time to obtain a third sample image. The subject in the third sample image has different degrees of missing. The actual subject integrity score of the sample of the third sample image is the second value, and the first value represents that the subject in the sample image is incomplete. More specifically, the second value is the ratio of the area occupied by the first subject mask image to the area occupied by the second subject mask image. The integrity of the subject in the third sample image can be quickly confirmed by the ratio of the transformed mask area to the original mask area.

[0133] See also Fig.10 , Fig.10An example schematic diagram of a completeness recognition model training method provided for an embodiment of the present application, wherein an image pool includes multiple first sample images, and the subject completeness score of the first sample image is 1 (label: 1), which indicates that the subject of the first sample image is complete. The first sample image label: 1 is subjected to first image processing, such as SD model local repainting, random scaling, and noise addition, and the first sample image and the first subject mask image are subjected to second image processing, such as random cropping, adding a template, and radioactive transformation. Finally, a second sample image is obtained without changing the subject mask and label. The second sample image includes a complete subject, so label: 1. There are different degrees of missing subjects in the third sample image, and the labels of different third sample images can be confirmed according to the ratio of the area occupied by the subject in the second subject mask image after transformation to the area occupied by the subject in the first subject mask image.

[0134] In the embodiment of the present application, by obtaining a first sample image including a complete subject and a first subject mask image corresponding to the first sample image, a first transformation process is performed on the first sample image to obtain a second sample image including a complete subject, and a second transformation process is performed on the first sample image and the first subject mask image to obtain a third sample image including a missing subject and a second subject mask image corresponding to the third sample image. By performing a variety of image processing on the sample images, the changes in subject integrity in real scenes can be simulated in all directions and from multiple angles, providing rich and high-quality training materials for the model.

[0135] The following will be combined with the attached Figure 11-Figure 12 , the integrity recognition model training device provided in the embodiment of the present application is introduced in detail. It should be noted that the attached Figure 11-Figure 12 The integrity recognition model training device in the embodiment of the present invention is used to execute the Figure 2-Figure 10 For the convenience of explanation, only the part related to the embodiment of the present application is shown. For the specific technical details not disclosed, please refer to this specification. Figure 2-Figure 10 The embodiment shown.

[0136] See also Fig.11, which shows a schematic diagram of the structure of a completeness recognition model training device provided by an exemplary embodiment of the present application. The completeness recognition model training device can be implemented as all or part of the device through software, hardware or a combination of both. The device 1 includes an initialization unit 11, an acquisition unit 12, a prediction unit 13, a contrast loss determination unit 14, a first regression loss determination unit 15 and a training unit 16.

[0137] An initialization unit 11, used for initializing a first initial structural parameter in the integrity recognition model;

[0138] The acquisition unit 12 is used to acquire sample image data; the sample image data includes a sample image, sample subject integrity description data and a sample actual subject integrity score;

[0139] The prediction unit 13 is used to input the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score; the first sample body integrity score is obtained based on the sample image features;

[0140] A contrast loss determination unit 14, configured to determine contrast loss values ​​corresponding to the sample text features and the sample image features using a contrast loss function;

[0141] A first regression loss determination unit 15, configured to determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score by using a first regression loss function;

[0142] The training unit 16 is used to adjust the first initial structure parameter based on the contrast loss value and the first regression loss value to obtain a first target structure parameter, and generate a trained integrity recognition model based on the first target structure parameter.

[0143] Optionally, the integrity recognition model includes a cutout module, a text encoder and an image encoder, and the prediction unit 13 is specifically used to input the sample body integrity description data into the text encoder to obtain sample text features;

[0144] Inputting the sample image into the cutout module to obtain a sample main body mask image;

[0145] The sample image and the sample body mask image are input into the image encoder to obtain sample image features and a first sample body completeness score.

[0146] Optionally, the image encoder includes an original image convolution layer, a main image convolution layer, a cross attention module, an encoding layer and a first fully connected layer, and the prediction unit 13 is specifically used to input the sample image into the original image convolution layer to obtain the sample original image features;

[0147] Inputting the sample subject mask image into the subject image convolution layer to obtain sample subject image features;

[0148] Inputting the sample original image features and the sample main image features into the cross attention module to obtain the sample fusion image features;

[0149] Inputting the sample fusion image features into the encoding layer to obtain sample image features;

[0150] The sample image features are input into the fully connected layer to obtain a first sample body completeness score.

[0151] Optionally, the first initial structural parameters include initial parameters of an image encoder, and the training unit 16 is specifically configured to adjust the initial parameters of the image encoder based on the contrast loss value and the first regression loss value to obtain target parameters of the image encoder;

[0152] The image encoder target parameters are used as the first target structure parameters, and a trained integrity recognition model is generated based on the first target structure parameters.

[0153] Optionally, the acquisition unit 12 is further configured to acquire a first sample image including a complete subject and a first subject mask image corresponding to the first sample image;

[0154] Performing a first transformation process on the first sample image to obtain a second sample image including a complete subject; the sample actual subject integrity score of the second sample image is a first value;

[0155] A second transformation process is performed on the first sample image and the first subject mask image to obtain a third sample image including a missing subject and a second subject mask image corresponding to the third sample image; the sample actual subject completeness score of the third sample image is a second value; and the second value is a ratio of an area occupied by the first subject mask image to an area occupied by the second subject mask image.

[0156] See also Fig.12 , which shows a schematic diagram of the structure of an integrity recognition model training device provided by an exemplary embodiment of the present application. The integrity recognition model training device can be implemented as all or part of the device through software, hardware or a combination of both. Fig.12As shown, the device 1 further comprises a distillation unit 17 and an application unit 18 .

[0157] The distillation unit 17 is used to initialize the second initial structure parameters of the convolutional neural network; generate a trained image encoder based on the image encoder target parameters; input the sample image into the convolutional neural network to obtain a second sample body integrity score; input the sample image into the cutout module to obtain a sample body mask image; input the sample body encoded image and the sample image into the trained image encoder to obtain a third sample body integrity score; use a second regression loss function to determine the second regression loss value corresponding to the second sample body integrity score and the sample actual body integrity score; use a distillation loss function to determine the distillation loss value corresponding to the second sample body integrity score and the third sample body integrity score; adjust the second initial structure parameters based on the second regression loss value and the distillation loss value to obtain second target structure parameters, and generate a trained integrity recognition model based on the second target structure parameters.

[0158] An application unit 18, configured to obtain a target image;

[0159] The target image is input into the trained integrity recognition model to obtain a target subject integrity score of the target image.

[0160] It should be noted that the integrity recognition model training device provided in the above embodiment only uses the division of the above functional modules as an example when executing the integrity recognition model training method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the integrity recognition model training device provided in the above embodiment and the integrity recognition model training method embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.

[0161] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0162] The present application also provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned Figure 2-Figure 10The method of the embodiment shown in the figure can be specifically executed by referring to Figure 2-Figure 10 The specific description of the illustrated embodiment will not be repeated here.

[0163] Please refer to Fig.13 , which shows a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. The electronic device in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.

[0164] The processor 110 may include one or more processing cores. The processor 110 uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions and processes data of the terminal 100 by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Optionally, the processor 110 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 110 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user pages, and applications; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 110, but may be implemented separately through a communication chip.

[0165] The memory 120 may include a Random Access Memory (RAM), or may also include a Read-Only Memory (ROM). Optionally, the memory 120 includes a Non-Transitory Computer-Readable Storage Medium. The memory 120 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc. The operating system can be an Android system, including a system developed based on the Android system in depth, an IOS system developed by Apple Inc., including a system developed based on the IOS system in depth, or other systems.

[0166] The memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party application programs run in the user space. In order to ensure that different third-party application programs can achieve better running effects, the operating system allocates corresponding system resources for different third-party application programs. However, there are also differences in the system resource requirements of different application scenarios in the same third-party application program. For example, in the local resource loading scenario, the third-party application program has a higher requirement for the disk reading speed; in the animation rendering scenario, the third-party application program has a higher requirement for the GPU performance. The operating system and the third-party application program are independent of each other, and the operating system often cannot timely perceive the current application scenario of the third-party application program, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application program.

[0167] In order to enable the operating system to distinguish the specific application scenarios of third-party application programs, it is necessary to establish data communication between the third-party application programs and the operating system, so that the operating system can obtain the current scenario information of the third-party application programs at any time, and then perform targeted system resource adaptation based on the current scenario.

[0168] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays. The touch screen display can be designed as a full screen, a curved screen or a special-shaped screen. The touch screen display can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of the present application.

[0169] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above drawings does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange the components differently. For example, the electronic device also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, a WiFi module, a power supply, and a Bluetooth module, which will not be described in detail here.

[0170] exist Fig.13 In the electronic device shown, the processor 110 may be used to call a computer application stored in the memory 120 and specifically perform the following operations:

[0171] Initializing the first initial structural parameters in the integrity identification model;

[0172] Acquire sample image data; the sample image data includes a sample image, sample subject integrity description data and a sample actual subject integrity score;

[0173] Inputting the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score; the first sample body integrity score is confirmed based on the sample image features;

[0174] Using a contrast loss function to determine the contrast loss values ​​corresponding to the sample text features and the sample image features;

[0175] Determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score using a first regression loss function;

[0176] The first initial structural parameters are adjusted based on the contrast loss value and the first regression loss value to obtain first target structural parameters, and a trained integrity recognition model is generated based on the first target structural parameters.

[0177] In one embodiment, the completeness recognition model includes a cutout module, a text encoder, and an image encoder layer. When the processor 110 inputs the sample image data into the completeness recognition model to obtain sample text features, sample image features, and a first sample subject completeness score, the following operations are specifically performed:

[0178] Inputting the sample body completeness description data into the text encoder to obtain sample text features;

[0179] Inputting the sample image into the cutout module to obtain a sample main body mask image;

[0180] The sample image and the sample body mask image are input into the image encoder to obtain sample image features and a first sample body completeness score.

[0181] In one embodiment, the image encoder includes an original image convolution layer, a subject image convolution layer, a cross attention module, an encoding layer, and a first fully connected layer. When the processor 110 inputs the sample image and the sample subject mask image into the image encoder to obtain the sample image feature and the first sample subject integrity, the following operations are specifically performed:

[0182] Inputting the sample image into the original image convolution layer to obtain sample original image features;

[0183] Inputting the sample subject mask image into the subject image convolution layer to obtain sample subject image features;

[0184] Inputting the sample original image features and the sample main image features into the cross attention module to obtain the sample fusion image features;

[0185] Inputting the sample fusion image features into the encoding layer to obtain sample image features;

[0186] The sample image features are input into the fully connected layer to obtain a first sample body completeness score.

[0187] In one embodiment, the first initial structural parameters include initial parameters of an image encoder, and the processor 110 adjusts the first initial structural parameters based on the contrast loss value and the first regression loss value to obtain first target structural parameters, and generates the trained integrity recognition model based on the first target structural parameters, specifically performing the following operations:

[0188] Adjusting the initial parameters of the image encoder based on the contrast loss value and the first regression loss value to obtain target parameters of the image encoder;

[0189] The image encoder target parameters are used as the first target structure parameters, and a trained integrity recognition model is generated based on the first target structure parameters.

[0190] In one embodiment, after adjusting the initial parameters of the image encoder based on the contrast loss value and the first regression loss value to obtain the target parameters of the image encoder, the processor 110 further performs the following operations:

[0191] Initialize the second initial structure parameters of the convolutional neural network;

[0192] generating a trained image encoder based on the image encoder target parameters;

[0193] Inputting the sample image into the convolutional neural network to obtain a second sample subject integrity score;

[0194] Inputting the sample image into the cutout module to obtain a sample main body mask image;

[0195] Inputting the sample body encoded image and the sample image into the trained image encoder to obtain a third sample body completeness score;

[0196] Determine a second regression loss value corresponding to the second sample subject integrity score and the sample actual subject integrity score using a second regression loss function;

[0197] Determine the distillation loss value corresponding to the second sample body integrity score and the third sample body integrity score using a distillation loss function;

[0198] The second initial structure parameters are adjusted based on the second regression loss value and the distillation loss value to obtain second target structure parameters, and a trained integrity recognition model is generated based on the second target structure parameters.

[0199] In one embodiment, the processor 110 further performs the following operations before acquiring the sample image data:

[0200] Acquire a first sample image including a complete subject and a first subject mask image corresponding to the first sample image;

[0201] Performing a first transformation process on the first sample image to obtain a second sample image including a complete subject; the sample actual subject integrity score of the second sample image is a first value;

[0202] A second transformation process is performed on the first sample image and the first subject mask image to obtain a third sample image including a missing subject and a second subject mask image corresponding to the third sample image; the sample actual subject completeness score of the third sample image is a second value; and the second value is a ratio of an area occupied by the first subject mask image to an area occupied by the second subject mask image.

[0203] In one embodiment, the processor 110 further performs the following operations:

[0204] Get the target image;

[0205] The target image is input into the trained integrity recognition model to obtain a target subject integrity score of the target image.

[0206] In the embodiment of the present application, by initializing the first initial structural parameter in the integrity recognition model, sample image data is obtained, and the sample image data is input into the integrity recognition model to obtain sample text features, sample image features and a first sample subject integrity score, the first sample subject integrity score is confirmed based on the sample image features, the contrast loss function is used to determine the contrast loss value corresponding to the sample text features and the sample image features, the first regression loss function is used to determine the first sample subject integrity score and the first regression loss value corresponding to the sample actual subject integrity score, the first initial structural parameter is adjusted based on the contrast loss value and the first regression loss value to obtain the first target structural parameter, and the trained integrity recognition model is generated based on the first target structural parameter. The integrity recognition model trained in this way can make a refined and continuous prediction of image integrity, and the trained subject integrity model is helpful to achieve accurate image selection and improve image quality in the process of material optimization and creative advertising synthesis, and ultimately serve higher levels of visual expression and communication needs.

[0207] Further, by acquiring sample image data, the sample subject integrity description data is input into a text encoder to obtain sample text features, the sample image is input into a cutout module to obtain a sample subject mask image, the sample image and the sample subject mask image are input into an image encoder to obtain sample image features and a first sample subject integrity score, a contrast loss function is used to determine the contrast loss value corresponding to the sample text features and the sample image features, a first regression loss function is used to determine the first sample subject integrity score and the first regression loss value corresponding to the sample actual subject integrity score, the image encoder initial parameters are adjusted based on the contrast loss value and the first regression loss value to obtain the image encoder target parameters, the image encoder target parameters are used as the first target structure parameters, and the trained integrity recognition model is generated based on the first target structure parameters. Adding a cutout model to the integrity recognition model and first accurately obtaining the subject mask information of the image helps the subsequent neural network to focus more on the subject part of the image, significantly reducing the difficulty of model prediction.

[0208] Furthermore, by obtaining sample image data, inputting the sample subject completeness description data into the text encoder to obtain sample text features, inputting the sample image into the cutout module to obtain the sample subject mask image, inputting the sample image into the original image convolution layer to obtain the sample original image features, inputting the sample subject mask image into the subject image convolution layer to obtain the sample subject image features, inputting the sample original image features and the sample subject image features into the cross attention module to obtain the sample fusion image features, the model's understanding of the image can be enhanced, the sample fusion image features are input into the encoding layer to obtain the sample image features, and the sample image features are input into the fully connected layer to obtain the first sample subject completeness score. Through this training method, the effect of processing and generating tasks can be improved.

[0209] Furthermore, by initializing the first initial structural parameters in the integrity recognition model, sample image data is obtained, and the sample image data is input into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score, a contrast loss function is used to determine the contrast loss values ​​corresponding to the sample text features and the sample image features, and a first regression loss function is used to determine the first sample body integrity score and the first regression loss value corresponding to the sample actual body integrity score. After adjusting the initial parameters of the image encoder based on the contrast loss value and the first regression loss value to obtain the target parameters of the image encoder, the second initial structural parameters of the convolutional neural network are initialized, and a trained image encoder is generated based on the target parameters of the image encoder. The sample image is input into the convolutional neural network to obtain the second sample subject integrity score, the sample image is input into the cutout module to obtain the sample subject mask image, the sample subject encoding image and the sample image are input into the trained image encoder to obtain the third sample subject integrity score, the second regression loss function is used to determine the second regression loss value corresponding to the second sample subject integrity score and the sample actual subject integrity score, the distillation loss function is used to determine the distillation loss value corresponding to the second sample subject integrity score and the third sample subject integrity score, the second initial structure parameter is adjusted based on the second regression loss value and the distillation loss value to obtain the second target structure parameter, and the trained integrity recognition model is generated based on the second target structure parameter. The convolutional neural network is used to distill the trained image encoder, which reduces the model's dependence on the mask image during the reasoning process, significantly shortens the model's reasoning time, and thus greatly improves the performance.

[0210] Furthermore, by obtaining a first sample image including a complete subject and a first subject mask image corresponding to the first sample image, a first transformation process is performed on the first sample image to obtain a second sample image including a complete subject, and a second transformation process is performed on the first sample image and the first subject mask image to obtain a third sample image including a missing subject and a second subject mask image corresponding to the third sample image. By performing various image processing on the sample images, the changes in subject integrity in real scenes can be simulated in all directions and from multiple angles, providing rich and high-quality training materials for the model.

[0211] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0212] The above disclosure is only the preferred embodiment of this specification, which certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.

Claims

1. A completeness recognition model training method, characterized in that: include: Initializing the first initial structural parameters in the integrity identification model; Get sample image data; The sample image data includes a sample image, sample subject integrity description data and a sample actual subject integrity score; Inputting the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score; the first sample body integrity score is confirmed based on the sample image features; Using a contrast loss function to determine the contrast loss values ​​corresponding to the sample text features and the sample image features; Determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score using a first regression loss function; The first initial structural parameters are adjusted based on the contrast loss value and the first regression loss value to obtain first target structural parameters, and a trained integrity recognition model is generated based on the first target structural parameters.

2. The method according to claim 1, characterized in that The completeness recognition model includes a cutout module, a text encoder and an image encoder; The step of inputting the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score includes: Inputting the sample body completeness description data into the text encoder to obtain sample text features; Inputting the sample image into the cutout module to obtain a sample main body mask image; The sample image and the sample body mask image are input into the image encoder to obtain sample image features and a first sample body completeness score.

3. The method according to claim 2, characterized in that The image encoder includes an original image convolution layer, a subject image convolution layer, a cross attention module, an encoding layer and a first fully connected layer; The step of inputting the sample image and the sample body mask image into the image encoder to obtain a sample image feature and a first sample body integrity score includes: Inputting the sample image into the original image convolution layer to obtain sample original image features; Inputting the sample subject mask image into the subject image convolution layer to obtain sample subject image features; Inputting the sample original image features and the sample main image features into the cross attention module to obtain the sample fusion image features; Inputting the sample fusion image features into the encoding layer to obtain sample image features; The sample image features are input into the fully connected layer to obtain a first sample body completeness score.

4. The method according to claim 1, characterized in that The first initial structure parameters include image encoder initial parameters; The first initial structural parameter is adjusted based on the contrast loss value and the first regression loss value to obtain a first target structural parameter, and the trained integrity recognition model is generated based on the first target structural parameter, including: Adjusting the initial parameters of the image encoder based on the contrast loss value and the first regression loss value to obtain target parameters of the image encoder; The image encoder target parameters are used as the first target structure parameters, and a trained integrity recognition model is generated based on the first target structure parameters.

5. The method according to claim 4, characterized in that After the initial parameters of the image encoder are adjusted based on the contrast loss value and the first regression loss value to obtain the target parameters of the image encoder, the method further includes: Initialize the second initial structure parameters of the convolutional neural network; generating a trained image encoder based on the image encoder target parameters; Inputting the sample image into the convolutional neural network to obtain a second sample subject integrity score; Inputting the sample image into the cutout module to obtain a sample main body mask image; Inputting the sample body encoded image and the sample image into the trained image encoder to obtain a third sample body completeness score; Determine a second regression loss value corresponding to the second sample subject integrity score and the sample actual subject integrity score using a second regression loss function; Determine the distillation loss value corresponding to the second sample body integrity score and the third sample body integrity score using a distillation loss function; The second initial structure parameters are adjusted based on the second regression loss value and the distillation loss value to obtain second target structure parameters, and a trained integrity recognition model is generated based on the second target structure parameters.

6. The method according to claim 1, characterized in that Before acquiring the sample image data, the method further includes: Acquire a first sample image including a complete subject and a first subject mask image corresponding to the first sample image; Performing a first transformation process on the first sample image to obtain a second sample image including a complete subject; the sample actual subject integrity score of the second sample image is a first value; A second transformation process is performed on the first sample image and the first subject mask image to obtain a third sample image including a missing subject and a second subject mask image corresponding to the third sample image; the sample actual subject completeness score of the third sample image is a second value; and the second value is a ratio of an area occupied by the first subject mask image to an area occupied by the second subject mask image.

7. The method according to claim 1, characterized in that The method further comprises: Get the target image; The target image is input into the trained integrity recognition model to obtain a target subject integrity score of the target image.

8. A completeness recognition model training device, characterized in that: The device comprises: An initialization unit, used for initializing a first initial structural parameter in the integrity recognition model; An acquisition unit, configured to acquire sample image data; the sample image data includes a sample image, sample subject integrity description data, and a sample actual subject integrity score; A prediction unit, configured to input the sample image data into the integrity recognition model to obtain sample text features, sample image features and a first sample body integrity score; the first sample body integrity score is obtained based on the sample image features; A contrast loss determination unit, used to determine the contrast loss values ​​corresponding to the sample text features and the sample image features using a contrast loss function; A first regression loss determining unit, configured to determine a first regression loss value corresponding to the first sample subject integrity score and the sample actual subject integrity score by using a first regression loss function; A training unit is used to adjust the first initial structure parameter based on the contrast loss value and the first regression loss value to obtain a first target structure parameter, and generate a trained integrity recognition model based on the first target structure parameter.

9. An electronic device, characterized in that: include: Processor and memory; The memory stores a computer program, wherein the computer program is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 7.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.