A data labeling method, device, apparatus, and readable storage medium

By automatically annotating a pre-trained aesthetic model and combining it with sample annotation scores from client feedback, the model selects the best composition boxes and generates a training dataset for retraining. This solves the problems of high cost and low accuracy of manual annotation, and achieves efficient and accurate aesthetic model training.

CN118379540BActive Publication Date: 2026-05-29SHENZHEN LINKRIC TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN LINKRIC TECH CO LTD
Filing Date
2024-03-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing aesthetic model training and fine-tuning require a large amount of manual intervention, resulting in high labor costs and significant deviations between the annotation results and the ideal situation, making it difficult to achieve optimal model performance.

Method used

The pre-trained aesthetic model automatically labels the composition scores of candidate boxes to obtain high-scoring composition images. It then uses the sample labeling scores fed back by the client to select the best ones. Combined with a small amount of manual labeling, it generates a labeled training set to retrain the aesthetic model, reducing manual costs and improving labeling accuracy.

Benefits of technology

This approach reduces manual labor costs while improving the accuracy of automatic annotation and composition capabilities of aesthetic models, generating high-quality training datasets and enhancing the model's aesthetic scoring and composition abilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118379540B_ABST
    Figure CN118379540B_ABST
Patent Text Reader

Abstract

The application discloses a data labeling method and device, equipment and a readable storage medium, which can be applied to the field of artificial intelligence technology, and automatically labels the composition frame score of each candidate frame based on a pre-trained aesthetic large model, performs first labeling to obtain a high-score composition image, performs second labeling based on the sample labeling score of the high-score composition image fed back by a client, obtains multiple optimal composition frames and corresponding labeling score results, the first labeling is realized based on automatic labeling, and the second optimal selection is realized based on a small amount of manual labeling, that is, it is not necessary to score each candidate frame manually, the pre-trained aesthetic large model and a small amount of manual intervention are used to obtain a labeling training set for retraining the aesthetic large model, the aesthetic large model is fine-tuned based on the labeling training set, the artificial cost of constructing the aesthetic large model is reduced, and the accuracy of automatic labeling of the aesthetic large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a data annotation method, apparatus, device, and readable storage medium. Background Technology

[0002] Automatic aesthetic data annotation strategies can use aesthetic models to perform aesthetic scoring and aesthetic composition annotation on unlabeled images without human intervention. Existing aesthetic models are typically trained and fine-tuned using publicly available datasets, such as aesthetic scoring datasets AVA, PARA, AADB, and TAD66K, and aesthetic composition datasets CPC, GAI CD, and FCDB.

[0003] Existing annotation methods for aesthetic composition datasets require manual annotation of optimal composition boxes and aesthetic scoring of each composition box. Therefore, when using public datasets to train and fine-tune aesthetic models, a large amount of manual intervention is required, resulting in extremely high labor costs. Furthermore, the high subjectivity and high degree of freedom of the aesthetic scoring task mean that the annotation results may deviate significantly from the ideal situation when the manual workload is high, making it difficult for aesthetic models to achieve optimal model performance.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a data annotation method, apparatus, device, and readable storage medium, as follows:

[0006] A data annotation method, comprising:

[0007] Obtain a set of candidate bounding boxes for each sample image in the sample set, wherein the set of candidate bounding boxes includes multiple candidate bounding boxes;

[0008] For each of the sample images, based on the bounding box scores of each of the candidate boxes output by the pre-trained aesthetic large model, multiple high-scoring bounding boxes of the sample image are obtained.

[0009] Multiple high-resolution composition images of the sample image and annotation instructions are sent to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-resolution composition images.

[0010] For each labeled account, the sample composition image corresponding to the labeled account is selected based on the sample labeling scores of multiple high-resolution composition images of the sample image;

[0011] Based on the sample composition images corresponding to each labeled account, the preferred composition frame of the sample image and the corresponding labeling score result are obtained;

[0012] The pre-trained aesthetic model is retrained using a labeled training set to obtain a trained aesthetic model. The labeled training set includes multiple retraining sample data, which includes corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

[0013] Optionally, the pre-trained aesthetic model can be retrained using a labeled training set to obtain a trained aesthetic model, including:

[0014] The sample images and candidate box sets in each of the retraining sample data are input into the pre-trained aesthetic model. The preferred composition boxes are used as the first annotation data, and the annotation scores are used as the second annotation data. The parameters of the pre-trained aesthetic model are fine-tuned until the preset retraining completion conditions are met, and the trained aesthetic model is obtained.

[0015] Optionally, data annotation methods also include:

[0016] Obtain the candidate bounding box set of the image to be labeled;

[0017] The image to be labeled and the set of candidate bounding boxes of the image to be labeled are input into the trained aesthetic model to obtain the labeled data of the image to be labeled output by the trained aesthetic model. The labeled data of the image to be labeled includes the preferred bounding box of the image to be labeled and the corresponding labeled score result.

[0018] The labeled data of each of the images to be labeled is collected to obtain the training data.

[0019] Optionally, based on the bounding box scores of each candidate box output by the pre-trained aesthetic large model, multiple high-scoring composition images of the sample image are obtained, including:

[0020] The sample image and each candidate box in the candidate box set are input into the pre-trained aesthetic large model to obtain the composition box score of each candidate box output by the pre-trained aesthetic large model.

[0021] The candidate boxes are sorted from high to low according to their frame scores, and the top n candidate boxes are selected as the high-scoring frames of the sample image, where n is the number of high-scoring frames pre-configured.

[0022] Based on each of the high-resolution composition frames, the sample image is cropped to obtain n high-resolution composition images of the sample image.

[0023] Optionally, the target image includes the sample image and the image to be labeled, and obtaining the candidate bounding box set of the target image includes:

[0024] Obtain a first type of candidate box in the target image. The first type of candidate box is a rectangular box with an aspect ratio equal to the first target ratio and an area ratio equal to the first target scale.

[0025] Obtain a second type of candidate box for the target image. The second type of candidate box includes a rectangular box whose target key points coincide with the target key points of the target image and whose aspect ratio is equal to the second target ratio value.

[0026] Obtain a third type of candidate box for the target image, the third type of candidate box including rectangular boxes whose composition satisfies the target composition rules and whose area ratio is equal to the third target scale;

[0027] Redundant candidate boxes in the first type, the second type, and the third type are deleted to generate a set of candidate boxes for the target image.

[0028] Optionally, based on the sample annotation scores of multiple high-resolution composition images of the sample image, the sample composition image corresponding to the annotation account is selected, including:

[0029] Based on the sample annotation scores of multiple high-resolution composition images of the sample image, the top m high-resolution composition images with the highest sample annotation scores are selected as candidate sample composition images of the sample image, where m is the pre-configured number of sample compositions.

[0030] Outlier test results are obtained by performing an outlier test on the labeled account. The outlier test results of the labeled account include the outlier test results of the labeled account for the sample image and / or the outlier test results of the labeled account for the sample set. The outlier test results of the labeled account for the sample image are used to indicate the degree of deviation between the labeled account's labeling results for the sample image and the labeling results of other labeled accounts for the sample image. The outlier test results of the labeled account for the sample set are used to indicate the degree of deviation between the labeled account's labeling results for the sample set and the labeling results of other labeled accounts for the sample set. The sample set includes multiple sample images.

[0031] If the outlier test result of the annotation account for the sample image indicates that the deviation of the annotation result of the annotation account for the sample image from the annotation result of other annotation accounts for the sample image is not higher than a preset first deviation threshold, and the outlier test result of the annotation account for the sample set indicates that the deviation of the annotation result of the annotation account for the sample set from the annotation result of other annotation accounts for the sample set is not higher than a preset second deviation threshold, then the candidate sample composition image of the sample image is used as the sample composition image of the sample image.

[0032] Optionally, based on the sample composition images corresponding to each labeled account, obtaining the preferred composition frame of the sample image and the corresponding labeling score results includes:

[0033] For each of the sample composition images of the sample image, the average of the annotation scores of each sample composition image is calculated as the annotation score result of the sample composition image;

[0034] Obtain the top m sample composition images sorted from largest to smallest by labeled score results, and obtain m optimal composition images;

[0035] Obtain candidate bounding boxes corresponding to each of the preferred composition images as preferred composition bounding boxes;

[0036] The annotation score of each of the preferred composition images is obtained as the annotation score of the corresponding preferred composition frame.

[0037] A data annotation device, comprising:

[0038] A candidate box generation unit is used to obtain a set of candidate boxes for each sample image in the sample set, wherein the set of candidate boxes includes multiple candidate boxes;

[0039] The sample automatic scoring unit is used to obtain multiple high-scoring composition images of each sample image based on the composition box scores of each candidate box output by the pre-trained aesthetic large model.

[0040] The sample manual scoring unit is used to send multiple high-scoring composition images of the sample image and annotation instructions to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-scoring composition images.

[0041] The sample composition selection unit is used to select the sample composition image corresponding to the labeled account for each labeled account based on the sample labeling score of multiple high-scoring composition images of the sample image;

[0042] The preferred composition selection unit is used to obtain the preferred composition frame of the sample image and the corresponding annotation score result based on the sample composition image corresponding to each of the labeled accounts;

[0043] The model retraining unit is used to retrain the pre-trained aesthetic large model using a labeled training set to obtain a trained aesthetic large model. The labeled training set includes multiple retraining sample data, which includes corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

[0044] A data annotation device includes: a memory and a processor;

[0045] The memory is used to store programs;

[0046] The processor is used to execute the program and implement the various steps of the data annotation method.

[0047] A readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the various steps of a data annotation method.

[0048] A computer program product includes a computer program that, when executed by a processor, implements the various steps of a data annotation method.

[0049] As can be seen from the above technical solutions, the data annotation method, apparatus, device, and readable storage medium provided in this application embodiment obtain a candidate box set for each sample image in a sample set, the candidate box set including multiple candidate boxes. For each sample image, based on the composition box scores of each candidate box output by the pre-trained aesthetic large model, multiple high-resolution composition images of the sample image are obtained. The multiple high-resolution composition images of the sample image and annotation instructions are sent to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-resolution composition images. For each annotation account, based on the sample annotation scores of the multiple high-resolution composition images of the sample image, the sample composition image corresponding to the annotation account is selected. Based on the sample composition images corresponding to each annotation account, the preferred composition box of the sample image and the corresponding annotation score result are obtained. The pre-trained aesthetic large model is retrained using the annotation training set to obtain a trained aesthetic large model. The annotation training set includes multiple retraining sample data, and the retraining sample data includes the corresponding sample image, candidate box set, preferred composition box, and annotation score result. This method automatically labels the bounding box scores of each candidate box based on a pre-trained aesthetic model, performing the first labeling to obtain high-scoring composition images. A second labeling is then performed based on the sample labeling scores of the high-scoring composition images fed back by the client, obtaining multiple preferred composition boxes and their corresponding labeling scores. The first labeling is achieved automatically, while the second selection is achieved with minimal manual labeling. This eliminates the need for manual scoring of each candidate box. By utilizing the pre-trained aesthetic model and minimal manual intervention, a labeled training set is obtained for retraining the aesthetic model. Fine-tuning of the aesthetic model is then achieved based on this labeled training set, reducing the manual cost of building the aesthetic model and improving the accuracy of automatic labeling. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart illustrating a specific implementation of a data annotation method provided in this application embodiment;

[0052] Figure 2 A schematic diagram illustrating a candidate box generation method provided in an embodiment of this application;

[0053] Figure 3 A flowchart illustrating a data annotation method provided in an embodiment of this application;

[0054] Figure 4 This is a schematic diagram of the structure of a data annotation device provided in an embodiment of this application;

[0055] Figure 5 This is a schematic diagram of the structure of a data annotation device provided in an embodiment of this application. Detailed Implementation

[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0057] This application provides a data annotation method applicable to, but not limited to, scenarios where a large-scale aesthetic model is used to automatically annotate training images (i.e., images to be annotated) to obtain training data. Specifically, the large-scale aesthetic model is retrained based on sample images to achieve fine-tuning of the model and improve the accuracy of automatic annotation. Optionally, this application is applicable to a data annotation system, which includes a server and multiple annotation clients. The multiple annotation clients are pre-bound to annotation accounts, and each annotation account is a unique identifier registered by the annotator. Specifically, the server is used to construct the large-scale aesthetic model and generate training data.

[0058] Figure 1 The specific implementation flow of a data annotation method provided in the embodiments of this application is as follows: Figure 1 As shown, this method is applied to the server side and specifically includes:

[0059] S101. Obtain the first type of candidate bounding box of the sample image.

[0060] In this embodiment, the sample image is any original image in the sample set, which includes multiple original images. The first type of candidate box is a rectangle with an aspect ratio equal to a first target ratio value and an area ratio equal to a first target scale. The first target ratio value is any preset ratio value in the first ratio value set, and the first target scale is any preset scale in the first scale set. Optionally, the first ratio value set is {1, 4 / 3, 3 / 4, 16 / 9, 9 / 16}, and the first scale set is {0.5, 0.4, 0.3}. Therefore, the first type of candidate box includes all rectangles with an aspect ratio equal to any one of {1, 4 / 3, 3 / 4, 16 / 9, 9 / 16} and a scale equal to any one of {0.5, 0.4, 0.3}. It should be noted that the aspect ratio refers to the ratio of the width to the height of the first type of candidate box, and the area ratio refers to the ratio of the area of ​​the first type of candidate box to the area of ​​the sample image. The first ratio value set and the first scale set are pre-configured based on the application scenario.

[0061] Optionally, the specific methods for obtaining the first type of candidate boxes include:

[0062] The first target ratio value is obtained from the first set of ratio values, and the first target scale is obtained from the first set of scale values. A rectangular box is generated based on the first target ratio value and the first target scale, using the target pixel in the sample image as the top-left corner. This rectangular box is then used as the first type of candidate box. In this embodiment, the target pixel is a sampled pixel in the sampled point set. The sampled point set includes multiple sampled pixels with a preset sampling interval. That is, to avoid candidate box redundancy caused by similar visual information of neighboring pixels, this step selects sampled pixels at equal intervals. The first type of candidate box is generated using each sampled pixel as the top-left corner, the first target ratio value as the aspect ratio, and the first target scale as the area ratio.

[0063] For example, if the preset sampling interval is equal to 1 / 10 of the short side of the sample image, and 9 / 16 is selected as the first target ratio value and 0.5 as the first target scale, then the first type of candidate box is generated by using 1 / 10 of the short side of the sample image as the sampling interval, 9 / 16 as the aspect ratio, and 0.5 as the area ratio.

[0064] S102, Obtain the second type of candidate box of the sample image.

[0065] In this embodiment, the second type of candidate bounding box includes a rectangular box whose target key point coincides with the target key point of the sample image, whose aspect ratio is equal to the second target ratio value, and whose area ratio is equal to the second target scale. The target key point is any preset key point in the key point set, the second target ratio value is any preset ratio value in the second ratio value set, and the second target scale is any preset scale in the second scale set. For example, the key point set includes the top-left corner, bottom-left corner, top-right corner, bottom-right corner, and center point.

[0066] It should be noted that the first set of scale values ​​and the second set of scale values ​​can be the same or different, and the second set of scale values ​​and the first set of scale values ​​can be the same or different. For example, the first set of scale values ​​and the second set of scale values ​​can be configured as the same set {1,4 / 3,3 / 4,16 / 9,9 / 16}, and the second set of scale values ​​can be configured as a different set {1,0.9,0.8,0.7,0.6}.

[0067] Optionally, specific methods for obtaining the second type of candidate boxes include:

[0068] This step involves selecting any preset keypoint from the keypoint set as the target keypoint, selecting any preset ratio value from the second ratio value set as the second target ratio value, and selecting any preset scale from the second scale set as the second target scale. Using the target keypoint on the sample image as the target keypoint, a bounding box is generated based on the second target ratio value and the second target scale, serving as the second type of candidate bounding box. This step uses the target keypoint as fixed vertices, the second target ratio value as the aspect ratio, and the second target scale as the area ratio to generate a second type of candidate bounding box with higher similarity to the sample image.

[0069] For example, if the top left corner is selected as the target key point, 9 / 16 is the second target ratio value, and 0.8 is the second target scale, then the top left corner of the sample image is used as the top left corner point, 9 / 16 is used as the aspect ratio, and 0.8 is used as the area ratio to generate the second type of candidate box.

[0070] S103. Obtain the third type of candidate box of the sample image.

[0071] In this embodiment, the third type of candidate box includes rectangular boxes whose composition satisfies the target composition rule and whose area ratio is equal to the third target scale. The target composition rule is any preset composition rule in the set of composition rules, and the third target scale is any preset scale in the set of third scales. Specifically, the set of composition rules is {central composition, rule of thirds, golden ratio composition}, and the set of third scales is {1, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3}.

[0072] Optionally, specific methods for obtaining the third type of candidate boxes include:

[0073] Take any preset composition rule from the set of composition rules as the target composition rule, take any preset scale from the set of third scales as the third target scale, use the target composition rule as the composition, and generate a rectangular box based on the third target scale as the third type of candidate box.

[0074] For example, if we choose center composition as the target composition rule and 0.6 as the third target scale, then using center composition as the composition rule and 0.6 as the area ratio, we can generate a third type of candidate box.

[0075] S104. Delete redundant candidate boxes in the first, second and third candidate boxes to generate a candidate box set for the sample image.

[0076] In this embodiment, the candidate box set includes multiple candidate boxes. Redundant candidate boxes include those that meet preset redundancy conditions. Redundancy conditions include an Intersection over Union (IoU) with at least one other candidate box that is greater than a preset threshold and a priority level no higher than other candidate boxes. The other candidate boxes are candidate boxes from the first, second, and third categories of candidate boxes. That is, when the IoU of a target candidate box with other candidate boxes is greater than the preset threshold and its priority is no higher than other candidate boxes, the target candidate box is deleted as a redundant candidate box. The target candidate box includes any one of the first, second, and third categories of candidate boxes. It should be noted that the priority is preset based on the candidate box category and attributes. Optionally, the priority is configured from low to high according to the candidate box category as: first category candidate boxes, second category candidate boxes, and third category candidate boxes.

[0077] It should be noted that S101 to S104 are optional methods for generating a set of candidate boxes for sample images based on a preset candidate box generation strategy, such as... Figure 2 As shown, the sample image is the first image 201. The candidate box generation strategy includes local region candidate box generation, vertex candidate box generation, fixed rule candidate box generation, and redundant candidate box deletion.

[0078] Image 202 illustrates the generation of local region candidate boxes, generating multiple local region candidate boxes, i.e., the first type of candidate boxes, based on each sampling point. Figure 2 The local region candidate box shown is candidate box h1 with sampling point o1 as the top-left vertex. Image 203 illustrates the generation of vertex candidate boxes. Based on preset vertices, i.e., key points, multiple vertex candidate boxes, i.e., second-type candidate boxes, are generated, such as... Figure 2As shown, the vertex candidate boxes include candidate box h2 with the top-left vertex o2 of the first image 201 as its top-left vertex, candidate box h3 with the bottom-right vertex o3 of the first image 201 as its bottom-right vertex, and candidate box h4 with the center point o4 of the first image 201 as its center point. Image 204 illustrates a schematic diagram of the generation of fixed-rule candidate boxes, generating multiple fixed-rule candidate boxes based on composition rules, which are also known as the third type of candidate boxes, such as... Figure 2 As shown, the fixed rule candidate boxes include candidate box h5 generated with center composition as the composition rule and candidate box h6 generated with third composition as the composition rule.

[0079] Furthermore, the candidate box generation strategy stipulates that when deleting redundant candidate boxes based on redundancy conditions, the redundancy conditions include a redundancy judgment condition IoU greater than 0.7 and a candidate box priority of fixed rule candidate boxes > vertex candidate boxes > local region candidate boxes.

[0080] As can be seen, this method generates composition boxes that better suit the aesthetic composition task through a candidate box generation strategy. It should be noted that the candidate boxes are stored in a preset data format, such as storing the coordinates of the four vertices of a candidate box or storing the coordinates of one vertex and its aspect ratio. For specific candidate box data formats, please refer to existing technologies.

[0081] S105. Input the sample image and each candidate box in the candidate box set into the pre-trained aesthetic big model to obtain the composition box score of each candidate box output by the aesthetic big model.

[0082] In this embodiment, the pre-training method for the aesthetic large model can refer to existing technologies.

[0083] S106. Sort the candidate boxes from highest to lowest according to their frame scores, and select the top n candidate boxes as the high-scoring frames of the sample image.

[0084] S107. Based on the high-resolution composition frame, crop the sample image to obtain n high-resolution composition images.

[0085] In this embodiment, n is the number of high-resolution images pre-configured.

[0086] S108. Send n high-resolution composition images and annotation instructions to the client corresponding to each annotation account, so that each client sends the sample annotation scores of the n high-resolution composition images.

[0087] In this embodiment, the annotation instruction is used to instruct aesthetic scores to be annotated for n high-scoring composition images. It should be noted that the annotation accounts are selected from a preset set of accounts, which includes multiple accounts, and each account corresponds to an annotator.

[0088] S109. For each labeled account, based on the sample labeling scores of the n high-resolution composition images sent by the client corresponding to the labeled account, select the top m high-resolution composition images with the highest sample labeling scores as candidate sample composition images corresponding to the labeled account.

[0089] In this embodiment, m is the pre-configured number of sample compositions. The annotation account is used to identify the account registered by the annotator based on their identity information. Specifically, n high-resolution composition images are sent to the interactive interface of the client logged into the annotation account, and the scores of each high-resolution composition image entered by the annotator in the interactive interface are received as the sample annotation scores of the high-resolution composition images.

[0090] In this embodiment, the score of the high-resolution composition image can be any value within a preset score range, for example, any value between 0 and 1. For each labeled account, the top m high-resolution composition images ranked by sample label scores are selected as candidate sample composition images, and the labeled account, candidate sample composition images, and sample label scores are stored accordingly.

[0091] S110. Calculate the selection probability of each candidate sample composition image, take the sum of the selection probabilities of the m candidate sample composition images with the highest selection probability as the statistical probability value, take the sum of the selection probabilities of the m candidate sample composition images corresponding to the labeled account as the target probability value, calculate the ratio of the target probability value to the statistical probability value, and take it as the outlier metric value of the labeled account for the sample image.

[0092] In this embodiment, the selection probability of each candidate sample composition image is calculated. The selection probability refers to the ratio of the number of corresponding annotation accounts to the total number of annotation accounts. In other words, the selection probability of a candidate sample composition image represents the probability that the candidate sample composition image is selected by the annotator.

[0093] Optionally, a histogram is constructed based on the selection probability of each candidate sample image. The horizontal axis of the histogram represents the selection probability of each candidate sample image, and the vertical axis represents the corresponding selection probability. The sum of the vertical axis values ​​of the m candidate sample images with the highest vertical axis values ​​is selected as the statistical probability value, and the sum of the vertical axis values ​​of the m candidate sample images corresponding to the labeled account is selected as the target probability value. The outlier metric of the labeled account is equal to the ratio of the target probability value to the statistical probability value. That is, the outlier metric of the labeled account is used to indicate the degree of outlier of the labeled account for the labeled score of the sample image. The smaller the outlier metric, the higher the degree of outlier. Taking n=20, the number of annotators k=10, and m=3 as an example... Figure 3 An example is provided: a histogram showing the probability of a candidate sample image being selected, such as... Figure 3As shown, the candidate sample images include image 1, image 2, image 3, image 5, image 6 and image 7. The selection probability of image 1 is 0.4, the selection probability of image 2 is 0.6, the selection probability of image 3 is 1, the selection probability of image 5 is 0.6, the selection probability of image 6 is 0.3 and the selection probability of image 7 is 0.1.

[0094] The statistical probability value is 1 + 0.6 + 0.6 = 2.2. The candidate sample composition images corresponding to the labeled account K1 include images 2, 3 and 5, so the target probability value of labeled account K1 is 0.6 + 1 + 0.6 = 2.2. The candidate sample composition images corresponding to the labeled account K2 include images 1, 3 and 7, so the target probability value of labeled account K2 is 0.4 + 1 + 0.1 = 1.6.

[0095] The outlier metric for the labeled account K1 on the sample image is 2.2 / 2.2 = 1, and the outlier metric for the labeled account K2 on the sample image is 1.6 / 2.2 = 0.7.

[0096] S111. If the outlier metric of the labeled account for the sample image is less than the preset metric threshold, then the outlier test result of the labeled account for the sample image will be set as unqualified.

[0097] For example, with a measurement threshold of 0.8, if the outlier metric of labeled account K1 for the sample image is greater than the measurement threshold, and the outlier metric of labeled account K2 for the sample image is less than the measurement threshold, then the outlier test result of labeled account K1 for the sample image is qualified, and the outlier test result of labeled account K2 for the sample image is unqualified.

[0098] In this embodiment, the outlier metric of the target annotation account (any annotation account) for the sample image is used to measure the degree of deviation between the annotation result of the target annotation account for the sample image and the annotation result of other annotation accounts (annotation accounts other than the target annotation account) for the sample image. The annotation result of the sample image refers to the selection result of m sample composition images. It can be understood that the lower the outlier metric, the higher the degree of deviation.

[0099] S112. If the failure rate of the outlier test results of the labeled account is greater than the preset elimination threshold, then the outlier test results of the labeled account for the sample set shall be set as unqualified.

[0100] In this embodiment, taking a sample set consisting of X sample images as an example, if the outlier test result of annotation account K1 for the X sample images is qualified, the failure rate of the outlier test result of annotation account K1 is (Xx) / X. If the failure rate is greater than the elimination threshold p, then the outlier test result of the annotation account for the sample set is unqualified, that is, the annotation accuracy of the annotator corresponding to the annotation account is low.

[0101] In this embodiment, the failure rate of the outlier test results of the target annotation account is used to indicate the degree of deviation between the annotation results of the target annotation account on the sample image and the annotation results of other annotation accounts on the sample image. It can be understood that the higher the failure rate, the greater the degree of deviation.

[0102] S113. Send a re-annotation instruction to the client corresponding to the first annotation account, so that the client resends the sample annotation scores of n high-resolution composition images, and re-executes S110 to S113 until the outlier test result of the first annotation account for the sample images is qualified, and the candidate sample composition image corresponding to the first annotation account is used as the sample composition image.

[0103] In this embodiment, the first labeled account refers to a labeled account that fails the outlier test for the sample image but passes the outlier test for the sample set.

[0104] For example, if the labeling account K2 fails the outlier test for sample image P, but passes the outlier test for the sample set, a relabeling instruction is sent to the client of labeling account K2. The relabeling instruction is used to instruct the n high-scoring composition images of P to be relabeled. Optionally, the n high-scoring composition images are reordered and displayed on the interactive interface.

[0105] S114. Send a substitute annotation instruction to the client corresponding to the second annotation account, so that the client sends the sample annotation scores of n high-resolution composition images.

[0106] In this embodiment, the second labeled account includes a substitute labeled account that fails the outlier test for the sample set (outlier unqualified labeled account). Optionally, the substitute labeled account is selected from the labeled account set.

[0107] S115. For the second labeled account, select the top m high-scoring composition images of the sample labeled scores as the candidate sample composition images corresponding to the second labeled account. Replace the candidate sample composition images corresponding to the labeled account with the outlier test results with the candidate sample composition images corresponding to the second labeled account. Repeat S110 to S115 until all labeled accounts pass the outlier test results for all sample images. Obtain the sample composition images corresponding to each labeled account.

[0108] S116. If all labeled accounts pass the validity test for all sample images, obtain the sample composition image and the corresponding label score result of the sample image.

[0109] In this embodiment, the validity test includes an outlier test. The sample composition images of the sample images include the sample composition images corresponding to each annotation account. The annotation score of the target sample composition image is obtained by weighted summation of the sample annotation scores of each annotation account for the target sample composition image. The weight of each annotation account for the sample annotation score of the target sample composition image can be determined based on the importance of each annotation account.

[0110] Optionally, the annotation score is the average of the annotation scores of all samples, meaning that each annotation account has the same level of importance. Continuing the previous example, Image 1 corresponds to 4 annotation accounts, meaning its selection probability is 0.4, and the sample annotation scores are 0.8, 0.9, 0.7, and 0.9 respectively. Therefore, the annotation score for Image 1 is (0.8 + 0.9 + 0.7 + 0.9) / 4 = 0.825. Image 3 corresponds to 10 annotation accounts, meaning its selection probability is 1, and the sample annotation scores are all 0.6. Therefore, the annotation score for Image 3 is 0.6.

[0111] It should be noted that this application improves the quality of manual annotation by conducting validity tests on aesthetic composition data annotation, including outlier tests. Based on the results of the outlier tests, sample composition images of sample images and corresponding annotation scores are obtained.

[0112] S117. Based on the sample composition images and corresponding annotation scores of each sample image, obtain the preferred composition frames and corresponding annotation scores of each sample image.

[0113] In this embodiment, the preferred bounding boxes include the high-scoring bounding boxes corresponding to the top m sample bounding images in the labeled score ranking. The labeled score of the preferred bounding boxes is the labeled score of the corresponding sample bounding images. It should be noted that S105 to S117 is a specific method for selecting preferred bounding boxes of sample images based on a pre-trained aesthetic model and a human intervention strategy, providing training data for the retraining of the aesthetic model.

[0114] S118. Retrain the aesthetic model using the labeled training set to obtain a trained aesthetic model.

[0115] In this embodiment, the labeled training set includes multiple retraining sample data, which includes the corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

[0116] The retraining process includes: inputting the sample images and candidate box sets from the retraining sample data into the pre-trained aesthetic large model, using the preferred composition boxes as the first annotation data and the annotation score results as the second annotation data, fine-tuning the various parameters of the aesthetic large model until the preset retraining completion conditions are met, and obtaining the trained aesthetic large model.

[0117] Optionally, the retraining completion conditions include reaching a preset number of iterations or a preset accuracy threshold. For methods to determine whether the aesthetic large model has met the retraining completion conditions, please refer to the prior art.

[0118] It should be noted that S101-118 above describes the specific implementation process of an optional aesthetic large-scale model construction method. As can be seen from the above technical solution, this method automatically generates the bounding box scores of multiple candidate boxes for the sample image based on a pre-trained aesthetic large-scale model. It then selects the n high-scoring bounding boxes with the highest scores, achieving the first optimization process for the candidate boxes of the sample image. This first optimization process requires no human intervention. Furthermore, based on a task distribution format, the client is instructed to provide the sample annotation scores of the n high-scoring bounding images obtained by cropping the sample image using the high-scoring bounding boxes. A second optimization process is then performed on the n high-scoring bounding boxes based on these sample annotation scores, resulting in multiple preferred bounding boxes and their corresponding annotation scores for the sample image. This second optimization process relies on a small amount of manual annotation. As can be seen, the multiple preferred bounding boxes and their corresponding annotation scores for the sample image, combining the model annotation results and the manual annotation results, demonstrate high accuracy in terms of both the preferred bounding boxes and their corresponding annotation scores.

[0119] Furthermore, a labeled training set is generated using the candidate bounding box set, multiple preferred composition boxes, and corresponding labeled scores of each sample image. This set is then used to retrain the aesthetic model, enabling fine-tuning of the aesthetic model. Since a single sample image can generate multiple candidate bounding boxes and corresponding labeled scores, a large amount of retraining sample data can be generated from a small number of sample images. Thus, the retraining sample data is characterized by its large quantity and high quality. Retraining the aesthetic model based on this retraining sample data improves the training efficiency and output accuracy of the aesthetic model, resulting in a finely tuned aesthetic model with better aesthetic scoring and composition capabilities.

[0120] S119. Obtain the set of candidate bounding boxes for the image to be labeled.

[0121] In this embodiment, the candidate box set of the image to be labeled includes multiple candidate boxes of the image to be labeled. The method for obtaining the candidate box set of the image to be labeled can be found in S101 to S104.

[0122] S120. Input the image to be labeled and the set of candidate boxes of the image to be labeled into the trained aesthetic model to obtain the labeling score of each candidate box of the image to be labeled output by the aesthetic model.

[0123] S121. Generate a training dataset based on the annotation scores of each candidate box in multiple images to be annotated.

[0124] In this embodiment, the training dataset includes multiple training datasets. Each training dataset includes a target sample image obtained by cropping the image to be labeled based on the candidate box and the labeling score of the image to be labeled corresponding to the candidate box.

[0125] It is understandable that a set of training data can be obtained from a single image to be labeled. Therefore, a first set of training data can be obtained based on a first number of images to be labeled. Each set of training data includes a second set of training data, where the second number is the number of candidate boxes for the images to be labeled corresponding to the training data set. Clearly, a well-trained aesthetic model can automatically label images, obtaining labeled target sample images. This labeling process requires no human intervention, reducing the manual cost of training data, improving labeling effectiveness, and further enhancing the performance of the aesthetic model trained on the training data.

[0126] It should be noted that S119-121 above describes the specific implementation process of the training data generation method. As can be seen from the above technical solution, the finely tuned aesthetic model possesses reliable aesthetic scoring and composition capabilities, and can automatically perform aesthetic scoring and composition annotation on unlabeled data. Based on the fully automated aesthetic model, the automatic annotation of unlabeled data is achieved, generating a training dataset. The annotation process requires no human intervention, saving significant manpower and time, greatly improving data annotation efficiency and accuracy, significantly reducing the difficulty and cost of aesthetic composition data annotation, standardizing annotation freedom, reducing annotation bias, and improving annotation efficiency. Based on this automatically annotated data, a lightweight aesthetic model can be trained, adapting to resource-limited application scenarios and enhancing the aesthetic evaluation capabilities of the lightweight aesthetic model.

[0127] In summary, the training data generation method provided in this application, based on an aesthetic composition candidate box generation strategy, can generate composition boxes that better meet the needs of aesthetic composition tasks, thereby improving the quality of aesthetic compositions. The aesthetic composition data annotation strategy based on minimal manual intervention can significantly reduce the difficulty and cost of aesthetic composition data annotation, and improve data annotation efficiency. Aesthetic composition data annotation based on minimal manual intervention allows for targeted fine-tuning of the large aesthetic model, making it more adaptable to current application scenarios. Furthermore, data can be collected based on new requirements, and the model can be iterated to adapt to changing needs. The appropriately fine-tuned large aesthetic model can automatically perform aesthetic scoring and composition data annotation, saving significant manpower and annotation time. This automatically annotated data can be used as training data to train a lightweight aesthetic model. Even with limited resources, a lightweight model can be used to complete aesthetic evaluation tasks.

[0128] It should be noted that S101 to S121 above is an optional specific implementation process of this application. The aesthetic large model construction method provided by this application also includes other specific implementation processes.

[0129] For example, S109 is only one optional method for obtaining sample composition images corresponding to labeled accounts. In another optional embodiment, the score of high-scoring composition images can be 0 or 1, and the number of high-scoring composition images with a score of 1 that the annotator inputs in the interactive interface is m. That is, the user can only annotate m high-scoring composition images with a score of 1. This is equivalent to the annotator selecting m high-scoring composition images from n high-scoring composition images and annotating these m high-scoring composition images with a score of 1. Then, these m high-scoring composition images are directly obtained as sample composition images, and the labeled account, sample composition images, and sample annotation scores are stored accordingly. This eliminates the step of manually selecting composition frames and also eliminates the need for precise aesthetic scoring of high-scoring composition images. Optionally, considering cost and operability, n=20, the number of annotators is k=10, and m=3. Based on this, the method for obtaining the preferred composition frames and corresponding annotation scores of each sample image includes selecting the top m sample composition images ranked from high to low as preferred composition frames.

[0130] For example, S111 to S116 are only one optional method for outlier testing. The outlier test results of the annotation account for the sample image are used to indicate the degree of deviation between the annotation results of the annotation account for the sample image and the annotation results of other annotation accounts for the sample image. The outlier test results of the annotation account for the sample set are used to indicate the degree of deviation between the annotation results of the annotation account for the sample set and the annotation results of other annotation accounts for the sample set.

[0131] If the outlier test result of the target annotation account for the sample image indicates that the deviation of the annotation result of the target annotation account for the sample image from the annotation result of other annotation accounts for the sample image is higher than the preset first deviation threshold, it means that the annotation result of the target annotation account for the sample image is unreasonable, and the target annotation account needs to be instructed to re-annotate n high-resolution composition images of the sample image to obtain m sample composition images.

[0132] If the outlier test result of the target annotation account for the sample set indicates that the deviation of the annotation result of the target annotation account for the sample set from the annotation result of other annotation accounts for the sample set is higher than the preset second deviation threshold, it means that the annotation result of the target annotation account for the sample set is unreasonable. It is necessary to indicate that the target annotation account should annotate n high-resolution composition images of the sample images of the replacement annotation account to obtain m sample composition images of the replacement target annotation account.

[0133] In other alternative embodiments, outlier testing can also be implemented through other specific methods. For example, the method for obtaining the outlier test results of the target labeled account for the sample image and / or the outlier test results of the target labeled account for the sample set can also include other specific implementation processes.

[0134] For example, S103 is an optional step. If there is no main target or no location label for the main target, then it is not necessary to obtain the third type of candidate box.

[0135] For example, in other optional embodiments, the validity test also includes probe data testing. After obtaining the sample image, the method further includes: obtaining a probe data set and probe sample image data; performing probe data testing on each labeled account based on the probe data set; and obtaining probe data test results.

[0136] In this embodiment, the probe data set includes multiple probe data, and each probe data includes a probe sample image and m corresponding preferred probe composition images.

[0137] Optionally, for any labeled account, the method for conducting probe data testing on the labeled account based on the probe dataset includes:

[0138] B1. Obtain the sample composition image set and probe preferred composition image set of the labeled account. The sample composition image set of the labeled account includes the sample composition images of all probe sample images corresponding to the labeled account, and the probe preferred composition image set includes the probe preferred composition images of all probe sample images.

[0139] B2. Calculate the overlap between the sample composition image set of the labeled account and the probe preferred composition image set. If the overlap is greater than the preset overlap threshold, the labeled account is determined to be a qualified labeled account for the probe. If the overlap is not greater than the overlap threshold, the labeled account is determined to be an unqualified labeled account for the probe.

[0140] B3. Send the substitute annotation instruction to the client corresponding to the third annotation account, so that the client sends the sample annotation scores of n high-resolution composition images.

[0141] In this embodiment, the third labeling account includes a substitute labeling account for the labeling account whose probe test result for the sample set is unqualified (probe unqualified labeling account). Optionally, the substitute labeling account is selected from the labeling account set.

[0142] B4. For the third labeled account, select the top m high-scoring composition images from the sample label scores and use them as the sample composition images corresponding to the third labeled account. Replace the sample composition images corresponding to the corresponding unqualified probe labeled accounts with the sample composition images corresponding to the third labeled account. Re-execute the method of testing probe data on labeled accounts based on probe data sets until the probe data test results of all labeled accounts are qualified.

[0143] It should be noted that this application conducts validity tests on aesthetic composition data annotation. The validity tests include probe data tests and outlier tests. The outlier test is used to detect the degree of deviation between each annotation account and other annotation accounts, which can unify the aesthetic annotation standards of multiple annotation accounts annotating the sample set. The probe data test is used to detect the overlap between each annotation account and the standard annotation. Since the standard annotation is authoritative and professional, the probe data test can detect the professionalism of each annotation account. It can be seen that obtaining sample composition images and corresponding annotation scores of sample images based on probe data tests and outlier tests can improve the quality of manual annotation, that is, improve the accuracy of manual annotation and avoid annotation deviations caused by subjective bias.

[0144] In summary, the data annotation method provided in the embodiments of this application can be summarized as follows: Figure 3 The process shown is as follows: Figure 3 As shown, this method includes:

[0145] S301. Obtain the candidate bounding box set for each sample image in the sample set.

[0146] In this embodiment, the candidate box set includes multiple candidate boxes, which are rectangular boxes generated based on a preset candidate box generation strategy for cropping sample images.

[0147] S302. For each sample image, based on the bounding box scores of each candidate box output by the pre-trained aesthetic large model, obtain multiple high-scoring bounding box images of the sample image.

[0148] S303. Send multiple high-resolution composite images of the sample image and annotation instructions to the client corresponding to each annotation account.

[0149] In this embodiment, the annotation instruction for the sample image is used to instruct each client to provide the sample annotation scores for multiple high-resolution composition images.

[0150] S304. For each labeled account, based on the sample labeling scores of multiple high-resolution composition images of the sample image, select the sample composition image corresponding to the labeled account.

[0151] S305. Based on the sample composition images corresponding to each labeled account, obtain the preferred composition frame of the sample image and the corresponding labeling score result.

[0152] S306. Use the labeled training set to retrain the pre-trained aesthetic model to obtain a well-trained aesthetic model.

[0153] In this embodiment, the labeled training set includes multiple retraining sample data, which includes the corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

[0154] As can be seen from the above technical solution, the data annotation method provided in this application provides a method for obtaining a set of candidate boxes for each sample image in a sample set, wherein the set of candidate boxes includes multiple candidate boxes. For each sample image, based on the composition box scores of each candidate box output by the pre-trained aesthetic large model, multiple high-resolution composition images of the sample image are obtained. The multiple high-resolution composition images of the sample image and annotation instructions are sent to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-resolution composition images. For each annotation account, based on the sample annotation scores of the multiple high-resolution composition images of the sample image, the sample composition image corresponding to the annotation account is selected. Based on the sample composition images corresponding to each annotation account, the preferred composition box of the sample image and the corresponding annotation score result are obtained. The pre-trained aesthetic large model is retrained using the annotation training set to obtain a trained aesthetic large model. The annotation training set includes multiple retraining sample data, which includes the corresponding sample image, candidate box set, preferred composition box, and annotation score result. This method automatically labels the bounding box scores of each candidate box based on a pre-trained aesthetic model, performing the first labeling to obtain high-scoring composition images. A second labeling is then performed based on the sample labeling scores of the high-scoring composition images fed back by the client, obtaining multiple preferred composition boxes and their corresponding labeling scores. The first labeling is achieved automatically, while the second selection is achieved with minimal manual labeling. This eliminates the need for manual scoring of each candidate box. By utilizing the pre-trained aesthetic model and minimal manual intervention, a labeled training set is obtained for retraining the aesthetic model. Fine-tuning of the aesthetic model is then achieved based on this labeled training set, reducing the manual cost of building the aesthetic model and improving the accuracy of automatic labeling.

[0155] Figure 4 This application provides a schematic diagram of the structure of a data annotation device according to an embodiment of the present application. Figure 4 As shown, the device may include:

[0156] The candidate box generation unit 401 is used to obtain a candidate box set for each sample image in the sample set, wherein the candidate box set includes multiple candidate boxes;

[0157] The sample automatic scoring unit 402 is used to obtain multiple high-scoring composition images of the sample image based on the composition box scores of each candidate box output by the pre-trained aesthetic large model for each sample image.

[0158] The sample manual scoring unit 403 is used to send multiple high-scoring composition images of the sample image and annotation instructions to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-scoring composition images.

[0159] The sample composition selection unit 404 is used to select the sample composition image corresponding to the labeled account for each labeled account based on the sample labeling score of multiple high-resolution composition images of the sample image.

[0160] The preferred composition selection unit 405 is used to obtain the preferred composition frame of the sample image and the corresponding annotation score result based on the sample composition image corresponding to each of the labeled accounts.

[0161] The model retraining unit 406 is used to retrain the pre-trained aesthetic large model using a labeled training set to obtain a trained aesthetic large model. The labeled training set includes multiple retraining sample data, which includes corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

[0162] Optionally, the model retraining unit is specifically used to: input the sample images and candidate box sets in each of the retraining sample data into the pre-trained aesthetic large model, use the preferred composition box as the first annotation data and the annotation score result as the second annotation data, fine-tune the various parameters of the pre-trained aesthetic large model until the preset retraining completion conditions are met, and obtain the trained aesthetic large model.

[0163] Optionally, the candidate box generation unit is also used to obtain a set of candidate boxes for the image to be labeled;

[0164] The data annotation device further includes: a data annotation unit, used to input the image to be annotated and the candidate bounding box set of the image to be annotated into the trained aesthetic model, to obtain the annotation data of the image to be annotated output by the trained aesthetic model, wherein the annotation data of the image to be annotated includes the preferred composition box of the image to be annotated and the corresponding annotation score result; and to collect the annotation data of each image to be annotated to obtain training data.

[0165] Optionally, the automatic sample scoring unit is specifically used for: inputting the sample image and each candidate box in the candidate box set into the pre-trained aesthetic large model to obtain the composition box score of each candidate box output by the pre-trained aesthetic large model; sorting each candidate box according to the composition box score from high to low, selecting the top n candidate boxes as high-scoring composition boxes of the sample image, where n is the pre-configured number of high-scoring composition boxes; and cropping the sample image based on each of the high-scoring composition boxes to obtain n high-scoring composition images of the sample image.

[0166] Optionally, the target image includes the sample image and the image to be labeled, and the candidate box generation unit is used to obtain a set of candidate boxes for the target image, specifically for:

[0167] Obtain a first type of candidate bounding box for the target image, wherein the first type of candidate bounding box is a rectangle with an aspect ratio equal to a first target ratio and an area ratio equal to a first target scale; obtain a second type of candidate bounding box for the target image, wherein the second type of candidate bounding box includes rectangles whose target key points coincide with the target key points of the target image and whose aspect ratio is equal to a second target ratio; obtain a third type of candidate bounding box for the target image, wherein the third type of candidate bounding box includes rectangles whose composition satisfies the target composition rules and whose area ratio is equal to a third target scale; delete redundant candidate bounding boxes in the first type of candidate bounding box, the second type of candidate bounding box, and the third type of candidate bounding box to generate a candidate bounding box set for the target image.

[0168] Optionally, the sample composition selection unit is specifically used for: selecting the top m high-scoring composition images based on the sample annotation scores of multiple high-scoring composition images of the sample image, as candidate sample composition images of the sample image, where m is a pre-configured number of sample compositions; performing outlier testing on the annotation account to obtain the outlier test result of the annotation account, wherein the outlier test result of the annotation account includes the outlier test result of the annotation account for the sample image and / or the outlier test result of the annotation account for the sample set, wherein the outlier test result of the annotation account for the sample image is used to indicate the degree of deviation between the annotation result of the annotation account for the sample image and the annotation result of other annotation accounts for the sample image, and the outlier test result of the annotation account for the sample image is used to indicate the degree of deviation between the annotation result of the annotation account for the sample image and the annotation result of other annotation accounts for the sample image. The outlier test results of the sample set are used to indicate the degree of deviation between the annotation results of the annotation account and the annotation results of other annotation accounts. The sample set includes multiple sample images. If the outlier test results of the annotation account for the sample image indicate that the degree of deviation between the annotation results of the annotation account for the sample image and the annotation results of other annotation accounts is not higher than a preset first deviation threshold, and the outlier test results of the annotation account for the sample set indicate that the degree of deviation between the annotation results of the annotation account for the sample image and the annotation results of other annotation accounts is not higher than a preset second deviation threshold, then the candidate sample composition image of the sample image is used as the sample composition image of the sample image.

[0169] Optionally, the preferred composition selection unit is specifically used for: for each of the sample composition images of the sample images, calculating the average of the sample annotation scores of each sample composition image as the annotation score result of the sample composition image; obtaining the top m sample composition images sorted from largest to smallest by annotation score results to obtain m preferred composition images; obtaining the candidate boxes corresponding to each of the preferred composition images as preferred composition boxes; and obtaining the annotation score results of each of the preferred composition images as the annotation score results of the corresponding preferred composition boxes.

[0170] Figure 5 A schematic diagram of the data annotation device is shown. The device may include: at least one processor 501, at least one communication interface 502, at least one memory 503, and at least one communication bus 504.

[0171] In this embodiment of the application, the number of processor 501, communication interface 502, memory 503 and communication bus 504 is at least one, and processor 501, communication interface 502 and memory 503 communicate with each other through communication bus 504.

[0172] The processor 501 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0173] The memory 503 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0174] The memory stores a program, and the processor can execute the program stored in the memory to implement the various steps of the data annotation method provided in this application embodiment, as follows:

[0175] Obtain a set of candidate bounding boxes for each sample image in the sample set, wherein the set of candidate bounding boxes includes multiple candidate bounding boxes;

[0176] For each of the sample images, based on the bounding box scores of each of the candidate boxes output by the pre-trained aesthetic large model, multiple high-scoring bounding boxes of the sample image are obtained.

[0177] Multiple high-resolution composition images of the sample image and annotation instructions are sent to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-resolution composition images.

[0178] For each labeled account, the sample composition image corresponding to the labeled account is selected based on the sample labeling scores of multiple high-resolution composition images of the sample image;

[0179] Based on the sample composition images corresponding to each labeled account, the preferred composition frame of the sample image and the corresponding labeling score result are obtained;

[0180] The pre-trained aesthetic model is retrained using a labeled training set to obtain a trained aesthetic model. The labeled training set includes multiple retraining sample data, which includes corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

[0181] Optionally, the pre-trained aesthetic model is retrained using a labeled training set to obtain a trained aesthetic model. This includes: inputting sample images and candidate box sets from each of the retraining sample data into the pre-trained aesthetic model, using preferred composition boxes as the first labeled data and labeled scores as the second labeled data, and fine-tuning various parameters of the pre-trained aesthetic model until the preset retraining completion conditions are met to obtain the trained aesthetic model.

[0182] Optionally, the data annotation method further includes: obtaining a set of candidate bounding boxes for the image to be annotated; inputting the image to be annotated and the set of candidate bounding boxes for the image to be annotated into the trained aesthetic model to obtain the annotation data of the image to be annotated output by the trained aesthetic model, wherein the annotation data of the image to be annotated includes the preferred bounding boxes of the image to be annotated and the corresponding annotation scores; and aggregating the annotation data of each image to be annotated to obtain training data.

[0183] Optionally, based on the bounding box scores of each candidate box output by the pre-trained aesthetic large model, multiple high-resolution composition images of the sample image are obtained, including: inputting the sample image and each candidate box in the candidate box set into the pre-trained aesthetic large model to obtain the bounding box scores of each candidate box output by the pre-trained aesthetic large model; sorting each candidate box according to the bounding box scores from high to low, selecting the top n candidate boxes as the high-resolution composition boxes of the sample image, where n is the pre-configured number of high-resolution compositions; and cropping the sample image based on each high-resolution composition box to obtain n high-resolution composition images of the sample image.

[0184] Optionally, the target image includes the sample image and the image to be labeled. Obtaining the candidate box set of the target image includes: obtaining a first type of candidate box of the target image, wherein the first type of candidate box is a rectangle with an aspect ratio equal to a first target ratio value and an area ratio equal to a first target scale; obtaining a second type of candidate box of the target image, wherein the second type of candidate box includes rectangles whose target key points coincide with the target key points of the target image and whose aspect ratio is equal to a second target ratio value; obtaining a third type of candidate box of the target image, wherein the third type of candidate box includes rectangles whose composition satisfies the target composition rules and whose area ratio is equal to a third target scale; deleting redundant candidate boxes in the first type of candidate box, the second type of candidate box, and the third type of candidate box to generate the candidate box set of the target image.

[0185] Optionally, selecting the sample composition image corresponding to the annotation account based on the sample annotation scores of multiple high-resolution composition images of the sample image includes: selecting the top m high-resolution composition images by sample annotation scores from the multiple high-resolution composition images of the sample image as candidate sample composition images of the sample image, where m is a pre-configured number of sample compositions; performing an outlier test on the annotation account to obtain the outlier test result of the annotation account, wherein the outlier test result of the annotation account includes the outlier test result of the annotation account for the sample image and / or the outlier test result of the annotation account for the sample set, wherein the outlier test result of the annotation account for the sample image is used to indicate the annotation result of the annotation account for the sample image compared with the annotation result of other annotation accounts for the sample image. The outlier test result of the annotation account for the sample set is used to indicate the degree of deviation between the annotation result of the annotation account for the sample set and the annotation result of other annotation accounts for the sample set. The sample set includes multiple sample images. If the outlier test result of the annotation account for the sample image indicates that the degree of deviation between the annotation result of the annotation account for the sample image and the annotation result of other annotation accounts for the sample image is not higher than a preset first deviation threshold, and the outlier test result of the annotation account for the sample set indicates that the degree of deviation between the annotation result of the annotation account for the sample set and the annotation result of other annotation accounts for the sample set is not higher than a preset second deviation threshold, then the candidate sample composition image of the sample image is used as the sample composition image of the sample image.

[0186] Optionally, based on the sample composition images corresponding to each labeled account, obtaining the preferred composition frames and corresponding labeling scores of the sample images includes: for each sample composition image of the sample images, calculating the average of the labeling scores of each sample composition image as the labeling score result of the sample composition image; obtaining the top m sample composition images sorted by labeling scores from largest to smallest to obtain m preferred composition images; obtaining the candidate frames corresponding to each preferred composition image as preferred composition frames; and obtaining the labeling scores of each preferred composition image as the labeling score result of the corresponding preferred composition frame.

[0187] This application also provides a readable storage medium that stores a computer program suitable for execution by a processor. When executed by the processor, the computer program implements the various steps of the data annotation method provided in this application, as follows:

[0188] Obtain a set of candidate bounding boxes for each sample image in the sample set, wherein the set of candidate bounding boxes includes multiple candidate bounding boxes;

[0189] For each of the sample images, based on the bounding box scores of each of the candidate boxes output by the pre-trained aesthetic large model, multiple high-scoring bounding boxes of the sample image are obtained.

[0190] Multiple high-resolution composition images of the sample image and annotation instructions are sent to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-resolution composition images.

[0191] For each labeled account, the sample composition image corresponding to the labeled account is selected based on the sample labeling scores of multiple high-resolution composition images of the sample image;

[0192] Based on the sample composition images corresponding to each labeled account, the preferred composition frame of the sample image and the corresponding labeling score result are obtained;

[0193] The pre-trained aesthetic model is retrained using a labeled training set to obtain a trained aesthetic model. The labeled training set includes multiple retraining sample data, which includes corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

[0194] Optionally, the pre-trained aesthetic model is retrained using a labeled training set to obtain a trained aesthetic model. This includes: inputting sample images and candidate box sets from each of the retraining sample data into the pre-trained aesthetic model, using preferred composition boxes as the first labeled data and labeled scores as the second labeled data, and fine-tuning various parameters of the pre-trained aesthetic model until the preset retraining completion conditions are met to obtain the trained aesthetic model.

[0195] Optionally, the data annotation method further includes: obtaining a set of candidate bounding boxes for the image to be annotated; inputting the image to be annotated and the set of candidate bounding boxes for the image to be annotated into the trained aesthetic model to obtain the annotation data of the image to be annotated output by the trained aesthetic model, wherein the annotation data of the image to be annotated includes the preferred bounding boxes of the image to be annotated and the corresponding annotation scores; and aggregating the annotation data of each image to be annotated to obtain training data.

[0196] Optionally, based on the bounding box scores of each candidate box output by the pre-trained aesthetic large model, multiple high-resolution composition images of the sample image are obtained, including: inputting the sample image and each candidate box in the candidate box set into the pre-trained aesthetic large model to obtain the bounding box scores of each candidate box output by the pre-trained aesthetic large model; sorting each candidate box according to the bounding box scores from high to low, selecting the top n candidate boxes as the high-resolution composition boxes of the sample image, where n is the pre-configured number of high-resolution compositions; and cropping the sample image based on each high-resolution composition box to obtain n high-resolution composition images of the sample image.

[0197] Optionally, the target image includes the sample image and the image to be labeled. Obtaining the candidate box set of the target image includes: obtaining a first type of candidate box of the target image, wherein the first type of candidate box is a rectangle with an aspect ratio equal to a first target ratio value and an area ratio equal to a first target scale; obtaining a second type of candidate box of the target image, wherein the second type of candidate box includes rectangles whose target key points coincide with the target key points of the target image and whose aspect ratio is equal to a second target ratio value; obtaining a third type of candidate box of the target image, wherein the third type of candidate box includes rectangles whose composition satisfies the target composition rules and whose area ratio is equal to a third target scale; deleting redundant candidate boxes in the first type of candidate box, the second type of candidate box, and the third type of candidate box to generate the candidate box set of the target image.

[0198] Optionally, selecting the sample composition image corresponding to the annotation account based on the sample annotation scores of multiple high-resolution composition images of the sample image includes: selecting the top m high-resolution composition images by sample annotation scores from the multiple high-resolution composition images of the sample image as candidate sample composition images of the sample image, where m is a pre-configured number of sample compositions; performing an outlier test on the annotation account to obtain the outlier test result of the annotation account, wherein the outlier test result of the annotation account includes the outlier test result of the annotation account for the sample image and / or the outlier test result of the annotation account for the sample set, wherein the outlier test result of the annotation account for the sample image is used to indicate the annotation result of the annotation account for the sample image compared with the annotation result of other annotation accounts for the sample image. The outlier test result of the annotation account for the sample set is used to indicate the degree of deviation between the annotation result of the annotation account for the sample set and the annotation result of other annotation accounts for the sample set. The sample set includes multiple sample images. If the outlier test result of the annotation account for the sample image indicates that the degree of deviation between the annotation result of the annotation account for the sample image and the annotation result of other annotation accounts for the sample image is not higher than a preset first deviation threshold, and the outlier test result of the annotation account for the sample set indicates that the degree of deviation between the annotation result of the annotation account for the sample set and the annotation result of other annotation accounts for the sample set is not higher than a preset second deviation threshold, then the candidate sample composition image of the sample image is used as the sample composition image of the sample image.

[0199] Optionally, based on the sample composition images corresponding to each labeled account, obtaining the preferred composition frames and corresponding labeling scores of the sample images includes: for each sample composition image of the sample images, calculating the average of the labeling scores of each sample composition image as the labeling score result of the sample composition image; obtaining the top m sample composition images sorted by labeling scores from largest to smallest to obtain m preferred composition images; obtaining the candidate frames corresponding to each preferred composition image as preferred composition frames; and obtaining the labeling scores of each preferred composition image as the labeling score result of the corresponding preferred composition frame.

[0200] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0201] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0202] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data annotation method, characterized in that, include: Obtain a set of candidate bounding boxes for each sample image in the sample set, wherein the set of candidate bounding boxes includes multiple candidate bounding boxes; For each of the sample images, based on the bounding box scores of each of the candidate boxes output by the pre-trained aesthetic large model, multiple high-scoring bounding boxes of the sample image are obtained. Multiple high-resolution composition images of the sample image and annotation instructions are sent to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-resolution composition images. For each labeled account, the sample composition image corresponding to the labeled account is selected based on the sample labeling scores of multiple high-resolution composition images of the sample image; Based on the sample composition images corresponding to each labeled account, the preferred composition frame of the sample image and the corresponding labeling score result are obtained; The pre-trained aesthetic model is retrained using a labeled training set to obtain a trained aesthetic model. The labeled training set includes multiple retraining sample data, which includes corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

2. The method according to claim 1, characterized in that, The process of retraining the pre-trained aesthetic model using a labeled training set to obtain a trained aesthetic model includes: The sample images and candidate box sets in each of the retraining sample data are input into the pre-trained aesthetic model. The preferred composition boxes are used as the first annotation data, and the annotation scores are used as the second annotation data. The parameters of the pre-trained aesthetic model are fine-tuned until the preset retraining completion conditions are met, and the trained aesthetic model is obtained.

3. The method according to claim 2, characterized in that, The data annotation method also includes: Obtain the candidate bounding box set of the image to be labeled; The image to be labeled and the set of candidate bounding boxes of the image to be labeled are input into the trained aesthetic model to obtain the labeled data of the image to be labeled output by the trained aesthetic model. The labeled data of the image to be labeled includes the preferred bounding box of the image to be labeled and the corresponding labeled score result. The labeled data of each of the images to be labeled is collected to obtain the training data.

4. The method according to claim 3, characterized in that, The bounding box scores of each candidate box output by the pre-trained aesthetic large model are used to obtain multiple high-scoring bounding box images of the sample image, including: The sample image and each candidate box in the candidate box set are input into the pre-trained aesthetic large model to obtain the composition box score of each candidate box output by the pre-trained aesthetic large model. The candidate boxes are sorted from high to low according to their frame scores, and the top n candidate boxes are selected as the high-scoring frames of the sample image, where n is the number of high-scoring frames pre-configured. Based on each of the high-resolution composition frames, the sample image is cropped to obtain n high-resolution composition images of the sample image.

5. The method according to claim 4, characterized in that, The target image includes the sample image and the image to be labeled. The candidate bounding box set for the target image includes: Obtain a first type of candidate box in the target image. The first type of candidate box is a rectangular box with an aspect ratio equal to the first target ratio and an area ratio equal to the first target scale. Obtain a second type of candidate box for the target image. The second type of candidate box includes a rectangular box whose target key points coincide with the target key points of the target image and whose aspect ratio is equal to the second target ratio value. Obtain a third type of candidate box for the target image, the third type of candidate box including rectangular boxes whose composition satisfies the target composition rules and whose area ratio is equal to the third target scale; Redundant candidate boxes in the first type, the second type, and the third type are deleted to generate a set of candidate boxes for the target image.

6. The method according to claim 1, characterized in that, The selection of the sample composition image corresponding to the labeled account, based on the sample annotation scores of multiple high-resolution composition images of the sample image, includes: Based on the sample annotation scores of multiple high-resolution composition images of the sample image, the top m high-resolution composition images with the highest sample annotation scores are selected as candidate sample composition images of the sample image, where m is the pre-configured number of sample compositions. Outlier test results are obtained by performing an outlier test on the labeled account. The outlier test results of the labeled account include the outlier test results of the labeled account for the sample image and / or the outlier test results of the labeled account for the sample set. The outlier test results of the labeled account for the sample image are used to indicate the degree of deviation between the labeled account's labeling results for the sample image and the labeling results of other labeled accounts for the sample image. The outlier test results of the labeled account for the sample set are used to indicate the degree of deviation between the labeled account's labeling results for the sample set and the labeling results of other labeled accounts for the sample set. The sample set includes multiple sample images. If the outlier test result of the annotation account for the sample image indicates that the deviation of the annotation result of the annotation account for the sample image from the annotation result of other annotation accounts for the sample image is not higher than a preset first deviation threshold, and the outlier test result of the annotation account for the sample set indicates that the deviation of the annotation result of the annotation account for the sample set from the annotation result of other annotation accounts for the sample set is not higher than a preset second deviation threshold, then the candidate sample composition image of the sample image is used as the sample composition image of the sample image.

7. The method according to claim 6, characterized in that, The step of obtaining the preferred frame of the sample image and the corresponding annotation score based on the sample composition image corresponding to each of the labeled accounts includes: For each of the sample composition images of the sample image, the average of the individual sample annotation scores of the sample composition image is calculated as the annotation score result of the sample composition image; Obtain the top m sample composition images sorted from largest to smallest by labeled score results, and obtain m optimal composition images; Obtain candidate bounding boxes corresponding to each of the preferred composition images as preferred composition bounding boxes; The annotation score of each of the preferred composition images is obtained as the annotation score of the corresponding preferred composition frame.

8. A data annotation device, characterized in that, include: A candidate box generation unit is used to obtain a set of candidate boxes for each sample image in the sample set, wherein the set of candidate boxes includes multiple candidate boxes; The sample automatic scoring unit is used to obtain multiple high-scoring composition images of each sample image based on the composition box scores of each candidate box output by the pre-trained aesthetic large model. The sample manual scoring unit is used to send multiple high-scoring composition images of the sample image and annotation instructions to the client corresponding to each annotation account. The annotation instructions of the sample image are used to instruct each client to provide the sample annotation scores of the multiple high-scoring composition images. The sample composition selection unit is used to select the sample composition image corresponding to the labeled account for each labeled account based on the sample labeling score of multiple high-scoring composition images of the sample image; The preferred composition selection unit is used to obtain the preferred composition frame of the sample image and the corresponding annotation score result based on the sample composition image corresponding to each of the labeled accounts; The model retraining unit is used to retrain the pre-trained aesthetic large model using a labeled training set to obtain a trained aesthetic large model. The labeled training set includes multiple retraining sample data, which includes corresponding sample images, candidate box sets, preferred composition boxes, and labeled score results.

9. A data annotation device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the data annotation method as described in any one of claims 1 to 7.

10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the data annotation method as described in any one of claims 1 to 7.