Image Processing Method, Apparatus, Electronic Device, and Readable Storage Medium

By connecting the model and composition and aesthetic modules, multiple cropping areas are determined and scored, the problem of insufficient aesthetics of image cropping in the prior art is solved, and an aesthetic and diverse cropping effect is achieved, which is suitable for Internet services.

CN114529558BActive Publication Date: 2025-07-11VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210122211.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-09
Publication Date
2025-07-11
Estimated Expiration
2042-02-09

AI Technical Summary

Technical Problem

The prior art image cropping method results in lower aesthetics of the cropped images.

Method used

Using the image processing method based on the connection model, multiple cropping areas are determined and scored through the composition model and aesthetic module branch, aesthetic score information is output, and aesthetic scores are provided for multiple cropable borders.

Benefits of technology

Ensure that the cropped images selected by users are more beautiful, and provide multiple scale cropped images to meet the diverse needs of Internet services, save machine resources, and simplify service deployment logic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529558B_ABST
    Figure CN114529558B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, apparatus, electronic device and readable storage medium, belonging to the field of electronic technology. Among them, the method includes: obtaining feature information of a first image based on a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model; determining N1 first regions in the first image according to the feature information, where N1 is a positive integer; determining first aesthetic score information corresponding to each of the first regions; outputting border lines corresponding to each of the first regions in the first image, and outputting first aesthetic score information corresponding to each of the first regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to an image processing method, apparatus, electronic device, and readable storage medium. Background Art

[0002] Currently, when users browse pages such as video recommendations, multiple pictures are presented on the page, and each picture is used to represent the brief information of the corresponding video. In this way, it helps users quickly browse the main content of each video on the recommendation page, so that users can select interesting pictures for operations such as clicking, and then the interface displays the video playback page corresponding to the picture.

[0003] Generally, for a certain video, a certain frame of the video is selected, and in this picture, a certain area is selected and cropped to be used as the cover of the video, and finally displayed on the video recommendation page. In the prior art, methods for cropping pictures include: cropping a certain area with the upper left corner of the picture as the reference point; cropping a certain area with the center of the picture as the reference point; and so on.

[0004] It can be seen that cropping pictures by the methods in the prior art results in pictures obtained by cropping having relatively low aesthetic value in terms of composition. Summary of the Invention

[0005] The objective of the embodiments of this application is to provide an image processing method, which can solve the problem that cropping pictures by the methods in the prior art results in pictures obtained by cropping having relatively low aesthetic value in terms of composition.

[0006] In a first aspect, the embodiments of this application provide an image processing method, which includes: obtaining the feature information of a first image based on a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model; determining N1 first regions in the first image according to the feature information, where N1 is a positive integer; determining the first aesthetic score information corresponding to each of the first regions; and outputting the border lines corresponding to each of the first regions in the first image, and outputting the first aesthetic score information corresponding to each of the first regions.

[0007] Second aspect, an embodiment of the present application provides an image processing apparatus, which includes: an acquisition module, configured to acquire feature information of a first image based on a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model; a first determination module, configured to determine N1 first regions in the first image according to the feature information, where N1 is a positive integer; a second determination module, configured to determine first aesthetic score information corresponding to each of the first regions; and an output module, configured to output border lines corresponding to each of the first regions in the first image, and output first aesthetic score information corresponding to each of the first regions.

[0008] Third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0009] Fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0010] Fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0011] Sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0012] In this way, in the embodiment of the present application, based on pre-training the connection model in advance, the first image can be input into the connection model. Thus, according to the feature information of the first image, the connection model uses the algorithm of the composition model to determine N1 first regions in the first image. The first regions can be used as the images to be retained after cropping. Further, after determining the N1 first regions, the connection model uses the algorithm of the aesthetic module branch to perform aesthetic scoring on each of the first regions respectively according to the N1 first regions, so as to determine first aesthetic score information corresponding to each of the first regions. Finally, in the first image, the border lines of each of the first regions are output, and at the same time, the first aesthetic score information corresponding to each of the first regions is output. It can be seen that based on the finally presented output effect, multiple aesthetic scores of the croppable borders can be provided for the user to refer to, so as to ensure that the cropped image selected by the user is relatively beautiful. Description of the Drawings

[0013] Figure 1It is one of the flowcharts of the image processing method according to the embodiments of the present application;

[0014] Figure 2 It is the output schematic diagram of the image processing method according to the embodiments of the present application;

[0015] Figure 3 It is the second flowchart of the image processing method according to the embodiments of the present application;

[0016] Figure 4 It is the third flowchart of the image processing method according to the embodiments of the present application;

[0017] Figure 5 It is the fourth flowchart of the image processing method according to the embodiments of the present application;

[0018] Figure 6 It is the fifth flowchart of the image processing method according to the embodiments of the present application;

[0019] Figure 7 It is the block diagram of the image processing device according to the embodiments of the present application;

[0020] Figure 8 It is one of the schematic diagrams of the hardware structure of the electronic device according to the embodiments of the present application;

[0021] Figure 9 It is the second schematic diagram of the hardware structure of the electronic device according to the embodiments of the present application. Detailed implementation manners

[0022] Next, the technical solutions of the embodiments of the present application will be clearly described in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0023] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0024] Next, the image processing method provided by the embodiments of the present application will be described in detail in conjunction with the accompanying drawings, through specific embodiments and their application scenarios.

[0025] Figure 1 The figure shows a flowchart of an image processing method according to an embodiment of the present application. The method is applied to an electronic device and includes:

[0026] Step 110: Obtain feature information of a first image based on a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model.

[0027] Optionally, the image processing method in the present application is implemented through a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model.

[0028] Correspondingly, in this step, the first image is input into the connection model as an initial image.

[0029] Optionally, the first image is a picture, which can be a certain frame picture in a video; it can be a certain long picture; and so on.

[0030] The feature information of the first image includes image feature information, which is used to reflect what objects are included in the first image, such as people, objects, scenes, etc.

[0031] Step 120: Determine N1 first regions in the first image according to the feature information, where N1 is a positive integer.

[0032] In this step, according to the feature information of the first image, the algorithm corresponding to the composition model used in the connection model obtains N1 first regions in the first image.

[0033] For example, according to the feature information of the first image, the main feature information of the first image can be identified, which can be a human face. Thus, taking the human face as a reference, the first region is divided in the first image.

[0034] Among them, in the present application, through the connection model, N1 first regions can be determined in the first image, where N1 is a positive integer, which can be one or more.

[0035] For example, for the determined N1 first regions, all include human faces.

[0036] Step 130: Determine the first aesthetic score information corresponding to each first region.

[0037] In this step, according to the sub-images in the determined N1 first regions, the algorithm corresponding to the aesthetic model branch used in the connection model obtains the aesthetic score information of each sub-image, so as to obtain the first aesthetic score information corresponding to each first region.

[0038] Optionally, the first aesthetic score information is a score from 1 to 10.

[0039] Step 140: Output the border lines corresponding to each first region in the first image, and output the first aesthetic score information corresponding to each first region.

[0040] In this step, relevant calculations are performed on the first image based on the connection model, and finally, the border lines of each first region and the first aesthetic score information corresponding to each first region are output in the first image.

[0041] Optionally, the border lines of different first regions are displayed separately, such as using lines of different colors or different thicknesses to achieve the effect of distinguishable display.

[0042] See Figure 2 , in the illustrated Picture 1, multiple boxes are output, each box is used to represent a first region, and below Picture 1, the schematic symbols of each box and the corresponding scores are displayed.

[0043] In this way, in the embodiment of the present application, based on pre-training the connection model, the first image can be input into the connection model. Thus, the connection model determines N1 first regions in the first image according to the feature information of the first image by using the algorithm of the composition model. The first regions can be the images retained after cropping. Further, after determining the N1 first regions, the connection model respectively performs aesthetic scoring on each first region according to the N1 first regions by using the algorithm of the aesthetic module branch to determine the first aesthetic score information corresponding to each first region. Finally, in the first image, the border lines of each first region are output, and at the same time, the first aesthetic score information corresponding to each first region is output. It can be seen that based on the final presented output effect, multiple aesthetic scores of the croppable borders are provided for the user's reference, so as to ensure that the cropped image selected by the user is relatively beautiful.

[0044] In the process of the image processing method in another embodiment of the present application, Step 120 includes:

[0045] Sub-step A1: Determine N2 second regions in the first image based on the composition model, where N2 is a positive integer.

[0046] Among them, the composition model is trained based on the first target image data.

[0047] Before this step, it is necessary to complete the preparation of the composition model so that in this step, N2 second regions in the first image can be determined through relevant calculations of the composition model.

[0048] See Figure 3For the flowchart shown, first, a public dataset for image cropping is collected and defined as the first target image data. Then, a large model is used as the model backbone to train the composition model. At this time, the labels of the composition model are the cropping box and the confidence level, where the confidence level label is 0 or 1, and the label format of the corresponding training set is: image - cropping box - type, that is, <imageA - x1, y1, x2, y2 - 1> or <imageA - x1, y1, x2, y2 - 0>. Among them, imageA is used to represent the name of a certain image sample for training, (x1, y1, x2, y2) is used to represent the coordinates of the cropping box in the image output by the composition model, 1 is used to indicate that there is a cropping box in the image, and 0 is used to indicate that there is no cropping box in the image.

[0049] It should be noted that it is default that the cropping box output by the composition model is a rectangle, and (x1, y1) and (x2, y2) are used to represent the coordinates of two opposite corners of the rectangle respectively.

[0050] For example, the data output by the composition model is: <image1 - 142, 0, 705, 426 - 1>, <image2 - 487, 94, 570, 165 - 1>.

[0051] See Figure 3 , the coordinates of the cropping box can be output by the locator, and "0" or "1" can be output by the classifier.

[0052] See Figure 3 , label 2 is used to represent an image sample in the first target image data. Thus, based on Figure 3 the process shown, a large number of images in the first target image data can be processed to obtain the output results corresponding to each image.

[0053] Among them, the first target image data includes: a large number of images in which cropping boxes are drawn for a certain feature information in the case where the participants are subjective factors. Therefore, taking these images as samples and through training, the composition model has the function of determining the cropping box in the image for a certain feature information. For example, the composition model determines the cropping box in the image for the portrait feature information; another example is that the composition model determines the cropping box in the image for the object feature information. It can be seen that through continuous training, the finally output cropping box image of the composition model can express the main content of the initial image and can be used in most image cropping scenarios.

[0054] In addition, see Figure 3 , before the locator and classifier process the features, there are also processes such as feature extraction and dimensionality reduction of the image, which will not be elaborated here.

[0055] Therefore, based on the above preparation of the composition model, in this embodiment, N2 second regions in the first image can be output by the composition model.

[0056] Sub-step A2: Determine N3 third regions in the first image based on the connection model, where N3 is a positive integer.

[0057] In this step, referring to Figure 4 the process shown, the connection model is based on the composition model, takes the first image 3 as input, and the locator outputs the box coordinates of N3 third regions.

[0058] It should be noted that the first region, the second region, and the third region in this application are all used to represent the cropping box.

[0059] Sub-step A3: When the overlap degree between the fourth region and at least one second region is greater than the first threshold, determine the fourth region as a first region, where the fourth region is a third region.

[0060] In this step, compare: the N3 third regions of the first image output by the connection model, and the N2 second regions of the first image output by the composition model.

[0061] As can be seen from the foregoing steps, the composition model is trained with a large number of images. Therefore, the accuracy of the N2 second regions of the first image output by the composition model is relatively high. Thus, when the matching rate between the N3 third regions of the first image output by the connection model and the N2 second regions of the first image output by the composition model is relatively high, some or all of the third regions of the first image output by the connection model can be used as the finally output first regions.

[0062] In this step, a comparison method is provided, that is, compare any one of the third regions (such as the fourth region) with the N2 second regions. If the overlap degree between the fourth region and at least one second region meets the condition of being greater than the first threshold, it is considered that the fourth region can be used as one of the finally output first regions. By analogy, N1 first regions can be output.

[0063] Among them, if the overlap degree between a certain third region and any one of the second regions cannot meet the condition of being greater than the first threshold, then this third region is excluded.

[0064] Correspondingly, N3≥N1.

[0065] It should be noted that this embodiment can be understood as the training process of the connection model. The training purpose is to make the matching rate between the N3 third regions output by the connection model and the N2 second regions output by the composition model relatively high, so as to complete the training. In subsequent applications, any image can be directly processed, and the third regions in the obtained image can be directly used as the first regions for output, without the need for a comparison process anymore.

[0066] In this embodiment, on the one hand, the composition model outputs N2 second regions of the first image. Since the composition model is a trained model, the second regions it outputs are more in line with the image cropping requirements. On the other hand, the connection model outputs N3 third regions of the first image, and the N3 third regions are compared with the N2 second regions to select some or all of the third regions that meet the comparison requirements in the N3 third regions as the first regions finally presented in the first image, so that the first regions output by the connection model in this application are more in line with the image cropping requirements.

[0067] In the process of the image processing method in another embodiment of this application, step 130 includes:

[0068] Sub-step B1: Determine the second aesthetic score information corresponding to the fifth region based on the aesthetic model, where the fifth region is a first region.

[0069] Among them, the aesthetic module is trained based on the second target image data.

[0070] Before this step, it is necessary to complete the preparation of the aesthetic model so that in this step, through relevant calculations of the aesthetic model, the second aesthetic score information corresponding to each first region can be determined.

[0071] In the preparation link, first collect the public dataset of picture ratings and define this dataset as the second target image data. For example, use the Comparative Photo Composition dataset, the aesthetic visual analysis dataset (referred to as the AVA dataset), and collect some news and video cover pictures, etc., and perform manual rating annotations according to the rules of the public dataset, with scores ranging from 0 to 10, so as to complete the preparation work of the second target image data. See Figure 5, then, take the image sample 2 in the second target image data as the input and input it into the aesthetics model. When trained by the regressor (sigmoid), it is mapped to the 0-1 interval. The corresponding training data format is: picture - aesthetics score: <imageA - B points>. Where imageA is used to represent the name of a certain image sample for training, and B points is used to represent the aesthetics score corresponding to this image sample.

[0072] Among them, the scores from 0 to 10 points have a one-to-one correspondence relationship in the 0-1 interval.

[0073] For example, the data output by the aesthetics model is: <image1 - 5.3 points>, <image2 - 3.1 points>.

[0074] Among them, the second target image data includes a large number of images with aesthetic scores marked under the condition that the participants are subjective factors. Therefore, taking these images as samples and through training, the aesthetics model has the function of determining the aesthetic score information in terms of aesthetic considerations such as the composition of the image.

[0075] In addition, refer to Figure 5 , before the regressor processes the features, it also includes processes such as feature extraction and dimensionality reduction of the image, which will not be elaborated here.

[0076] Sub-step B2: Based on the connection model, obtain a 1024-dimensional feature vector of the fifth region.

[0077] Sub-step B3: After performing a one-dimensional convolution operation on the 1024-dimensional feature vector of the fifth region, obtain a 256-dimensional feature vector of the fifth region.

[0078] Sub-step B4: Input the 256-dimensional feature vector of the fifth region into the objective function and output the third aesthetic score information corresponding to the fifth region.

[0079] In the above steps, refer to Figure 4 the shown process. In the connection model, combine the branches in the aesthetics model so that when taking the first image 3 as the input, after a series of processes, the obtained N1 first regions are used as the input of the regressor, and thus the regressor outputs the third aesthetic score information of the N1 first regions.

[0080] Refer to Figure 4, in combination with the processing of branches in the aesthetic model, after passing through the backbone, the connection model passes through a Region of Interest (ROI) extraction layer to extract the feature information of the cropping frame, and then transfers to the full connect (FC) layer 1 (1024 dimensions) and FC2 (256 dimensions) of the deep learning network and a regressor: where the operation from FC1 to FC2 is a one-dimensional convolution operation that compresses the 1024-dimensional feature vector into 256 dimensions for learning the features of the aesthetic score; then the 256-dimensional features are input into the regressor function, and the output is an aesthetic score, that is, the third aesthetic score information.

[0081] Sub-step B5: Calculate the loss function based on the second aesthetic score information and the third aesthetic score information.

[0082] In this step, calculate the loss function for the third aesthetic score information and the second aesthetic score information obtained from the aesthetic model.

[0083] Among them, use the second aesthetic score information as the ground truth that the model should output in the ideal state for calculating the loss function.

[0084] Sub-step B6: When the loss function is minimized, determine the third aesthetic score information as the first aesthetic score information corresponding to the fifth region.

[0085] In this step, compare the N1 third aesthetic score information of the first image output by the connection model and the N1 second aesthetic score information of the first image output by the aesthetic model until the loss function between the third aesthetic score information and the second aesthetic score information is minimized and the matching rate reaches the highest, and it is considered that the matching rate is relatively high.

[0086] As can be seen from the foregoing steps, the aesthetic model is trained with a large number of images. Therefore, the accuracy of the N1 second aesthetic score information of the first image output by the aesthetic model is relatively high. Thus, when the matching rate between the N1 third aesthetic score information of the first image output by the connection model and the N1 second aesthetic score information of the first image output by the aesthetic model is relatively high, the N1 third aesthetic score information of the first image output by the connection model can be used as the first aesthetic score information finally presented to the user.

[0087] Reference may be made to comparing the third aesthetic score information of any one of the first regions (such as the fifth region) with its corresponding second aesthetic score information. If the matching rate between the two satisfies the condition of being greater than the second threshold, it is considered that the third aesthetic score information of the fifth region can finally be presented to the user as the first aesthetic score information. Among them, the condition that the matching rate between the two satisfies being greater than the second threshold can correspond to that the loss between the third aesthetic score information and the second aesthetic score information reaches the minimum.

[0088] Among them, if the matching rate between the third aesthetic score information of a certain first region and its corresponding second aesthetic score information cannot satisfy the condition of being greater than the second threshold, the third aesthetic score information of this first region will be eliminated.

[0089] It should be noted that this embodiment can be understood as the training process of the connection model. The training purpose is to make the matching rate between the N1 third aesthetic score information output by the connection model and the N1 second aesthetic score information output by the aesthetic model relatively high, so as to complete the training. In subsequent applications, any image can be directly processed, and the third aesthetic score information of the obtained first region can be directly output as the first aesthetic score information without the need for a comparison process.

[0090] In this embodiment, on the one hand, the aesthetic model outputs the N1 second aesthetic score information of the first image. Since the aesthetic model is a trained model, the second aesthetic score information it outputs is more in line with the image aesthetic scoring standard. On the other hand, the connection model outputs the N1 third aesthetic score information of the first image, and the third aesthetic score information is compared with the corresponding second aesthetic score information to retain the third aesthetic score information with a relatively high matching rate as the first aesthetic score information finally presented to the user, so that the first aesthetic score information output by the connection model in this application is more in line with the image aesthetic requirements.

[0091] In another embodiment of the present application, refer to Figure 4 , which shows the general process of processing an image by a connection model formed by combining a composition model and an aesthetic model branch.

[0092] Optionally, for the connection model, the model body still adopts the intelligent cropping architecture, but it is replaced with a small network of resNet50 (a deep learning model with 50 convolutional layers) to improve the model operation efficiency. After detecting features and extracting regions of interest, 1024-dimensional feature values are obtained. One branch obtains the frame coordinates (one or many) of the cropping frame through a locator (softmax). Another branch has changed compared with the aesthetic model. A new fully connected layer with 256 dimensions of FC2 is added to learn the quality score features, and then a regressor is used for aesthetic score prediction, so that the cropping frame and the aesthetic score are output simultaneously.

[0093] Based on the connection model formed by the above combination, a training method of knowledge distillation can be adopted to complete the training of the model.

[0094] As Figure 6 shown, the input of the composition model 601 (the processing flow is shown in the first dashed box) and the connection model (the processing flow is shown in the second dashed box) is the same, which is the original image P. During the knowledge distillation process, the cropping frame a needs to be output as the true position label of the connection model.

[0095] Furthermore, taking the true position label of the connection model as the standard, the connection model outputs a cropping frame b that matches the cropping frame a.

[0096] The input of the aesthetics model 602 (the processing flow is shown in the third dashed box) is the picture P2, which is a small picture cropped from the original image P by the cropping frame b, and the output is the aesthetics score a, which serves as the true score label of the connection model. After passing through the backbone, the aesthetics model directly goes through FC, generally a 1024-dimensional feature vector, and then the score is output through the calculation of the regressor; while the connection model, after passing through the backbone, will also go through an ROI extraction layer to extract the feature information of the cropping frame, and then transfer to FC1 (1024 dimensions), FC2 (256 dimensions) and the regressor: among them, the operation from FC1 to FC2 is a one-dimensional convolution operation, which compresses the 1024-dimensional feature vector into 256 dimensions for learning the features of aesthetic scoring; then the 256-dimensional features are input into the regressor function, and the output is an aesthetics score, and the loss function is calculated with the ground truth of the aesthetics score obtained by the aesthetics model. Therefore, the connection model actually performs aesthetic scoring on the cropping frame image, rather than on the input original image P.

[0097] Among them, taking the true score label of the connection model as the standard, the connection model outputs an aesthetics score b that matches the aesthetics score a.

[0098] See Figure 6 , there are mainly two loss functions (loss) in the training of the connection model, namely the coordinate loss function (SiteLoss) and the scoring loss function (ScoreLoss). Among them, SiteLoss is mainly the intersection over union loss function (IoUloss) between the true box and the predicted box, and ScoreLoss is mainly the mean squared error loss function (MSELoss) between the true value and the predicted value of the scores between the quality scores. The total loss function TotalLoss = α * SiteLoss + β * ScoreLoss.

[0099] Referentially, the training process of the connection model is as follows:

[0100] The training image P passes through the connection model, and the main body of the resNet50 model to obtain a 16*16 feature layer (featureMap). Then, small 7*7*1024 blocks that may be the cropping frames are extracted through ROI, and then a one-dimensional vector of 1*1*1024 is obtained through a convolution (conv) operation. Subsequently, position and score predictions are performed: one branch directly performs position prediction to obtain the cropping frame b, and the other branch first performs conv calculation to compress to 256-dimensional features, and then obtains the aesthetic score b through the sigmoid regression function calculation. That is to say, the score b is the aesthetic score for the cropping frame, while the separate aesthetic model can only score the complete image; at the same time, the image P passes through the composition model to obtain the cropping frame a.

[0101] Among them, the training image P is cropped using the cropping frame b to obtain the cropped small image P2, and then passed through the aesthetic model to obtain the aesthetic score a.

[0102] The cropping frames a and b are substituted into the SiteLoss formula, and the aesthetic scores are substituted into the ScoreLoss formula. Finally, the hyperparameters α and β are adjusted to calculate the TotalLoss and perform model parameter update training.

[0103] In this embodiment, the composition model can output the frame coordinates (x1, y1, x2, y2) of the cropping frame and the probability value of the object category within the frame, and the aesthetic model can output the aesthetic score of the cropping frame. Since the cropping frame and the aesthetic score are not linearly related, that is, after the cropping frame changes, the change of the aesthetic score cannot be predicted. And the knowledge distillation method adopted in this embodiment can combine the two models to perfectly solve the correlation problem between the cropping frame and the aesthetic score, that is, use the aesthetic model to provide training labels for the connection model, so that the model training can be realized by modifying the output branch of the composition model.

[0104] In the process of the image processing method in another embodiment of this application, step 140 includes:

[0105] Sub-step C1: Output the border line corresponding to the first region in the first image according to the coordinate information of each first position point in the first region.

[0106] In this embodiment, the locator in the model can be used to obtain the coordinate information of each first position point in the first region, so as to output the border lines of each first region in the first image according to the coordinate information of these points.

[0107] Optionally, in this embodiment, the shape of the first region, that is, the shape of the cropping frame, can be preset, so that the connection between points can be performed according to the relative position relationship between each first position point in the first region, and finally the border line of the first region is formed.

[0108] Optionally, each first position point in the first region includes key points of the first region, such as points at all corner positions.

[0109] Optionally, with the upper left corner of the first image as the origin, a coordinate system is established, extending horizontally to the right of the origin as the positive direction of the X-axis, and extending vertically downward from the origin as the positive direction of the Y-axis.

[0110] In this embodiment, with reference to the coordinate information of each position point in the first region, border lines corresponding to different first regions are presented in the first image, enabling the user to clearly see each cropping frame in the first image.

[0111] In the process of the image processing method according to another embodiment of the present application, step 120 includes:

[0112] Sub-step D1: Determine N1 first regions in the first image according to the first shape and N4 preset ratios corresponding to the first shape, where N4 is a positive integer and N1≥N4.

[0113] Optionally, the first shape and the corresponding N4 preset ratios are automatically set by the model.

[0114] Optionally, the user inputs the first shape and the corresponding N4 preset ratios, so that the model can set the first shape and the corresponding N4 preset ratios according to the user input.

[0115] Further, based on the preset first shape, determine the first region in the first image according to the first shape.

[0116] For example, the first shape is a rectangle.

[0117] Further, based on the preset ratio, limit the first shape according to the preset ratio.

[0118] For example, the preset ratios include at least one of the length-width ratios of a rectangle: 1:1, 1:2, 2:1, 3:4, 4:3, 9:16, 16:9, etc.

[0119] Optionally, the number of preset ratios can be multiple, and each preset ratio can correspondingly output at least one first region.

[0120] In this embodiment, based on the shape and proportional dimensions of the preset cropping frame, more cropping frames can be output, facilitating subsequent application selection.

[0121] In more embodiments of the present application, in practical applications, when there is a need to call for image cropping, first, the connection model outputs the top m cropped images with their corresponding aesthetic scores; second, the caller filters according to the required image size type and selects n (m > n) images that meet the required size from the m cropped images; finally, based on the aesthetic scores of the n cropped images, the optimal cropped image is selected as the result image.

[0122] In summary, the existing image cropping methods lack research on compositional aesthetics. Aesthetic evaluation refers to using traditional image features or deep learning features for learning, simulating human subjective feelings, and finally assigning an aesthetic score to an image through an algorithm. Based on this, the present application utilizes knowledge distillation, that is, a learning method in which the teacher models (composition model and aesthetic model) provide high-quality labels to the student model (connection model), and trains the small model network through an excellent large model network, enabling the small model to also have recognition and detection capabilities close to those of the large model. Finally, based on the method for aesthetic composition cropping of images based on knowledge distillation provided by the present application, when an image is input, the top N cropping frames and the aesthetic scores corresponding to the cropping frames can be output to simultaneously achieve compositional cropping and aesthetic evaluation, providing multiple scale cropped image selections for the demand side. Common scenarios include providing diverse selections for web pages, news, and video cover images.

[0123] Among them, the improvement points of the present application are as follows: building a teacher model for compositional and aesthetic scoring tasks using a large model as the basis for training the connection model, rather than combining them to complete the cropping task; for ordinary compositional cropping networks, due to the difficulty of defining aesthetics, only a small number of high-quality cropped images can be used as labels, resulting in a relatively low reference significance for the confidence of the compositional network. In response to this, the knowledge distillation method can be used to train the connection model, using the cropping frames and scores output by the teacher model as the true labels for training the model, solving the label problem, and thus ensuring the accuracy of the connection model's cropping and scoring; providing diverse cropping size labels to facilitate the subsequent connection model to output results of multiple scales and provide diverse selections; only one algorithm service is required to meet multiple requirements on the Internet business side, saving machine resources and simplifying the service deployment logic.

[0124] It can be seen that the present application proposes a method for aesthetic composition cropping of images based on knowledge distillation. In the absence of large-scale manual simultaneous annotation of cropping frames and corresponding aesthetic quality scores, a knowledge distillation scheme is introduced. First, the large models for compositional cropping and aesthetic evaluation are learned separately, and then these two large models are used as teacher models to conduct supervised learning on compositional scoring. Finally, a model that can both intelligently crop and provide aesthetic scoring for cropped images is obtained, which can simultaneously output cropped images of multiple scales and corresponding aesthetic scores, providing diverse selections for the demand side.

[0125] Reference can be made. Based on the current serial solutions of multiple deep learning models, many of them can draw on the method of knowledge distillation to integrate multiple tasks into one task, which not only simplifies the service deployment logic but also saves machine resources.

[0126] In the image processing method provided by the embodiments of the present application, the execution subject can be an image processing device. In the embodiments of the present application, taking the image processing device executing the image processing method as an example, the image processing device provided by the embodiments of the present application is described.

[0127] Figure 7 The block diagram of the image processing device according to another embodiment of the present application is shown. The device includes:

[0128] An acquisition module 10, configured to obtain feature information of a first image based on a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model;

[0129] A first determination module 20, configured to determine N1 first regions in the first image according to the feature information, where N1 is a positive integer;

[0130] A second determination module 30, configured to determine first aesthetic score information corresponding to each first region;

[0131] An output module 40, configured to output border lines corresponding to each first region in the first image, and output first aesthetic score information corresponding to each first region.

[0132] In this way, in the embodiments of the present application, based on pre-training the connection model in advance, the first image can be input into the connection model. Thus, the connection model determines N1 first regions in the first image according to the feature information of the first image by using the algorithm of the composition model. The first regions can be used as the images to be retained after cropping. Further, after determining the N1 first regions, the connection model uses the algorithm of the aesthetic module branch to perform aesthetic scoring on each first region according to the N1 first regions to determine first aesthetic score information corresponding to each first region. Finally, in the first image, the border lines of each first region are output, and at the same time, the first aesthetic score information corresponding to each first region is output. It can be seen that based on the finally presented output effect, aesthetic scores of multiple croppable borders are provided for the user's reference, so as to ensure that the cropped image selected by the user is relatively beautiful.

[0133] Optionally, the first determination module 20 includes:

[0134] A first determination unit, configured to determine N2 second regions in the first image based on the composition model, where N2 is a positive integer;

[0135] A second determination unit, configured to determine N3 third regions in the first image based on connections, where N3 is a positive integer;

[0136] A third determination unit, configured to determine the fourth region as a first region when the overlap degree between the fourth region and at least one second region is greater than a first threshold, where the fourth region is a third region;

[0137] Wherein, the composition model is trained based on first target image data.

[0138] Optionally, the second determination module 30 includes:

[0139] A fourth determination unit, configured to determine second aesthetic score information corresponding to each first region based on an aesthetic model, where the fifth region is a first region;

[0140] An acquisition unit, configured to acquire a 1024-dimensional feature vector of the fifth region based on a connection model;

[0141] An operation unit, configured to obtain a 256-dimensional feature vector of the fifth region after performing a one-dimensional convolution operation on the 1024-dimensional feature vector of the fifth region;

[0142] A first output unit, configured to input the 256-dimensional feature vector of the fifth region into an objective function and output third aesthetic score information corresponding to the fifth region;

[0143] A calculation unit, configured to calculate a loss function based on the second aesthetic score information and the third aesthetic score information;

[0144] A fifth determination unit, configured to determine the third aesthetic score information as the first aesthetic score information corresponding to the fifth region when the loss function is minimized;

[0145] Wherein, the aesthetic model is trained based on second target image data.

[0146] Optionally, the output module 40 includes:

[0147] A second output unit, configured to output a border line corresponding to the first region in the first image according to the coordinate information of each first position point of the first region.

[0148] Optionally, the first determination module 20 includes:

[0149] A sixth determination unit, configured to determine N1 first regions in the first image according to a first shape and N4 preset ratios corresponding to the first shape, where N4 is a positive integer and N1 ≥ N4.

[0150] The image processing device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0151] The image processing device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0152] The image processing device provided in the embodiments of the present application can implement each process implemented in the above method embodiments. To avoid repetition, it will not be elaborated here.

[0153] Optionally, as Figure 8 shown, the embodiments of the present application further provide an electronic device 100, including a processor 101, a memory 102, a program or instruction stored on the memory 102 and executable on the processor 101. When the program or instruction is executed by the processor 101, it implements each step of any of the above image processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0154] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0155] Figure 9 Schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0156] The electronic device 1000 includes, but is not limited to, components such as a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010, etc.

[0157] Those skilled in the art can understand that the electronic device 1000 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1010 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 9 The structure of the electronic device shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0158] Among them, the processor 1010 is used to obtain the feature information of the first image based on a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model; according to the feature information, determine N1 first regions in the first image, where N1 is a positive integer; determine the first aesthetic score information corresponding to each of the first regions; output the border lines corresponding to each of the first regions in the first image, and output the first aesthetic score information corresponding to each of the first regions.

[0159] In this way, in the embodiment of the present application, based on pre-training the connection model, the first image can be input into the connection model. Thus, the connection model determines N1 first regions in the first image according to the feature information of the first image, using the algorithm of the composition model. The first regions can be used as the images to be retained after cropping. Further, after determining the N1 first regions, the connection model respectively performs aesthetic scoring on each of the first regions according to the N1 first regions, using the algorithm of the aesthetic module branch, to determine the first aesthetic score information corresponding to each of the first regions. Finally, in the first image, the border lines corresponding to each of the first regions are output, and at the same time, the first aesthetic score information corresponding to each of the first regions is output. It can be seen that based on the final presented output effect, multiple aesthetic scores of the croppable borders are provided for the user's reference, so as to ensure that the cropped image selected by the user is relatively beautiful.

[0160] Optionally, the processor 1010 is further configured to determine N2 second regions in the first image based on the composition model, where N2 is a positive integer; determine N3 third regions in the first image based on the connection model, where N3 is a positive integer; when the overlap degree between the fourth region and at least one of the second regions is greater than a first threshold, determine the fourth region as one of the first regions, and the fourth region is one of the third regions; wherein, the composition model is trained based on first target image data.

[0161] Optionally, the processor 1010 is further configured to determine second aesthetic score information corresponding to a fifth region based on the aesthetic model, where the fifth region is one of the first regions; obtain a 1024-dimensional feature vector of the fifth region based on the connection model; after performing a one-dimensional convolution operation on the 1024-dimensional feature vector of the fifth region, obtain a 256-dimensional feature vector of the fifth region; input the 256-dimensional feature vector of the fifth region into an objective function, and output third aesthetic score information corresponding to the fifth region; calculate a loss function based on the second aesthetic score information and the third aesthetic score information; when the loss function is minimized, determine the third aesthetic score information as the first aesthetic score information corresponding to the fifth region; wherein, the aesthetic module is trained based on second target image data.

[0162] Optionally, the processor 1010 is further configured to output a border line corresponding to the first region in the first image according to the coordinate information of each first position point of the first region.

[0163] Optionally, the processor 1010 is further configured to determine N1 first regions in the first image according to a first shape and N4 preset ratios corresponding to the first shape, where N4 is a positive integer and N1≧N4.

[0164] In summary, the existing image cropping methods lack research on compositional aesthetics. Aesthetic evaluation refers to using traditional picture features or deep learning features for learning, simulating human subjective feelings, and finally giving a picture an aesthetic score through an algorithm. Based on this, this application uses knowledge distillation, that is, a learning method in which a teacher model (composition model and aesthetic model) provides high-quality labels to a student model (connection model), and trains a small model network through an excellent large model network, so that the small model can also have an identification and detection ability close to that of the large model. Finally, based on the picture aesthetic composition cropping method based on knowledge distillation provided by this application, when inputting a picture, the first N cropping frames and the aesthetic scores corresponding to the cropping frames can be output to simultaneously achieve compositional cropping and aesthetic evaluation, providing multiple scale cropped picture selections for the demand side. Common scenarios include providing diversified selections for web pages, news, and video cover pictures.

[0165] Among them, the improvement of this application lies in: building a teacher model for composition and aesthetic scoring tasks using a large model as the basis for connecting model training, rather than combining it to complete the cropping task; for an ordinary composition cropping network, due to the difficulty of aesthetic definition, only a small number of high-quality cropped images can be used as labels, resulting in a relatively low reference significance for the confidence of the composition network. In response to this, the knowledge distillation method can be used for connecting model training, taking the cropping frames and scores output by the teacher model as the true labels of the training model, solving the label problem, and thus ensuring the accuracy of the cropping and scoring of the connecting model; providing diverse cropping size labels to facilitate the subsequent connecting model to output results of multiple scales and provide diverse choices; only one algorithm service is required to meet multiple requirements on the Internet business side, saving machine resources and simplifying the service deployment logic.

[0166] It can be seen that this application proposes a method for picture aesthetic composition cropping based on knowledge distillation. In the case of the lack of large-scale manual simultaneous annotation of cropping frames and corresponding aesthetic quality scores, a knowledge distillation scheme is introduced. First, the composition cropping and aesthetic evaluation large models are learned separately, and then these two large models are used as teacher models to conduct supervised learning on composition scoring. Finally, a model that can not only perform intelligent cropping but also provide aesthetic scoring for cropped pictures is obtained, which can simultaneously output cropped pictures of multiple scales and corresponding aesthetic scores, providing diverse choices for the demand side.

[0167] It should be understood that in the embodiments of this application, the input unit 1004 may include a Graphics Processing Unit (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes the image data of static pictures or video images obtained by an image capture device (such as a camera) in the video image capture mode or the image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an action lever, which will not be elaborated here. The memory 1009 can be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 1010 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1010.

[0168] The memory 1009 can be used to store software programs and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 may include a volatile memory or a non-volatile memory, or the memory 1009 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0169] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1010 either.

[0170] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the image processing method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0171] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0172] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiment of the image processing method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0173] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0174] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above embodiment of the image processing method and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0175] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be executed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0176] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0177] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining feature information of a first image based on a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model; Determining N1 first regions in the first image according to the feature information, where N1 is a positive integer; Determining first aesthetic score information corresponding to each of the first regions; Outputting border lines corresponding to each of the first regions in the first image, and outputting first aesthetic score information corresponding to each of the first regions; The determining N1 first regions in the first image according to the feature information includes: Determining N2 second regions in the first image based on the composition model, where N2 is a positive integer; Determining N3 third regions in the first image based on the connection model, where N3 is a positive integer; When the degree of overlap between a fourth region and at least one of the second regions is greater than a first threshold, determining the fourth region as one of the first regions, where the fourth region is one of the third regions; Wherein, the composition model is trained based on first target image data.

2. The method according to claim 1, wherein The determining first aesthetic score information corresponding to each of the first regions includes: Determining second aesthetic score information corresponding to a fifth region based on the aesthetic module, where the fifth region is one of the first regions; Obtaining a 1024-dimensional feature vector of the fifth region based on the connection model; After performing a one-dimensional convolution operation on the 1024-dimensional feature vector of the fifth region, obtaining a 256-dimensional feature vector of the fifth region; Inputting the 256-dimensional feature vector of the fifth region into an objective function, and outputting third aesthetic score information corresponding to the fifth region; Calculating a loss function based on the second aesthetic score information and the third aesthetic score information; When the loss function is minimized, determining the third aesthetic score information as the first aesthetic score information corresponding to the fifth region; Wherein, the aesthetic module is trained based on second target image data.

3. The method according to claim 1, wherein The outputting border lines corresponding to each of the first regions in the first image includes: Outputting the border line corresponding to the first region in the first image according to the coordinate information of each first position point of the first region.

4. The method according to claim 1, characterized in that The determining N1 first regions in the first image includes: Determining N1 first regions in the first image according to a first shape and N4 preset ratios corresponding to the first shape, where N4 is a positive integer and N1≥N4.

5. An image processing apparatus, characterized in that, The apparatus includes: An obtaining module, configured to obtain feature information of a first image based on a connection model, where the connection model is obtained by connecting an aesthetic module branch on the basis of a composition model; A first determining module, configured to determine N1 first regions in the first image according to the feature information, where N1 is a positive integer; A second determining module, configured to determine first aesthetic score information corresponding to each of the first regions; An output module, configured to output border lines corresponding to each of the first regions in the first image, and output first aesthetic score information corresponding to each of the first regions; The first determination module includes: A first determination unit, configured to determine N2 second regions in the first image based on the composition model, where N2 is a positive integer; A second determination unit, configured to determine N3 third regions in the first image based on the connection model, where N3 is a positive integer; A third determination unit, configured to determine the fourth region as one of the first regions when the degree of overlap between the fourth region and at least one of the second regions is greater than a first threshold, where the fourth region is one of the third regions; Wherein, the composition model is trained based on first target image data.

6. The device according to claim 5, characterized in that, The second determination module includes: A fourth determination unit, configured to determine second aesthetic score information corresponding to each fifth region based on the aesthetics module, where the fifth region is one of the first regions; An acquisition unit, configured to acquire a 1024-dimensional feature vector of the fifth region based on the connection model; An operation unit, configured to obtain a 256-dimensional feature vector of the fifth region after performing a one-dimensional convolution operation on the 1024-dimensional feature vector of the fifth region; A first output unit, configured to input the 256-dimensional feature vector of the fifth region into an objective function and output third aesthetic score information corresponding to the fifth region; A calculation unit, configured to calculate a loss function based on the second aesthetic score information and the third aesthetic score information; A fifth determination unit, configured to determine the third aesthetic score information as the first aesthetic score information corresponding to the fifth region when the loss function is minimized; Wherein, the aesthetics module is trained based on second target image data.

7. The device according to claim 5, wherein The output module includes: A second output unit, configured to output a border line corresponding to the first region in the first image according to the coordinate information of each first position point of the first region.

8. The device according to claim 5, characterized in that, The first determination module includes: A sixth determination unit, configured to determine N1 first regions in the first image according to a first shape and N4 preset ratios corresponding to the first shape, where N4 is a positive integer and N1≥N4.

9. An electronic device, characterized in that, It includes a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the image processing method according to any one of claims 1 to 4 are implemented.

10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the image processing method according to any one of claims 1 to 4 are implemented.