Image processing method and device, electronic equipment and storage medium

By determining and utilizing the shooting style of objects in the image for composition, the problems of low efficiency and lack of targeted image composition in the prior art are solved, and a more efficient and high-quality composition effect is achieved.

CN120070656APending Publication Date: 2025-05-30BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311610678.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has low efficiency and lacks targeting in image post-processing, especially in composition processing, making it difficult to effectively utilize the shooting characteristics of the object in the initial image.

Method used

By obtaining the shooting styles of objects in the initial image, including shooting angles and crop positions, these styles are determined using image recognition models (such as deep learning networks or multi-task models), and composing based on these styles, adjusting the position and crop positions of objects in the target image, and adding presets to complete composition.

Benefits of technology

It improves the efficiency and pertinence of image composition, can generate more reasonable and high-quality composition images, and avoids the tedious process of trying different template designs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070656A_ABST
    Figure CN120070656A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring an initial image of an object; determining a shooting style of an object in the initial image; and performing composition on the object based on the shooting style of the object to obtain a target image after composition. Through the method, the composition efficiency can be improved, the composition is more targeted, and a more reasonable high-quality composition can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image processing method, an apparatus, an electronic device, and a storage medium. Background Art

[0002] With the vigorous development of computer technologies, post-processing of images is quite common. For example, in daily life, a user may need to add elements such as the sun or add text descriptions to an image; in commercial advertisements, it is also necessary to process product images to generate some product advertisement images, so as to attract consumers through advertising creativity and promote brand development. Therefore, it is particularly important to be able to efficiently perform post-composition processing on images. Summary of the Invention

[0003] The present disclosure provides an image processing method, an apparatus, an electronic device, and a storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:

[0005] Obtaining an initial image of an object;

[0006] Determining a shooting style of the object in the initial image;

[0007] Composing the object based on the shooting style of the object to obtain a target image after composition.

[0008] In some embodiments, the composing the object based on the shooting style of the object to obtain a target image after composition includes:

[0009] Based on the shooting style of the object, determining a target position of the object in the target image;

[0010] Setting the object based on the target position, and adding preset content to positions other than the target position in the target image to obtain the target image after composition.

[0011] In some embodiments, the shooting style includes: a shooting angle and / or a cropping position of the object, and the determining a target position of the object in the target image based on the shooting style of the object includes:

[0012] Based on the shooting angle of the object, determining a target area of the object in the target image;

[0013] And / or,

[0014] Based on the cropping position of the object, determining a target cropping position of the object in the target image.

[0015] In some embodiments, composing the object based on the shooting style of the object to obtain a target image after composition includes:

[0016] Composing the object at the target position while maintaining the shooting angle of the object to obtain the target image after composition.

[0017] In some embodiments, composing the object based on the shooting style of the object to obtain a target image after composition includes:

[0018] Composing the object at the target cropping position to obtain the target image after composition; wherein, the target cropping position is the cropping position of the object in the initial image.

[0019] In some embodiments, determining the shooting style of the object in the initial image includes:

[0020] Determining the shooting style of the object in the initial image based on an image recognition model; wherein, the image recognition model is obtained by training a deep learning network.

[0021] In some embodiments, the image recognition model is trained from an initial multi-task model, and the tasks include a shooting angle recognition task and a cropping position recognition task. The method further includes:

[0022] Obtaining a training sample data set; wherein, the training sample data set includes a plurality of training samples related to each task, and labels of each task corresponding to each training sample;

[0023] For each training sample, using the initial multi-task model to determine the prediction results of each training sample for each task;

[0024] For each task, calculating the loss value of the task based on the prediction results of each training sample for the task and the labels of each training sample for the task;

[0025] Determining the total loss of the initial multi-task model based on the loss values of each task;

[0026] Adjusting the parameters of the initial multi-task model based on the total loss to obtain the trained image recognition model.

[0027] According to a second aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, including:

[0028] A first acquisition module configured to acquire an initial image of an object;

[0029] A first determination module configured to determine the shooting style of the object in the initial image;

[0030] A composition module configured to compose the object based on the shooting style of the object to obtain a target image after composition.

[0031] In some embodiments, the composition module is further configured to determine a target position of the object in the target image based on the shooting style of the object; set the object based on the target position, and add preset content to positions other than the target position in the target image to obtain the target image after composition.

[0032] In some embodiments, the shooting style includes: a shooting angle and / or a cropping position of the object. The first determination module is further configured to determine a target area of the object in the target image based on the shooting angle of the object; and / or determine a target cropping position of the object in the target image based on the cropping position of the object.

[0033] In some embodiments, the composition module is further configured to compose the object while maintaining the shooting angle of the object at the target position to obtain the target image after composition.

[0034] In some embodiments, the composition module is further configured to compose the object with the target cropping position to obtain the target image after composition; wherein the target cropping position is the cropping position of the object in the initial image.

[0035] In some embodiments, the first determination module is further configured to determine the shooting style of the object in the initial image based on an image recognition model; wherein the image recognition model is obtained by training a deep learning network.

[0036] In some embodiments, the image recognition model is trained from an initial multi-task model. The tasks include a shooting angle recognition task and a cropping position recognition task. The apparatus further includes:

[0037] A second acquisition module configured to acquire a training sample data set; wherein the training sample data set includes a plurality of training samples related to each task, and labels of each task corresponding to each training sample;

[0038] A second determination module configured to, for each training sample, use the initial multi-task model to determine the prediction result of each training sample for each task;

[0039] A calculation module configured to, for each task, calculate the loss value of the task based on the prediction result of each training sample for the task and the label of each training sample for the task;

[0040] A third determination module, configured to determine the total loss of the initial multi-task model based on the loss value of each task;

[0041] A parameter tuning module, configured to adjust the parameters of the initial multi-task model based on the total loss to obtain the trained image recognition model.

[0042] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0043] A processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in the first aspect above.

[0044] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium, including:

[0045] When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in the first aspect above.

[0046] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0047] In the embodiments of the present disclosure, the electronic device determines the shooting style of the object in the initial image and composes the picture autonomously based on the shooting style of the object, without trying different template designs, thus improving the efficiency of picture composition; in addition, since the electronic device composes the picture based on the shooting style of the object, that is, composes the picture by using the shooting characteristics of the object in the initial image, the picture composition is more targeted and a more reasonable and high-quality picture composition can be obtained.

[0048] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0050] Figure 1 It is a flowchart example of an image processing method shown in the embodiments of the present disclosure.

[0051] Figure 2 It is an example of an initial image including an object in the embodiments of the present disclosure Figure 1 .

[0052] Figure 3 It is an example of an initial image including an object in the embodiments of the present disclosure Figure 2 .

[0053] Figure 4 In the embodiments of the present disclosureFigure 2 An example of the target image after corresponding composition Figure 1 。

[0054] Figure 5 In the embodiments of the present disclosure Figure 3 An example of the target image after corresponding composition Figure 2 。

[0055] Figure 6 It is an example diagram of a rotation angle in the embodiments of the present disclosure.

[0056] Figure 7 It is an example diagram of an inclination angle in the embodiments of the present disclosure.

[0057] Figure 8 It is an example diagram of determining a target area based on an object's shooting angle in the embodiments of the present disclosure.

[0058] Figure 9 It is an example diagram of determining a target cropping position based on an object's cropping position in the embodiments of the present disclosure.

[0059] Figure 10 It is an example of the target image after composition in the embodiments of the present disclosure Figure 1 。

[0060] Figure 11 It is an example of the target image after composition in the embodiments of the present disclosure Figure 2 。

[0061] Figure 12 It is an example diagram of the training architecture of the image recognition model in the embodiments of the present disclosure.

[0062] Figure 13 It is a diagram of an image processing device shown according to an exemplary embodiment.

[0063] Figure 14 It is a block diagram of an electronic device shown in the embodiments of the present disclosure. Detailed implementation manners

[0064] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0065] Embodiments of the present disclosure provide an image processing method, and the execution subject thereof may be an electronic device capable of running an application, such as a user equipment (UE), a mobile device, a user terminal, a mobile phone, a tablet computer, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the image processing method may be implemented by a processor in an electronic device calling computer-readable instructions stored in a memory.

[0066] Figure 1 is a flowchart example of an image processing method shown in embodiments of the present disclosure. As can be seen from Figure 1 it, the method includes the following steps:

[0067] S11. Obtain an initial image of an object;

[0068] S12. Determine the shooting style of the object in the initial image;

[0069] S13. Compose the object based on the shooting style of the object to obtain a target image after composition.

[0070] In step S11 of the embodiments of the present disclosure, an electronic device obtains an initial image of an object. The object may be a single object such as a person, an animal, or an item, or a combined object. For example, the object is a mobile phone or a headset, the object is a combination of a mobile phone and a person, or the object is a combination of a person and a puppy, etc. The initial image of the object may be a daily life image taken by a user, or an image for a specific purpose, such as a product image of a product such as a mobile phone or a headset.

[0071] Figure 2 is an example of an initial image of an object in embodiments of the present disclosure Figure 1 The figure includes a combined object of a mobile phone and a human finger; Figure 3 is an example of an initial image of an object in embodiments of the present disclosure Figure 2 The figure includes a single object headset.

[0072] In step S12, the electronic device determines the shooting style of the object in the initial image. Among them, the shooting style of the object may include the shooting angle of the object, whether there is cropping and the cropping position of the object, and may also include the proportion of the object in the initial image, the color of the object, etc.

[0073] In the embodiments of the present disclosure, an electronic device may determine the shooting style of an object based on traditional image processing methods. For example, the initial image is subjected to image segmentation and detection, and the type of the object, the edge position of the object, etc. are determined based on the detection and segmentation results. Then, based on the preset rules corresponding to the object, the shooting angle, the cropping position, the screen ratio, etc. of the object are determined. Exemplarily, for a mobile phone, which is usually of a type with a size within a certain range and is flat, based on some judgment rules corresponding to the mobile phone, the shooting style of the mobile phone can be determined; for another example, for a human face, which also has fixed facial feature distribution characteristics, based on some judgment rules of the human face, the shooting style of the human face can also be determined.

[0074] In addition, in the embodiments of the present disclosure, the electronic device may also determine the shooting style of the object based on a trained image recognition model. For example, the model can be trained based on a deep learning network.

[0075] In step S13, the electronic device composes the object based on the shooting style of the object. For example, according to the shooting angle of the object, the area and / or shooting angle of the object in the target image after composition can be determined; according to the cropping position of the object, the cropping position of the object in the target image after composition can be determined; according to the screen ratio of the object in the initial image, the screen ratio of the object in the target image can be adjusted; according to the color of the object, the color of the content other than the object in the target object can be determined, etc.

[0076] It should be noted that in the embodiments of the present disclosure, when the electronic device composes the object based on the shooting style of the object, it may obtain a single target image after composition, or may obtain multiple target images after composition. Taking the generation of a mobile phone advertisement picture as an example, the electronic device may determine the area of the object in the target image after composition based on the shooting angle of the mobile phone, and add information such as text or logo in the area outside the mobile phone area; alternatively, the electronic device may perform an affine transformation on the mobile phone in the initial image based on the shooting angle of the mobile phone to obtain a target image with a preset tilt angle and / or rotation angle; or, design the color of the advertisement background according to the color of the mobile phone, etc.

[0077] In addition, in the embodiments of the present disclosure, if the object is a combined object, the electronic device may also determine the main object according to the screen ratio of each object in the combined object. For example, the object with the largest screen ratio is determined as the main object, and the foregoing method is used to compose the main object.

[0078] In the related art, when performing post-processing on an image, such as when designing an advertisement picture, it is necessary to place the product at different positions in each preset design template and try different composition designs in the design templates, and then manually select a satisfactory composition, which is a cumbersome process and has low efficiency.

[0079] In the embodiments of the present disclosure, the electronic device determines the shooting style of the object in the initial image and composes the picture autonomously based on the shooting style of the object, without attempting different template designs, thus improving the efficiency of picture composition. In addition, since the electronic device composes the picture based on the shooting style of the object, that is, composes the picture by using the shooting characteristics of the object in the initial image, the picture composition is more targeted and a more reasonable and high-quality picture composition can be obtained.

[0080] In some embodiments, composing the picture of the object based on the shooting style of the object to obtain the target image after composition includes:

[0081] Determining the target position of the object in the target image based on the shooting style of the object;

[0082] Setting the object based on the target position, and adding preset content to the positions other than the target position in the target image to obtain the target image after composition.

[0083] In the embodiments of the present disclosure, when the electronic device composes the picture of the object based on the shooting style of the object, it first determines the target position of the object in the target image based on the shooting style of the object. Among them, the target position includes the area occupied by the object in the target image, and may also include the cropping position of the object in the target image, etc. After determining the target position of the object, the electronic device sets the object according to the target position during picture composition, and adds content such as text and logo to the positions other than the target position, so as to obtain the target image after composition.

[0084] Figure 4 For the embodiments of the present disclosure Figure 2 The corresponding example of the target image after composition Figure 1 , Figure 5 For the embodiments of the present disclosure Figure 3 The corresponding example of the target image after composition Figure 2 , such as Figure 4 and Figure 5 shown, the original angle of the object is retained in the image after composition, but the cropping position of the hand is adjusted in Figure 4 , and the position of the earphone in the image is adjusted in Figure 5 . In addition, some text information is added to the positions other than the target position where the object is located.

[0085] It can be understood that in the embodiments of the present disclosure, the electronic device first determines the target position of the object in the target image based on the shooting style of the object and then sets the object, and then adds preset content to the positions other than the target position to obtain the target image after composition, that is, realizes picture composition through the layout of the position relationship. The scheme is simple and the picture composition efficiency is high.

[0086] In some embodiments, the shooting style includes: the shooting angle and / or the cropping position of the object. Determining the target position of the object in the target image based on the shooting style of the object includes:

[0087] Determining the target area of the object in the target image based on the shooting angle of the object;

[0088] And / or,

[0089] Determining the target cropping position of the object in the target image based on the cropping position of the object.

[0090] In the embodiments of the present disclosure, the electronic device can determine the target area of the object in the target image based on the shooting angle of the object; wherein, the shooting angle can include the rotation angle and / or the tilt angle of the object. Taking a mobile phone as an example, the rotation angle can be the angle formed after rotating clockwise or counterclockwise around a certain central axis of the mobile phone, and the tilt angle can be expressed as the degree of tilt of the mobile phone relative to the horizon.

[0091] Figure 6 FIG. is an example diagram of a rotation angle in the embodiments of the present disclosure, Figure 7 FIG. is an example diagram of a tilt angle in the embodiments of the present disclosure. As Figure 6 shown, the angle marked by β is the rotation angle, which is the angle formed after the mobile phone rotates counterclockwise along the vertical central axis; as Figure 7 shown, the angle marked by α is the tilt angle, which is the angle formed after the mobile phone tilts counterclockwise.

[0092] It should be noted that in the embodiments of the present disclosure, the mapping relationship between the shooting angle and the target area can be stored in the electronic device. For example, when it is determined that the mobile phone rotates clockwise along the vertical central axis, or tilts clockwise by 0 to 70 degrees, and the mobile phone is not tilted (90° in the vertical direction), the corresponding target area of the mobile phone is the right area in the target image; when it is determined that the mobile phone rotates counterclockwise by 0 to 70° or tilts counterclockwise by 0 to 70 degrees, the corresponding target area of the mobile phone is the left area in the target image; and when it is determined that the mobile phone is in the horizontal direction (180°), the corresponding target area of the mobile phone is the lower area in the target image.

[0093] In addition, in the embodiments of the present disclosure, the electronic device can also determine the target cropping position of the object in the target image based on the cropping position of the object. As above Figure 2 and Figure 4 , there is a cropping of the hand in the original Figure 2 , and in the Figure 4 target image after composition, the cropping of the hand is still retained, but the cropping position is offset towards the mobile phone side relative to the original Figure 2 .

[0094] It should be noted that in the embodiments of the present disclosure, the mapping relationship between the cutting position and the target cutting position can also be stored in the electronic device. For example, the target cutting position is the same as the initial cutting position (i.e., the cutting position of the object in the initial image), or the target cutting position is offset by a predetermined distance based on the initial cutting position, or if the initial cutting position is a right cut, the target cutting position is an upper cut and a right cut, etc. In addition, if the object is not cut in the initial image, the object may not be cut in the target image either.

[0095] In the embodiments of the present disclosure, considering that the visual effects presented by different shooting angles of the object and / or the cutting position of the object are different, then different positions or cutting positions of the object in the target image after composition will also bring different visual effects. Therefore, in the embodiments of the present disclosure, determining the target area of the object in the target image based on the shooting angle of the object, and the target cutting position of the initial cutting position of the object in the target image can obtain a target image with better visual effects and improve the composition quality.

[0096] As described above, text can be added to the positions outside the target position in the target image to obtain the target image after composition. Exemplarily, Figure 8 FIG. is an exemplary diagram for determining the target area based on the shooting angle of the object in the embodiments of the present disclosure. Figure 9 FIG. is an exemplary diagram for determining the target cutting position based on the cutting position of the object in the embodiments of the present disclosure.

[0097] As Figure 8 shown, the product diagram marked with L0 is the initial image of the embodiments of the present disclosure, and the algorithm-recognized product angle marked with L1 is the shooting angle of the object in the initial image determined by the electronic device based on the algorithm. If the shooting angle corresponds to the angle range of the object marked with L3 / L4 / L5, the object is set in the right area of the target image, and text is added in the left area, so as to obtain the target image with left text and right picture marked with L9; if the shooting angle corresponds to the angle of the object marked with L6, the object is set in the lower area of the target image, and text is added in the upper area, so as to obtain the target image with upper text and lower picture marked with L10; if the shooting angle corresponds to the angle of the object marked with L7 / L8, the object is set in the left area of the target image, and text is added in the right area, so as to obtain the target image with right text and left picture marked with L11.

[0098] As Figure 9As shown, the algorithm identified by L2 determines whether there is a crop, that is, the electronic device determines the crop position of the object in the initial image based on the algorithm. If the object in the initial image is a top-edge crop identified by L12, then the object is also a top-edge crop in the target image, so as to obtain the target image with the crop edge sucking the top identified by L16; if the object in the initial image is a bottom-edge crop identified by L13, then the object is also a bottom-edge crop in the target image, so as to obtain the target image with the crop edge sucking the bottom identified by L17; if the object in the initial image is a left-edge crop identified by L14, then the object is also a left-edge crop in the target image, so as to obtain the target image with the crop edge sucking the left part identified by L18; if the object in the initial image is a right-edge crop identified by L15, then the object is also a right-edge crop in the target image, so as to obtain the target image with the crop edge sucking the right part identified by L19.

[0099] In some embodiments, composing the object based on the shooting style of the object to obtain the composed target image includes:

[0100] Composing the object at the target position while maintaining the shooting angle of the object to obtain the composed target image.

[0101] In some embodiments, composing the object based on the shooting style of the object to obtain the composed target image includes:

[0102] Composing the object with the target crop position to obtain the composed target image; wherein, the target crop position is the crop position of the object in the initial image.

[0103] In the embodiments of the present disclosure, when composing an object, the original shooting angle and / or crop position of the object in the initial image are maintained. On the basis of reasonably setting the target position of the object and adding preset content for composition at positions other than the target position, the shooting characteristics of the object in the initial image are maintained as much as possible. For example, for an advertisement picture of a product, the product department may be more familiar with the characteristics of the product and thus make a more targeted shooting to obtain the initial image, while the advertisement department needs to compose the product to obtain the target image. The advertisement composition focuses on the style of the advertisement. Therefore, while composing the advertisement, maintaining the shooting characteristics of the product as much as possible can make the composed target image better display the product while not losing the beauty of the style. It can be understood that through this method, a target image with high-quality composition can be obtained.

[0104] Figure 10 is an example of the composed target image in the embodiments of the present disclosure Figure 1 , such as Figure 10As shown, the target image marked with M1 is obtained by composing a mobile phone that is rotated 45° clockwise along the central axis and has no cropping. This target image is the left-side view of the product with text on the left and a picture on the right; the target image marked with M2 is obtained by composing a mobile phone that is rotated 45° counterclockwise along the central axis and has no cropping. This target image is the right-side view of the product with text on the right and a picture on the left; the target image marked with M3 is obtained by composing a mobile phone that has no cropping at 180°. This target image is the top view of the product with text on the top and a picture on the bottom; the target image marked with M4 is obtained by composing a mobile phone that has no cropping at 90°. This target image is the front view of the product with text on the left and a picture on the right.

[0105] Figure 11 These are examples of the target images after composition in the embodiments of the present disclosure Figure 2 , such as Figure 11 As shown, the target image marked with M5 is obtained by composing a mobile phone that is tilted 45° clockwise and has its lower edge cropped. This target image is the right-tilted view of the product with text on the left and a picture on the right and the cropped edge adsorbed to the bottom; the target image marked with M6 is obtained by composing a mobile phone that is tilted 45° counterclockwise and has its lower edge cropped. This target image is the left-tilted view of the product with text on the right and a picture on the left and the cropped edge adsorbed to the bottom.

[0106] In some embodiments, determining the shooting style of the object in the initial image includes:

[0107] Determining the shooting style of the object in the initial image based on an image recognition model; wherein, the image recognition model is obtained by training a deep learning network.

[0108] As mentioned above, the electronic device can determine the shooting style of the object based on the trained image recognition model. For example, based on the training sample data and the label value, after training and tuning parameters of networks such as Convolutional Neural Networks (CNN) and Deep Neural Networks (DNN), the image recognition model is obtained, where the label value is the shooting style of the object in the training sample. Based on the trained image recognition model, inputting the initial image into the model can obtain the shooting style of the object in the initial image.

[0109] It can be understood that the method of determining the shooting style of the object in the initial image based on the trained image recognition model is simple and effective.

[0110] In some embodiments, the image recognition model is trained from an initial multi-task model, and the tasks include a shooting angle recognition task and a cropping position recognition task. The method further includes:

[0111] Obtain a training sample dataset; wherein, the training sample dataset includes a plurality of training samples related to each task, and labels of each task corresponding to each training sample;

[0112] For each training sample, use the initial multi-task model to determine the prediction results of each training sample for each task;

[0113] For each task, based on the prediction results of each training sample for the task and the labels of each training sample for the task, calculate the loss value of the task;

[0114] Based on the loss values of each task, determine the total loss of the initial multi-task model;

[0115] Based on the total loss, adjust the parameters of the initial multi-task model to obtain the trained image recognition model.

[0116] In the embodiments of the present disclosure, the shooting styles as described above may simultaneously include the shooting angle and the cropping position of the object. In this regard, the embodiments of the present disclosure may train an image recognition model based on an initial multi-task model, wherein the tasks in the initial multi-task model include a shooting angle recognition task and a cropping position recognition task.

[0117] In the embodiments of the present disclosure, the loss value of each type of task may be calculated according to the difference between the prediction result and the label of each training sample for each type of task. For example, the loss value corresponding to each type of task may be calculated based on logistic regression or mean squared error. The specific implementation manner of calculating the loss value of each type of task is not limited in the embodiments of the present disclosure.

[0118] It should be noted that the loss value of each type of task is based on all the training samples in the training sample dataset. The loss value of a single task can be characterized as the average of the differences between the prediction results and the label values of all the training samples in the training sample dataset. The smaller the loss value, the smaller the difference. Correspondingly, the smaller the total loss determined based on the loss values of each type of task, the better the convergence degree of the model, that is, the better the accuracy of the model. Therefore, the embodiments of the present disclosure may adjust the parameters of the initial multi-task model based on the total loss to obtain an image recognition model with better accuracy than the initial multi-task model.

[0119] In the embodiments of the present disclosure, when determining the total loss based on the loss values of each type of task, it may be based on a predetermined loss function, with the loss value of each type of task as the independent variable and the total loss as the dependent variable to determine the total loss. For example, the sum of the loss values of each type of task may be used as the total loss; or different weights may be assigned to each type of task and then summed and the sum value may be used as the total loss. The embodiments of the present disclosure do not limit the determination method of the total loss.

[0120] It can be understood that in the embodiments of the present disclosure, the total loss is determined based on the losses of each type of task, and the parameters in the initial multi-task model are adjusted based on the total loss, rather than adjusting the parameters based on a single task. This enables the multi-task model to balance the processing of different tasks and train the multi-task as a whole, so that the obtained multi-task model can obtain more comprehensive and accurate multi-task recognition results.

[0121] Figure 12 This is an example diagram of the training architecture of the image recognition model in the embodiments of the present disclosure. As Figure 12 shown, according to the task objective marked by N1, the classification system is disassembled, that is, the tasks included in the initial multi-task model are determined (that is, the shooting styles to be recognized are determined). Subsequently, the training sample data set can be sorted based on the method marked by N2. The content for annotating the pictures can include the shooting angle and cropping position of the objects in the pictures. For example, the annotation includes a 30° clockwise rotation and upper edge cropping. The annotation of each image is the label of each training sample for the task. After determining the training sample data set, the classification model can be trained as marked by N3, and the final trained image recognition model can be obtained through model iteration and optimization marked by N4.

[0122] It should be noted that in the embodiments of the present disclosure, before training, the labeled data set can be randomly split in a ratio of 8:1:1. 80% of the data is divided into the training sample data set, 10% of the data is divided into the validation data set, and 10% of the data is divided into the test data set to evaluate the effect of model training. At the same time, it is ensured that the data in the test data set does not appear in the training data, ensuring no data leakage problem. Through this method, the accuracy of model training can be improved, thereby improving the accuracy of recognizing the shooting styles of objects in the initial images. In addition, in the embodiments of the present disclosure, the training process of the image recognition model can also be executed on a device other than the electronic device that executes the aforementioned image processing method.

[0123] Figure 13 This is a diagram of an image processing device shown according to an exemplary embodiment. Referring to Figure 13 , the device includes:

[0124] A first acquisition module 101 configured to acquire an initial image of an object;

[0125] A first determination module 102 configured to determine the shooting style of the object in the initial image;

[0126] A composition module 103 configured to compose the object based on the shooting style of the object to obtain a composed target image.

[0127] In some embodiments, the composition module 103 is further configured to determine a target position of the object in the target image based on the shooting style of the object; set the object based on the target position, and add preset content to positions other than the target position in the target image to obtain the composed target image.

[0128] In some embodiments, the shooting style includes: shooting angle and / or the cropping position of the object. The first determination module 102 is further configured to determine a target area of the object in the target image based on the shooting angle of the object; and / or determine a target cropping position of the object in the target image based on the cropping position of the object.

[0129] In some embodiments, the composition module 103 is further configured to compose the object while maintaining the shooting angle of the object at the target position to obtain the composed target image.

[0130] In some embodiments, the composition module 103 is further configured to compose the object at the target cropping position to obtain the composed target image; wherein the target cropping position is the cropping position of the object in the initial image.

[0131] In some embodiments, the first determination module 102 is further configured to determine the shooting style of the object in the initial image based on an image recognition model; wherein the image recognition model is obtained by training a deep learning network.

[0132] In some embodiments, the image recognition model is trained from an initial multi-task model. The tasks include a shooting angle recognition task and a cropping position recognition task. The apparatus further includes:

[0133] A second acquisition module, configured to acquire a training sample data set; wherein the training sample data set includes a plurality of training samples related to each task, and labels of each task corresponding to each training sample;

[0134] A second determination module, configured to, for each training sample, use the initial multi-task model to determine a prediction result of each training sample for each task;

[0135] A calculation module, configured to, for each task, calculate a loss value of the task based on the prediction result of each training sample for the task and the label of each training sample for the task;

[0136] A third determination module, configured to determine the total loss of the initial multi-task model based on the loss value of each task;

[0137] The parameter adjustment module is configured to adjust the parameters of the initial multi-task model based on the total loss to obtain the trained image recognition model.

[0138] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0139] Figure 14 It is a block diagram of an electronic device 800 shown in an embodiment of the present disclosure. For example, the electronic device 800 may be a mobile phone or a computer, etc.

[0140] Refer to Figure 14 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0141] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0142] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0143] The power supply component 806 provides power to various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0144] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0145] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice determination mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0146] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0147] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and the keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0148] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a communication standard-based wireless network, such as Wi-Fi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0149] In an exemplary embodiment, the electronic device 800 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0150] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0151] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the foregoing image processing method.

[0152] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0153] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An image processing method, characterized in that, the method includes: Obtaining an initial image of an object; Determining the shooting style of the object in the initial image; Composing the object based on the shooting style of the object to obtain a target image after composition.

2. The method according to claim 1, characterized in that, the composing the object based on the shooting style of the object to obtain a target image after composition includes: Based on the shooting style of the object, determining the target position of the object in the target image; Setting the object based on the target position and adding preset content to positions other than the target position in the target image to obtain the target image after composition.

3. The method according to claim 2, characterized in that, the shooting style includes: shooting angle and / or the cropping position of the object, and the determining the target position of the object in the target image based on the shooting style of the object includes: Based on the shooting angle of the object, determining the target area of the object in the target image; and / or, Based on the cropping position of the object, determining the target cropping position of the object in the target image.

4. The method according to claim 3, characterized in that, the composing the object based on the shooting style of the object to obtain a target image after composition includes: Composing the object while maintaining the shooting angle of the object at the target position to obtain the target image after composition.

5. The method according to claim 3, characterized in that, the composing the object based on the shooting style of the object to obtain a target image after composition includes: Composing the object with the target cropping position to obtain the target image after composition; wherein the target cropping position is the cropping position of the object in the initial image.

6. The method according to claim 1, characterized in that, the determining the shooting style of the object in the initial image includes: Determining the shooting style of the object in the initial image based on an image recognition model; wherein the image recognition model is obtained by training a deep learning network.

7. The method according to claim 1, characterized in that, the image recognition model is trained from an initial multi-task model, the tasks include a shooting angle recognition task and a cropping position recognition task, and the method further includes: Obtaining a training sample data set; wherein the training sample data set includes a plurality of training samples related to each task, and labels of each task corresponding to each training sample; For each training sample, using the initial multi-task model to determine the prediction result of each training sample for each task; For each task, based on the prediction result of each training sample for the task and the label of each training sample for the task, calculating the loss value of the task; Based on the loss value of each task, determining the total loss of the initial multi-task model; Based on the total loss, adjusting the parameters of the initial multi-task model to obtain the trained image recognition model.

8. An image processing device, characterized in that, the device includes: A first acquisition module, configured to acquire an initial image of an object; A first determination module, configured to determine a shooting style of the object in the initial image; A composition module, configured to compose the object based on the shooting style of the object to obtain a target image after composition.

9. An electronic device, characterized in that it includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 7.