Image processing method and device, electronic equipment and storage medium

The preset model trained by deep learning networks performs composition processing on the preview image, and filters out the target composition area, solving the problem of insufficient composition skills of ordinary photographers and improving the aesthetics and diversity of the image.

CN120034745APending Publication Date: 2025-05-23BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311562832.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to help ordinary photographers improve composition skills, resulting in insufficient aesthetics and diversity of images captured by electronic devices.

Method used

The preset model obtained through deep learning network training is used to compose the preview image, filter out the target composition area, and improve the pertinence and diversity of composition.

Benefits of technology

Automatic cropping or focusing is realized to generate images, reducing the errors of photographers' artificial adjustment of composition positions, and improving the stability and aesthetic quality of image shooting and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034745A_ABST
    Figure CN120034745A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, electronic equipment and a storage medium. The method comprises the following steps: determining a shooting scene of a preview image; performing composition on the preview image by using a preset model to obtain an initial composition area of the preview image; wherein the preset model is obtained by training a deep learning network; and screening the initial composition area according to the shooting scene to obtain a target composition area of the shooting scene, and performing composition processing according to the target composition area. Through the method, on one hand, pertinence and diversity of composition recommendation of the electronic equipment are improved, and the application range is wide; and on the other hand, the image processing stability of the electronic equipment is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence and image processing technology, and in particular to an image processing method and device, an electronic device, and a storage medium. Background Art

[0002] In daily photography, the aesthetics and diversity of images captured by electronic devices are limited by factors such as shooting time, shooting space, and user shooting skills, which leads to a lot of room for improvement. Usually, professional photographers can improve the shooting effect of electronic devices by adjusting the position of the subject, the relationship between the foreground and the background, and the color ratio. However, professional photographers need to learn photography knowledge for a long time to learn composition skills.

[0003] Therefore, it is necessary to provide an image processing method to help ordinary photographers improve their composition skills so that ordinary photographers can also obtain aesthetic, diverse and high-quality images through electronic devices. Summary of the invention

[0004] The present disclosure provides an image processing method and device, an electronic device, and a storage medium.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:

[0006] Determine the shooting scene of the preview image;

[0007] Composing the preview image using a preset model to obtain an initial composition area of ​​the preview image; wherein the preset model is obtained by training a deep learning network;

[0008] The initial composition area is screened according to the shooting scene to obtain a target composition area of ​​the shooting scene, and composition processing is performed according to the target composition area.

[0009] In some embodiments, the method further comprises:

[0010] Determine the segmented areas included in the preview image and the contents of the segmented areas to segment and detect the preview image, and determine the segmentation and detection results of the preview image;

[0011] The screening of the initial composition area according to the shooting scene to obtain a target composition area of ​​the shooting scene includes:

[0012] Determining a composition type corresponding to the shooting scene;

[0013] For each composition type, based on the initial composition area, the segmentation of the preview image, each segmented area of ​​the detection result, and the content of the segmented area, a target composition area matching the composition type is screened out from the initial composition area.

[0014] In some embodiments, the preset model is a multi-task model obtained by training a composition sub-network, a segmentation sub-network, and a detection sub-network based on deep learning;

[0015] The segmenting and detecting the preview image, and determining the segmented areas and contents of the segmented areas included in the preview image based on the segmentation and detection results of the preview image, include:

[0016] Determine the segmentation and detection results of the preview image using the segmentation subnetwork and the detection subnetwork of the preset model;

[0017] Determine the segmented area included in the preview image using the segmentation subnetwork of the preset model, and determine the content of the segmented area using the detection subnetwork of the preset model;

[0018] The step of composing the preview image by using a preset model to obtain an initial composition area of ​​the preview image includes:

[0019] The preview image is composed using the composition subnetwork of the preset model to obtain an initial composition area of ​​the preview image.

[0020] In some embodiments, determining the shooting scene of the preview image includes:

[0021] The shooting scene of the preview image is determined based on the segmented areas and the contents of the segmented areas included in the segmentation and detection results of the preview image.

[0022] In some embodiments, the method further comprises:

[0023] Obtain each sample image to be trained and the composition label area corresponding to each sample image;

[0024] Using the deep learning network to process each sample image, obtaining a composition prediction area corresponding to each sample image;

[0025] Based on the difference between the composition label area and the composition prediction area corresponding to each sample image, the parameters of the deep learning network are adjusted to obtain the preset model.

[0026] In some embodiments, the step of obtaining each sample image to be trained and a composition label area corresponding to each sample image includes:

[0027] For each sample image, determining a shooting scene of the sample image, and a segmentation area included in the segmentation and detection results of the sample image;

[0028] Based on the segmentation of the sample image and the segmented area included in the detection result, determine an alternative composition corresponding to the shooting scene of the sample image; wherein the alternative composition is a sub-image of the sample image;

[0029] Obtaining aesthetic scoring results for each candidate composition;

[0030] For each sample image, based on each candidate composition and the aesthetic score result of the candidate composition, a composition screening strategy associated with the shooting scene of the sample image is used to screen out a target composition of the sample image from the candidate compositions;

[0031] The area corresponding to the target composition is determined as the composition label area corresponding to the sample image.

[0032] In some embodiments, the step of selecting a target composition of the sample image from the candidate compositions based on each candidate composition and the aesthetic scoring result of the candidate composition and using a composition screening strategy associated with the shooting scene of the sample image includes:

[0033] Based on the aesthetic score result of each candidate composition, select aesthetic candidate compositions having an aesthetic score greater than a threshold value corresponding to the shooting scene from the candidate compositions;

[0034] From the aesthetic candidate compositions, a composition that satisfies a composition rule corresponding to the shooting scene of the sample image is screened out as the target composition.

[0035] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:

[0036] A first determining module, used to determine a shooting scene of a preview image;

[0037] A first obtaining module is used to compose the preview image using a preset model to obtain an initial composition area of ​​the preview image; wherein the preset model is obtained by training a deep learning network;

[0038] The second obtaining module is used to screen the initial composition area according to the shooting scene, obtain the target composition area of ​​the shooting scene, and perform composition processing according to the target composition area.

[0039] In some embodiments, the apparatus further comprises:

[0040] A second determination module is used to determine the segmented areas included in the preview image and the contents of the segmented areas to segment and detect the preview image, and determine the segmentation and detection results of the preview image;

[0041] The second obtaining module is also used to determine the composition type corresponding to the shooting scene; for each composition type, based on the initial composition area, the segmentation of the preview image, each segmented area of ​​the detection result and the content of the segmented area, a target composition area matching the composition type is screened out from the initial composition area.

[0042] In some embodiments, the preset model is a multi-task model obtained by training a composition sub-network, a segmentation sub-network, and a detection sub-network based on deep learning;

[0043] The second determination module is further used to determine the segmentation and detection results of the preview image using the segmentation subnetwork and the detection subnetwork of the preset model; determine the segmented area included in the preview image using the segmentation subnetwork of the preset model, and determine the content of the segmented area using the detection subnetwork of the preset model;

[0044] The first obtaining module is further used to compose the preview image using the composition subnetwork of the preset model to obtain an initial composition area of ​​the preview image.

[0045] In some embodiments, the first determination module is further used to determine the shooting scene of the preview image based on the segmented areas and content of the segmented areas included in the segmentation and detection results of the preview image.

[0046] In some embodiments, the apparatus further comprises:

[0047] A first acquisition module is used to acquire each sample image to be trained and a composition label area corresponding to each sample image;

[0048] A third obtaining module is used to process each sample image using the deep learning network to obtain a composition prediction area corresponding to each sample image;

[0049] The first adjustment module is used to adjust the parameters of the deep learning network based on the difference between the composition label area and the composition prediction area corresponding to each sample image to obtain the preset model.

[0050] In some embodiments, the first acquisition module is further used to determine, for each sample image, the shooting scene of the sample image and the segmentation area included in the segmentation and detection results of the sample image; determine the alternative composition corresponding to the shooting scene of the sample image based on the segmentation area included in the segmentation and detection results of the sample image; wherein the alternative composition is a sub-image of the sample image; obtain the aesthetic scoring result of each alternative composition; for each sample image, based on each alternative composition and the aesthetic scoring result of the alternative composition, use the composition screening strategy associated with the shooting scene of the sample image to screen out the target composition of the sample image from the alternative compositions; and determine the area corresponding to the target composition as the composition label area corresponding to the sample image.

[0051] In some embodiments, the first acquisition module is further used to screen out aesthetic alternative compositions having an aesthetic score greater than an aesthetic score threshold corresponding to the shooting scene from the alternative compositions based on the aesthetic score result of each alternative composition; and to screen out a composition that satisfies the composition rule corresponding to the shooting scene of the sample image from the aesthetic alternative compositions as the target composition.

[0052] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0053] A processor; a memory for storing processor executable instructions; wherein the processor is configured to execute the image processing method as described in the first aspect above.

[0054] According to a fourth aspect of an embodiment of the present disclosure, there is provided a storage medium, including:

[0055] When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image processing method as described in the first aspect above.

[0056] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0057] In the disclosed embodiment, on the one hand, the electronic device obtains the target composition area, so that the electronic device can automatically crop or focus to generate an image directly based on the target composition area, reducing the situation where the photographer needs to manually adjust the composition position and miss the best shooting opportunity, thereby improving the stability of image shooting and processing; on the other hand, the preset model is obtained by training the deep learning network, so the model can be trained by sample images covering a variety of shooting scenes such as portraits, buildings, and green plants, so that the preset model can predict the initial composition area for the preview image of any shooting scene, and then can filter out the target composition area corresponding to the shooting scene of the preview image, and is not limited to the processing of portraits, and has a wide range of applications; on the other hand, the prediction model can directly predict the initial composition area based on the preview image, and the composition efficiency is high; in addition, based on the initial composition area of ​​the preview image, based on the shooting scene screening, it is beneficial to obtain target composition areas of various composition types for the shooting scene, improve the pertinence and diversity of composition recommendations, and facilitate the electronic device to perform image processing to obtain rich and diverse images.

[0058] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0060] Figure 1 This is an example of an image processing method flow shown in the embodiment of the present disclosure Figure 1 .

[0061] Figure 2 This is an example of an image processing method flow shown in the embodiment of the present disclosure Figure 2 .

[0062] Figure 3 It is a module framework diagram of an intelligent composition system shown in an embodiment of the present disclosure.

[0063] Figure 4 is a diagram of an image processing device shown in an embodiment of the present disclosure.

[0064] Figure 5 It is a block diagram of an electronic device 800 shown in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0065] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0066] In related technology 1, multiple shooting modes are preset, and the best reference composition is provided in each shooting mode to guide the user to adjust the composition position through the viewfinder when shooting. In related technology 2, multiple composition templates are preset, and the target composition template required for shooting is determined based on the subject-object relationship of shooting, and then the target composition template is used to guide the user to adjust the composition position through the viewfinder when shooting. However, in related technologies 1-2, preset reference compositions are used to guide users to adjust the composition position during shooting, and the composition position needs to be adjusted manually, which makes the operation cumbersome and easy to miss the best shooting time, thereby causing low stability in image shooting and processing.

[0067] In related technology 3, a composition library is established for portraits, and the similarity between the portrait image to be photographed and the composition samples in the composition library is used to match a composition sample with the highest similarity to guide the user to shoot. However, related technology 3 is only applicable to processing portrait images, resulting in a small scope of application of image processing.

[0068] In the related art 4, a statistical composition law model is built based on a set of high-quality images, and the image is randomly cropped during composition, and the composition that best conforms to the statistical composition law is output as the target composition. However, in the related art 4, since the candidate composition is generated by random cropping, the number of candidate compositions is large and the correlation with the image is low, and the process of selecting the target composition from the candidate compositions is complicated and inefficient.

[0069] In addition, related technologies 1-2 obtain the target composition based on the preset composition, which results in the composition methods that can be referenced being limited to a few preset composition templates and lacking diversity; related technologies 3-4 only output an optimal target composition, and the composition result is single. The above related technologies have not fully considered the diversity of composition recommendations.

[0070] In this regard, the present disclosure provides an image processing method. Figure 1 This is an example of an image processing method flow shown in the embodiment of the present disclosure Figure 1 ,Depend on Figure 1 It can be seen that the following steps are included:

[0071] S11, determining a shooting scene of the preview image;

[0072] S12. Use a preset model to compose the preview image to obtain an initial composition area of the preview image; wherein, the preset model is obtained by training a deep learning network.

[0073] S13. Screen the initial composition area according to the shooting scene to obtain a target composition area of the shooting scene, and perform composition processing according to the target composition area.

[0074] In the embodiments of the present disclosure, the image processing method can be executed by a terminal device with a shooting function. For example, the preset model and the image processing method can be deployed on the terminal device. The terminal device acquires a preview image and executes the steps S11 to S13 above to obtain an image processed by composing according to the composition method recommended by the preview image. In addition, the terminal device can also collect the preview image and send it to other devices such as a server. For example, after the server executes the steps S11 to S13 above, it returns the target composition area to the terminal device. The terminal device, for example, crops the preview image according to the target composition area to display the recommended composition, or the terminal device automatically adjusts the lens position according to the target composition area to collect an image with the corresponding composition method for recommendation. Among them, the terminal device can be a mobile phone, a camera, a tablet computer, a vehicle-mounted device, a wearable device, etc. with a shooting function. In the embodiments of the present disclosure, taking the image processing method applied to an electronic device as an example, the electronic device includes the aforementioned terminal device or other devices such as a server.

[0075] In the embodiments of the present disclosure, the preview image can be an image acquired in any shooting scene. For example, the shooting scene can be a portrait scene, an architectural scene, a green plant scene, etc. In step S11, there are multiple ways for the electronic device to determine the shooting scene of the preview image. It can segment and detect the preview image and determine the shooting scene according to the content results of the segmentation and detection; it can also be to directly process the preview image using a trained scene detection model to determine the shooting scene.

[0076] For example, the preview image is subjected to panoramic segmentation or saliency detection segmentation. If it is detected that the preview image includes a person in the segmented area or the saliency area, and the area to which the person belongs is relatively large, it is determined to be a portrait scene; if it is detected that the preview image includes line features, corner features, and color features of a building scene in the segmented area or the saliency area, and the area size is relatively large, it is determined to be a building scene; if it is detected that the preview image includes plant-specific shape features and green color features in the segmented area or the saliency area, and the area size is relatively large, it can be determined to belong to a green plant scene. For another example, the scene detection model can be obtained by training a neural network model based on a large number of sample images collected in different shooting scenes and the scene labels corresponding to the sample images. After the preview image is input into the model, the shooting scene corresponding to the preview image can be obtained.

[0077] It should be noted that, in the embodiment of the present disclosure, the portrait scene can be further subdivided into a single-person scene or a multi-person scene. The embodiment of the present disclosure can distinguish between a single-person scene and a multi-person scene based on a face detection algorithm.

[0078] In the disclosed embodiments, different composition types may be used in different shooting scenes. Among them, the composition type may be a close-up composition for scaling a local area (detail shooting), or a composition type based on a primary-secondary relationship. For example, in a portrait scene, the composition type may include a facial close-up composition, a complete portrait composition, or a fusion composition of a portrait and the background, etc.; in an architectural scene, the composition type may include an architectural close-up composition, a composition of the main building with the sun as the background, a character check-in composition with the building as the main body, etc.; in a green plant scene, the composition type may include a leaf close-up composition, a green plant overall composition, etc. It should be noted that the composition size may be different under different composition types.

[0079] In step S12, the preset model is obtained by training the deep learning network, and the electronic device uses the preset model to compose the preview image to obtain the initial composition area of ​​the preview image. The initial composition area may be the corresponding area coordinates on the preview image. In the disclosed embodiment, the preset model may be a model associated with the shooting scene; or it may be a model not associated with the shooting scene. For example, for each shooting scene, based on the sample image under the shooting scene and the composition label area of ​​the sample image, a preset model associated with the shooting scene is trained. Based on the model, the electronic device can input the shooting scene and the preview image into the model to obtain the initial composition area of ​​the associated shooting scene; for another example, the preset model may also be a model not associated with the shooting scene obtained by training based on the sample image that does not distinguish the shooting scene and the composition label area of ​​the sample image. Based on the model, the electronic device can obtain the initial composition area by inputting the preview image.

[0080] In the disclosed embodiment, the preset model may be a single-task model or a multi-task model, but no matter it is a single-task model or a multi-task model, the preset model includes a composition task, through which the initial composition area of ​​the preview image can be obtained. Taking the multi-task preset model as an example, the tasks other than the composition task may be a segmentation task and / or a detection task, and the electronic device may use the segmentation task of the model to perform segmentation processing on the preview image to obtain a segmentation result, use the detection task to perform detection processing on the preview image to obtain a detection result, detect the image content of each segmentation result, and determine the aforementioned shooting scene based on the image content of the segmentation result and / or the detection result.

[0081] In step S13, the electronic device screens the initial composition area according to the shooting scene to obtain a target composition area, the purpose of which is to screen out a target composition area that better matches the shooting scene, or to screen out a variety of compositions that match the shooting scene.

[0082] In some embodiments, if the initial composition area is a composition area obtained based on the shooting scene guidance, a target composition area including multiple composition sizes may be selected based on the composition size information corresponding to the composition type corresponding to the shooting scene.

[0083] In other embodiments, the initial composition area is a composition area that is not obtained based on the shooting scene guidance. Then, based on the shooting scene, the content in the initial composition area can be screened to select a target composition area that meets the shooting scene. Generally, since there may be many initial composition areas obtained based on the preset model, there are also many composition areas that meet the shooting scene that are screened. In this regard, the embodiment of the present disclosure is also based on the diversity of composition recommendations. From the composition areas of multiple composition types that meet the shooting scene, a predetermined number of composition areas are screened as target composition areas for each composition type.

[0084] Exemplarily, in an architectural scene, the electronic device can provide multiple local close-up compositions of different sizes, for example, composition areas of architectural windows of different sizes and composition areas of architectural top surfaces; in addition, multiple compositions with the building as the main body can also be provided, for example, a composition with the building as the foreground and the lake as the background, a small-sized portrait as the foreground and a large-sized building as the background, etc.; in a portrait scene, the electronic device can provide multiple local close-up compositions of different sizes, for example, a composition area of ​​the face, a composition area of ​​the upper body of the human body; in addition, multiple compositions with the person as the main body can also be provided, for example, a composition with the portrait as the foreground and the sky as the background, a composition with the cake as the foreground and a large-area portrait as the background, etc.; in a green plant scene, the electronic device can provide multiple local close-up compositions of different sizes, for example, a composition area of ​​green plant leaves, a composition area of ​​green plant flowers; in addition, multiple compositions with green plants as the main body can also be provided, for example, a composition with green plants as the foreground and mountains as the background, a composition with dewdrops as the foreground and green plants as the background, etc.

[0085] In the disclosed embodiments, there are multiple ways for the electronic device to process the image according to the target composition area. The preview image may be cropped according to the target composition area to obtain the image, or the lens position may be automatically adjusted according to the target composition area to compose the image in a corresponding composition manner to obtain the image.

[0086] It should be noted that in the disclosed embodiment of the present invention, when the electronic device performs screening according to the content of the initial composition area, it can guide the screening of the target composition area by combining the content of the shooting scene determined by the segmentation and detection processing in the process of determining the shooting scene in the aforementioned step S11, or it can re-segment and detect the preview sub-image corresponding to the initial composition area to confirm the content of the shooting scene to guide the screening of the target composition area.

[0087] As mentioned above, the related technologies have problems such as low stability of image capture and processing, small scope of application, low composition efficiency, and lack of diversity in composition recommendations. In this regard, in the embodiments of the present disclosure, on the one hand, the electronic device obtains the target composition area, so that the electronic device can automatically crop or focus to generate an image based on the target composition area, thereby reducing the situation where the photographer needs to manually adjust the composition position and miss the best shooting opportunity, thereby improving the stability of image shooting and processing; on the other hand, the preset model is obtained by training the deep learning network, so the model can be trained by sample images covering a variety of shooting scenes such as portraits, buildings, and green plants, so that the preset model can predict the initial composition area for the preview image of any shooting scene, and then can filter out the target composition area corresponding to the shooting scene of the preview image, without being limited to the processing of portraits, and has a wide range of applications; on the other hand, the prediction model can directly predict the initial composition area based on the preview image, and the composition efficiency is high; in addition, based on the initial composition area of ​​the preview image, based on the shooting scene screening, it is beneficial to obtain target composition areas of various composition types for the shooting scene, thereby improving the pertinence and diversity of the composition recommendation, and facilitating the electronic device to perform image processing to obtain rich and diverse images.

[0088] In some embodiments, the method further comprises:

[0089] Segmenting and detecting the preview image to determine the segmentation and detection results of the preview image; determining the segmented areas included in the preview image and the contents of the segmented areas;

[0090] The screening of the initial composition area according to the shooting scene to obtain a target composition area of ​​the shooting scene includes:

[0091] Determining a composition type corresponding to the shooting scene;

[0092] For each composition type, based on the initial composition area, the segmentation of the preview image, each segmented area of ​​the detection result, and the content of the segmented area, a target composition area matching the composition type is screened out from the initial composition area.

[0093] Figure 2 This is an example of an image processing method flow shown in the embodiment of the present disclosure Figure 2 ,Depend on Figure 2 It can be seen that the following steps are included:

[0094] S10, segmenting and detecting the preview image to determine the segmentation and detection results of the preview image; determining the segmented areas included in the preview image and the contents of the segmented areas;

[0095] S11, determining a shooting scene of the preview image;

[0096] S12, composing the preview image using a preset model to obtain an initial composition area of ​​the preview image; wherein the preset model is obtained by training a deep learning network;

[0097] The aforementioned step S13 includes:

[0098] S131, determining a composition type corresponding to the shooting scene;

[0099] S132 . For each composition type, based on the initial composition area, the segmentation of the preview image, each segmented area of ​​the detection result, and the content of the segmented area, filter out a target composition area matching the composition type from the initial composition area.

[0100] In step S10, the electronic device may determine the segmented area included in the segmentation result of the preview image by methods such as panorama segmentation, saliency segmentation, and edge segmentation; and may determine the content of the segmented area included in the detection result of the preview image by methods such as color feature detection, shape feature detection, and texture feature detection. The segmentation and detection results may be the contents in the preview image and the segmented area where the contents are located. For example, by segmenting and detecting the preview image of the building, it may be determined that the preview image includes multiple contents such as the building, the sky, and the sun, as well as the segmented area information where the contents are located.

[0101] As mentioned above, there are different composition types in different shooting scenes. In the disclosed embodiment, the electronic device can preset one or more composition types corresponding to the shooting scene, and then after executing the aforementioned step S11 to determine the shooting scene of the preview image, the electronic device can use step S131 to drive the composition type corresponding to the shooting scene, and use step S132 to screen the composition based on each composition type of the shooting scene.

[0102] In step S132, for each composition type, the electronic device selects a target composition area matching the composition type from the initial composition area based on the initial composition area, the segmentation and detection results (content and the area where the content is located) of the preview image, and the content of the segmented areas. For example, in a portrait scene, for a facial close-up composition, the electronic device may compare an initial composition area with the content of the preview image and the segmented areas where the content is located. If the initial composition segmented area belongs to a segmented area of ​​the preview image, and the content corresponding to the segmented area is a face, it can be determined as a facial close-up composition. At this time, the electronic device may retain the initial composition area of ​​the facial close-up composition for subsequent composition recommendations. For another example, in an architectural scene, for a character punch-in composition with a building as the main body, the electronic device may compare an initial composition area with the content of the preview image and each segmented area where the content is located. If the initial segmented composition area includes multiple segmented areas of the preview image, and the content corresponding to the segmented area includes portraits and buildings, it can be determined as a character punch-in composition with a building as the main body. At this time, the electronic device can retain the initial composition area of ​​the character punch-in composition with the building as the main body for subsequent composition recommendations. For another example, in a green plant scene, for an overall composition of green plants, the electronic device may compare an initial composition area with the content of the preview image and each segmented area where the content is located. If the initial composition segmented area includes multiple segmented areas of the preview image, and the content of the segmented area includes leaves and flowers, it can be determined as an overall composition of green plants. At this time, the electronic device can retain the initial composition area of ​​the overall composition of green plants for subsequent composition recommendations.

[0103] It should be noted that, as mentioned above, there may be a particularly large number of initial composition areas. For example, there may be multiple initial composition areas belonging to facial close-up compositions. In the disclosed embodiment, the electronic device may also retain a preset number of initial composition areas matching the composition type for each composition type. For example, for facial close-up compositions, the initial composition areas of two facial close-up compositions are retained as part of the target composition area.

[0104] In the disclosed embodiment, the electronic device performs segmentation and detection on the preview image to obtain the content of the preview image and the content of the segmented areas where the contents are located, and can recommend one or more target compositions for each composition type of the shooting scene, thereby reducing the occurrence of homogeneous recommended compositions and improving the diversity of composition recommendations.

[0105] In some embodiments, the preset model is a multi-task model obtained by training a composition sub-network, a segmentation sub-network, and a detection sub-network based on deep learning;

[0106] The segmenting and detecting the preview image, and determining the segmented areas and contents of the segmented areas included in the preview image based on the segmentation and detection results of the preview image, include:

[0107] Determine the segmentation and detection results of the preview image using the segmentation subnetwork and the detection subnetwork of the preset model; determine the segmentation area included in the preview image using the segmentation subnetwork of the preset model, and determine the content of the segmentation area using the detection subnetwork of the preset model;

[0108] The step of composing the preview image by using a preset model to obtain an initial composition area of ​​the preview image includes:

[0109] The preview image is composed using the composition subnetwork of the preset model to obtain an initial composition area of ​​the preview image.

[0110] As mentioned above, the preset model can be a multi-task model. In the embodiment of the present disclosure, the preset model includes a composition task, a segmentation task and a detection task. The electronic device can use the segmentation subnetwork and the detection subnetwork of the preset model to determine the segmentation and detection result area of ​​the preview image (the content of the preview image and the area where the content is located), use the detection subnetwork to determine the content of each segmented area, and use the composition subnetwork to determine the initial composition area of ​​the preview image, thereby further based on the initial composition area, the content of the preview image and the area where the content is located, each segmented area and the content of the segmented area, filter out the target composition area from the initial composition area.

[0111] In the disclosed embodiment, the composition subnetwork, detection subnetwork, and segmentation subnetwork of the preset model may share a backbone network. Based on the features extracted by the backbone network, each subnetwork may further perform feature processing to obtain the corresponding results of each subnetwork based on the processed features. Among them, the backbone network may be ResNet, DarkNet, etc., the composition subnetwork and the detection subnetwork may be the neck and head parts of networks such as YOLO and R-CNN, and the segmentation subnetwork may be the Decoder part of networks such as SegNet and U-Net.

[0112] In the disclosed embodiment, when the electronic device performs multi-task model training on the preset model, each type of task (each sub-network) may correspond to a loss function, and the electronic device may determine the total loss based on the loss of each type of task, and adjust the model parameters based on the total loss to obtain the preset model, or may adjust the model parameters corresponding to each type of task based on the loss of each type of task to obtain the preset model.

[0113] In addition, in the embodiment of the present disclosure, the processing results obtained by the segmentation subnetwork and / or the detection subnetwork in the preset model can also assist in determining the shooting scene of the preview image in the aforementioned step S11.

[0114] In the disclosed embodiment, the electronic device uses a preset model to be set as a multi-task model. Since the multi-task model can share the extracted features during the learning process, and share and supplement each other with the learned information related to each task, the generalization ability of the model can be improved; and since multiple tasks share one model, less memory is occupied, which can save device resources relative to multiple single-task models; in addition, the results obtained by the segmentation sub-network and the detection sub-network can also assist in determining the shooting scene, thereby improving the reusability of the preset model and facilitating the lightweighting of the preset model as a whole.

[0115] In some embodiments, determining the shooting scene of the preview image includes:

[0116] The shooting scene of the preview image is determined based on the segmented areas and the contents of the segmented areas included in the segmentation and detection results of the preview image.

[0117] In the disclosed embodiment, there are multiple ways for the electronic device to determine the shooting scene based on the segmentation and detection results of the preview image, including the segmented areas and the contents of the segmented areas. The shooting scene may be determined based on the proportion of the content in each segmented area; or the shooting scene may be determined based on the relative position relationship of the content in each segmented area.

[0118] For example, if it is detected that buildings account for the highest proportion in each segmented area, the shooting scene is determined to be a building scene; if it is detected that the central area is a plant and the edge area is a road, the shooting scene is determined to be a green plant scene.

[0119] It should be noted that when detecting the shooting scene directly based on the entire preview image, the shooting scene is usually determined directly based on the content corresponding to the central area of ​​the preview image or the area with the highest color saturation. Due to the lack of analysis and judgment of the entire preview image, it is easy to cause scene judgment errors.

[0120] In the disclosed embodiment, the multi-task results of the electronic device based on the multi-task model can not only guide the screening of the target composition area, but also be used to determine the shooting scene of the preview image, thereby improving the reusability of the output results of the preset model and saving computing resources.

[0121] In some embodiments, the method further comprises:

[0122] Obtain each sample image to be trained and the composition label area corresponding to each sample image;

[0123] Processing each sample image using the deep learning network to obtain a composition prediction area corresponding to each sample image;

[0124] Based on the difference between the composition label area and the composition prediction area corresponding to each sample image, the parameters of the deep learning network are adjusted to obtain the preset model.

[0125] As mentioned above, the preset model may be a model associated with the shooting scene; or a model not associated with the shooting scene.

[0126] In some embodiments, the preset model is a model that is not associated with a shooting scene. The electronic device can obtain each sample image to be trained and the composition label area corresponding to each sample image, and use a deep learning network to process each sample image to obtain a composition prediction area corresponding to each sample image. Based on the difference between the composition label area corresponding to each sample image and the composition prediction area, the parameters of the deep learning network are adjusted to obtain the preset model.

[0127] In other embodiments, the preset model is a model of an associated shooting scene, and the electronic device can obtain each sample image to be trained and a composition label area corresponding to each sample image, wherein each sample image is also accompanied by shooting scene information. During model training, the parameters of the deep learning network are adjusted based on the difference between the composition prediction area and the composition label area of ​​the associated shooting scene, thereby obtaining a preset model of the associated shooting scene.

[0128] In addition, as mentioned above, the preset model can be a single-task model or a multi-task model. In the disclosed embodiment, the electronic device can also obtain the content label corresponding to the sample image, the label of the area where the segmented content is located, and the content label of the segmented area, so as to simultaneously train the aforementioned segmentation task and detection task. For example, the segmentation subnetwork can be used to process each sample image to obtain the segmentation prediction area corresponding to each sample image, and the parameters of the segmentation subnetwork can be adjusted based on the difference between the segmentation label area and the segmentation prediction area corresponding to each sample image.

[0129] In the embodiments of the present disclosure, the composition label area may be an area based on manual annotation or may be obtained by other automatic methods. In some embodiments, obtaining each sample image to be trained and the composition label area corresponding to each sample image includes:

[0130] For each sample image, determining a shooting scene of the sample image, and a segmentation area included in the segmentation and detection results of the sample image;

[0131] Based on the segmentation of the sample image and the segmented area included in the detection result, determine an alternative composition corresponding to the shooting scene of the sample image; wherein the alternative composition is a sub-image of the sample image;

[0132] Obtaining aesthetic scoring results for each candidate composition;

[0133] For each sample image, based on each candidate composition and the aesthetic score result of the candidate composition, a composition screening strategy associated with the shooting scene of the sample image is used to screen out a target composition of the sample image from the candidate compositions;

[0134] The area corresponding to the target composition is determined as the composition label area corresponding to the sample image.

[0135] In the embodiments of the present disclosure, there are multiple ways for the electronic device to determine the shooting scene of the sample image. The method may be to use a trained scene detection model to directly process the sample image to determine the shooting scene. Alternatively, the sample image may be segmented and detected to first obtain the segmentation and detection results (i.e., the content of the sample image and the area where the content is located) to determine the segmented area of ​​the sample image, and then determine the shooting scene of the sample image based on the content of the sample image and the image detection content of the segmented area where the content is located.

[0136] It should be noted that in the embodiments of the present disclosure, the segmented areas obtained by the electronic device through segmentation of the sample image can also be used as segmented area labels for training the segmentation sub-network in the preset model; in addition, the image detection content of the segmented area can also be used as segmented area content labels for training the detection sub-network in the preset model.

[0137] In the disclosed embodiment, the electronic device determines the alternative composition corresponding to the shooting scene of the sample image based on the content of the sample image and the segmented area included in the area where the content is located. Among them, the electronic device can compose according to the shooting scene of the sample image and for each composition type in the shooting scene using the preset composition rules of each composition type, and the composition method can be to screen or combine according to the segmented area of ​​the sample image, and determine the composition that meets the preset composition rules as the alternative composition. For example, in a building scene, when composing the main body of the building, the two segmented areas of the building and the sun in the segmented content of the sample image can be combined to obtain an alternative composition with the building as the main body and the sun as the background; when composing a close-up of the building, the multiple segmented areas including the building windows, the building top surface, the building lighting, etc. obtained by segmenting the sample image can be used to obtain multiple alternative images of the building close-ups of different parts of the building.

[0138] In the disclosed embodiment, there are multiple ways for the electronic device to obtain the aesthetic scoring result of each candidate composition. The aesthetic scoring result of each candidate composition can be determined based on a preset aesthetic scoring rule; or each candidate composition can be input into an aesthetic scoring model to obtain the aesthetic score of each candidate composition. Among them, the aesthetic scoring rules may include the color of the composition, the proportion of white space in the composition, the relative position relationship of each scene in the composition, etc. In the disclosed embodiment, the aesthetic scoring model does not need to be deployed in the electronic device.

[0139] In the disclosed embodiment, the electronic device will select a target composition of the sample image from the candidate compositions based on the composition screening strategy associated with the shooting scene of the sample image, wherein the composition screening strategy associated with the shooting scene may be at least one of the following: a scoring standard associated with the shooting scene, and may also be a size standard and a composition matching standard associated with the shooting scene. Exemplarily, the electronic device may select as the target composition an alternative composition whose aesthetic score result is higher than the scoring standard associated with the shooting scene; may also select as the target composition an alternative composition whose size meets the size standard associated with the shooting scene; may also select as the target composition an alternative composition whose composition elements meet the composition matching standard associated with the shooting scene.

[0140] In the disclosed embodiment, the electronic device determines the area corresponding to the target composition as the composition label area corresponding to the sample image, so as to facilitate training of a preset model including the composition task.

[0141] It should be noted that in the disclosed embodiment, the trained preset model is deployed on the electronic device side for image processing to directly predict the target composition area for the preview image. However, the training task of the preset model does not need to be deployed on the electronic device side for image processing, and the aesthetic scoring model or other models involved in the training process do not need to be deployed on the electronic device side, so as to make the electronic device side more lightweight.

[0142] In the disclosed embodiment, the electronic device determines the shooting scene of the sample image and the segmentation area of ​​the segmentation and detection results, further obtains alternative compositions of the sample image, and further screens out the target composition based on the aesthetic scoring results of the alternative compositions to determine the composition label, and guides the composition based on the shooting scene and the aesthetic scoring results, which is beneficial to enhancing the correlation between the target composition and the shooting scene while taking into account the aesthetic evaluation of the composition, so that a composition label area with better quality can be obtained, thereby improving the accuracy of the preset model training.

[0143] In some embodiments, the step of selecting a target composition of the sample image from the candidate compositions based on each candidate composition and the aesthetic scoring result of the candidate composition and using a composition screening strategy associated with the shooting scene of the sample image includes:

[0144] Based on the aesthetic score result of each candidate composition, select aesthetic candidate compositions having an aesthetic score greater than a threshold value corresponding to the shooting scene from the candidate compositions;

[0145] From the aesthetic candidate compositions, a composition that satisfies a composition rule corresponding to the shooting scene of the sample image is screened out as the target composition.

[0146] In the disclosed embodiment, the electronic device may preset an aesthetic score threshold for the shooting scene, and may set the aesthetic score thresholds of the shooting scenes to be the same or different. In the disclosed embodiment, the electronic device determines, based on the aesthetic score result of each candidate composition, a candidate composition having an aesthetic score result greater than the aesthetic score threshold corresponding to the shooting scene as an aesthetic candidate composition.

[0147] In the disclosed embodiment, the electronic device can filter out the target composition of the sample image that meets the shooting scene rules from the aesthetic alternative compositions, wherein the composition rules can be the size standard, composition matching standard, etc. associated with the shooting scene. For example, in a portrait scene, a portrait composition that meets the preset size is filtered out from the alternative images and determined as the target composition in the portrait scene, wherein the size of the facial close-up composition and the portrait complete composition in the portrait scene can be set to be different. For another example, in a building scene, for a character punch-in composition type with a building as the main body, from multiple building main body alternative compositions, the alternative composition with the character close to the center of the building is determined as the target composition, and the alternative composition with the character deviating from the center of the building is removed. For another example, in a green plant scene, for a leaf close-up composition type, from multiple close-up alternative compositions, the alternative composition that can highlight the characteristics of the plant is determined as the target composition, and the alternative composition that cannot highlight the characteristics of the plant is removed.

[0148] In the disclosed embodiment, the electronic device first determines the alternative compositions that meet the aesthetic conditions, and then uses the composition screening strategy associated with the shooting scene to screen the alternative compositions, which is conducive to adopting different screening rules for each shooting scene, improving the correlation between the target image and the sample image, and also helping to improve the aesthetics of the composition.

[0149] Figure 3 is a module framework diagram of an intelligent composition system shown in an embodiment of the present disclosure, such as Figure 3 As shown, the intelligent composition system includes a recommendation and scoring module L301, a scene rule module L302, and a prediction module L303.

[0150] In this embodiment, the sample image can be input into the intelligent composition system, and the candidate composition that meets the aesthetic rules in each shooting scene (i.e., the alternative composition in the embodiment of the present disclosure) can be obtained through the recommendation and scoring module L301, and then the candidate composition can be screened using the scene rule module L302 to obtain the target composition that meets the scene rule (i.e., the composition screening strategy associated with the shooting scene of the sample image in the embodiment of the present disclosure) and perform composition annotation (i.e., the composition area label in the embodiment of the present disclosure).

[0151] In this embodiment, the prediction module L303 (i.e., the preset model in the embodiment of the present disclosure) uses a deep learning network to process each sample image, obtains the predicted composition annotation corresponding to each sample image (i.e., the composition prediction area in the embodiment of the present disclosure), and combines the composition annotation obtained by the aforementioned scene rule module L302 to adjust the model parameters of the deep learning network so that the loss function converges, and finally obtains the preset model corresponding to the prediction module L303. Based on the trained preset model, the electronic device can use the prediction module L303 to predict one or more recommended compositions corresponding to each composition type in the shooting scene for any input image. Among them, the prediction module L303 can directly predict the target composition area corresponding to the recommended composition (for example, the specific coordinates of the composition).

[0152] In this embodiment, the recommendation and scoring module L301 includes: a detection and segmentation submodule L3011, a candidate composition generation submodule L3012, and a scoring submodule L3013. Among them, the detection and segmentation submodule L3011 is used to perform saliency detection, panoramic segmentation and other processing on the sample image to obtain the image content and the location of the content, and determine the shooting scene, that is, in the embodiment of the present disclosure, the electronic device determines the shooting scene of the sample image. In addition, the detection and segmentation submodule L3011 can retain the saliency detection results and the panoramic segmentation results, which are used for composition screening in the scene rule module L302 and training of the prediction model in the prediction module L303.

[0153] In this embodiment, the candidate composition generation submodule L3012 is used to generate multiple candidate compositions for the scene based on shooting scene information, saliency detection, panoramic segmentation and other results. That is, in the embodiment of the present disclosure, based on the segmentation of the sample image and the segmented area included in the detection results, the alternative composition corresponding to the shooting scene of the sample image is determined.

[0154] In this embodiment, the scoring submodule L3013 is used to perform aesthetic scoring on the candidate compositions in the shooting scene, that is, in the embodiment of the present disclosure, obtain the aesthetic scoring result of each candidate composition. The aesthetic scoring can be implemented using an aesthetic scoring model trained on a large amount of data.

[0155] In the disclosed embodiments, on the one hand, the intelligent composition system obtains the target composition area, and enables the electronic device to automatically crop or focus to generate an image directly based on the target composition area, thereby reducing the need for the photographer to manually adjust the composition position and causing the best shooting opportunity to be missed, thereby improving the stability of image capture and processing; on the other hand, the prediction model is obtained by training the deep learning network, and thus the model can be trained through sample images covering a variety of shooting scenes, so that the preset model can predict the initial composition area for the input image of any shooting scene, and then the target composition area corresponding to the shooting scene of the input image can be screened out, and it is not limited to the processing of portraits, and has a wide range of applications; on the other hand, the prediction model can directly predict the initial composition area based on the input image, and the composition efficiency is high; in addition, based on the initial composition area of ​​the input image, based on the shooting scene screening, it is beneficial to obtain target composition areas of various composition types for the shooting scene, thereby improving the pertinence and diversity of the composition recommendation, and facilitating the electronic device to perform image processing to obtain rich and diverse images.

[0156] Figure 4 is a diagram of an image processing device shown in an embodiment of the present disclosure, comprising Figure 4 It is known that, including:

[0157] A first determining module 401 is used to determine a shooting scene of a preview image;

[0158] A first obtaining module 402 is used to compose the preview image using a preset model to obtain an initial composition area of ​​the preview image; wherein the preset model is obtained by training a deep learning network;

[0159] The second obtaining module 403 is used to screen the initial composition area according to the shooting scene, obtain a target composition area of ​​the shooting scene, and perform composition processing according to the target composition area.

[0160] In some embodiments, the apparatus further comprises:

[0161] A second determination module is used to determine the segmented areas included in the preview image and the contents of the segmented areas to segment and detect the preview image, and determine the segmentation and detection results of the preview image;

[0162] The second obtaining module 403 is also used to determine the composition type corresponding to the shooting scene; for each composition type, based on the initial composition area, the segmentation of the preview image, each segmented area of ​​the detection result and the content of the segmented area, a target composition area matching the composition type is screened out from the initial composition area.

[0163] In some embodiments, the preset model is a multi-task model obtained by training a composition sub-network, a segmentation sub-network, and a detection sub-network based on deep learning;

[0164] The second determination module is further used to determine the segmentation and detection results of the preview image using the segmentation subnetwork and the detection subnetwork of the preset model; determine the segmented area included in the preview image using the segmentation subnetwork of the preset model, and determine the content of the segmented area using the detection subnetwork of the preset model;

[0165] The first obtaining module 402 is further configured to compose the preview image using the composition subnetwork of the preset model to obtain an initial composition area of ​​the preview image.

[0166] In some embodiments, the first determination module 401 is further configured to determine a shooting scene of the preview image based on the segmented areas and contents of the segmented areas included in the segmentation and detection results of the preview image.

[0167] In some embodiments, the apparatus further comprises:

[0168] A first acquisition module is used to acquire each sample image to be trained and a composition label area corresponding to each sample image;

[0169] A third obtaining module is used to process each sample image using the deep learning network to obtain a composition prediction area corresponding to each sample image;

[0170] The first adjustment module is used to adjust the parameters of the deep learning network based on the difference between the composition label area and the composition prediction area corresponding to each sample image to obtain the preset model.

[0171] In some embodiments, the first acquisition module is further used to determine, for each sample image, the shooting scene of the sample image and the segmentation area included in the segmentation and detection results of the sample image; determine the alternative composition corresponding to the shooting scene of the sample image based on the segmentation area included in the segmentation and detection results of the sample image; wherein the alternative composition is a sub-image of the sample image; obtain the aesthetic scoring result of each alternative composition; for each sample image, based on each alternative composition and the aesthetic scoring result of the alternative composition, use the composition screening strategy associated with the shooting scene of the sample image to screen out the target composition of the sample image from the alternative compositions; and determine the area corresponding to the target composition as the composition label area corresponding to the sample image.

[0172] In some embodiments, the first acquisition module is further used to screen out aesthetic alternative compositions having an aesthetic score greater than an aesthetic score threshold corresponding to the shooting scene from the alternative compositions based on the aesthetic score result of each alternative composition; and to screen out a composition that satisfies the composition rule corresponding to the shooting scene of the sample image from the aesthetic alternative compositions as the target composition.

[0173] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0174] Figure 5 8 is a block diagram of an electronic device 800 shown in an embodiment of the present disclosure. For example, the electronic device 800 may be a mobile phone, a mobile computer, a camera, etc.

[0175] Reference Figure 5 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0176] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0177] The memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0178] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0179] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0180] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0181] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0182] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of the components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0183] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as Wi-Fi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0184] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0185] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of an electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0186] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the aforementioned image processing method.

[0187] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims above.

[0188] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image processing method, It is characterized in that The method comprises: Determine the shooting scene of the preview image; Composing the preview image using a preset model to obtain an initial composition area of ​​the preview image; wherein the preset model is obtained by training a deep learning network; The initial composition area is screened according to the shooting scene to obtain a target composition area of ​​the shooting scene, and composition processing is performed according to the target composition area.

2. The method according to claim 1, It is characterized in that The method further comprises: Segmenting and detecting the preview image, and determining segmented areas and contents of the segmented areas included in the segmentation and detection results of the preview image; The screening of the initial composition area according to the shooting scene to obtain a target composition area of ​​the shooting scene includes: Determining a composition type corresponding to the shooting scene; For each composition type, based on the initial composition area, the segmentation of the preview image, each segmented area of ​​the detection result, and the content of the segmented area, a target composition area matching the composition type is screened out from the initial composition area.

3. The method according to claim 2, It is characterized in that The preset model is a multi-task model obtained by training a composition sub-network, a segmentation sub-network and a detection sub-network based on deep learning; The segmenting and detecting the preview image, and determining the segmented areas and contents of the segmented areas included in the preview image based on the segmentation and detection results of the preview image, include: Determine the segmentation and detection results of the preview image using the segmentation subnetwork and the detection subnetwork of the preset model; Determine the segmented area included in the preview image using the segmentation subnetwork of the preset model, and determine the content of the segmented area using the detection subnetwork of the preset model; The step of composing the preview image by using a preset model to obtain an initial composition area of ​​the preview image includes: The preview image is composed using the composition subnetwork of the preset model to obtain an initial composition area of ​​the preview image.

4. The method according to claim 3, It is characterized in that The step of determining the shooting scene of the preview image includes: The shooting scene of the preview image is determined based on the segmented areas and the contents of the segmented areas included in the segmentation and detection results of the preview image.

5. The method according to any one of claims 1 to 4, It is characterized in that The method further comprises: Obtain each sample image to be trained and the composition label area corresponding to each sample image; Using the deep learning network to process each sample image, obtaining a composition prediction area corresponding to each sample image; Based on the difference between the composition label area and the composition prediction area corresponding to each sample image, the parameters of the deep learning network are adjusted to obtain the preset model.

6. The method according to claim 5, It is characterized in that The step of obtaining each sample image to be trained and a composition label area corresponding to each sample image includes: For each sample image, determining a shooting scene of the sample image, and a segmentation area included in the segmentation and detection results of the sample image; Based on the segmentation of the sample image and the segmented area included in the detection result, determine an alternative composition corresponding to the shooting scene of the sample image; wherein the alternative composition is a sub-image of the sample image; Obtaining aesthetic scoring results for each candidate composition; For each sample image, based on each candidate composition and the aesthetic score result of the candidate composition, a composition screening strategy associated with the shooting scene of the sample image is used to screen out a target composition of the sample image from the candidate compositions; The area corresponding to the target composition is determined as the composition label area corresponding to the sample image.

7. The method according to claim 6, It is characterized in that The method of selecting a target composition of the sample image from the candidate compositions based on each candidate composition and the aesthetic scoring result of the candidate composition and using a composition screening strategy associated with the shooting scene of the sample image comprises: Based on the aesthetic score result of each candidate composition, select aesthetic candidate compositions having an aesthetic score greater than a threshold value corresponding to the shooting scene from the candidate compositions; From the aesthetic candidate compositions, a composition that satisfies a composition rule corresponding to the shooting scene of the sample image is screened out as the target composition.

8. An image processing device, It is characterized in that The device comprises: A first determining module, used to determine a shooting scene of a preview image; A first obtaining module is used to compose the preview image using a preset model to obtain an initial composition area of ​​the preview image; wherein the preset model is obtained by training a deep learning network; The second obtaining module is used to screen the initial composition area according to the shooting scene, obtain the target composition area of ​​the shooting scene, and perform composition processing according to the target composition area.

9. An electronic device, It is characterized in that include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the image processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, It is characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image processing method as claimed in any one of claims 1 to 7.