Image processing method, device, electronic device and storage medium

By calculating the similarity of image attributes and combining the aesthetic features of the reference image for feature alignment and fusion, the problem of single image beautification processing in the prior art is solved, and beautified images that are more in line with the aesthetic needs of users.

CN114140315BActive Publication Date: 2025-05-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111285143.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-01
Publication Date
2025-05-06
Estimated Expiration
2041-11-01

AI Technical Summary

Technical Problem

In the prior art, image beautification processing adopts the same beautification algorithm, resulting in a single target beautification image and it is difficult to meet the aesthetic needs of different users.

Method used

By obtaining the image to be processed and its attribute information, including aesthetic feature information, temporal information, spatial information and social information, the attribute similarity between the candidate image and the image to be processed is calculated, the reference image with the highest similarity is selected, its aesthetic feature information is extracted and the feature information of the image to be processed for feature alignment and fusion, and decoding is performed to generate the target beautification image.

Benefits of technology

It realizes the search for reference images with high similarity based on the aesthetic characteristics, time, space and social information of the image, and combines their aesthetic characteristics to perform image beautification processing to generate target beautification images that are more in line with users' aesthetic needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140315B_ABST
    Figure CN114140315B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image processing method, device, electronic device and storage medium. The method includes: obtaining an image to be processed and attribute information of the image to be processed, including aesthetic feature information, time information, spatial information and social information of the user to which the image belongs; calculating the attribute similarity between the image to be processed and the candidate image, and taking the candidate image with the highest attribute similarity as the reference image; extracting the features of the aesthetic feature information of the image to be processed and the features of the aesthetic feature information of the reference image, respectively obtaining the feature information to be processed and the reference feature information; performing feature alignment and feature fusion on the feature information to be processed and the reference feature information, decoding the obtained fused aesthetic feature information, and obtaining a target beautified image. In this way, by determining a reference image with a high attribute similarity with the image to be processed, and processing the image to be processed in combination with the reference image, the obtained target beautified image can better meet the aesthetic needs of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of the Internet, and in particular to an image processing method, device, electronic device and storage medium. Background Art

[0002] Image beautification is a very important function in image and video application software. For example, image beautification can include beautification of portraits, beautification of landscapes, beautification of food, etc.

[0003] Taking portrait beautification as an example, in the prior art, the following steps are usually adopted to beautify the portrait: first, at least one feature area on the face in the image is identified, then, based on the attribute feature information of the feature area, the beautification change parameters of the feature area are determined, and then, based on the determined beautification change parameters, the image is beautified to obtain a beautified portrait.

[0004] However, in the above-mentioned image processing method, the same beautification algorithm is used for all images. It is understandable that users' perception of beauty varies in terms of region, time and sociality. In other words, different regions, different periods and different groups of people have different perceptions of beauty. Therefore, when the same beautification algorithm is used to beautify images, the target beautified images obtained are very single and difficult to meet the aesthetic needs of different users. Summary of the invention

[0005] The present disclosure provides an image processing method, device, electronic device and storage medium to at least solve the problem that the same beautification algorithm is used to beautify the image in the related art, and the target beautified image obtained is very single and difficult to meet the aesthetic needs of different users. The technical solution of the present disclosure is as follows:

[0006] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:

[0007] Acquire an image to be processed and attribute information of the image to be processed, wherein the attribute information includes any one or more of aesthetic feature information, time information, space information, and social information of a user;

[0008] Calculating the attribute similarity between the image to be processed and the candidate image based on the attribute information of each candidate image acquired in advance and the attribute information of the image to be processed, and taking the candidate image with the highest attribute similarity as a reference image;

[0009] Extracting features of the aesthetic feature information of the image to be processed to obtain feature information to be processed, and extracting features of the aesthetic feature information of the reference image to obtain reference feature information;

[0010] Performing feature alignment and feature fusion on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information;

[0011] The fused aesthetic feature information is decoded to obtain a target beautified image.

[0012] Optionally, when the attribute information includes multiple items, calculating the attribute similarity between the image to be processed and the candidate image according to the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed includes:

[0013] Calculating the sub-item similarity between the image to be processed and the candidate image for the attribute information according to the attribute information of each candidate image acquired in advance and the attribute information of the image to be processed;

[0014] According to the weight of each item of attribute information obtained in advance, the sub-item similarities are weighted and summed to obtain the attribute similarity between the candidate image and the image to be processed.

[0015] Optionally, when the attribute information includes social information of the user, calculating the sub-item similarity of the attribute information between the image to be processed and the candidate image according to the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed includes:

[0016] According to the pre-acquired social information of the user to which each candidate image belongs and the social information of the user to which the image to be processed belongs, social network information is constructed, wherein each node in the social network represents the image to be processed and the user to which the candidate image belongs, and each connection line between nodes represents a communication between the image to be processed and the user to which the candidate image belongs;

[0017] According to the social network information, the social similarity between the image to be processed and the candidate images is calculated, and the social similarity is used as the sub-item similarity of the social information of the user to which the image to be processed and each candidate image belongs.

[0018] Optionally, calculating the social similarity between the image to be processed and each candidate image according to the social network information includes:

[0019] In the case where there is a target link between the user of the image to be processed and the nodes corresponding to the user of any candidate image, the social similarity between the user of the image to be processed and the candidate image is calculated according to the number of the target links, where the nodes at both ends of the target link are the nodes corresponding to the user of the image to be processed and the user of any candidate image respectively;

[0020] When the target link does not exist between the nodes corresponding to the user of the image to be processed and any of the candidate images, the product of the weights of each link in the shortest path between the user of the image to be processed and the candidate images in the social network is calculated as the social similarity between the user of the image to be processed and the candidate images, and the weight of each link is the social similarity between the image to be processed and the user of the candidate images represented by the corresponding two nodes.

[0021] Optionally, when the attribute information includes time information, calculating the sub-item similarity between the image to be processed and the candidate image for the attribute information according to the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed includes:

[0022] The absolute value of the difference between the time information of each candidate image and the time information of the image to be processed is calculated, and the absolute value of the natural base e is raised to the power to obtain the time similarity between the image to be processed and the candidate image, and the time similarity is used as the sub-item similarity between the image to be processed and the candidate image for the time information.

[0023] Optionally, when the attribute information includes spatial information, calculating the sub-item similarity between the image to be processed and the candidate image for the attribute information according to the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed includes:

[0024] The 2-norm Euclidean distance between the spatial information of each candidate image and the spatial information of the image to be processed is calculated, the 2-norm Euclidean distance is normalized to obtain the spatial similarity between the image to be processed and the candidate image, and the spatial similarity is used as the sub-item similarity between the image to be processed and the candidate image for the spatial information.

[0025] Optionally, when the attribute information includes aesthetic feature information, calculating the sub-item similarity between the image to be processed and the candidate image for the attribute information based on the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed includes:

[0026] The cosine distance between the aesthetic feature information of each candidate image and the aesthetic feature information of the image to be processed is calculated to obtain the aesthetic feature similarity between the image to be processed and the candidate image, and the aesthetic feature similarity is used as the sub-item similarity for the aesthetic feature information between the image to be processed and the candidate image.

[0027] Optionally, the extracting features of the aesthetic feature information of the image to be processed to obtain the feature information to be processed, and extracting features of the aesthetic feature information of the reference image to obtain reference feature information, includes:

[0028] Inputting the aesthetic feature information of the reference image into a neural network model to obtain first feature information and second feature information of the aesthetic feature information of the reference image as the feature information to be processed;

[0029] Inputting the aesthetic feature information of the image to be processed into a neural network model to obtain third feature information of the aesthetic feature information of the image to be processed as the reference feature information, wherein the third feature information, the first feature information and the second feature information have the same size and correspond to different feature spaces respectively;

[0030] The step of performing feature alignment and feature fusion on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information includes:

[0031] A matrix multiplication operation is performed on the third feature information and the second feature information, and a matrix multiplication operation is performed on the obtained operation result and the first feature information to obtain fused aesthetic feature information.

[0032] Optionally, after decoding the fused aesthetic feature information to obtain a target beautified image, the method further includes:

[0033] Obtaining user evaluation information of the target beautified image;

[0034] The weight of the attribute information is updated according to the user evaluation information.

[0035] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:

[0036] An acquisition unit is configured to acquire the image to be processed and attribute information of the image to be processed, wherein the attribute information includes any one or more of aesthetic feature information, time information, space information, and social information of the user to which the image belongs;

[0037] a calculation unit configured to calculate the attribute similarity between the image to be processed and the candidate image according to the attribute information of each candidate image acquired in advance and the attribute information of the image to be processed, and use the candidate image with the highest attribute similarity as a reference image;

[0038] an extraction unit configured to extract the aesthetic feature information of the image to be processed to obtain the feature information to be processed, and to extract the aesthetic feature information of the reference image to obtain the reference feature information;

[0039] A fusion unit is configured to perform feature alignment and feature fusion on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information;

[0040] The processing unit is configured to perform decoding processing on the fused aesthetic feature information to obtain a target beautified image.

[0041] Optionally, when the attribute information includes multiple items, the computing unit is configured to execute:

[0042] According to the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed, respectively calculating the sub-item similarity between the image to be processed and each candidate image for the attribute information;

[0043] According to the weight of each item of attribute information obtained in advance, the sub-item similarities are weighted and summed to obtain the attribute similarity between the candidate image and the image to be processed.

[0044] Optionally, when the attribute information includes social information of the user, the computing unit is configured to execute:

[0045] According to the pre-acquired social information of the user to which each candidate image belongs and the social information of the user to which the image to be processed belongs, social network information is constructed, wherein each node in the social network represents the image to be processed and the user to which the candidate image belongs, and each connection line between nodes represents a communication between the image to be processed and the user to which the candidate image belongs;

[0046] According to the social network information, the social similarity between the image to be processed and the candidate images is calculated, and the social similarity is used as the sub-item similarity of the social information of the user to which the image to be processed and each candidate image belongs.

[0047] Optionally, the computing unit is specifically configured to execute:

[0048] In the case where there is a target link between the user of the image to be processed and the nodes corresponding to the user of any candidate image, the social similarity between the user of the image to be processed and the candidate image is calculated according to the number of the target links, where the nodes at both ends of the target link are the nodes corresponding to the user of the image to be processed and the user of any candidate image respectively;

[0049] When the target link does not exist between the nodes corresponding to the user of the image to be processed and any of the candidate images, the product of the weights of each link in the shortest path between the user of the image to be processed and the candidate images in the social network is calculated as the social similarity between the user of the image to be processed and the candidate images, and the weight of each link is the social similarity between the image to be processed and the user of the candidate images represented by the corresponding two nodes.

[0050] Optionally, when the attribute information includes time information, the computing unit is configured to execute:

[0051] The absolute value of the difference between the time information of each candidate image and the time information of the image to be processed is calculated, and the absolute value of the natural base e is raised to the power to obtain the time similarity between the image to be processed and the candidate image, and the time similarity is used as the sub-item similarity between the image to be processed and the candidate image for the time information.

[0052] Optionally, when the attribute information includes spatial information, the computing unit is configured to execute:

[0053] The 2-norm Euclidean distance between the spatial information of each candidate image and the spatial information of the image to be processed is calculated, the 2-norm Euclidean distance is normalized to obtain the spatial similarity between the image to be processed and the candidate image, and the spatial similarity is used as the sub-item similarity between the image to be processed and the candidate image for the spatial information.

[0054] Optionally, when the attribute information includes aesthetic feature information, the computing unit is configured to execute:

[0055] The cosine distance between the aesthetic feature information of each candidate image and the aesthetic feature information of the image to be processed is calculated to obtain the aesthetic feature similarity between the image to be processed and the candidate image, and the aesthetic feature similarity is used as the sub-item similarity for the aesthetic feature information between the image to be processed and the candidate image.

[0056] Optionally, the extraction unit is configured to execute:

[0057] Inputting the aesthetic feature information of the reference image into a neural network model to obtain first feature information and second feature information of the aesthetic feature information of the reference image as the feature information to be processed;

[0058] Inputting the aesthetic feature information of the image to be processed into a neural network model to obtain third feature information of the aesthetic feature information of the image to be processed as the reference feature information, wherein the third feature information, the first feature information and the second feature information have the same size and correspond to different feature spaces respectively;

[0059] The fusion unit is configured to perform:

[0060] A matrix multiplication operation is performed on the third feature information and the second feature information, and a matrix multiplication operation is performed on the obtained operation result and the first feature information to obtain fused aesthetic feature information.

[0061] Optionally, the device further includes an updating unit configured to execute:

[0062] Obtaining user evaluation information of the target beautified image;

[0063] The weight of the attribute information is updated according to the user evaluation information.

[0064] According to a third aspect of an embodiment of the present disclosure, there is provided an image processing electronic device, including:

[0065] processor;

[0066] a memory for storing instructions executable by the processor;

[0067] Wherein, the processor is configured to execute the instructions to implement the image processing method described in the first item above.

[0068] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an image processing electronic device, the image processing electronic device is enabled to perform the image processing method described in the first item above.

[0069] According to a fifth aspect of an embodiment of the present disclosure, there is provided a computer program product, comprising a computer program, wherein the computer program implements the image processing method described in the first item above when executed by a processor.

[0070] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0071] The image to be processed and its attribute information are obtained, wherein the attribute information includes any one or more of aesthetic feature information, time information, space information and social information of the user to which it belongs; the attribute similarity between the image to be processed and the candidate images is calculated based on the attribute information of each candidate image obtained in advance and the attribute information of the image to be processed, and the candidate image with the highest attribute similarity is used as the reference image; the aesthetic feature information of the image to be processed is extracted to obtain the features of the feature information to be processed, and the features of the aesthetic feature information of the reference image are extracted to obtain the reference feature information; feature alignment and feature fusion are performed on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information; the fused aesthetic feature information is decoded to obtain the target beautified image.

[0072] In this way, the reference image with the highest attribute similarity to the image to be processed can be found based on the attribute information such as aesthetic feature information, time information, space information and social information of the user to which the image to be processed belongs. The higher the attribute similarity, the more similar the aesthetic features of the reference image and the image to be processed, the closer the time and space distances are, and the higher the social intimacy of the users, then the more likely it is that the reference image and the user to which the image to be processed belong have consistent aesthetic orientations. Then, the reference image is combined with the image to be processed for feature alignment and fusion, and the target beautified image is obtained after decoding, which improves the beautification effect of the target beautified image, and the processing method is more accurate and more likely to meet the aesthetic needs of the user to which the image to be processed belongs.

[0073] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0075] Figure 1 The figure is a flowchart of an image processing method according to an exemplary embodiment.

[0076] Figure 2 The diagram is a logical diagram showing social network information according to an exemplary embodiment.

[0077] Figure 3 The diagram is a logic diagram of feature alignment and feature fusion according to an exemplary embodiment.

[0078] Figure 4 is a block diagram of an image processing apparatus according to an exemplary embodiment.

[0079] Figure 5A block diagram of an image processing electronic device is shown according to an exemplary embodiment.

[0080] Figure 6 The invention is a block diagram of a device for image processing according to an exemplary embodiment. DETAILED DESCRIPTION

[0081] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0082] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0083] Figure 1 is a flowchart of an image processing method according to an exemplary embodiment. The image processing method includes the following steps.

[0084] In step S11, the image to be processed and its attribute information are obtained, wherein the attribute information includes any one or more of aesthetic feature information, time information, space information and social information of the user to which the image belongs.

[0085] Among them, the image to be processed can be a portrait, a landscape or other types of images, without specific limitation. Aesthetic feature information can include information such as color features, texture features, light and shadow features, depth of field features, etc. of the image, which can reflect the aesthetic needs of the user to whom the image belongs. Time information is the time information corresponding to the image, which can be the acquisition time or upload time of the image, or the time reflected by the image content, etc. Spatial information is the spatial information corresponding to the image, which can be the acquisition location or upload location of the image, or the geographical location reflected by the image content, etc. The social information of the user to whom the image belongs can reflect the social situation of the user to whom the image belongs. The more frequent the social interaction between two users, the closer the relationship between the two users.

[0086] Generally, users' aesthetic orientations are continuous, temporal, regional and social. That is to say, the aesthetic needs of the same user for multiple images may be similar, and the aesthetic needs of the same era may be even more similar. For example, the aesthetic needs of users 10 years ago are completely different from those of users now, and users' aesthetic needs are easily affected by recent hot events. In addition, the aesthetic needs of users in different regions are also likely to be different, and the aesthetic needs between close friends tend to be similar.

[0087] In step S12, the attribute similarity between the image to be processed and the candidate images is calculated based on the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed, and the candidate image with the highest attribute similarity is used as the reference image.

[0088] In this step, the attribute information of the image to be processed may be one item or multiple items. When the attribute information includes multiple items, the sub-item similarity between the image to be processed and the candidate image for each item of attribute information can be calculated based on the attribute information of each candidate image obtained in advance and the attribute information of the image to be processed; then, the sub-item similarities are weighted and summed according to the weight of each item of attribute information obtained in advance to obtain the attribute similarity between the candidate image and the image to be processed.

[0089] In the present disclosure, the higher the attribute similarity between the image to be processed and the candidate image, the more likely it is that the aesthetic orientation of the user to whom the candidate image and the image to be processed belong is consistent. In this way, when the attribute information includes multiple information such as aesthetic feature information, time information, spatial information, and social information of the user, the sub-item similarities of each attribute information are weighted and summed, and the obtained attribute similarity is more comprehensive, and the aesthetic orientation of the user to whom the reference image belongs and the user to whom the image to be processed belongs is more likely to be consistent.

[0090] For example, if the attribute information may include time information, space information, social information of the user, and aesthetic feature information, then the above process may adopt the following formula:

[0091] score=alpha_t*score_t+alpha_loc*score_loc+alpha_social*score_social+alpha_aes*score_aes

[0092] Among them, score_t represents the sub-item similarity corresponding to time information, alpha_t represents the weight corresponding to time information, score_loc represents the sub-item similarity corresponding to spatial information, alpha_loc represents the weight corresponding to spatial information, score_social represents the sub-item similarity corresponding to the social information of the user, alpha_social represents the weight corresponding to the social information of the user, score_aes represents the sub-item similarity corresponding to aesthetic feature information, alpha_aes represents the weight corresponding to aesthetic feature information. The weight of each attribute information is a decimal in [0,1].

[0093] In one implementation, after obtaining the target beautified image, user evaluation information of the target beautified image may be obtained, and then the weight of the attribute information may be updated according to the user evaluation information.

[0094] Among them, user evaluation information can be reflected through the user's operation on the target beautified image. For example, when the user forwards, comments, likes or collects the target beautified image, it means that the user has a high evaluation of the target beautified image. Through the user's operation, the user's evaluation information on the target beautified image can be determined, and then, the value of the weight of the attribute information is adjusted through the gradient descent algorithm. In this way, the weight of each user corresponding to each attribute information is different, which can further realize the personalization of image processing.

[0095] In this step, if the attribute information includes the social information of the user, then social network information can be constructed based on the pre-acquired social information of the user of each candidate image and the social information of the user of the image to be processed, where each node in the social network represents the image to be processed and the user of the candidate image, and each line between the nodes represents an exchange between the image to be processed and the user of the candidate image; then, based on the social network information, the social similarity between the image to be processed and the candidate images is calculated, and the social similarity is used as the sub-item similarity of the social information of the user between the image to be processed and each candidate image.

[0096] In this way, through social network information, the social situation between the users to whom the image to be processed and the candidate images belong can be analyzed. Generally, the more frequent the social interactions, the closer the relationship, and the aesthetic orientation consistency between users with close relationships is higher. Therefore, based on the social similarity between the image to be processed and the candidate images, a reference image that is closer to the aesthetic needs of the user to whom the image to be processed belongs can be determined.

[0097] Among them, according to the social network information, calculating the social similarity between the image to be processed and each candidate image, which may specifically include: when there is a target link between the user of the image to be processed and the node corresponding to the user of any candidate image, according to the number of target links, calculating the social similarity between the user of the image to be processed and any candidate image, the two end nodes of the target link are respectively the nodes corresponding to the user of the image to be processed and the user of any candidate image; when there is no target link between the user of the image to be processed and the node corresponding to the user of any candidate image, calculating the product of the weight of each link in the shortest path between the user of the image to be processed and the candidate image in the social network as the social similarity between the user of the image to be processed and the candidate image, the weight of each link is the social similarity between the image to be processed and the user of the candidate image represented by the corresponding two nodes.

[0098] That is to say, the social similarity between the image to be processed and the candidate images is determined based on the number of communications between the users to which they belong. The more frequent the communications between the users to which they belong, the higher the social similarity between the image to be processed and the candidate images. When there is no communication between the users to which the image to be processed and the candidate images belong, the social similarity between the image to be processed and the candidate images can also be determined based on the shortest path between the two nodes.

[0099] For example, Figure 2 As shown in FIG. 1 , it is a logical diagram of social network information, where each node in the social network represents a user, and the user can be a user of the image to be processed or the candidate image, including user 1, user 2, user 3, user 4 and user 5, and each line represents a communication between two users corresponding to the two end nodes. Each user shoots or uploads an image to be processed or a candidate image, and a series of associated data can be established for the image to be processed or the candidate image. For example, an image uploaded by user 1 can be represented as image 1 (aesthetic feature information, time information, spatial information), and each image has a unique identifier image_id. In addition, the set of all other users with social relationships corresponding to each user can be named user_friend.

[0100] When calculating social similarity, it is calculated by the frequency of communication between users. For example, if the number of target connections between two users exceeds 100, the social similarity is considered to be 1. If the number of target connections is less than 100, the social similarity is the number of target connections divided by 100. If there is no target connection between two users, it is necessary to calculate the product of the weights of the shortest paths of the two users on the social network.

[0101] If the attribute information includes time information, then the step of calculating the sub-item similarity between the image to be processed and the candidate images for the time information may include: calculating the absolute value of the difference between the time information of each candidate image and the time information of the image to be processed, and calculating the absolute value of the natural base e to the power to obtain the time similarity between the image to be processed and the candidate images, and using the time similarity as the sub-item similarity between the image to be processed and the candidate images for the time information.

[0102] In this way, through the time information, the time similarity between the image to be processed and the candidate images can be analyzed. Usually, users have more consistent aesthetic orientations towards images with higher time similarity. For example, they may have similar aesthetic trends in the same period. Therefore, based on the time similarity between the image to be processed and the candidate images, a reference image that is closer to the aesthetic needs of the user to whom the image to be processed belongs can be determined.

[0103] If the attribute information includes spatial information, then the step of calculating the sub-item similarity between the image to be processed and the candidate images for the spatial information may include: calculating the 2-norm Euclidean distance between the spatial information of each candidate image and the spatial information of the image to be processed, normalizing the 2-norm Euclidean distance to obtain the spatial similarity between the image to be processed and the candidate images, and using the spatial similarity as the sub-item similarity between the image to be processed and the candidate images for the spatial information.

[0104] In this way, the spatial similarity between the image to be processed and the candidate images can be analyzed through spatial information. The spatial information can be the longitude and latitude of the image collection location. Usually, the aesthetic orientation of users in the same region is relatively consistent. For example, users in different countries may have different aesthetic orientations, while users in the same country may have relatively consistent aesthetic orientations due to various factors. Therefore, based on the spatial similarity between the image to be processed and the candidate images, a reference image that is closer to the aesthetic needs of the user to whom the image to be processed belongs can be determined.

[0105] If the attribute information includes aesthetic feature information, then the step of calculating the sub-item similarity of the aesthetic feature information between the image to be processed and the candidate images may include: calculating the cosine distance between the aesthetic feature information of each candidate image and the aesthetic feature information of the image to be processed, obtaining the aesthetic feature similarity between the image to be processed and the candidate images, and using the aesthetic feature similarity as the sub-item similarity of the aesthetic feature information between the image to be processed and the candidate images.

[0106] In this way, the aesthetic feature similarity between the image to be processed and the candidate images can be analyzed through the aesthetic feature information. The aesthetic feature information can be information such as the color harmony, rule of thirds, lighting, depth of field, visual balance, object emphasis, etc. of the image. When the aesthetic feature information between images is similar in phase, the user's processing of the images is also likely to be similar. For example, for images that lack exposure, it is usually necessary to increase the exposure, and for images with low contrast, it is usually necessary to increase the contrast, and so on. Therefore, based on the aesthetic feature similarity between the image to be processed and the candidate images, a reference image that is closer to the aesthetic needs of the user to whom the image to be processed belongs can be determined.

[0107] In step S13, features of the aesthetic feature information of the image to be processed are extracted to obtain the feature information to be processed, and features of the aesthetic feature information of the reference image are extracted to obtain the reference feature information.

[0108] In this step, the aesthetic feature information of the reference image can be input into the neural network model to obtain the first feature information and the second feature information of the aesthetic feature information of the reference image as the feature information to be processed; and the aesthetic feature information of the image to be processed can be input into the neural network model to obtain the third feature information of the aesthetic feature information of the image to be processed as the reference feature information. The third feature information, the first feature information and the second feature information have the same size and correspond to different feature spaces respectively.

[0109] For example, the value feature and key feature of the aesthetic feature information of the reference image can be extracted as the first feature information and the second feature information, respectively, and the query feature of the aesthetic feature information of the image to be processed can be extracted as the third feature information. Among them, the value feature is the feature of the aesthetic feature information of the reference image in the value feature space, the key feature is the feature of the aesthetic feature information of the reference image in the key feature space, and the query feature is the feature of the aesthetic feature information of the image to be processed in the query feature space. The value, key and query represent different feature spaces respectively. The neural network model can map the input information to different feature spaces, extract the features of the input information in different feature spaces, and obtain the feature information corresponding to different feature spaces. These feature information have the same size.

[0110] In step S14, feature alignment and feature fusion are performed on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information.

[0111] In this step, feature alignment and feature fusion of the feature information to be processed and the reference feature information may specifically include: performing a matrix multiplication operation on the third feature information and the second feature information, performing a matrix multiplication operation on the result of the operation and the first feature information, and obtaining fused aesthetic feature information.

[0112] In this way, by performing feature extraction, feature alignment and feature fusion on the aesthetic feature information of the image to be processed and the reference image, the aesthetic features of the reference image can be combined with the aesthetic features of the image to be processed. The obtained fused aesthetic feature information can reflect the aesthetic features of the reference image to a certain extent, and is therefore more likely to meet the aesthetic orientation of the user to whom the image to be processed belongs.

[0113] For example, Figure 3 The figure shows a logic diagram of feature alignment and feature fusion.

[0114] The aesthetic feature information of the reference image and the aesthetic feature information of the image to be processed are input into a CNN (Convolutional Neural Networks). Assuming that the sizes of the two aesthetic feature information are both HxWxC, CNN extracts the value features of the aesthetic feature information of the reference image, and then reshapes them into a size of (HxW)xC, extracts the key features of the aesthetic feature information of the reference image, and then reshapes them into a size of (HxW)xC, extracts the query features of the aesthetic feature information of the image to be processed, and then reshapes them into a size of (HxW)xC.

[0115] Then, the query feature and the key feature are matrix multiplied, and the size of the matrix multiplication result is (HxW)x(HxW). The matrix multiplication result of the query feature and the key feature is matrix multiplied with the value feature to obtain the result of the feature alignment and feature fusion module, that is, the fused aesthetic feature information of size HxWxC.

[0116] In step S15, the fused aesthetic feature information is decoded to obtain a target beautified image.

[0117] The decoding process is the process of reconstructing an image based on feature information. In this step, the aesthetic decoder can be used to decode the fused aesthetic feature information to obtain a target beautified image. Since the aesthetic features of the reference image with a high attribute similarity to the image to be processed are referenced in the beautification process of the target beautified image, the target beautified image is more likely to conform to the aesthetic orientation of the user to whom the image to be processed belongs.

[0118] From the above, it can be seen that the technical solution provided by the embodiments of the present disclosure can find the reference image with the highest attribute similarity with the image to be processed based on the attribute information such as aesthetic feature information, time information, spatial information of the image to be processed and social information of the user to which it belongs. The higher the attribute similarity, the more similar the aesthetic features of the reference image and the image to be processed, the closer the time and space distance, the higher the social intimacy of the user to which they belong, and the more likely it is that the reference image and the user to whom the image to be processed belong have the same aesthetic orientation. Then, the reference image is combined with the reference image to perform feature alignment and fusion on the image to be processed, and the target beautified image is obtained after decoding. The target beautified image is more likely to meet the aesthetic needs of the user to whom the image to be processed belongs.

[0119] Figure 4 is a block diagram of an image processing device according to an exemplary embodiment, the device comprising:

[0120] An acquisition unit 201 is configured to acquire an image to be processed and attribute information of the image to be processed, wherein the attribute information includes any one or more of aesthetic feature information, time information, space information, and social information of a user;

[0121] The calculation unit 202 is configured to calculate the attribute similarity between the image to be processed and the candidate image according to the attribute information of each candidate image acquired in advance and the attribute information of the image to be processed, and use the candidate image with the highest attribute similarity as a reference image;

[0122] The extraction unit 203 is configured to extract features of the aesthetic feature information of the image to be processed to obtain the feature information to be processed, and extract features of the aesthetic feature information of the reference image to obtain reference feature information;

[0123] A fusion unit 204 is configured to perform feature alignment and feature fusion on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information;

[0124] The processing unit 205 is configured to perform decoding processing on the fused aesthetic feature information to obtain a target beautified image.

[0125] In one implementation, when the attribute information includes multiple items, the calculation unit 202 is configured to execute:

[0126] Calculating the sub-item similarity between the image to be processed and each candidate image for each item of attribute information based on the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed;

[0127] According to the weight of each item of attribute information obtained in advance, the sub-item similarities are weighted and summed to obtain the attribute similarity between the candidate image and the image to be processed.

[0128] In one implementation, when the attribute information includes social information of the user, the computing unit 202 is configured to execute:

[0129] According to the pre-acquired social information of the user to which each candidate image belongs and the social information of the user to which the image to be processed belongs, social network information is constructed, wherein each node in the social network represents the image to be processed and the user to which the candidate image belongs, and each connection line between nodes represents a communication between the image to be processed and the user to which the candidate image belongs;

[0130] According to the social network information, the social similarity between the image to be processed and the candidate images is calculated, and the social similarity is used as the sub-item similarity of the social information of the user to which the image to be processed and each candidate image belongs.

[0131] In one implementation, the computing unit 202 is specifically configured to execute:

[0132] In the case where there is a target link between the user of the image to be processed and the nodes corresponding to the user of any candidate image, the social similarity between the user of the image to be processed and the candidate image is calculated according to the number of the target links, where the nodes at both ends of the target link are the nodes corresponding to the user of the image to be processed and the user of any candidate image respectively;

[0133] When the target link does not exist between the nodes corresponding to the user of the image to be processed and any of the candidate images, the product of the weights of each link in the shortest path between the user of the image to be processed and the candidate images in the social network is calculated as the social similarity between the user of the image to be processed and the candidate images, and the weight of each link is the social similarity between the image to be processed and the user of the candidate images represented by the corresponding two nodes.

[0134] In one implementation, when the attribute information includes time information, the calculation unit 202 is configured to execute:

[0135] The absolute value of the difference between the time information of each candidate image and the time information of the image to be processed is calculated, and the absolute value of the natural base e is raised to the power to obtain the time similarity between the image to be processed and the candidate image, and the time similarity is used as the sub-item similarity between the image to be processed and the candidate image for the time information.

[0136] In one implementation, when the attribute information includes spatial information, the computing unit is configured to execute:

[0137] The 2-norm Euclidean distance between the spatial information of each candidate image and the spatial information of the image to be processed is calculated, the 2-norm Euclidean distance is normalized to obtain the spatial similarity between the image to be processed and the candidate image, and the spatial similarity is used as the sub-item similarity between the image to be processed and the candidate image for the spatial information.

[0138] In one implementation, when the attribute information includes aesthetic feature information, the computing unit 202 is configured to execute:

[0139] The cosine distance between the aesthetic feature information of each candidate image and the aesthetic feature information of the image to be processed is calculated to obtain the aesthetic feature similarity between the image to be processed and the candidate image, and the aesthetic feature similarity is used as the sub-item similarity for the aesthetic feature information between the image to be processed and the candidate image.

[0140] In one implementation, the extracting unit 203 is configured to execute:

[0141] Inputting the aesthetic feature information of the reference image into a neural network model to obtain first feature information and second feature information of the aesthetic feature information of the reference image as the feature information to be processed;

[0142] Inputting the aesthetic feature information of the image to be processed into a neural network model to obtain third feature information of the aesthetic feature information of the image to be processed as the reference feature information, wherein the third feature information, the first feature information and the second feature information have the same size and correspond to different feature spaces respectively;

[0143] The fusion unit 204 is configured to perform:

[0144] A matrix multiplication operation is performed on the third feature information and the second feature information, and a matrix multiplication operation is performed on the obtained operation result and the first feature information to obtain fused aesthetic feature information.

[0145] In one implementation, the apparatus further includes an updating unit configured to execute:

[0146] Obtaining user evaluation information of the target beautified image;

[0147] The weight of the attribute information is updated according to the user evaluation information.

[0148] From the above, it can be seen that the technical solution provided by the embodiments of the present disclosure can find the reference image with the highest attribute similarity with the image to be processed based on the attribute information such as aesthetic feature information, time information, spatial information of the image to be processed and social information of the user to which it belongs. The higher the attribute similarity, the more similar the aesthetic features of the reference image and the image to be processed, the closer the time and space distance, the higher the social intimacy of the user to which they belong, and the more likely it is that the reference image and the user to whom the image to be processed belong have the same aesthetic orientation. Then, the reference image is combined with the reference image to perform feature alignment and fusion on the image to be processed, and the target beautified image is obtained after decoding. The target beautified image is more likely to meet the aesthetic needs of the user to whom the image to be processed belongs.

[0149] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0150] Figure 5 It is a block diagram of an image processing electronic device according to an exemplary embodiment.

[0151] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by a processor of an electronic device to perform the above method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0152] In an exemplary embodiment, a computer program product is also provided. When the computer program product is executed on a computer, the computer is enabled to implement the above-mentioned image processing method.

[0153] From the above, it can be seen that the technical solution provided by the embodiments of the present disclosure can find the reference image with the highest attribute similarity with the image to be processed based on the attribute information such as aesthetic feature information, time information, spatial information of the image to be processed and social information of the user to which it belongs. The higher the attribute similarity, the more similar the aesthetic features of the reference image and the image to be processed, the closer the time and space distance, the higher the social intimacy of the user to which they belong, and the more likely it is that the reference image and the user to whom the image to be processed belong have the same aesthetic orientation. Then, the reference image is combined with the reference image to perform feature alignment and fusion on the image to be processed, and the target beautified image is obtained after decoding. The target beautified image is more likely to meet the aesthetic needs of the user to whom the image to be processed belongs.

[0154] Figure 6 is a block diagram of a device 800 for image processing according to an exemplary embodiment.

[0155] For example, apparatus 800 may be a mobile phone, a computer, a digital broadcast electronic device, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0156] Reference Figure 6 , the device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0157] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0158] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0159] The power supply component 807 provides power to the various components of the device 800. The power supply component 807 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.

[0160] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0161] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the device 800 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 404 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0162] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0163] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800, and the sensor assembly 814 can also detect the position change of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0164] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0165] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the methods described in the first and second aspects.

[0166] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the device 800 to complete the above method. Optionally, for example, the storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0167] In an exemplary embodiment, a computer program product containing instructions is also provided. When the computer program product is run on a computer, the computer is enabled to perform the image processing method described in the above embodiment.

[0168] From the above, it can be seen that the technical solution provided by the embodiments of the present disclosure can find the reference image with the highest attribute similarity with the image to be processed based on the attribute information such as aesthetic feature information, time information, spatial information of the image to be processed and social information of the user to which it belongs. The higher the attribute similarity, the more similar the aesthetic features of the reference image and the image to be processed, the closer the time and space distance, the higher the social intimacy of the user to which they belong, and the more likely it is that the reference image and the user to whom the image to be processed belong have the same aesthetic orientation. Then, the reference image is combined with the reference image to perform feature alignment and fusion on the image to be processed, and the target beautified image is obtained after decoding. The target beautified image is more likely to meet the aesthetic needs of the user to whom the image to be processed belongs.

[0169] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0170] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image processing method, characterized in that: include: Acquire an image to be processed and attribute information of the image to be processed, wherein the attribute information includes aesthetic feature information, time information, space information, and social information of a user; Calculating the sub-item similarity between the image to be processed and the candidate image for the attribute information according to the attribute information of each candidate image acquired in advance and the attribute information of the image to be processed; According to the weight of each attribute information obtained in advance, weighted sum calculation is performed on the sub-item similarities to obtain the attribute similarity between the candidate image and the image to be processed, and the candidate image with the highest attribute similarity is used as the reference image; Extracting features of the aesthetic feature information of the image to be processed to obtain feature information to be processed, and extracting features of the aesthetic feature information of the reference image to obtain reference feature information; Performing feature alignment and feature fusion on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information; Decoding the fused aesthetic feature information to obtain a target beautified image; The calculating of the sub-item similarity of the attribute information between the image to be processed and the candidate image according to the attribute information of each candidate image acquired in advance and the attribute information of the image to be processed includes: According to the pre-acquired social information of the user to which each candidate image belongs and the social information of the user to which the image to be processed belongs, constructing social network information, wherein each node in the social network represents the user to which the image to be processed belongs and the user to which the candidate image belongs, and each connection line between the nodes represents a communication between the user to which the image to be processed belongs and the user to which the candidate image belongs; According to the social network information, the social similarity between the image to be processed and the candidate images is calculated, and the social similarity is used as the sub-item similarity of the social information of the user to which the image to be processed and each candidate image belongs.

2. The image processing method according to claim 1, characterized in that: The step of calculating the social similarity between the image to be processed and each candidate image according to the social network information includes: In the case where there is a target link between the user of the image to be processed and the nodes corresponding to the user of any candidate image, the social similarity between the user of the image to be processed and the candidate image is calculated according to the number of the target links, where the nodes at both ends of the target link are the nodes corresponding to the user of the image to be processed and the user of any candidate image respectively; When the target link does not exist between the nodes corresponding to the user of the image to be processed and any of the candidate images, the product of the weights of each link in the shortest path between the user of the image to be processed and the candidate images in the social network is calculated as the social similarity between the user of the image to be processed and the candidate images, and the weight of each link is the social similarity between the image to be processed and the user of the candidate images represented by the corresponding two nodes.

3. The image processing method according to claim 1, characterized in that: The calculating, based on the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed, the sub-item similarity between the image to be processed and the candidate image for the attribute information includes: The absolute value of the difference between the time information of each candidate image and the time information of the image to be processed is calculated, and the absolute value of the natural base e is raised to the power to obtain the time similarity between the image to be processed and the candidate image, and the time similarity is used as the sub-item similarity between the image to be processed and the candidate image for the time information.

4. The image processing method according to claim 1, characterized in that: The calculating, based on the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed, the sub-item similarity between the image to be processed and the candidate image for the attribute information includes: The 2-norm Euclidean distance between the spatial information of each candidate image and the spatial information of the image to be processed is calculated, the 2-norm Euclidean distance is normalized to obtain the spatial similarity between the image to be processed and the candidate image, and the spatial similarity is used as the sub-item similarity between the image to be processed and the candidate image for the spatial information.

5. The image processing method according to claim 1, characterized in that: The calculating, based on the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed, the sub-item similarity between the image to be processed and the candidate image for the attribute information includes: The cosine distance between the aesthetic feature information of each candidate image and the aesthetic feature information of the image to be processed is calculated to obtain the aesthetic feature similarity between the image to be processed and the candidate image, and the aesthetic feature similarity is used as the sub-item similarity for the aesthetic feature information between the image to be processed and the candidate image.

6. The image processing method according to claim 1, characterized in that: The step of extracting the features of the aesthetic feature information of the image to be processed to obtain the feature information to be processed, and extracting the features of the aesthetic feature information of the reference image to obtain the reference feature information, comprises: Inputting the aesthetic feature information of the reference image into a neural network model to obtain first feature information and second feature information of the aesthetic feature information of the reference image as the feature information to be processed; Inputting the aesthetic feature information of the image to be processed into a neural network model to obtain third feature information of the aesthetic feature information of the image to be processed as the reference feature information, wherein the third feature information, the first feature information and the second feature information have the same size and correspond to different feature spaces respectively; The step of performing feature alignment and feature fusion on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information includes: A matrix multiplication operation is performed on the third feature information and the second feature information, and a matrix multiplication operation is performed on the obtained operation result and the first feature information to obtain fused aesthetic feature information.

7. The image processing method according to claim 1, characterized in that: After decoding the fused aesthetic feature information to obtain a target beautified image, the method further includes: Obtaining user evaluation information of the target beautified image; The weight of the attribute information is updated according to the user evaluation information.

8. An image processing device, characterized in that: include: An acquisition unit is configured to acquire an image to be processed and attribute information of the image to be processed, wherein the attribute information includes aesthetic feature information, time information, space information, and social information of a user; A calculation unit configured to respectively calculate the sub-item similarity between the image to be processed and the candidate image for each item of attribute information based on the pre-acquired attribute information of each candidate image and the attribute information of the image to be processed; According to the weight of each attribute information obtained in advance, weighted sum calculation is performed on the sub-item similarities to obtain the attribute similarity between the candidate image and the image to be processed, and the candidate image with the highest attribute similarity is used as the reference image; an extraction unit configured to extract features of the aesthetic feature information of the image to be processed to obtain the feature information to be processed, and to extract features of the aesthetic feature information of the reference image to obtain reference feature information; A fusion unit is configured to perform feature alignment and feature fusion on the feature information to be processed and the reference feature information to obtain fused aesthetic feature information; A processing unit is configured to perform decoding processing on the fused aesthetic feature information to obtain a target beautified image; Wherein, the computing unit is configured to execute: According to the pre-acquired social information of the user to which each candidate image belongs and the social information of the user to which the image to be processed belongs, social network information is constructed, wherein each node in the social network represents the image to be processed and the user to which the candidate image belongs, and each connection line between nodes represents a communication between the image to be processed and the user to which the candidate image belongs; According to the social network information, the social similarity between the image to be processed and the candidate images is calculated, and the social similarity is used as the sub-item similarity of the social information of the user to which the image to be processed and each candidate image belongs.

9. The image processing device according to claim 8, characterized in that: The computing unit is specifically configured to execute: In the case where there is a target link between the user of the image to be processed and the nodes corresponding to the user of any candidate image, the social similarity between the user of the image to be processed and the candidate image is calculated according to the number of the target links, where the nodes at both ends of the target link are the nodes corresponding to the user of the image to be processed and the user of any candidate image respectively; When the target link does not exist between the nodes corresponding to the user of the image to be processed and any of the candidate images, the product of the weights of each link in the shortest path between the user of the image to be processed and the candidate images in the social network is calculated as the social similarity between the user of the image to be processed and the candidate images, and the weight of each link is the social similarity between the image to be processed and the user of the candidate images represented by the corresponding two nodes.

10. The image processing device according to claim 8, characterized in that: The computing unit is configured to execute: The absolute value of the difference between the time information of each candidate image and the time information of the image to be processed is calculated, and the absolute value of the natural base e is raised to the power to obtain the time similarity between the image to be processed and the candidate image, and the time similarity is used as the sub-item similarity between the image to be processed and the candidate image for the time information.

11. The image processing device according to claim 8, characterized in that: The computing unit is configured to execute: The 2-norm Euclidean distance between the spatial information of each candidate image and the spatial information of the image to be processed is calculated, the 2-norm Euclidean distance is normalized to obtain the spatial similarity between the image to be processed and the candidate image, and the spatial similarity is used as the sub-item similarity between the image to be processed and the candidate image for the spatial information.

12. The image processing device according to claim 8, characterized in that: The computing unit is configured to execute: The cosine distance between the aesthetic feature information of each candidate image and the aesthetic feature information of the image to be processed is calculated to obtain the aesthetic feature similarity between the image to be processed and the candidate image, and the aesthetic feature similarity is used as the sub-item similarity for the aesthetic feature information between the image to be processed and the candidate image.

13. The image processing device according to claim 8, characterized in that: The extraction unit is configured to execute: Inputting the aesthetic feature information of the reference image into a neural network model to obtain first feature information and second feature information of the aesthetic feature information of the reference image as the feature information to be processed; Inputting the aesthetic feature information of the image to be processed into a neural network model to obtain third feature information of the aesthetic feature information of the image to be processed as the reference feature information, wherein the third feature information, the first feature information and the second feature information have the same size and correspond to different feature spaces respectively; The fusion unit is configured to perform: A matrix multiplication operation is performed on the third feature information and the second feature information, and a matrix multiplication operation is performed on the obtained operation result and the first feature information to obtain fused aesthetic feature information.

14. The image processing device according to claim 8, characterized in that: The apparatus further comprises an updating unit configured to execute: Obtaining user evaluation information of the target beautified image; The weight of the attribute information is updated according to the user evaluation information.

15. An image processing electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image processing method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an image processing electronic device, the image processing electronic device is enabled to execute the image processing method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Image processing method and device, interactive display device and electronic equipment

    CN111986076A

  • Video recognition method and device and computer readable storage medium

    CN112580599A