Image cutting method and device, equipment, storage medium and computer program product
Through the methods of object detection, semantic segmentation, visual evaluation and dynamic cropping, the problem of poor display effect of hotel appearance pictures due to different sizes and aspect ratios is solved, achieving better visual experience and unified display effect.
Patent Information
- Application Number
- CN202510133344.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-23
AI Technical Summary
In the hotel industry, due to different shooting angle, equipment and time, the picture size and aspect ratio are different, and it is difficult to unify the display, resulting in poor visual effects.
The initial image is subjected to detect the initial image through the object detection model to determine the first cropped image; the first cropped image is semantically segmented, and the second cropped image is obtained; the second cropped image is visually evaluated, and the third cropped image is generated; based on the size of the image display container and the size of the initial image, the third cropped image is dynamically cropped to generate the target cropped image.
It improves the display effect of the image, avoids visual losses caused by stretching or compression of the image, and reduces the investment in manpower and material resources, improving the timeliness and uniformity of display.
Smart Images

Figure CN120032129A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image cropping method, apparatus, device, storage medium and computer program product. Background Art
[0002] In the hotel industry, in order to improve online display effects and user booking experience, hotels usually need to upload their exterior pictures to various online booking platforms or self-built websites. However, due to the differences in shooting angles, shooting equipment, shooting time and other factors of hotel exterior pictures, the uploaded pictures often have different sizes and aspect ratios, which brings challenges to the display of pictures.
[0003] Currently, there are two main types of technologies on the market to address this challenge:
[0004] 1. Fixed aspect ratio display technology. By setting a fixed aspect ratio, no matter what the size of the uploaded image is, it will be displayed in this ratio. Although this method solves the problem of uniformity in image display to a certain extent, the visual effect of the image is greatly reduced because images of different sizes are rigidly stretched or compressed to fit the fixed aspect ratio, and the image may be deformed, thus affecting the user's visual experience.
[0005] 2. Multi-size image production technology. According to the requirements of different scenes, the original images are cropped in advance and saved as images of different sizes. This method effectively avoids the problems caused by image stretching or compression and improves the visual effect. However, this solution requires a lot of manpower and material resources for image preprocessing and storage, which is costly and time-consuming. In addition, as screen size and resolution continue to change, the image library needs to be constantly updated to adapt to new display requirements, which greatly increases the complexity of the solution.
[0006] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention
[0007] The main purpose of the present application is to provide an image cropping method, device, equipment, storage medium and computer program product, aiming to solve the technical problem of poor image display effect.
[0008] To achieve the above object, the present application proposes an image cropping method, which comprises:
[0009] Performing subject detection on the initial image by using an object detection model to determine a first cropped image of the initial image;
[0010] performing semantic segmentation on the first cropped image, and cropping the first cropped image based on a result of the semantic segmentation to obtain a second cropped image of the first cropped image;
[0011] performing a visual assessment on the second cropped image, and cropping the second cropped image based on a result of the visual assessment to generate a third cropped image of the second cropped image;
[0012] Based on the size of the image display container and the size of the initial image, the third cropped image is dynamically cropped to generate a target cropped image of the third cropped image.
[0013] In one embodiment, the step of performing subject detection on the initial image by using the target detection model to determine the first cropped image of the initial image comprises:
[0014] Conduct target subject detection and recognition training on the target detection model;
[0015] Performing subject detection on the initial image by detecting and recognizing the trained target detection model to obtain a detection result of the target subject in the initial image;
[0016] Based on the detection result of the target subject, a first cropped image of the initial image is determined.
[0017] In one embodiment, the step of performing semantic segmentation on the first cropped image, cropping the first cropped image based on a result of the semantic segmentation, and obtaining a second cropped image of the first cropped image comprises:
[0018] Perform semantic category recognition and segmentation training on the semantic segmentation model;
[0019] Performing semantic segmentation on the first cropped image by using a semantic segmentation model after recognition segmentation training to obtain a semantic segmentation result of the first cropped image;
[0020] Based on the semantic segmentation result of the first cropped image, the first cropped image is cropped to generate a second cropped image of the first cropped image.
[0021] In one embodiment, the steps of visually evaluating the second cropped image, cropping the second cropped image based on a result of the visual evaluation, and generating a third cropped image of the second cropped image include:
[0022] Training visual evaluation models to evaluate aesthetic composition rules;
[0023] Performing a visual evaluation on the second cropped image by evaluating the trained visual evaluation model to generate a visual score map of the second cropped image;
[0024] The second cropped image is cropped based on the visual score map to generate a third cropped image of the second cropped image.
[0025] In one embodiment, the step of cropping the second cropped image based on the visual score map to generate a third cropped image of the second cropped image comprises:
[0026] Based on the second cropped image, generating a plurality of candidate cropping frames of the second cropped image;
[0027] Based on the visual scoring graphs of the plurality of candidate cropping frames, visually scoring the second cropped image using the visual scoring graphs to obtain a plurality of visual scoring results;
[0028] Based on the visual scoring result, a third cropped image of the second cropped image is determined.
[0029] In one embodiment, the step of dynamically cropping the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image includes:
[0030] Obtaining the real-time size of the image display container and the aspect ratio of the initial image;
[0031] Based on the real-time size of the image display container and the aspect ratio of the initial image, the third cropped image is dynamically cropped and adjusted to generate a target cropped image of the third cropped image.
[0032] In addition, to achieve the above-mentioned purpose, the present application also proposes an image cropping device, the image cropping device comprising:
[0033] A subject detection module, used to perform subject detection on the initial image through an object detection model to determine a first cropped image of the initial image;
[0034] a semantic segmentation module, configured to perform semantic segmentation on the first cropped image, and crop the first cropped image based on a result of the semantic segmentation to obtain a second cropped image of the first cropped image;
[0035] a visual assessment module, configured to perform a visual assessment on the second cropped image, crop the second cropped image based on a result of the visual assessment, and generate a third cropped image of the second cropped image;
[0036] The dynamic cropping module is used to dynamically crop the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes an image cropping device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image cropping method described above.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image cropping method described above are implemented.
[0039] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the image cropping method described above are implemented.
[0040] One or more technical solutions proposed in this application have at least the following technical effects:
[0041] The embodiments of the present application propose an image cropping method, apparatus, device, storage medium and computer program product, which performs subject detection on an initial image through a target detection model to determine a first cropped image of the initial image; performs semantic segmentation on the first cropped image, crops the first cropped image based on the result of the semantic segmentation, and obtains a second cropped image of the first cropped image; performs visual evaluation on the second cropped image, crops the second cropped image based on the result of the visual evaluation, and generates a third cropped image of the second cropped image; and dynamically crops the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image. The present application can improve the display effect of an image by performing subject detection, semantic segmentation, visual evaluation and dynamic cropping on an initial image. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0044] Figure 1 A schematic diagram of a process flow provided for the first embodiment of the image cropping method of the present application;
[0045] Figure 2 This is a schematic diagram of the module structure of the image cropping device according to an embodiment of the present application;
[0046] Figure 3 Schematic diagram of the device structure of the hardware operating environment involved in the image cropping method in the embodiment of the present application. DETAILED DESCRIPTION
[0047] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0048] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0049] The main solution of the embodiment of the present application is: perform subject detection on the initial image through the target detection model to determine the first cropped image of the initial image; perform semantic segmentation on the first cropped image, crop the first cropped image based on the result of the semantic segmentation, and obtain a second cropped image of the first cropped image; perform visual evaluation on the second cropped image, crop the second cropped image based on the result of the visual evaluation, and generate a third cropped image of the second cropped image; based on the size of the image display container and the size of the initial image, dynamically crop the third cropped image to generate a target cropped image of the third cropped image.
[0050] In this embodiment, for the convenience of description, the following description is made by taking the image recognition and cropping device as the execution subject.
[0051] In the hotel industry, in order to improve online display effects and user booking experience, hotels usually need to upload their exterior pictures to various online booking platforms or self-built websites. However, due to the differences in shooting angles, shooting equipment, shooting time and other factors of hotel exterior pictures, the uploaded pictures often have different sizes and aspect ratios, which brings challenges to the display of pictures.
[0052] Currently, there are two main types of technologies on the market to address this challenge:
[0053] 1. Fixed aspect ratio display technology. By setting a fixed aspect ratio, no matter what the size of the uploaded image is, it will be displayed in this ratio. Although this method solves the problem of uniformity in image display to a certain extent, the visual effect of the image is greatly reduced because images of different sizes are rigidly stretched or compressed to fit the fixed aspect ratio, and the image may be deformed, thus affecting the user's visual experience.
[0054] 2. Multi-size image production technology. According to the requirements of different scenes, the original images are cropped in advance and saved as images of different sizes. This method effectively avoids the problems caused by image stretching or compression and improves the visual effect. However, this solution requires a lot of manpower and material resources for image preprocessing and storage, which is costly and time-consuming. In addition, as screen size and resolution continue to change, the image library needs to be constantly updated to adapt to new display requirements, which greatly increases the complexity of the solution.
[0055] The present application provides a solution that can improve the display effect of an image by performing subject detection, semantic segmentation, visual evaluation, and dynamic cropping on an initial image.
[0056] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, an image cropping device, etc. The following takes the image cropping device as an example to illustrate this embodiment and the following embodiments.
[0057] Based on this, the present application embodiment provides an image cropping method, referring to Figure 1 , Figure 1 This is a flowchart diagram of the first embodiment of the image cropping method of the present application.
[0058] In this embodiment, the image cropping method includes steps S11 to S14:
[0059] Step S11, performing subject detection on the initial image by using an object detection model to determine a first cropped image of the initial image.
[0060] It should be noted that the target detection model refers to an algorithm model that can detect and locate key target areas (such as faces, objects, etc.) in the initial image, by analyzing the content of the initial image, determining whether there are main objects of predefined categories in the image, and returning the locations of these objects (usually in the form of bounding boxes). In one embodiment of the present application, the target detection model refers to the YOLOv5 algorithm model.
[0061] In addition, it should be noted that the initial image refers to an original image that has not been processed or cropped. The first cropped image refers to an image obtained after the initial image is subjected to subject detection by the object detection model.
[0062] Specifically, the target detection model is trained to detect and identify the target subject; the subject detection is performed on the initial image by the target detection model after the detection and identification training to obtain the detection result of the target subject in the initial image; based on the detection result of the target subject, the first cropping area of the initial image is determined.
[0063] Step S12: performing semantic segmentation on the first cropped image, and cropping the first cropped image based on a result of the semantic segmentation to obtain a second cropped image of the first cropped image.
[0064] It should be noted that the second cropped image is a further cropped image after semantic segmentation is performed on the first cropped image.
[0065] Specifically, the semantic segmentation model is trained for recognition and segmentation of semantic categories; the first cropped image is semantically segmented by the semantic segmentation model after the recognition and segmentation training to obtain a semantic segmentation result of the first cropped image; based on the semantic segmentation result of the first cropped image, the first cropped image is cropped to generate a second cropped image of the first cropped image.
[0066] Step S13: Perform visual evaluation on the second cropped image, and crop the second cropped image based on the result of the visual evaluation to generate a third cropped image of the second cropped image.
[0067] It should be noted that the third cropped image is an image that is further cropped after visual evaluation of the second cropped image.
[0068] Specifically, the visual evaluation model is trained to evaluate aesthetic composition rules; the second cropped image is visually evaluated by evaluating the trained visual evaluation model to generate a visual score map of the second cropped image; the second cropped image is cropped based on the visual score map to generate a third cropped image of the second cropped image.
[0069] Step S14: dynamically cropping the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image.
[0070] It should be noted that the image display container refers to an interface element or area used to display an image, and its size is fixed. In a web page, application or other graphical interface, the image display container defines the maximum size that an image can be displayed.
[0071] Specifically, the real-time size of the image display container and the aspect ratio of the initial image are obtained; based on the real-time size of the image display container and the aspect ratio of the initial image, the third cropped image is dynamically cropped and adjusted to generate a target cropped image of the third cropped image.
[0072] Through the above scheme, this embodiment performs subject detection on the initial image through the target detection model to determine the first cropped image of the initial image; performs semantic segmentation on the first cropped image, crops the first cropped image based on the result of the semantic segmentation, and obtains the second cropped image of the first cropped image; performs visual evaluation on the second cropped image, crops the second cropped image based on the result of the visual evaluation, and generates a third cropped image of the second cropped image; based on the size of the image display container and the size of the initial image, dynamically crops the third cropped image to generate a target cropped image of the third cropped image. This application can improve the display effect of the image by performing subject detection, semantic segmentation, visual evaluation, and dynamic cropping on the initial image.
[0073] Based on the above implementation scheme, in a feasible implementation manner, the step of performing subject detection on the initial image through the target detection model to determine the first cropped image of the initial image includes S21 to S23:
[0074] Step S21, performing target subject detection and recognition training on the target detection model.
[0075] It should be noted that the target subject refers to the subject object that is selected and expected to be accurately detected and identified by the trained target detection model in the image subject detection and identification. This subject can be various, such as people, animals, vehicles, buildings, specific items, etc., depending on the actual application scenario and needs. In the training process, the target subject is the learning focus of the target detection model, and the model learns its characteristics and attributes so that it can accurately identify and locate the subject from the image in subsequent processing. In the obtained detection result, the target subject is the key object in the image confirmed and marked by the model. In one embodiment of the present application, the initial image is an image of a hotel, and the target object refers to the hotel appearance (doors and windows of the hotel, signs, etc.) and the hotel interior (beds, toilets, etc.) in the hotel image.
[0076] Specifically, an image set containing a target subject is obtained, and the target subject in the image is annotated by an annotation tool (such as LabelImg), and an annotation file suitable for the target detection model format is generated, and the annotation file contains the category and bounding box information of the target subject. The image set is divided into a training set and a validation set, and the target detection model (YOLOv5 algorithm model) is initialized according to the characteristics of the image set and the training requirements, such as setting the number of target categories, the size of the input initial image, etc. The model is trained based on the training set by the YOLOv5 algorithm model, and the network parameters of the model are continuously updated during the training process to obtain the target detection model with the minimum loss function, and the model training is completed. Then, the target detection model that has completed the model training is verified by the validation set, the performance of the model is evaluated, and the detection and recognition training of the target subject of the target detection model is completed.
[0077] Step S22, performing subject detection on the initial image by using the trained target detection model to obtain a detection result of the target subject in the initial image.
[0078] It should be noted that the detection result of the target subject refers to the target bounding box obtained after subject detection is performed on the initial image, including the coordinates of the target bounding box.
[0079] Specifically, the initial image is preprocessed, such as adjusting the image size to meet the input requirements of the model, normalizing it, etc., to meet the input requirements of the target detection model. The preprocessed image is input into the trained target detection model, and forward propagation calculation is performed to calculate the bounding box coordinates, category label and confidence score of the output target body. According to the confidence threshold and confidence score, the qualified detection results are screened out. The bounding box coordinates in the detection result are converted from the normalized form to the pixel coordinates of the original image. The specific formula is as follows:
[0080]
[0081] Among them, x center y is the normalized horizontal coordinate of the center point of the detected target bounding box on the image width; center is the normalized ordinate of the center point of the detected target bounding box at the image height; width is the width of the detected target bounding box, normalized to the ratio of the image width; height is the height of the detected target bounding box, normalized to the ratio of the image height; image width is the actual width of the preprocessed image, in pixels; image height is the actual height of the preprocessed image, in pixels; x 1 is the pixel coordinate of the left edge of the bounding box on the width of the image; x 2 is the pixel coordinate of the right edge of the bounding box on the width of the image; 1y is the pixel coordinate of the upper edge of the bounding box at the image height; 2 is the pixel coordinate of the bottom edge of the bounding box in terms of the image height.
[0082] Step S23: determining a first cropped image of the initial image based on the detection result of the target subject.
[0083] Specifically, the initial image is cropped according to the coordinate information of the target bounding box to obtain a first cropped image containing the target body. Furthermore, necessary post-processing operations are performed on the cropped first cropped image according to specific needs, such as adjusting the image size, removing black edges, etc.
[0084] Through the above scheme, this embodiment performs subject detection on the initial image, so that the target subject can be more accurately identified and located during the detection process, thereby improving the accuracy of target subject detection; by performing subject detection on the initial image and then cropping it, it is possible to avoid indiscriminate processing of the entire image, thereby improving the efficiency of image processing.
[0085] Based on the above embodiment, in a feasible implementation manner, the step of performing semantic segmentation on the first cropped image, cropping the first cropped image based on the result of the semantic segmentation, and obtaining a second cropped image of the first cropped image includes:
[0086] Step S31, performing semantic category recognition and segmentation training on the semantic segmentation model.
[0087] It should be noted that semantic segmentation models are used for image segmentation tasks, where the model classifies each pixel of the input image into a predefined semantic category. This model can identify different objects or regions in the image and assign a corresponding category label to each pixel, thereby achieving fine image segmentation.
[0088] Specifically, collect a large number of annotated image data sets, where the pixels in each image are marked as specific semantic categories, and divide the data sets into training sets and validation sets; select a suitable semantic segmentation model architecture, such as fully convolutional neural network FCN, U-net network U-Net, etc., and determine the specific parameters and configuration of the model according to the model architecture. Use the training set to train the model, and adjust the model parameters through the back propagation algorithm to minimize the difference between the predicted label and the true label; during the training process, regularly use the validation set to evaluate the performance of the model, and adjust the training strategy based on the evaluation results. Use the test set to evaluate the trained model, and calculate indicators such as segmentation accuracy, recall rate, and F1 score; based on the evaluation results, determine whether the model meets the application requirements. When it meets the application requirements, complete the model training.
[0089] Step S32: performing semantic segmentation on the first cropped image by using the semantic segmentation model after recognition segmentation training to obtain a semantic segmentation result of the first cropped image.
[0090] Specifically, the first cropped image is preprocessed, such as normalization and scaling, to ensure that it meets the input requirements of the semantic segmentation model. The preprocessed first cropped image is input into the trained semantic segmentation model. The model classifies the image pixel by pixel to generate a semantic segmentation result. The semantic segmentation result is post-processed, such as smoothing and removing small areas, to improve the accuracy and usability of the segmentation result.
[0091] Step S33: cropping the first cropped image based on the semantic segmentation result of the first cropped image to generate a second cropped image of the first cropped image.
[0092] Specifically, according to the semantic segmentation result, the semantic category area to be retained is determined; based on the semantic category area to be retained, the first cropped image is cropped through an image processing library (such as OpenCV) to generate a second cropped image.
[0093] Through the above scheme, this embodiment can make the cropped image more focused on the target subject through semantic segmentation, thereby further reducing the interference of background information and improving the accuracy of semantic segmentation; through fine cropping, a clearer image that is more focused on the target semantic category can be generated, which is beneficial to subsequent image analysis, processing or application.
[0094] Based on the above embodiment, in a feasible implementation manner, the steps of visually evaluating the second cropped image, cropping the second cropped image based on the result of the visual evaluation, and generating a third cropped image of the second cropped image include S41 to S43:
[0095] Step S41, performing evaluation training of aesthetic composition rules on the visual evaluation model.
[0096] It should be noted that a visual assessment model is a computer program based on machine learning or deep learning technology that is specially trained to automatically analyze and evaluate the aesthetic composition rules of an image. The model is able to receive image input and score the image based on predefined aesthetic criteria (such as the rule of thirds, balance, contrast, color harmony, etc.). These scores reflect the aesthetic quality of the image in terms of composition, color, subject expression, etc.
[0097] Specifically, a large amount of image data containing different composition styles is collected, ensuring that the dataset contains high-quality images taken by professional photographers, as well as images taken by ordinary users to cover a wide range of composition styles and quality. The collected image set is divided into a test set and a validation set. The collected images are annotated and scored according to aesthetic composition rules (such as the rule of thirds, symmetry, balance, foreground-background contrast, etc.). A suitable deep learning model is selected as the basis of the visual evaluation model, which can process image data and output aesthetic scores; the selected visual evaluation model is trained using the annotated training set. During the training process, the model will learn to identify aesthetic composition rules in the image and give an aesthetic score of the image based on these rules. After the model completes the evaluation, it outputs a score map of the same size as the input image, where the value of each pixel represents the aesthetic score at that position. After the training is completed, the performance of the model is evaluated using the validation set to ensure that it can accurately predict the aesthetic score of the image.
[0098] Step S42: Perform visual evaluation on the second cropped image by evaluating the trained visual evaluation model to generate a visual score map of the second cropped image.
[0099] It should be noted that the visual score map is used to show the aesthetic scores of different regions of the image by the visual evaluation model, and is presented in the form of an image, where the color, brightness or other visual attributes of each pixel or image region represent the aesthetic score of the region. The visual score map can help users intuitively understand which parts of the image meet the aesthetic composition rules and which parts need to be improved.
[0100] Specifically, the second cropped image to be evaluated is input into the trained visual evaluation model, and the model extracts features of the second cropped image. Based on the extracted features, the model predicts the aesthetic score of the second cropped image and generates a visual score map of the second cropped image. The size of the visual score map is the same as that of the second cropped image, and each pixel has a corresponding aesthetic score.
[0101] Step S43: cropping the second cropped image based on the visual score map to generate a third cropped image of the second cropped image.
[0102] Specifically, based on the second cropped image, several candidate cropping frames of the second cropped image are generated; based on the visual scoring maps of the several candidate cropping frames, the second cropped image is visually scored through the visual scoring maps to obtain several visual scoring results; based on the visual scoring results, a third cropped image of the second cropped image is determined.
[0103] Through the above scheme, this embodiment can ensure that the cropped image is more in line with the aesthetic standards in composition by evaluating and training the visual evaluation model on the rules of aesthetic composition, thereby further improving the aesthetic quality of the image; the image is automatically evaluated by the visual evaluation model and a visual scoring chart is generated, which can greatly improve the efficiency and accuracy of cropping. Compared with the traditional manual cropping method, this method is more objective, fast and consistent.
[0104] Based on the above implementation scheme, in a feasible implementation manner, the step of cropping the second cropped image based on the visual score map to generate a third cropped image of the second cropped image includes S51 to S53:
[0105] Step S51: generating a plurality of candidate cropping frames of the second cropping image based on the second cropping image.
[0106] It should be noted that the candidate cropping box refers to a series of rectangular bounding boxes preset or dynamically generated on the second cropped image in order to find the best cropping area during the image processing or editing process. Each candidate cropping box defines a specific area in the image, which may have different sizes, positions, and aspect ratios, so as to explore the best expression of image composition, perspective, or aesthetic value while maintaining the integrity of the image content.
[0107] Specifically, according to the size of the initial image, composition rules (such as the rule of thirds, the golden section, etc.), and the real-time size of the image display container, the parameters of the candidate cropping frame are defined, including the size range, aspect ratio, position offset, etc.; according to the defined parameters, a series of candidate cropping frames are generated on the image, and the candidate cropping frames can be evenly distributed or can be generated by giving priority to the main object and high aesthetic score area of the image. In one embodiment of the present application, the candidate cropping frame gives priority to the main object and high aesthetic score area.
[0108] Step S52: Based on the visual scoring maps of the plurality of candidate cropping frames, visual scoring is performed on the second cropped image using the visual scoring maps to obtain a plurality of visual scoring results.
[0109] Specifically, for each candidate cropping frame, the corresponding image area is extracted from the second cropped image; the extracted image area is aligned with the visual score map to ensure that each pixel or image area corresponds to the corresponding position in the score map; based on the visual score map, the average score and the highest score of the image area in each candidate cropping frame are calculated as the visual score result of the cropping frame.
[0110] Step S53: determining a third cropped image of the second cropped image based on the visual scoring result.
[0111] Specifically, according to the visual scoring results, the candidate cropping frames are sorted, and the cropping frames with higher scores are given priority. According to the sorting results, the candidate cropping frame with the highest score is selected as the best cropping solution. If there are multiple cropping frames with similar scores, other factors (such as composition balance, subject prominence, etc.) can be further considered for screening; according to the selected best cropping frame, the corresponding image area is extracted from the second cropped image to generate a third cropped image.
[0112] Based on the above implementation scheme, in a feasible implementation manner, the step of dynamically cropping the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image includes S61-S62:
[0113] Step S61, obtaining the real-time size of the image display container and the aspect ratio of the initial image.
[0114] Specifically, the real-time size of the image display container in the page is monitored in real time. For example, the current width and height of the image display container can be obtained through JavaScript technology; the original width and height of the initial image are obtained, and the aspect ratio is calculated, that is, the width divided by the height (or the height divided by the width).
[0115] Step S62: dynamically cropping and adjusting the third cropped image based on the real-time size of the image display container and the aspect ratio of the initial image to generate a target cropped image of the third cropped image.
[0116] Specifically, according to the real-time size of the image display container and the aspect ratio of the initial image, the width and height of the target cropped image are calculated, the aspect ratio of the target cropped image is kept the same as that of the initial image, and at the same time, the cropped image is ensured to fit the image display container without stretching or compression. Based on the calculated width and height of the target cropped image, the third cropped image is cropped and adjusted to generate the target cropped image.
[0117] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the image cropping method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0118] This application also provides an image cropping device, please refer to Figure 2 , the image cropping device comprises:
[0119] A subject detection module 201 is used to perform subject detection on an initial image by using an object detection model to determine a first cropped image of the initial image;
[0120] A semantic segmentation module 202, configured to perform semantic segmentation on the first cropped image, and crop the first cropped image based on a result of the semantic segmentation to obtain a second cropped image of the first cropped image;
[0121] A visual evaluation module 203 is used to perform a visual evaluation on the second cropped image, crop the second cropped image based on a result of the visual evaluation, and generate a third cropped image of the second cropped image;
[0122] The dynamic cropping module 204 is configured to dynamically crop the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image.
[0123] The image cropping device provided by the present application adopts the image cropping method in the above embodiment, which can solve the technical problem of poor image display effect. Compared with the prior art, the beneficial effects of the image cropping device provided by the present application are the same as the beneficial effects of the image cropping method provided by the above embodiment, and other technical features in the image cropping device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0124] The present application provides an image cropping device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image cropping method in the above-mentioned embodiment 1.
[0125] Reference below Figure 3 , which shows a schematic diagram of the structure of an image cropping device suitable for implementing an embodiment of the present application. The image cropping device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The image cropping device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0126] like Figure 3As shown, the image cropping device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the image cropping device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image cropping device to communicate with other devices wirelessly or by wire to exchange data. Although the image cropping device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0127] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0128] The image cropping device provided by the present application adopts the image cropping method in the above embodiment to solve the technical problem of poor image display effect. Compared with the prior art, the beneficial effects of the image cropping device provided by the present application are the same as the beneficial effects of the image cropping method provided by the above embodiment, and the other technical features in the image cropping device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0129] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0130] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0131] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, wherein the computer-readable program instructions are used to execute the image cropping method in the above-mentioned embodiment.
[0132] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0133] The computer-readable storage medium may be included in the image cropping device; or may exist independently without being assembled into the image cropping device.
[0134] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the image cropping device, the image cropping device: performs subject detection on the initial image through a target detection model to determine a first cropped image of the initial image; performs semantic segmentation on the first cropped image, crops the first cropped image based on the result of the semantic segmentation, and obtains a second cropped image of the first cropped image; performs visual evaluation on the second cropped image, crops the second cropped image based on the result of the visual evaluation, and generates a third cropped image of the second cropped image; dynamically crops the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image.
[0135] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0136] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0137] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0138] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned image cropping method, and can solve the technical problem of poor image display effect. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the image cropping method provided in the above-mentioned embodiment, and will not be repeated here.
[0139] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned image cropping method when executed by a processor.
[0140] The computer program product provided by this application can solve the technical problem of poor image display effect. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as the beneficial effects of the image cropping method provided by the above embodiment, which will not be repeated here.
[0141] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. An image cropping method, characterized in that: The method comprises: Performing subject detection on the initial image by using an object detection model to determine a first cropped image of the initial image; performing semantic segmentation on the first cropped image, and cropping the first cropped image based on a result of the semantic segmentation to obtain a second cropped image of the first cropped image; performing a visual assessment on the second cropped image, and cropping the second cropped image based on a result of the visual assessment to generate a third cropped image of the second cropped image; Based on the size of the image display container and the size of the initial image, the third cropped image is dynamically cropped to generate a target cropped image of the third cropped image.
2. The method according to claim 1, characterized in that The step of performing subject detection on the initial image by using the target detection model to determine the first cropped image of the initial image comprises: Conduct target subject detection and recognition training on the target detection model; Performing subject detection on the initial image by detecting and recognizing the trained target detection model to obtain a detection result of the target subject in the initial image; Based on the detection result of the target subject, a first cropped image of the initial image is determined.
3. The method according to claim 1, characterized in that The step of performing semantic segmentation on the first cropped image, cropping the first cropped image based on the result of the semantic segmentation, and obtaining a second cropped image of the first cropped image comprises: Perform semantic category recognition and segmentation training on the semantic segmentation model; Performing semantic segmentation on the first cropped image by using a semantic segmentation model after recognition segmentation training to obtain a semantic segmentation result of the first cropped image; Based on the semantic segmentation result of the first cropped image, the first cropped image is cropped to generate a second cropped image of the first cropped image.
4. The method according to claim 1, characterized in that The steps of visually evaluating the second cropped image, cropping the second cropped image based on a result of the visual evaluation, and generating a third cropped image of the second cropped image include: Training visual evaluation models to evaluate aesthetic composition rules; Performing a visual evaluation on the second cropped image by evaluating the trained visual evaluation model to generate a visual score map of the second cropped image; The second cropped image is cropped based on the visual score map to generate a third cropped image of the second cropped image.
5. The method according to claim 4, characterized in that The step of cropping the second cropped image based on the visual score map to generate a third cropped image of the second cropped image comprises: Based on the second cropped image, generating a plurality of candidate cropping frames of the second cropped image; Based on the visual scoring graphs of the plurality of candidate cropping frames, visually scoring the second cropped image using the visual scoring graphs to obtain a plurality of visual scoring results; Based on the visual scoring result, a third cropped image of the second cropped image is determined.
6. The method according to claim 1, characterized in that The step of dynamically cropping the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image comprises: Obtaining the real-time size of the image display container and the aspect ratio of the initial image; Based on the real-time size of the image display container and the aspect ratio of the initial image, the third cropped image is dynamically cropped and adjusted to generate a target cropped image of the third cropped image.
7. An image cropping device, characterized in that: The device comprises: A subject detection module, used to perform subject detection on the initial image through an object detection model to determine a first cropped image of the initial image; a semantic segmentation module, configured to perform semantic segmentation on the first cropped image, and crop the first cropped image based on a result of the semantic segmentation to obtain a second cropped image of the first cropped image; a visual assessment module, configured to perform a visual assessment on the second cropped image, crop the second cropped image based on a result of the visual assessment, and generate a third cropped image of the second cropped image; The dynamic cropping module is used to dynamically crop the third cropped image based on the size of the image display container and the size of the initial image to generate a target cropped image of the third cropped image.
8. An image cropping device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image cropping method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image cropping method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the image cropping method according to any one of claims 1 to 6 are implemented.