Image cutting method and device based on semantic perception
Through semantic perception technology, combining the visual characteristics of the image and deep semantic information, the direction of the cropping box is dynamically adjusted, which solves the problem of lack of personalization and flexibility of image cropping results in the prior art, and achieves a more accurate and flexible image cropping effect.
Patent Information
- Application Number
- CN202411872824.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing image cropping methods ignore deep semantic information in images, cannot fully tap the potential content of the image, and the cropping method lacks personalization and flexibility.
The instance pixel mask and semantic labels in the image are obtained through the segmentation detection model, a weighted undirected graph and an adjacency matrix are constructed, and the spectral clustering is used for graph convolutional networks, candidate center of gravity group is obtained, and the crop angle is determined based on the rotation angle and weight. Finally, the minimum external box is obtained through composition for cropping.
Effectively combine visual features and semantic information, dynamically adjust the direction of the cropping box, improve the accuracy and flexibility of the cropping results, reduce manual intervention, and improve the robustness of the model.
Smart Images

Figure CN119991695A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image cropping, and in particular relates to a method and device for image cropping based on semantic perception. Background Art
[0002] The image cropping task aims to accurately capture the boundaries of the original image to retain the most informative and visually attractive parts of the image, thereby optimizing content display and improving visual effects. With the continuous development of digital image processing technology and the increasing diversification and personalization of image processing needs, image cropping has gradually become a key task in many fields such as computer vision and automated design, and has begun to be widely used in scenarios such as social media, advertising design, and image retrieval.
[0003] At present, most mainstream image cropping methods are based on image detection information, salient regions and other visual features, combined with composition principles and photography theory for cropping. However, many traditional methods ignore the deep semantic information contained in the image during application and fail to fully explore the potential content of the image. At the same time, most existing cropping techniques use fixed cropping frame angles and ratios, failing to take into account the diversity of instances in the image and the complexity of the layout, resulting in a lack of personalization and flexibility in the cropping results. How to organically combine the visual features and semantic information of the image and flexibly adjust the cropping method according to the specific content is still a key issue that needs to be solved in the field of image cropping. Summary of the invention
[0004] The purpose of the present invention is to provide an image cropping method and device based on semantic perception.
[0005] In a first aspect, the present invention provides a semantically-aware image cropping method, which comprises the following steps: Step 1: Get the original image ; Use the segmentation detection model to obtain the original image The pixel mask corresponding to each instance in and semantic tags ; Based on the pixel mask and semantic tags Obtaining joint feature representation ;in, ; For the original image The number of instances in ; Step 2: Original image As nodes, we construct a weighted undirected graph and the adjacency matrix ; Step 3: Extract pixel mask The boundary point set ; Through the rotation matrix Boundary point collection Rotate to different angles and get the rotation angle . Rotation Angle Make the boundary point set The area of the positive rectangular rotation boundary is the smallest.
[0006] Step 4: Use graph convolutional networks to jointly represent features and weighted undirected graph Construct a Laplace matrix; use spectral clustering to perform data segmentation on the eigenvalues and eigenvectors of the Laplace matrix, and use each cluster obtained by segmentation as a set of candidate centroid groups; Step 5: According to the weight of the nodes in each candidate centroid group and rotation angle Determine the corresponding cropping angle ;in, ; For the The number of nodes in the candidate centroid group; ; is the number of candidate centroid groups; according to the clipping angle Get the coverage The minimum bounding rectangle of all nodes in the candidate centroid group is used as the minimum bounding box, and the final candidate box is obtained by composition; the final candidate box is used to Perform cropping and obtain the cropping result.
[0007] Preferably, in step 1, the feature joint representation The method to obtain is as follows: Use the concept dictionary to obtain each semantic label Detailed definition of ; with semantic tags And the corresponding definition A collection of text inputs ; Use image encoder pixel mask Local visual features ; Use text encoder to extract text input Text features ; Based on local visual features and text features The composed set is used as a joint representation of features .
[0008] As a preference, in the step 2, the undirected graph The relationship weights between different nodes in The method to obtain is as follows: in, for structural similarity; for semantic similarity; To combine global visual features and the spatial relationship between the local features of the two nodes; , and is a hyperparameter.
[0009] As a preference, in step 5, the weight The method to obtain is as follows: in, is the feature extraction function; is the minimum enclosing rectangle of the boundary point set; is the minimum positive circumscribed rectangle of the boundary point set; is the area ratio; , and is a hyperparameter.
[0010] As a preferred embodiment, the minimum circumscribed rectangle The method to obtain is: Extract the boundary point set of the pixel mask through the convex hull algorithm, using the rotation matrix Rotate the boundary point set to obtain the minimum horizontal and vertical coordinates and the maximum horizontal and vertical coordinates in the rotated boundary point set; use the minimum horizontal and vertical coordinates and the maximum horizontal and vertical coordinates as the two endpoints of the diagonal of the minimum circumscribed rectangle, and according to the rotation angle Get the The minimum bounding rectangle of the nodes .
[0011] As a preferred embodiment, in the step 5, the cutting angle The method to obtain is as follows: Get the rotation angle of the nodes in the candidate centroid group Standard Deviation ; If the standard deviation is less than the set threshold, the clipping angle Equal to the mean of the rotation angles of the nodes in the candidate centroid group ; If the standard deviation is greater than or equal to the set threshold, the clipping angle The expression is: in, is the rotation angle of the zth node in the candidate centroid group.
[0012] Preferably, in step 5, the method for obtaining the composition mode is as follows: Set up a composition template according to the composition method; calculate the similarity between the instance distribution of the candidate centroid group and different composition templates , whose expression is: in, ; For the The visual area that the composition template focuses on; ; is the number of composition templates; is the node centroid position; All similarities corresponding to each candidate centroid group In the above example, select the similarity greater than the preset similarity threshold. The corresponding composition method is used as the composition method of the candidate centroid group .
[0013] As a preference, the node centroid position The expression is: in, is the pixel mask of the zth node; Indicates that it belongs to the pixel mask Pixels.
[0014] Preferably, in step 5, the method for obtaining the final candidate box is as follows: Match the crop ratio closest to the smallest bounding box , crop ratio The expression is: in, To include The ratio of the minimum bounding box of all nodes in the candidate centroid group; Enlarge the length or width of the minimum bounding rectangle to obtain the initial candidate frame, so that the aspect ratio of the initial candidate frame is consistent with the cropping ratio The same; the size and position of the initial candidate frame are adjusted according to the golden ratio and the composition method to obtain the final candidate frame.
[0015] In the second aspect, the present invention provides an image cropping device based on semantic perception, which is used to implement the above-mentioned image cropping method based on semantic perception, including a feature extraction module, a center of gravity group screening module and an image cropping module; the feature extraction module is used to extract image features and text features of the original image; the center of gravity group screening module is used to obtain candidate center of gravity groups based on image features and text features; the image cropping module confirms the final candidate frame based on the characteristics of the candidate center of gravity group and crops the original image.
[0016] The present invention has the following beneficial effects: 1. The present invention effectively makes up for the limitation of traditional methods that rely only on visual features. It enriches semantic information by introducing a concept dictionary and further explores potential semantic relationships through semantic similarity, which helps to fully explore and utilize semantic information, thereby establishing a richer and more unified conceptual space between images and semantic labels, further improving the accuracy of cropping results and the robustness of the model.
[0017] 2. The present invention breaks through the limitation of fixed angles of traditional methods. By comprehensively considering the angle information of each instance, the direction of the cropping frame is dynamically adjusted, so that the cropping result can adapt to the layout and structure of the image content more flexibly and accurately.
[0018] 3. The present invention deeply integrates photography theory and computer vision. By comprehensively considering the geometric properties, directional information, density distribution and other characteristics of the instance, it automatically selects the appropriate cropping method and composition ratio, greatly reducing the need for manual intervention and has strong versatility and extensibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of Example 1 of the present invention.
[0020] Figure 2 This is a schematic diagram of the framework of Example 1 of the present invention.
[0021] Figure 3 This is a schematic diagram of node weighting in Example 1 of the present invention.
[0022] Figure 4 This is a comparison diagram of the minimum circumscribed rectangle and the minimum positive circumscribed rectangle in Example 1 of the present invention.
[0023] Figure 5 Schematic diagram of the focus area of different composition methods in Example 1 of the present invention.
[0024] Figure 6 The figure is a schematic diagram comparing the cropping results of the present invention and the existing image cropping method.
[0025] Figure 7 This is a schematic diagram of the structure of Example 2 of the present invention. DETAILED DESCRIPTION
[0026] The present invention will be further described below in conjunction with the accompanying drawings.
[0027] Example 1 like Figure 1 and Figure 2 As shown, a semantic-aware image cropping method includes the following steps: Step 1: Get the original image ;in, and are the height and width of the original image respectively; is the number of channels of the original image. The segmentation detection model is used to obtain the pixel mask corresponding to each instance in the original image. and semantic tags , ;in, ; is the number of instances in the original image; It is a predefined set of semantic categories built based on concept databases such as Things. To ensure universality, concepts not defined in the concept dictionary WordNet will be filtered out from the semantic category set.
[0028] Based on semantic tags , using the concept dictionary, obtain each semantic label Detailed definition of , whose expression is: in, A concept dictionary.
[0029] The concept dictionary includes classic language knowledge bases such as WordNet and BabelNet, which can obtain the definition of each tag based on semantic tags, and provide prior knowledge for tags through richer text representations, which helps to represent complex semantic information in the form of feature embedding and generate a more unified and rich concept space. This enhancement process not only improves the model's ability to understand the semantics of tags, but also provides a basis for mining semantic relationships between different categories.
[0030] Semantic tags And the corresponding definition A collection of text inputs .
[0031] Step 2: Extract the global visual features of the original image through the image encoder and pixel mask Local visual features , global visual features and local visual features It is expressed as: in, is an image encoder; and is the spatial resolution of global visual features; is the dimension of the global visual feature.
[0032] Extract text input via text encoder Text features , text features It is expressed as: in, A text encoder.
[0033] Based on local visual features and text features Construct a joint representation of the features of each instance , which facilitates the subsequent graph structure construction and analysis.
[0034] Step 3: Construct a weighted undirected graph and the adjacency matrix , which is expressed as: in, is the node set, ; For the Instances; is the set of edges between nodes; For the Instances and The relationship weight between instances; .
[0035] The method for obtaining the relationship weights between different instances is as follows: a. Obtain an undirected graph based on the brightness, contrast, and structure information of the pixel masks corresponding to different instances The structural similarity between two different nodes in , whose expression is: sim a,b struct =[ l( M a , M b ) ] α 1 [ c( M a , M b ) ] α 2 [ s( M a , M b ) ] α 3 in, is the brightness contrast function; is the contrast comparison function; is the structural comparison function; , and is a hyperparameter, , used to control the weights of contributions from different modules; are the pixel masks corresponding to different instances.
[0036] b. Text input corresponding to different instances Get an undirected graph The semantic similarity between two different nodes in , whose expression is: in, Enter the text corresponding to different instances.
[0037] c. Based on the semantic similarity between two nodes , structural similarity Calculate the relationship weight between nodes with spatial relationships , whose expression is: in, , and is a hyperparameter; To combine global visual features And the spatial relationship between the local features of the two nodes.
[0038] Hyperparameters and Used to control the weight of each similarity, hyperparameter Used to control the weight of position information, global visual features Local features can be and To guide and adjust the distance calculation between local features, it can better capture the spatial correlation between nodes.
[0039] Step 4: Get the weight corresponding to the node 4-1. If Figure 3 and Figure 4 As shown, get the centroid position of each node respectively , minimum positive bounding rectangle and minimum bounding rectangle; the centroid position of each node The expression is: in, Indicates that it belongs to a pixel mask Pixels.
[0040] The two endpoints of the diagonal of the minimum positive circumscribed rectangle are expressed as and ;in, ; Get the minimum positive circumscribed rectangle based on the coordinates of the two endpoints of the diagonal .
[0041] The method to obtain the minimum enclosing rectangle is as follows: a. Extract pixel mask through convex hull algorithm The boundary point set , which is expressed as: in, represents the convex hull algorithm; is the coordinate of the kth boundary point; For the The number of boundary points of an instance; .
[0042] b. Define the rotation matrix ;in, For the The rotation angle of each instance; through the rotation matrix Boundary point collection Rotate to different angles and obtain the set of rotated boundary points for: in, is the coordinate of the kth rotated boundary point.
[0043] c. Based on the boundary point set Get the rotated positive rectangle rotation boundary, which is expressed as: in, and The boundary point sets The minimum horizontal and vertical coordinates in and The boundary point sets The maximum abscissa and ordinate in ; and The boundary point set The horizontal and vertical coordinates of the middle boundary point.
[0044] d. Based on the positive rectangle rotation boundary, calculate the area of the rotated rectangle , whose expression is: e. Select the rotation angle that minimizes the rotation boundary area of the positive rectangle as the rotation angle of the corresponding instance , which is expressed as: in, The input value corresponding to the value of the minimized variable.
[0045] f. The two endpoints of the diagonal of the minimum circumscribed rectangle are expressed as and ; Based on the two endpoints of the diagonal of the minimum circumscribed rectangle and the minimum rotation angle Get the The minimum bounding rectangle of instances .
[0046] 4-2. To take size into account, calculate the area ratio of each instance , whose expression is: in, is the pixel mask Total number of midpoints; .
[0047] 4-3. Comprehensively consider visual attributes such as angle and size to obtain the weight corresponding to each node , whose expression is: in, is the feature extraction function; , and is a hyperparameter used to control the weight of the contribution of different visual attributes.
[0048] The feature extraction function is used to normalize and fuse the areas of the minimum enclosing rectangle or the minimum positive enclosing rectangle, and then map them into quantitative feature values related to weight calculation.
[0049] Step 5: Use graph convolutional network (GCN) to jointly represent the features of nodes Modeling; for each node, it is necessary to consider all its neighboring nodes and the feature information contained in itself. The graph convolutional network processes edges and nodes as follows: in, It is The node feature matrix of the layer; is the node feature after graph convolution; is a nonlinear activation function; is the degree matrix; is the adjacency matrix after processing; For the The weight matrix of the layer; is a hyperparameter, when When the value of is 1, it means that the importance of the node's own features is the same as that of its neighbors; for dimensional identity matrix; For the node The sum of the connected edge weights; For Node and nodes The edge weight between nodes and nodes If there is no connection, the weight is 0.
[0050] Based on the adjacency matrix and degree matrix calculated by the graph convolutional network, construct the Laplacian matrix; calculate the minimum Laplacian matrix The eigenvalues corresponding to the eigenvectors are arranged into a matrix by column. Q min = [ q 1 , q 2 ,..., q K ] ;in, is the number of candidate centroid groups; For the Eigenvectors corresponding to small eigenvalues; . Use K-means algorithm to calculate the matrix Cluster all rows of , obtain cluster division, and construct a set of candidate centroid groups with the divided clusters ;in, For the The instance information contained in each candidate centroid group.
[0051] Step 6: Crop the original image 6-1. Get the clipping angle The standard deviation of the rotation angles of the nodes in the candidate centroid group Determine the consistency of node direction, standard deviation The expression is: in, For the The number of nodes in the candidate centroid group; is the rotation angle of the zth node in the candidate centroid group; For the The mean of the rotation angles of the candidate centroid group; .
[0052] If the standard deviation is less than the set threshold, the node directions are considered to be highly consistent, and the clipping angle of the candidate centroid group is is the mean of the rotation angles of the candidate centroid group ; The threshold of the standard deviation is set as [10° ,20 °] .
[0053] If the standard deviation is greater than or equal to the set threshold, it means that the directions of the nodes are very different, and the overall angle cannot be used directly as the direction of the cropping frame. In this case, the node with the largest weight is selected as the dominant node, and the influence of other nodes on the angle of the cropping frame is determined according to its weight ratio. The expression is:
[0054] in, is the weight of the zth node in the candidate centroid group.
[0055] 6-2. Get the crop ratio According to the cutting angle Get the minimum bounding rectangle covering all nodes in the tth group of candidate centroids as the minimum bounding box, and automatically match the closest cropping ratio for: in, To include The ratio of the minimum bounding box of all nodes in the group candidate centroid.
[0056] 6-3. Obtaining the composition method like Figure 5 As shown in the figure, a series of composition templates are set up according to the composition method. Each composition template focuses on different visual areas. For example, in the horizontal line, vertical line and diagonal line composition methods, the visual center of gravity is distributed near different auxiliary lines; in the center composition method, the visual center of gravity is concentrated in the center to present a sense of balanced symmetry; in the thirds composition method and the golden ratio composition method, the visual center of gravity is distributed in the intersection area of the auxiliary lines.
[0057] Calculate the similarity between each candidate centroid group and different composition templates based on its instance distribution , whose expression is: in, ; For the The visual area that the composition template focuses on; ; is the number of composition templates; The first The centroid position of the nodes.
[0058] All similarities corresponding to each candidate centroid group In the above example, select the similarity greater than the preset similarity threshold. The corresponding composition method is used as the composition method of the candidate centroid group .
[0059] 6-4. Enlarge the length or width of the minimum bounding box to obtain the initial candidate box, so that the aspect ratio of the initial candidate box is consistent with the cropping ratio The size of the initial candidate box is adjusted according to the golden ratio and composition method (content compactness ) and position, obtain the final candidate box, and adjust it as follows:
[0060] in, and are the width and height of the final candidate box respectively; and are the width and height of the initial candidate box respectively; ; For composition The number of
[0061] Use the final candidate box to Perform cropping and obtain the cropping result.
[0062] Step 7: Method evaluation like Figure 6As shown, the cropping results of the present invention are compared with the cropping results of the existing image cropping method. The existing image cropping method ignores the directional information of the instance in the image, adopts a fixed cropping frame angle, ignores the directional difference between the cropping results, and causes the cropping result to lack flexibility. The present invention calculates the minimum circumscribed rectangle and the minimum positive circumscribed rectangle of each instance, and introduces the angle factor for calculation. When the rotation angles in the candidate centroid group are highly consistent, the overall angle is directly used as the direction of the cropping frame. In the case of a large angle difference, the cropping frame angle is adjusted according to the node with the largest weight, thereby ensuring that the cropping frame is more in line with the layout and structure of the actual content. At the same time, the present invention optimizes the image cropping and composition process by combining the aesthetic principles in photography theory. When processing the candidate centroid group, the system selects an appropriate composition method, such as a centralized composition or a symmetrical composition, according to the centroid distribution density and the geometric layout to ensure visual focus and overall harmony. The present invention also takes into account the compactness of the content, and the system can intelligently adjust the cropping method to ensure the balance and tension of the image. In addition, the system also dynamically adjusts the aspect ratio according to the size of the image and the proportional rules in photography theory to ensure the best visual effect. By combining aesthetics with computer vision technology, the present invention achieves an organic fusion of artistry and technology, which helps to enhance the aesthetic value of the cutting results.
[0063] Example 2 like Figure 7 As shown, a semantically-aware image cropping device is used to implement the cropping method in Example 1, which includes a feature extraction module, a center of gravity group screening module and an image cropping module; the feature extraction module is used to extract image features and text features of the original image; the center of gravity group screening module is used to obtain candidate center of gravity groups based on image features and text features; the image cropping module confirms the final candidate frame based on the characteristics of the candidate center of gravity group, such as the standard deviation of the rotation angle of the candidate center of gravity group, the weight of the elements within the group, the bounding box size, and the content compactness, and crops the original image.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; obviously, the drawings are only some examples or embodiments of the present application, and ordinary technicians in this field can also apply the present application to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in this field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.
[0065] Although the present invention has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and all of these belong to the protection scope of this application. Therefore, the protection scope of this application shall be based on the attached claims.
Claims
1. A semantic-aware image cropping method, characterized in that: The following steps are involved: Step 1: Get the original image ; Use the segmentation detection model to obtain the original image The pixel mask corresponding to each instance in and semantic tags ; Based on the pixel mask and semantic tags Obtaining joint feature representation ;in, ; For the original image The number of instances in ; Step 2: Original image As nodes, we construct a weighted undirected graph and the adjacency matrix ; Step 3: Extract pixel mask The boundary point set ; Through the rotation matrix Boundary point collection Rotate to different angles and get the rotation angle ; Rotation angle Make the area of the positive rectangular rotation boundary as small as possible; Step 4: Use graph convolutional networks to jointly represent features and weighted undirected graph Construct a Laplace matrix; use spectral clustering to perform data segmentation on the eigenvalues and eigenvectors of the Laplace matrix, and use each cluster obtained by segmentation as a set of candidate centroid groups; Step 5: According to the weight of the nodes in each candidate centroid group and rotation angle Determine the corresponding cropping angle ;in, ; For the The number of nodes in the candidate centroid group; ; is the number of candidate centroid groups; according to the clipping angle Get the coverage The minimum bounding rectangle of all nodes in the candidate centroid group is used as the minimum bounding box, and the final candidate box is obtained by composition; the final candidate box is used to Perform cropping and obtain the cropping result.
2. The semantic-aware image cropping method according to claim 1, characterized in that: In the step 1, the feature joint representation The method to obtain is as follows: Use the concept dictionary to obtain each semantic label Detailed definition of ; with semantic tags And the corresponding definition A collection of text inputs ; Use image encoder pixel mask Local visual features ; Extract text input using text encoder Text features ; Based on local visual features and text features The composed set is used as a joint representation of features .
3. The semantic-aware image cropping method according to claim 1, characterized in that: In the step 2, the undirected graph The relationship weights between different nodes in The method to obtain is as follows: in, for structural similarity; for semantic similarity; To combine global visual features and the spatial relationship between the local features of the two nodes; , and is a hyperparameter.
4. The semantic-aware image cropping method according to claim 1, characterized in that: In step 5, the weight The method to obtain is as follows: in, is the feature extraction function; is the minimum enclosing rectangle of the boundary point set; is the minimum positive circumscribed rectangle of the boundary point set; is the area ratio; , and is a hyperparameter.
5. The semantic-aware image cropping method according to claim 4, characterized in that: The minimum bounding rectangle The method to obtain is: Extract the boundary point set of the pixel mask through the convex hull algorithm, using the rotation matrix Rotate the boundary point set to obtain the minimum horizontal and vertical coordinates and the maximum horizontal and vertical coordinates in the rotated boundary point set; use the minimum horizontal and vertical coordinates and the maximum horizontal and vertical coordinates as the two endpoints of the diagonal of the minimum circumscribed rectangle, and according to the rotation angle Get the The minimum bounding rectangle of the nodes .
6. The semantic-aware image cropping method according to claim 1, characterized in that: In step 5, the cutting angle The method to obtain is as follows: Get the rotation angle of the nodes in the candidate centroid group Standard Deviation ; If the standard deviation is less than the set threshold, the clipping angle Equal to the mean of the rotation angles of the nodes in the candidate centroid group ; If the standard deviation is greater than or equal to the set threshold, the clipping angle The expression is: in, is the rotation angle of the zth node in the candidate centroid group.
7. The semantic-aware image cropping method according to claim 1, characterized in that: In the step 5, the method for obtaining the composition mode is as follows: Set up a composition template according to the composition method; calculate the similarity between the instance distribution of the candidate centroid group and different composition templates , whose expression is: in, ; For the The visual area that the composition template focuses on; ; is the number of composition templates; is the node centroid position; All similarities corresponding to each candidate centroid group In the above example, select the similarity greater than the preset similarity threshold. The corresponding composition method is used as the composition method of the candidate centroid group .
8. The semantic-aware image cropping method according to claim 7, characterized in that: The node centroid position The expression is: in, is the pixel mask of the zth node; Indicates that it belongs to the pixel mask Pixels.
9. The semantic-aware image cropping method according to claim 1, characterized in that: In step 5, the method for obtaining the final candidate box is as follows: Match the crop ratio closest to the smallest bounding box , crop ratio The expression is: in, To include The ratio of the minimum bounding box of all nodes in the candidate centroid group; Enlarge the length or width of the minimum bounding rectangle to obtain the initial candidate frame, so that the aspect ratio of the initial candidate frame is consistent with the cropping ratio The same; the size and position of the initial candidate frame are adjusted according to the golden ratio and the composition method to obtain the final candidate frame.
10. An image cropping device based on semantic perception, characterized in that: Used to implement the semantic-aware image cropping method described in claim 1, which includes a feature extraction module, a center of gravity group screening module and an image cropping module; the feature extraction module is used to extract image features and text features of the original image; the center of gravity group screening module is used to obtain candidate center of gravity groups based on image features and text features; the image cropping module confirms the final candidate frame based on the characteristics of the candidate center of gravity group and crops the original image.
Citation Information
Patent Citations
Person image composition cropping method and apparatus, device and storage medium
CN108009998A
Image cutting method and device based on semantic content
CN111612004A
Intelligent image clipping method and system based on visual element relationship
CN113763391A
Image cutting method and device
CN116309627A
Method of automatic cropping
US20100073402A1