Image cutting method and device based on text guidance, equipment and storage medium

Through the image cropping method based on text-guided deep learning and reinforcement learning, the problem of insufficient understanding of user intent is solved, cropping results that are consistent with aesthetics and semantics are generated, and efficient image cropping effects are achieved.

CN120707845APending Publication Date: 2025-09-26INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376196.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing image cropping techniques fail to fully understand user intent, resulting in inaccurate and unattractive cropping results, especially when dealing with complex images and small objects on the edges.

Method used

Through a text-guided method, deep learning is used to extract image and text features, generate an initial cropping frame, and combine it with a reinforcement learning algorithm for dynamic optimization. By integrating aesthetic evaluation, semantic similarity and boundary constraints, multiple candidate cropping images are generated, and finally the best cropping result is selected through a multi-dimensional scoring mechanism.

Benefits of technology

It achieves automatic cropping of images that meet human aesthetic evaluation based on user intention, ensuring that the cropping results achieve the best balance between aesthetic quality and semantic consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707845A_ABST
    Figure CN120707845A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and provides an image cutting method and device based on text guidance, equipment and a storage medium, and the method comprises the steps: carrying out the feature extraction of a to-be-cut image and a text description, and generating an initial cutting frame which is in semantic matching with the text description based on the extracted image features and text features; dynamically optimizing the initial cutting frame by adopting a reinforcement learning algorithm to obtain a plurality of candidate cutting images; and scoring each candidate cutting image by adopting a multi-dimensional scoring mechanism, and determining a target cutting image based on the score of each candidate cutting image. According to the method, the initial cutting frame is determined based on the image features and the text features, then the cutting frame is optimized by designing the composite reward function of the aesthetic evaluation result, the semantic similarity score and the boundary constraint through the reinforcement learning algorithm, and finally the optimal image aesthetic cutting result conforming to the user intention is automatically and efficiently generated. And the optimal balance between the aesthetic quality and the semantic consistency is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a text-guided image cropping method, apparatus, device and storage medium. Background Art

[0002] With the rapid development of internet technology, social media has become an integral part of people's lives, filled with a vast amount of images and text. Image cropping, a common image processing method on these social media platforms, plays a crucial role in improving image visual quality and meeting users' personalized needs. Cropping is particularly important during image upload and dissemination, particularly for generating appropriate thumbnails or meeting specific display requirements.

[0003] However, most existing image cropping techniques ignore user intent. With the rise of vision-language models in recent years, some researchers have begun to leverage these models to achieve image cropping based on user intent. However, existing image cropping techniques based on vision-language models still suffer from two major flaws: first, they rely on single-step cropping, which lacks asymptotic and iterative properties, resulting in inaccurate cropping results; second, they overly focus on semantic similarity and ignore aesthetic quality, resulting in cropping results that conform to the description but lack aesthetic appeal. Therefore, designing a method that can automatically crop images based on user intent to meet human aesthetic standards is a pressing technical challenge in this field. Summary of the Invention

[0004] The present invention provides a text-guided image cropping method, apparatus, device and storage medium, which are used to solve the defects of existing image cropping technology, such as insufficient understanding of user intention, difficulty in retaining small objects, and lack of aesthetics in cropped images.

[0005] The present invention provides a text-guided image cropping method, comprising: Extracting features of the image to be cropped and the text description, and generating an initial cropping frame that semantically matches the text description based on the extracted image features and text features; Dynamically optimizing the initial cropping frame using a reinforcement learning algorithm to obtain a plurality of candidate cropped images, wherein a reward function of the reinforcement learning algorithm is determined based on an aesthetic evaluation result, a semantic similarity score, and a boundary constraint condition, wherein the aesthetic evaluation result is used to characterize the difference between a current candidate cropped image and a previous candidate cropped image, the semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description, and the boundary constraint condition is used to constrain the cropping frame to not exceed the boundary of the image to be cropped; A multi-dimensional scoring mechanism is used to score each candidate cropped image, and a target cropped image is determined based on the scores of each candidate cropped image. The multi-dimensional scoring mechanism includes an aesthetic quality score and a semantic similarity score. The aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

[0006] According to a text-guided image cropping method provided by the present invention, feature extraction is performed on the image to be cropped and the text description, and based on the extracted image features and text features, an initial cropping frame semantically matching the text description is generated, including: Based on an image encoder, feature extraction is performed on the image to be cropped to obtain the image features, and based on a text encoder, feature extraction is performed on the text description to obtain the text features; Using a cross-attention mechanism, the image features and the text features are enhanced and fused to obtain fused features; Based on the decoder, the fusion features are applied to generate a candidate detection frame, and the minimum bounding rectangle of the candidate detection frame is used as the initial cropping frame.

[0007] According to a text-guided image cropping method provided by the present invention, the initial cropping frame is dynamically optimized using a reinforcement learning algorithm to obtain multiple candidate cropped images, including: Based on an observation state and a cropping strategy, selecting an action from a predefined action space and performing the selected action on the initial cropping frame to obtain a new cropping frame, wherein the observation state is determined based on the image to be cropped and the initial cropping frame; Cropping the image to be cropped based on the new cropping frame to obtain a current candidate cropped image, and applying the current candidate cropped image to update the observation state and the cropping strategy; Based on the updated observation state and cropping strategy, continue to select an action from the predefined action space to perform the selected action on the new cropping frame, and repeat the above steps until multiple candidate cropping images are obtained.

[0008] According to a text-guided image cropping method provided by the present invention, applying the current candidate cropping image to update the observation state and the cropping strategy includes: Determining current observation information based on the new cropping frame, the current candidate cropping image, and the image to be cropped, and updating the observation state according to historical observation information and the current observation information to obtain an updated observation state; Based on the current candidate cropped image and the reward function, a reward value is calculated, and the cropping strategy is updated according to the reward value to obtain an updated cropping strategy.

[0009] According to a text-guided image cropping method provided by the present invention, the predefined action space includes a plurality of predefined geometric transformation operations.

[0010] According to a text-guided image cropping method provided by the present invention, the multi-dimensional scoring mechanism is used to score each candidate cropped image, including: Performing semantic similarity scoring on any candidate cropped image and the text description to obtain a semantic similarity score for the candidate cropped image; Performing an aesthetic quality score on any candidate cropped image to obtain an aesthetic score for the candidate cropped image; The semantic similarity score and the aesthetic score of any candidate cropped image are weightedly fused to obtain the score of any candidate cropped image.

[0011] According to a text-guided image cropping method provided by the present invention, the semantic similarity scoring of any candidate cropped image and the text description is performed to obtain the semantic similarity score of any candidate cropped image, including: Extracting an image feature vector of any candidate cropped image and a text feature vector of the text description based on a text-image multimodal model; A similarity calculation is performed based on the image feature vector and the text feature vector, and a semantic similarity score of any candidate cropped image is determined according to the calculation result.

[0012] The present invention also provides an image cropping device based on text guidance, comprising: A cropping frame generating unit is used to extract features of the image to be cropped and the text description, and generate an initial cropping frame that semantically matches the text description based on the extracted image features and text features; a cropping frame optimization unit, configured to dynamically optimize the initial cropping frame using a reinforcement learning algorithm to obtain a plurality of candidate cropped images, wherein a reward function of the reinforcement learning algorithm is determined based on an aesthetic evaluation result, a semantic similarity score, and a boundary constraint condition, wherein the aesthetic evaluation result is used to characterize the difference between a current candidate cropped image and a previous candidate cropped image, the semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description, and the boundary constraint condition is used to constrain the cropping frame to not exceed the boundary of the image to be cropped; A cropped image determination unit is configured to score each candidate cropped image using a multi-dimensional scoring mechanism and determine a target cropped image based on the scores of each candidate cropped image, wherein the multi-dimensional scoring mechanism includes an aesthetic quality score and a semantic similarity score, and the aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, it implements any of the above-described text-guided image cropping methods.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described text-guided image cropping methods.

[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned text-guided image cropping methods.

[0016] The text-guided image cropping method, apparatus, device, and storage medium provided by the present invention can generate an initial cropping frame that matches the semantics of the text description by extracting the image features of the image to be cropped and the text features of the text description; a reinforcement learning algorithm is used to collaboratively optimize the initial cropping frame through a composite reward function that integrates multiple objectives such as aesthetic evaluation results, semantic similarity scores, and boundary constraints, thereby obtaining multiple candidate cropped images; a multi-dimensional scoring mechanism is further used to rank and analyze the candidate cropped images, thereby automatically and efficiently generating high-quality cropping results that meet the user's intentions, achieving an optimal balance between aesthetic quality and semantic consistency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is one of the flow charts of the text-guided image cropping method provided by the present invention; Figure 2 This is the second flow chart of the text-guided image cropping method provided by the present invention; Figure 3 This is the image result automatically cropped based on the text description provided by the present invention; Figure 4 is a schematic diagram of the ablation experiment results provided by the present invention; Figure 5 Schematic diagram of the structure of the text-guided image cropping device provided by the present invention; Figure 6It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0020] With the rapid development of internet technology, social media has become an integral part of people's daily lives. Images and text information are ubiquitous on these platforms. When users upload images from their cameras to social media or share them across social media platforms, they often need to crop them to enhance their visual quality, such as by generating thumbnails. The textual content presented on social media is often closely related to the user's intended use of the image. Therefore, the results of image cropping are significantly influenced by user intent. Different users may have different cropping requirements and preferences for the same image.

[0021] However, most current image cropping methods fail to fully consider user intent, which is crucial for achieving ideal cropping results. Therefore, designing a method that can automatically crop images based on user intent to meet human aesthetic evaluation has become a challenging and highly practical problem.

[0022] In recent years, vision-language models have made significant progress in object detection. Some researchers have begun to explore leveraging these models to implement image cropping based on user intent. The basic idea of ​​this approach is to use the user-entered text as a query and use a vision-language model to evaluate the semantic similarity between the image and the query text. However, when the image composition is complex or the part mentioned in the user's text is located at the edge of the image, the expected thumbnail should match the user's description rather than simply redirecting or cropping the entire image, which would lead to serious aesthetic issues.

[0023] Currently, there is very little research that simultaneously considers user text descriptions and the aesthetic quality of cropped images, and there are several limitations and deficiencies. First, most existing image cropping methods adopt a single-step cropping strategy, that is, generating the final cropping result in one go, ignoring the gradual and iterative nature of image cropping, which may lead to inaccurate and unstable cropping results. Second, these methods are mainly based on a single cropping objective, that is, they mainly consider the semantic similarity between image and text, while to some extent ignoring the aesthetic quality of the image, which may lead to a lack of aesthetics in the cropping results. Therefore, there is an urgent need for a cropping method that can comprehensively consider the semantic similarity between image and text and the aesthetic quality of the image to determine a more accurate cropping target, thereby obtaining a more natural and beautiful cropping result.

[0024] To address these issues, this paper proposes a text-guided aesthetic cropping optimization framework to address existing image cropping techniques, including insufficient understanding of user intent, difficulty preserving small objects, and a lack of aesthetic appeal in cropped results. By constructing a joint text-visual space and leveraging deep reinforcement learning to optimize cropping decisions, this framework achieves aesthetic cropping results that align with user intent, thus overcoming these drawbacks.

[0025] Figure 1 This is one of the flow charts of the text-guided image cropping method provided by the present invention. Figure 1 As shown, the method includes: Step 110 : extracting features from the image to be cropped and the text description, and generating an initial cropping frame that semantically matches the text description based on the extracted image features and text features.

[0026] It should be noted that the image to be cropped refers to the original image to be cropped. This image will be cropped based on the text description provided by the user to highlight or display the portion related to the text description. The text description is a brief description of what the user wants the image to appear after cropping. The user guides the image cropping algorithm to help it determine which portions of the image to retain or emphasize. It should be understood that the image to be cropped can be uploaded or selected by the user, while the text description can be obtained based on user input (such as text input or voice input).

[0027] Specifically, for images to be cropped, deep learning models (such as convolutional neural networks (CNNs) and Transformer-based visual models) can be used to extract image features. These features can capture the image's visual content, such as color, texture, and shape. For text descriptions, natural language processing techniques, such as word embeddings or pre-trained models like BERT, can be used to extract text features. These features can capture the semantic information in the text description. It should be understood that the extracted image features are digital representations of the image data, reflecting the image's visual content and structure; whereas the extracted text features are digital representations of the text description, reflecting the semantic information and context in the text description.

[0028] In one embodiment, after extracting the image features of the image to be cropped and the text features of the text description, a cross-attention mechanism can be used to enhance and fuse them. The cross-attention mechanism allows the model to establish connections between image and text features, emphasizing the image regions most relevant to the text description by calculating attention weights, thereby generating an initial cropping box that matches the semantics of the text description.

[0029] In another embodiment, after extracting the image features of the image to be cropped and the text features of the text description, a matching algorithm (such as cosine similarity or inner product) can be used to calculate the similarity between the image features and the text features. Based on this similarity, the region on the image that best matches the text description can be located, and an initial cropping box containing this region can be generated.

[0030] Here, the initial cropping box is a rectangular region that locates the area of ​​the image to be cropped that is most relevant to the text description based on the matching results of image and text features. This cropping box is the starting point of the cropping process and is subsequently further optimized and adjusted to obtain the final cropping result.

[0031] It is understood that step 110 is the foundation of the entire cropping process. Through feature extraction and matching, it provides an initial cropping region related to the text description for subsequent cropping operations. This step ensures that the cropping operation is closely aligned with the user's intent, helping to improve the accuracy and pertinence of cropping.

[0032] Step 120: Dynamically optimize the initial cropping frame using a reinforcement learning algorithm to obtain multiple candidate cropped images. The reward function of the reinforcement learning algorithm is determined based on an aesthetic evaluation result, a semantic similarity score, and a boundary constraint. The aesthetic evaluation result is used to characterize the difference between the current candidate cropped image and the previous candidate cropped image. The semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description. The boundary constraint is used to constrain the cropping frame to not exceed the boundary of the image to be cropped.

[0033] It's important to note that reinforcement learning (RL) is a branch of machine learning that focuses on learning how to take actions in an environment to maximize a cumulative reward through trial and error. In the image cropping task, a reinforcement learning algorithm is used to dynamically adjust the position and size of the cropping box to find the optimal cropping result. This process is iterative, with the algorithm trying different cropping box configurations and evaluating each configuration based on a reward function.

[0034] Specifically, the reinforcement learning algorithm takes an initial cropping box as a starting point and attempts to optimize the cropping result through a series of actions (such as moving the cropping box's boundaries and resizing it). These actions are selected based on an evaluation of the current state (i.e., the current cropping box configuration) and the expected reward. The algorithm continuously tries new actions and updates its cropping policy based on the reward function to better select actions in future iterations. This process continues until the algorithm finds a satisfactory cropping result or reaches a predetermined number of iterations.

[0035] It's understandable that during the iterative process of reinforcement learning, the algorithm generates multiple different cropping configurations, each corresponding to a candidate crop image. These candidate crops are generated as the algorithm attempts to find the optimal cropping result. By comparing the performance of these candidate images on the reward function, the algorithm can gradually approach the optimal cropping result.

[0036] The reward function is a core component of the reinforcement learning algorithm, used to evaluate the quality of each cropping box configuration (i.e., each candidate cropped image). In this embodiment of the present invention, the reward function is determined based on the aesthetic evaluation results, the semantic similarity score, and the boundary constraints. The aesthetic evaluation results reflect the difference between the current candidate cropped image and the previous candidate cropped image; the semantic similarity score measures the semantic consistency between the candidate cropped image and the text description; and the boundary constraints ensure that the cropping box does not exceed the boundaries of the image to be cropped.

[0037] In this embodiment of the present invention, a reinforcement learning algorithm plays a crucial role in the text-guided image cropping method. It dynamically adjusts the configuration of the cropping box to find the optimal cropping result, continuously learning and improving its strategy in the process. The reward function, serving as the algorithm's guiding principle, comprehensively considers multiple objectives, including aesthetic evaluation results, semantic similarity scores, and boundary constraints. This ensures that the cropping result not only conforms to the user's text description but also achieves a high level of aesthetics, while also preventing the cropping box from exceeding the image boundaries.

[0038] Step 130 , scoring each candidate cropped image using a multi-dimensional scoring mechanism, and determining a target cropped image based on the scores of each candidate cropped image, wherein the multi-dimensional scoring mechanism includes an aesthetic quality score and a semantic similarity score, and the aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

[0039] It should be noted that step S120 is a process of dynamically optimizing the initial cropping frame based on a reinforcement learning algorithm. In this process, the reinforcement learning algorithm performs multi-objective collaborative optimization on the initial cropping frame by designing a composite reward function that includes aesthetic evaluation, semantic similarity scoring, and boundary constraints. This means that although aesthetic evaluation and semantic similarity evaluation are also taken into account in step S120, the optimal cropping result is gradually approached by adjusting the cropping area. However, this process is dynamic and relies on the iteration and optimization strategy of reinforcement learning.

[0040] To address this issue, the present invention, based on step 120, further provides a more refined evaluation method through step S130, thereby more accurately measuring the aesthetic and semantic performance of each candidate cropped image. Furthermore, the reinforcement learning process itself has a certain degree of randomness. The multi-dimensional scoring mechanism of step S130 allows for a more objective and comprehensive evaluation of each candidate cropped image, thereby mitigating the randomness of the reinforcement learning process and providing a more reliable basis for the final cropping solution selection.

[0041] Specifically, after obtaining multiple candidate crop images through the reinforcement learning algorithm process, a multi-dimensional scoring mechanism can be used to score and rank each candidate crop image, thereby selecting the candidate crop image with the highest score as the final target crop image. Here, the multi-dimensional scoring mechanism is an evaluation method that comprehensively considers multiple different dimensions or criteria to score the candidate objects. In step 130, the multi-dimensional scoring mechanism is used to evaluate the quality of each candidate crop image. These dimensions can include aesthetic quality, semantic similarity, etc. By comprehensively considering these dimensions, the mechanism can more comprehensively and objectively evaluate the quality of each candidate crop image.

[0042] It is understood that the multi-dimensional scoring can include an aesthetic quality score and a semantic similarity score. The aesthetic quality score is used to evaluate the aesthetic performance of the candidate cropped image. Aesthetic quality can be judged based on factors such as image composition, color matching, and lighting effects. A high-quality cropped image should have an attractive visual effect and resonate with the viewer. The semantic similarity score dimension measures the semantic consistency between the candidate cropped image and the text description. It reflects whether the image content accurately conveys the information or intent in the text description.

[0043] Specifically, the multi-dimensional scoring mechanism scores candidate cropping images by defining a scoring criterion or weight for each dimension. For key dimensions such as aesthetic quality and semantic similarity, specialized algorithms or models can be designed to calculate the scores. These algorithms or models can be based on technologies such as deep learning, computer vision, or natural language processing. When calculating the scores, the algorithms consider factors such as the visual features of the image, the semantic information in the text description, and the interactions between the dimensions. Ultimately, each candidate cropping image receives a scoring vector composed of scores from multiple dimensions.

[0044] After obtaining the score vectors for each candidate cropping image, a strategy can be used to determine the target cropping image. This strategy can be based on methods such as score sorting, weighted averaging, and threshold screening. For example, the aesthetic quality score and semantic similarity score can be weighted averaged to obtain a comprehensive score, and the candidate cropping image with the highest comprehensive score can be selected as the target cropping image. Alternatively, a score threshold can be set, and only candidate cropping images with scores exceeding this threshold are considered potential target cropping images, and the best one is selected from them. Here, the target cropping image refers to the final output of the entire text-guided image cropping method. It is the cropping result that is considered to best meet the user's text description and aesthetic requirements after feature extraction, reinforcement learning optimization, and multi-dimensional scoring mechanism evaluation.

[0045] The method provided by the embodiments of the present invention can generate an initial cropping frame that matches the semantics of the text description by extracting image features of the image to be cropped and text features of the text description. A reinforcement learning algorithm is used to collaboratively optimize the initial cropping frame using a composite reward function that integrates multiple objectives, such as aesthetic evaluation results, semantic similarity scores, and boundary constraints, to obtain multiple candidate cropped images. Furthermore, a multi-dimensional scoring mechanism is used to rank and analyze the candidate cropped images, automatically and efficiently generating high-quality cropping results that meet the user's intent, achieving an optimal balance between aesthetic quality and semantic consistency.

[0046] Based on the above embodiment, step 110 specifically includes: Step 111: performing feature extraction on the image to be cropped based on an image encoder to obtain the image features, and performing feature extraction on the text description based on a text encoder to obtain the text features; Step 112: Use a cross-attention mechanism to enhance and fuse the image features and the text features to obtain fused features.

[0047] Specifically, a Transformer-based visual model (such as Swin Transformer) can be used as an image encoder to extract image features of the image to be cropped. At the same time, the pre-trained language model BERT is used as the text encoder to extract the text features of the text description .

[0048] After extracting the image and text features, they can be enhanced and fused using a cross-attention mechanism to produce fused features. The cross-attention mechanism is a deep learning attention mechanism that allows a model to focus on relevant information in one sequence (e.g., a sequence of image features) while processing another sequence (e.g., a sequence of text features). In step 112, the cross-attention mechanism is used to enhance and fuse the image and text features. This allows the model to identify the image features most relevant to the text description and the text features most relevant to the image content, thereby enabling cross-modal information interaction.

[0049] Specifically, a cross-attention mechanism is used to enhance and fuse image and text features. This can be achieved through the following steps: First, the model calculates the attention weights between each element in the image features and each element in the text features. These weights reflect the correlation between the image and text elements. Then, the model performs a weighted sum of the image and text features based on the calculated attention weights. In this way, each image element is enhanced based on its correlation with the text element, and each text element is enhanced based on its correlation with the image element. Finally, the model fuses the weighted sum of the image and text features to obtain fused features. These fused features contain both visual information from the image and semantic information from the text, providing comprehensive guidance for the subsequent cropping task, enabling the model to more accurately understand the relationship between the text description and the image content and produce more accurate cropping results.

[0050] Step 113: Based on the decoder, the fusion feature is applied to generate a candidate detection frame, and the minimum bounding rectangle of the candidate detection frame is used as the initial cropping frame.

[0051] Specifically, a decoder is a deep learning model that is typically used to convert encoded features or representations into an output sequence. In embodiments of the present invention, the decoder is used to generate candidate detection boxes based on the fused features. The decoder can be built based on architectures such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), or transformers.

[0052] Specifically, the process of the decoder generating candidate detection frames based on the fused features can include the following steps: First, the decoder decodes the fused features and converts them into a preliminary representation of a series of candidate detection frames. Then, the decoder applies a regression model to predict these preliminary representations to obtain the precise coordinates of the candidate detection frames. In addition, the decoder can also apply techniques such as non-maximum suppression to remove overlapping or redundant candidate detection frames to obtain the final candidate detection frames. Here, the candidate detection frame refers to a rectangular area that may contain the target object in the image to be cropped. It is generated by the decoder based on the fused features and represents the image area that is most relevant to the text description.

[0053] After obtaining the candidate detection frame, the model will select the smallest bounding rectangle of the detection frame as the initial cropping frame. This bounding rectangle completely encompasses all pixels within the candidate detection frame and has the smallest area. Using the candidate detection frame's minimum bounding rectangle as the initial cropping rectangle ensures that the cropped image contains the key information from the text description while minimizing unnecessary background information. Subsequent cropping optimization steps can then be performed based on this initial cropping rectangle, resulting in a more accurate and aesthetically pleasing cropping result.

[0054] Based on any of the above embodiments, step 120 specifically includes: Step 121 , based on an observation state and a cropping strategy, select an action from a predefined action space, and perform the selected action on the initial cropping frame to obtain a new cropping frame, wherein the observation state is determined based on the image to be cropped and the initial cropping frame.

[0055] It's important to note that the reinforcement learning framework includes key components such as the agent, environment, observation state, action space, and reward function. The observation state refers to the information the agent receives from the environment, which guides its decision-making process. The clipping policy is the basis for the agent to select an action given the observation state. It is typically a function or model whose input is the observation state and whose output is the probability distribution of a particular action in the action space.

[0056] Reinforcement learning typically involves multiple iterations. In the first iteration, the observation state is typically determined by extracting features from the image to be cropped and the initial cropping box. These features can include visual features such as the image's color, texture, and shape, as well as geometric features such as the coordinates and size of the cropping box. The agent can use these features to construct an initial observation state and begin its decision-making process based on this state. This initial observation state serves as the basis for the agent to learn the cropping strategy, which is continuously updated and optimized in subsequent iterations.

[0057] Specifically, the agent first receives the current observation state as input. Then, the agent applies a clipping policy function or model to map the observation state to a probability distribution in a predefined action space. Finally, the agent selects an action based on the probability distribution. This can be achieved through random sampling, a greedy policy (selecting the action with the highest probability), or some exploration strategy (such as the ε-greedy policy).

[0058] After selecting an action, the selected action is performed on the initial cropping frame, which typically involves adjusting the initial cropping frame's position, size, or shape. These adjustments can be based on the selected action type (e.g., zoom in, zoom out, translate, rotate, etc.) and action parameters (e.g., the size and direction of the adjustment). After performing the action, the agent obtains a new cropping frame. Here, the new cropping frame refers to the cropping frame obtained after the agent adjusts the initial cropping frame according to the selected action. This new cropping frame may have a different position, size, or shape, which is used for subsequent image cropping and observation state updates.

[0059] It is understandable that the predefined action space may include multiple (e.g., 19) predefined geometric transformation operations and their combinations and variations, such as scaling, translation, rotation, flipping, and termination actions, etc., so that windows of all specified sizes and positions starting from the initial cropping frame can be covered.

[0060] Step 122 : Crop the image to be cropped based on the new cropping frame to obtain a current candidate cropping image, and apply the current candidate cropping image to update the observation state and the cropping strategy.

[0061] Specifically, after obtaining a new cropping frame, the coordinate information of the new cropping frame (including the coordinates of the upper left and lower right corners, or parameters such as the center point, width, and height) can be used to extract a sub-image from the image to be cropped. This sub-image is the image area obtained based on the new cropping frame, that is, the current candidate cropping image. Here, the current candidate cropping image refers to the image that is extracted from the image to be cropped based on the current new cropping frame during each iteration. This image will serve as input for subsequent steps (such as observation state updates, cropping policy updates, and reward value calculations) to guide the agent's decision-making process. As the number of iterations increases, the current candidate cropping image will continue to approach the user's desired cropping result.

[0062] It is understandable that the current candidate crop image provides new visual information, which is used to update the observation state. The observation state can be composed of information such as the coordinates and size of the current cropping box, the historical action sequence, and the feature representation of the image, which is used to guide the subsequent decision-making process. At the same time, the current candidate crop image is also used to calculate the reward value, which reflects the quality of the current cropping result. Based on the reward value and the historical performance of the cropping policy, the cropping policy can be updated according to the policy gradient. In this way, the agent can gradually learn how to choose the optimal action based on the observation state to maximize the long-term reward.

[0063] It's important to note that the purpose of updating the observation state and cropping strategy is to continuously optimize the agent's decision-making process, gradually bringing it closer to the user's desired cropping result. By continuously updating the observation state, the agent acquires more comprehensive and accurate image information; and by continuously updating the cropping strategy, the agent gradually learns how to choose the optimal action based on this information. As the number of iterations increases, the agent's cropping results become increasingly accurate and meet user expectations.

[0064] Furthermore, in step 122, applying the current candidate cropping image to update the observation state and the cropping strategy includes: Step 1221: determining current observation information based on the new cropping frame, the current candidate cropping image, and the image to be cropped, and updating the observation state according to historical observation information and the current observation information to obtain an updated observation state; Step 1222: Calculate a reward value based on the current candidate cropped image and the reward function, and update the cropping strategy according to the reward value to obtain an updated cropping strategy.

[0065] Specifically, starting from the second iteration, the observation state can be composed of historical observation information and current observation information. The current observation information refers to the information extracted from the new cropping frame, the current candidate cropping image, and the original image (i.e., the image to be cropped) in the current iteration. For example, the current observation information may include the original image, the new cropping frame coordinates, and the score difference between the cropped image (i.e., the current candidate cropping image) and the original image. Historical observation information refers to the observation information accumulated in previous iterations and can be recorded by the LSTM (Long Short-Term Memory) unit of the reinforcement learning model. After each iteration, the LSTM unit can record the current observation information. In this way, in subsequent iterations, this historical observation information can be used to assist the agent's decision-making process.

[0066] After obtaining the current observation information and historical observation information, these information are fused or spliced ​​to form a more comprehensive and accurate representation of the observation state, that is, to obtain the updated observation state.

[0067] The current candidate crop is also used to calculate a reward, which is then used to update the cropping strategy. The reward is calculated based on the current candidate crop and a predefined reward function. This reward function comprises multiple components, including aesthetic evaluation, semantic similarity, and boundary constraints. The aesthetic evaluation is performed by using a View Finding Network (VFN) model to evaluate the difference between the output of the new crop and the previous crop. The semantic similarity score is achieved using a semantic similarity network, which assesses the semantic similarity between the current candidate crop and the user-entered text description. Boundary constraints restrict the position and size of the crop to ensure that the cropped image meets specific requirements or specifications. For example, the crop cannot extend beyond the boundaries of the original image, or the cropped image must meet a certain aspect ratio. These constraints influence the final reward by calculating a penalty term.

[0068] When calculating the reward value, the scores of the above components can be weighted and summed or otherwise combined to obtain a comprehensive reward value. Here, the reward value is used to evaluate how well the agent performs an action in a given state. In image cropping tasks, the reward value can be considered a quantitative evaluation of the current cropping result. A higher reward value means that the cropping result is more in line with user expectations or preset standards; a lower reward value means that the cropping result needs improvement.

[0069] Once the reward is calculated, the parameters of the clipping policy can be updated based on the reward. For example, a gradient or update can be calculated based on the reward and the clipping policy's historical performance and applied to the clipping policy parameters. This way, over more iterations, the clipping policy will gradually learn how to choose the optimal action based on the observed state to maximize the long-term reward.

[0070] Step 123 : Based on the updated observation state and cropping strategy, continue to select an action from the predefined action space to perform the selected action on the new cropping frame, and repeat the above steps until a plurality of candidate cropped images are obtained.

[0071] Specifically, step 123 is an iterative step in the image cropping process, which continues to select an action from the predefined action space based on the updated observation state and cropping strategy, and performs the selected action on the new cropping box. This process is repeated until multiple candidate cropped images are obtained.

[0072] In each iteration, the agent selects an optimal action (such as adjusting the size, position, or shape of the cropping box) based on its current observation state and cropping strategy. It then performs the selected action on the new cropping box, generating a new candidate cropped image. The agent then updates its observation state and cropping strategy based on this new candidate cropped image. This process continues until a stopping condition is met (such as reaching a preset number of iterations or the cropping result meeting a certain quality standard).

[0073] Based on any of the above-mentioned embodiments, the reinforcement learning algorithm provided by the embodiments of the present invention employs an actor-critic architecture. This architecture improves learning efficiency and stability by simultaneously training an actor (policy network) and a critic (value network). The actor is responsible for selecting the probability distribution of actions in a given state, that is, selecting actions based on the current policy. This typically uses a neural network to approximate the policy function, with the state as input and the possible action probability distribution as output. The critic is responsible for evaluating the quality of the actor's actions, that is, estimating the value of the state or state-action pair. Similarly, the critic uses another neural network to approximate the value function, with the state (or state and action) as input and a scalar value representing the estimated value. The critic's role is to provide optimization direction for the actor, telling the actor which actions are good and which are bad.

[0074] Specifically, in the Actor-Critic architecture, the current state space It is composed of the current cropping frame coordinates, historical action sequence and image features. Given a text description and obtaining the initial cropping frame from the original image , the agent selects actions based on the highest expected reward, taking into account both historical experience (i.e., historical observation information) and current observation information , where the current observation information Including the original image, cropping frame coordinates, the score difference between the cropped image and the original image; the historical observation information and current observation information memorized by the LSTM unit Together they constitute the current observation state : The action space contains 19 predefined geometric transformation operations, including scaling, translation, and terminal actions, which allow the agent to cover all windows of specified sizes and positions starting from the initial crop box: in represents the probability distribution of the action at step t+1, is the probability distribution of the policy output, is the raw score of the output, which can be stopped when the score no longer increases to obtain the best cropping box.

[0075] Reward Function Defined as: in is the output difference between the aesthetic evaluation model VFN of the new cropping frame and the previous cropping frame, is the similarity value between the cropped image and the text description, and these indicators are used to calculate the reward for each action. In addition, the agent can also be assigned Negative rewards, and is a hyperparameter that adjusts their relative importance.

[0076] Based on any of the above embodiments, in step 130, the multi-dimensional scoring mechanism is used to score each candidate cropped image, including: Step 131 : performing semantic similarity scoring on any candidate cropped image and the text description to obtain a semantic similarity score of the candidate cropped image.

[0077] Specifically, the process of scoring the semantic similarity between any candidate cropped image and text description is essentially a multimodal information matching process, which usually involves converting the image and text into feature vectors respectively and calculating the similarity between the two feature vectors.

[0078] Furthermore, step 131 specifically includes: Step 1311: extracting an image feature vector of any candidate cropped image and a text feature vector of the text description based on a text-image multimodal model; Step 1312 : performing similarity calculation based on the image feature vector and the text feature vector, and determining a semantic similarity score of any candidate cropped image according to the calculation result.

[0079] Specifically, a pre-trained text-image multimodal model can be used to extract the image feature vector and the text feature vector of any candidate cropped image, respectively. Then, based on these two feature vectors, the similarity between them can be calculated using cosine similarity, Euclidean distance, or other measurement methods. This similarity value can be used as the semantic similarity score of the candidate cropped image. Here, the semantic similarity score of any candidate cropped image refers to the degree of semantic match between the candidate cropped image and the text description. This score can be a value between 0 and 1, where a higher score indicates that the candidate cropped image is semantically closer or matches the text description.

[0080] It is understood that the text-image multimodal model can be a CLIP (Contrastive Language-Image Pre-training) model, a deep learning model that can learn the correspondence between images and text. This model uses contrastive learning to project images and text into a common embedding space, so that similar image and text pairs are closer in this space, while dissimilar pairs are farther apart.

[0081] During the training process, the CLIP model receives a large number of image-text pairs as input and learns how to map the images and text in these pairs into a shared embedding space. In this way, after training is completed, the CLIP model can be used to extract feature vectors of images and texts and calculate the similarity between them.

[0082] Step 132 : performing an aesthetic quality score on any candidate cropped image to obtain an aesthetic score of any candidate cropped image.

[0083] Specifically, the aesthetic quality score of any candidate cropped image can be achieved using a pre-trained aesthetic evaluation model (i.e., VFN model). This model can be a deep learning network that can perform aesthetic analysis on the image and give an aesthetic quality score.

[0084] It is understandable that the aesthetic evaluation model has been trained on a large number of aesthetic images and is therefore able to more accurately assess the aesthetic quality of an image in terms of composition, color, contrast, clarity, etc. During the scoring process, a candidate cropped image can be input into the aesthetic evaluation model, and the aesthetic quality score output by the model is obtained as the aesthetic score of the candidate cropped image.

[0085] Here, the aesthetic score of any candidate cropped image refers to the score of the candidate cropped image in terms of aesthetic quality. This score is usually a numerical value that reflects the quality of the image in terms of composition, color, contrast, etc.

[0086] Step 133 : performing weighted fusion on the semantic similarity score and the aesthetic score of any candidate cropped image to obtain a score of any candidate cropped image.

[0087] Specifically, for any candidate cropped image, after calculating its semantic similarity score and aesthetic score, a weighted sum of the semantic similarity score and aesthetic score is performed based on pre-set weights to obtain the final score for the candidate cropped image. It should be understood that the weights corresponding to the semantic similarity score and aesthetic score can be pre-set based on actual needs or experimental results.

[0088] For example, through the reinforcement learning algorithm process of the above embodiment, several candidate cropped images can be obtained. Then, the CLIP model can be used to extract the image feature vector of each candidate cropped image. and the text feature vector of the text description , and the two are expressed as follows: These vectors are normalized unit vectors. According to the matching mechanism of the CLIP model, the cosine similarity is used to measure the vector clamping angle to reflect their similarity: where s∈R, For the aesthetic quality scoring task, the VFN network is used to achieve the goal of maximizing the aesthetic score output by the model. The cropped image is input into the VFN network to obtain the image features, which are multiplied by the pre-trained weight matrix to obtain the aesthetic score: in is a feature vector, W Is a trainable weight matrix, represents matrix multiplication. Given a text description T The cropping result (i.e. cropped image) under The final score is expressed as: where λ is a hyperparameter.

[0089] After calculating the score of each candidate cropped image, the candidate cropped image can be sorted according to the score, and the candidate cropped image with the highest score is selected as the final target cropped image.

[0090] Based on any of the above embodiments, Figure 2 This is the second flow chart of the text-guided image cropping method provided by the present invention. Figure 2 As shown, an embodiment of the present invention provides a method for automatically generating an image intelligent cropping that meets the user's intention and the public's aesthetic taste based on a given text description. The method can be implemented based on an image cropping model, which includes three modules, namely a candidate box generation module, a reinforcement learning optimization module, and a multi-dimensional scoring module. Among them, the candidate box generation module is used to determine the initial cropping box based on the visual language model; the reinforcement learning optimization module is used to optimize the cropping box based on the reinforcement learning algorithm, designing a composite reward function of an aesthetic scoring network, a semantic similarity evaluation network, and boundary constraints; the multi-dimensional scoring module is used to score each candidate cropping result to determine the final image aesthetic cropping result. The method specifically includes the following steps: Step S1: The candidate box generation module based on multimodal feature fusion extracts cross-modal features of the input image (i.e., the image to be cropped) and the text description through a pre-trained visual language model, and generates an initial cropping box that semantically matches the text description.

[0091] Specifically, the Swin Transformer can be used as an image encoder to extract image features; BERT can be used as a text encoder to extract text features; then, a cross-attention mechanism is used for feature enhancement and cross-modal feature fusion; finally, a dynamic query decoder is used to generate candidate detection boxes and calculate the minimum enclosing rectangle as the initial cropping box. It should be understood that this step, by introducing text as input to initially reflect user intent, introduces the characteristics and advantages of large-scale open vocabulary knowledge into the image cropping algorithm.

[0092] Step S2: Dynamically optimize the initial cropping frame based on the reinforcement learning framework, perform multi-objective collaborative optimization of a composite reward function that includes an aesthetic scoring network, a semantic similarity evaluation network, and boundary constraints, and adjust the cropping area through a predefined action space to gradually approach the optimal cropping result.

[0093] It's important to note that the reinforcement learning framework uses an actor-critic architecture, which includes multiple FC (Fully Connected) layers, multiple ReLU (an activation function), LSTM units, and linear layers. The linear layers are divided into the actor (policy network) linear layer and the critic (value network) linear layer. The LSTM unit is used to record observation information during each iteration to accumulate historical observation information.

[0094] Specifically, in the initial step, the initial cropping frame is input into a convolutional network for feature extraction. The extracted features (such as cropping frame coordinates, size, and image characteristics) are used to determine the current observation information and, combined with historical observation information, form the current observation state. Based on the current observation state and cropping strategy, the agent selects an optimal action from a predefined action space. It then executes the selected action on the initial cropping frame, generating a new candidate cropped image. The agent then updates the observation state based on this new candidate cropped image and calculates a reward value using a reward function to update the cropping strategy. A new round of iteration begins based on the updated observation state and cropping strategy. This process continues until a stopping condition is met (such as reaching a preset number of iterations or the cropping result meeting a certain quality standard).

[0095] Step S3: After obtaining several candidate cropping images through the above reinforcement learning algorithm process, these candidate cropping images are scored and ranked using a multi-dimensional scoring mechanism. The scoring mechanism includes a text-image semantic similarity score based on the CLIP model and an aesthetic quality score based on a deep neural network. The final cropping scheme is determined through weighted fusion, that is, the target cropping image is obtained. This can reduce the randomness of the reinforcement learning process and achieve an optimal balance between aesthetic quality and semantic consistency.

[0096] In the method provided by an embodiment of the present invention, a candidate frame generation module extracts features of images and text through a pre-trained visual language model to generate an initial cropping frame that conforms to the text description; a reinforcement learning framework is adopted to collaboratively optimize the initial cropping frame through a composite reward function, comprehensively considering multiple objectives such as aesthetic score, semantic similarity, and boundary constraints; a multi-dimensional scoring mechanism is used to rank and analyze the candidate cropping results to achieve an optimal balance between aesthetic quality and semantic consistency, thereby automatically and efficiently generating the best image aesthetic cropping result that meets the user's intention.

[0097] Based on any of the above embodiments, this embodiment of the present invention selected two representative annotated datasets, the GAIC dataset and the Horanyi-PR dataset, for experiments. The GAIC dataset consists of 280,000 candidate detection boxes across 3,336 images, with 2,636 images used for training, 200 images for validation, and 500 images for testing. Each cropped box is assigned a quality score between 1 and 5. To apply the text-guided cropping algorithm, the GAIC images are enhanced by incorporating text descriptions for each annotated image.

[0098] The Horanyi-PR dataset randomly selects 100 images from the MS-COCO test set and creates a caption for each image. It then adds eight ground truth bounding box annotations for each caption. The purpose of this dataset is to evaluate the ability to automatically crop images with captions.

[0099] In addition, during the experiment, some complex images were annotated with different text descriptions to qualitatively evaluate the adaptability of the model of the embodiment of the present invention to complex scenes and different texts. Figure 3 This is the image result automatically cropped based on the text description provided by the present invention, such as Figure 3 As shown, by utilizing text input as a control condition, different cropping results can be produced on the same image.

[0100] Figure 4 is a schematic diagram of the ablation experiment results provided by the present invention, such as Figure 4 As shown, the embodiment of the present invention also creates some baselines to verify the effectiveness of the method. Figure 4 (a) corresponds to the original image, and (b) corresponds to the input text description. The ablation experiment results are as follows: ① Lack of candidate box generation module: Figure 4 As shown in (c), the second row does not prioritize the pensive man, but instead focuses on two individuals, resulting in a cropped image that does not match the text description.

[0101] ② Lack of aesthetic rewards: exclude aesthetic quality evaluation from the reward function within the reinforcement learning module and maintain semantic similarity. Figure 4 As shown in (d), the cropped image is closely aligned with the text description but lacks visual appeal or aesthetic quality.

[0102] ③ Lack of semantic similarity reward: Modify the reinforcement learning module so that it only considers aesthetic quality in the reward function. Figure 4 As shown in (e), the cropped image shows visual appeal but fails to faithfully represent the user-specified preferences expressed in the text.

[0103] ④ Lack of sampling and sorting: The variation of cropping results is due to the inherent uncertainty of the reinforcement learning process. By sampling multiple cropping results and selecting the best one based on aesthetics and semantic similarity with the text. Figure 4 As shown in (f), removing the sampling sorting process will cause the final cropping result to be unstable.

[0104] Based on any of the above embodiments, a user survey was conducted in accordance with an embodiment of the present invention, as specifically described below: The results of the method provided by the present invention were compared with those of several existing image cropping techniques. A total of 50 participants participated in the study, including 22 computer graphics or computer vision researchers aged between 20 and 50. Each participant evaluated 40 randomly selected text-image pairs, comparing the cropping results of the method proposed by the present invention with those of another method. Participants were briefly introduced to the goals and parameters of the motion customization task and were asked to evaluate the cropping results based on text relevance, aesthetic quality, and user satisfaction. The results of the user survey are shown in Table 1 below:

[0105] As shown in the table above, G-DINO, GAIC, A2RL, ReIC, and Ground Truth, shown in the first column, are all existing image cropping technologies. This shows that the intelligent image cropping method provided by the present invention has achieved higher user preference, particularly in terms of aesthetic quality and user satisfaction. Because the visual-based algorithm focuses on accurately cropping bounding boxes, although the method provided by the present invention does not perform as well as some direct object detection algorithms (such as G-DINO and Ground Truth) in terms of text relevance, the performance difference is not significant. Furthermore, in all other aspects, the method provided by the present invention outperforms other algorithms. User studies have shown that the intelligent image cropping method provided by the present invention excels in maintaining description quality and aesthetic appeal, achieving the highest user preference and superior cropping results.

[0106] The text-guided image cropping device provided by the present invention is described below. The text-guided image cropping device described below and the text-guided image cropping method described above can refer to each other.

[0107] Based on any of the above embodiments, Figure 5 Schematic diagram of the structure of the text-guided image cropping device provided by the present invention. Figure 5 As shown, the device includes: The cropping frame generating unit 510 is configured to extract features of the image to be cropped and the text description, and generate an initial cropping frame that semantically matches the text description based on the extracted image features and text features; a cropping frame optimization unit 520 configured to dynamically optimize the initial cropping frame using a reinforcement learning algorithm to obtain a plurality of candidate cropped images, wherein a reward function of the reinforcement learning algorithm is determined based on an aesthetic evaluation result, a semantic similarity score, and a boundary constraint condition, wherein the aesthetic evaluation result is used to characterize the difference between the current candidate cropped image and the previous candidate cropped image, the semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description, and the boundary constraint condition is used to constrain the cropping frame to not exceed the boundary of the image to be cropped; The cropped image determination unit 530 is configured to score each candidate cropped image using a multi-dimensional scoring mechanism and determine a target cropped image based on the scores of each candidate cropped image. The multi-dimensional scoring mechanism includes an aesthetic quality score and a semantic similarity score. The aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

[0108] The device provided by the embodiment of the present invention can generate an initial cropping frame that matches the semantics of the text description by extracting image features of the image to be cropped and text features of the text description. A reinforcement learning algorithm is used to collaboratively optimize the initial cropping frame through a composite reward function that integrates multiple objectives such as aesthetic evaluation results, semantic similarity scores, and boundary constraints, thereby obtaining multiple candidate cropped images. Furthermore, a multi-dimensional scoring mechanism is used to rank and analyze the candidate cropped images, thereby automatically and efficiently generating high-quality cropping results that meet the user's intent, achieving an optimal balance between aesthetic quality and semantic consistency.

[0109] Based on any of the above embodiments, the cropping frame generating unit 510 is specifically configured to: Based on an image encoder, feature extraction is performed on the image to be cropped to obtain the image features, and based on a text encoder, feature extraction is performed on the text description to obtain the text features; Using a cross-attention mechanism, the image features and the text features are enhanced and fused to obtain fused features; Based on the decoder, the fusion features are applied to generate a candidate detection frame, and the minimum bounding rectangle of the candidate detection frame is used as the initial cropping frame.

[0110] Based on any of the above embodiments, the cropping frame optimization unit 520 includes: an execution subunit, configured to select an action from a predefined action space based on an observation state and a cropping strategy, and execute the selected action on the initial cropping frame to obtain a new cropping frame, wherein the observation state is determined based on the image to be cropped and the initial cropping frame; an updating subunit, configured to crop the image to be cropped based on the new cropping frame to obtain a current candidate cropped image, and apply the current candidate cropped image to update the observation state and the cropping strategy; The iterative subunit is configured to continue selecting an action from the predefined action space based on the updated observation state and cropping strategy to perform the selected action on the new cropping frame, and repeat the above steps until a plurality of candidate cropped images are obtained.

[0111] Based on any of the above embodiments, the updating subunit is specifically configured to: Determining current observation information based on the new cropping frame, the current candidate cropping image, and the image to be cropped, and updating the observation state according to historical observation information and the current observation information to obtain an updated observation state; Based on the current candidate cropped image and the reward function, a reward value is calculated, and the cropping strategy is updated according to the reward value to obtain an updated cropping strategy.

[0112] Based on any of the above embodiments, the predefined action space includes multiple predefined geometric transformation operations.

[0113] Based on any of the above embodiments, the cropped image determining unit 530 includes: a similarity score calculation subunit, configured to perform semantic similarity scoring on any candidate cropped image and the text description to obtain a semantic similarity score of the candidate cropped image; an aesthetic score calculation subunit, configured to perform an aesthetic quality score on any of the candidate cropped images to obtain an aesthetic score for the candidate cropped image; The fusion determination subunit is used to perform weighted fusion on the semantic similarity score and the aesthetic score of any candidate cropped image to obtain the score of any candidate cropped image.

[0114] Based on any of the above embodiments, the similarity score calculation subunit is specifically configured to: Extracting an image feature vector of any candidate cropped image and a text feature vector of the text description based on a text-image multimodal model; A similarity calculation is performed based on the image feature vector and the text feature vector, and a semantic similarity score of any candidate cropped image is determined according to the calculation result.

[0115] Figure 6 An example of a physical structure diagram of an electronic device is shown below. Figure 6As shown, the electronic device may include: a processor 610 , a communication interface 620 , a memory 630 and a communication bus 640 , wherein the processor 610 , the communication interface 620 , and the memory 630 communicate with each other via the communication bus 640 . The processor 610 can call logic instructions in the memory 630 to execute a text-guided image cropping method, the method including: extracting features of the image to be cropped and the text description, and generating an initial cropping frame that semantically matches the text description based on the extracted image features and text features; dynamically optimizing the initial cropping frame using a reinforcement learning algorithm to obtain multiple candidate cropped images, wherein a reward function of the reinforcement learning algorithm is determined based on an aesthetic evaluation result, a semantic similarity score, and a boundary constraint condition, wherein the aesthetic evaluation result is used to characterize the difference between the current candidate cropped image and the previous candidate cropped image, the semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description, and the boundary constraint condition is used to constrain the cropping frame to not exceed the boundary of the image to be cropped; and scoring each candidate cropped image using a multi-dimensional scoring mechanism, and determining a target cropped image based on the score of each candidate cropped image, wherein the multi-dimensional scoring includes an aesthetic quality score and a semantic similarity score, and the aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

[0116] In addition, the logic instructions in the aforementioned memory 630 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the relevant art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0117] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the text-guided image cropping method provided by the above methods, the method comprising: extracting features of the image to be cropped and the text description, and generating an initial cropping frame that semantically matches the text description based on the extracted image features and text features; using a reinforcement learning algorithm to dynamically optimize the initial cropping frame to obtain multiple candidate cropped images, and the reward function of the reinforcement learning algorithm is based on An aesthetic evaluation result, a semantic similarity score, and a boundary constraint condition are determined. The aesthetic evaluation result is used to characterize the difference between the current candidate cropped image and the previous candidate cropped image. The semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description. The boundary constraint condition is used to constrain the cropping box to not exceed the boundary of the image to be cropped. A multi-dimensional scoring mechanism is used to score each candidate cropped image, and a target cropped image is determined based on the score of each candidate cropped image. The multi-dimensional scoring includes an aesthetic quality score and a semantic similarity score. The aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

[0118] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the text-guided image cropping method provided by the above methods, the method comprising: extracting features of an image to be cropped and a text description, and generating an initial cropping frame that semantically matches the text description based on the extracted image features and text features; dynamically optimizing the initial cropping frame using a reinforcement learning algorithm to obtain multiple candidate cropped images, wherein a reward function of the reinforcement learning algorithm is determined based on an aesthetic evaluation result, a semantic similarity score, and boundary constraints, wherein the aesthetic evaluation result is used to characterize the difference between a current candidate cropped image and a previous candidate cropped image, the semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description, and the boundary constraints are used to constrain the cropping frame to not exceed the boundary of the image to be cropped; and scoring each candidate cropped image using a multi-dimensional scoring mechanism, and determining a target cropped image based on the scores of each candidate cropped image, wherein the multi-dimensional scoring includes an aesthetic quality score and a semantic similarity score, and the aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

[0119] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0120] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A text-guided image cropping method, characterized in that: include: Extracting features of the image to be cropped and the text description, and generating an initial cropping frame that semantically matches the text description based on the extracted image features and text features; Dynamically optimizing the initial cropping frame using a reinforcement learning algorithm to obtain a plurality of candidate cropped images, wherein a reward function of the reinforcement learning algorithm is determined based on an aesthetic evaluation result, a semantic similarity score, and a boundary constraint condition, wherein the aesthetic evaluation result is used to characterize the difference between a current candidate cropped image and a previous candidate cropped image, the semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description, and the boundary constraint condition is used to constrain the cropping frame to not exceed the boundary of the image to be cropped; A multi-dimensional scoring mechanism is used to score each candidate cropped image, and a target cropped image is determined based on the scores of each candidate cropped image. The multi-dimensional scoring mechanism includes an aesthetic quality score and a semantic similarity score. The aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

2. The text-guided image cropping method according to claim 1, characterized in that: The step of extracting features from the image to be cropped and the text description, and generating an initial cropping frame that semantically matches the text description based on the extracted image features and text features, includes: Based on an image encoder, feature extraction is performed on the image to be cropped to obtain the image features, and based on a text encoder, feature extraction is performed on the text description to obtain the text features; Using a cross-attention mechanism, the image features and the text features are enhanced and fused to obtain fused features; Based on the decoder, the fusion features are applied to generate a candidate detection frame, and the minimum bounding rectangle of the candidate detection frame is used as the initial cropping frame.

3. The text-guided image cropping method according to claim 1, characterized in that: The reinforcement learning algorithm is used to dynamically optimize the initial cropping frame to obtain multiple candidate cropping images, including: Based on an observation state and a cropping strategy, selecting an action from a predefined action space and performing the selected action on the initial cropping frame to obtain a new cropping frame, wherein the observation state is determined based on the image to be cropped and the initial cropping frame; Cropping the image to be cropped based on the new cropping frame to obtain a current candidate cropped image, and applying the current candidate cropped image to update the observation state and the cropping strategy; Based on the updated observation state and cropping strategy, continue to select an action from the predefined action space to perform the selected action on the new cropping frame, and repeat the above steps until multiple candidate cropping images are obtained.

4. The text-guided image cropping method according to claim 3, characterized in that: The applying the current candidate cropping image to update the observation state and the cropping strategy includes: Determining current observation information based on the new cropping frame, the current candidate cropping image, and the image to be cropped, and updating the observation state according to historical observation information and the current observation information to obtain an updated observation state; Based on the current candidate cropped image and the reward function, a reward value is calculated, and the cropping strategy is updated according to the reward value to obtain an updated cropping strategy.

5. The text-guided image cropping method according to claim 3, characterized in that: The predefined action space includes a plurality of predefined geometric transformation operations.

6. The text-guided image cropping method according to any one of claims 1 to 5, characterized in that: The multi-dimensional scoring mechanism is used to score each candidate cropped image, including: Performing semantic similarity scoring on any candidate cropped image and the text description to obtain a semantic similarity score for the candidate cropped image; Performing an aesthetic quality score on any candidate cropped image to obtain an aesthetic score for the candidate cropped image; The semantic similarity score and the aesthetic score of any candidate cropped image are weightedly fused to obtain the score of any candidate cropped image.

7. The text-guided image cropping method according to claim 6, characterized in that: The step of performing semantic similarity scoring on any candidate cropped image and the text description to obtain a semantic similarity score of any candidate cropped image includes: Extracting an image feature vector of any candidate cropped image and a text feature vector of the text description based on a text-image multimodal model; A similarity calculation is performed based on the image feature vector and the text feature vector, and a semantic similarity score of any candidate cropped image is determined according to the calculation result.

8. A text-guided image cropping device, characterized in that: include: A cropping frame generating unit is used to extract features of the image to be cropped and the text description, and generate an initial cropping frame that semantically matches the text description based on the extracted image features and text features; a cropping frame optimization unit, configured to dynamically optimize the initial cropping frame using a reinforcement learning algorithm to obtain a plurality of candidate cropped images, wherein a reward function of the reinforcement learning algorithm is determined based on an aesthetic evaluation result, a semantic similarity score, and a boundary constraint condition, wherein the aesthetic evaluation result is used to characterize the difference between a current candidate cropped image and a previous candidate cropped image, the semantic similarity score is used to characterize the semantic similarity between the candidate cropped image and the text description, and the boundary constraint condition is used to constrain the cropping frame to not exceed the boundary of the image to be cropped; A cropped image determination unit is configured to score each candidate cropped image using a multi-dimensional scoring mechanism and determine a target cropped image based on the scores of each candidate cropped image, wherein the multi-dimensional scoring mechanism includes an aesthetic quality score and a semantic similarity score, and the aesthetic quality score is used to characterize the aesthetic quality of the candidate cropped image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the text-guided image cropping method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the text-guided image cropping method according to any one of claims 1 to 7 is implemented.