Composition Bounding Box Proposal Method Unrestricted by Image Boundaries

By adopting a combination method of frame regression module and feature extraction module in image processing, the problem of composition cropping in the prior art is solved, and the composition bounding box recommendation is achieved without being restricted by image boundary is improved, and the aesthetic quality of the photo is improved.

CN116563518BActive Publication Date: 2025-06-06HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310454466.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-25
Publication Date
2025-06-06
Estimated Expiration
2043-04-25

AI Technical Summary

Technical Problem

Existing composition cropping methods are limited by image boundaries and cannot obtain the best compositional photos without being restricted by image boundaries.

Method used

Using the method based on the box regression module, missing view feature extraction module, full view feature extraction module and feature completion module, the composition bounding box that is not limited by image boundaries is predicted through feature merging and loss function optimization.

Benefits of technology

It realizes the recommendation of composition bounding box without being restricted by image boundaries, helping users take photos with higher aesthetic quality and improving the quality and freedom of composition recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116563518B_ABST
    Figure CN116563518B_ABST
Patent Text Reader

Abstract

A method for recommending a composition bounding box without being restricted by an image boundary, belonging to the technical field of image processing of the present invention. It solves the problem that the existing composition cropping method is restricted by the image boundary. Each sample of the present invention includes a full-view image I, any missing-view image I obtained by cropping the full-view image I init , and the collection of cropping boxes corresponding to the missing-view image I init in the full-view image I. A sample is randomly selected from the sample set to train the missing-view feature extraction module, the feature completion module, and the box regression module. During the training process, through I init the missing-view feature map Z vis is obtained. Using Z vis the extended feature map Z pad is predicted. After merging the features of Z pad and Z vis , they are sent to the box regression module for predicting the composition bounding box, and the total loss value is calculated to update the network parameters of the missing-view feature extraction module, the box regression module, and the feature completion module. The present invention is used to recommend a composition bounding box for the camera view.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing. Background Art

[0002] With the popularity of electronic devices such as mobile phones and cameras, taking photos has become a common activity in daily life. However, due to the lack of photography experience and skills of non-professional users, it is difficult to take photos with harmonious composition and high aesthetic quality. Therefore, how to help users improve image composition is an important research direction.

[0003] In recent years, the image cropping method based on neural network has achieved good results. The image cropping method removes redundant content through cropping operation to obtain a sub-image with good composition. However, this method is a post-processing of the image. When the optimal cropping is not completely within the image range, it will inevitably get a suboptimal solution due to the image boundary restriction. Therefore, how to help users get photos with the best composition without being restricted by the image boundary is an urgent problem to be solved. Summary of the invention

[0004] The purpose of the present invention is to solve the problem that the existing composition cropping method is limited by the image boundary. The present invention provides a composition boundary box recommendation method that is not limited by the image boundary.

[0005] A composition bounding box recommendation method that is not restricted by image boundaries is implemented based on a box regression module, a missing view feature extraction module, a full view feature extraction module, and a feature completion module. The internal structures of the missing view feature extraction module and the full view feature extraction module are the same. The method includes the following steps:

[0006] S1. Construct sample set:

[0007] Each sample in the sample set includes a full-view image I and any missing view image I obtained by cropping the full-view image I. init , and a cropping frame collection obtained by cropping the full-view image I; the cropping frame collection includes N groups of calibration results, each group of calibration results includes a cropping frame position information c i And the confidence p of the crop box i ; Among them, c i is the position information of the cropping box in the i-th group calibration result in the cropping box collection, p i is the confidence of the cropping box in the i-th group of calibration results in the cropping box collection, i = 1, 2, ... N; p 1 to p N The values ​​are all 1, where each cropping frame is located in the area covered by the image content of the full-view image I, and each cropping frame is a missing view image I init The ideal composition bounding box;

[0008] S2, initializing the frame regression module, the missing view feature extraction module, the full view feature extraction module and the feature completion module, and the parameter values ​​of the missing view feature extraction module and the full view feature extraction module after initialization are exactly the same;

[0009] S3. Randomly extract a sample from the sample set and train the box regression module, the missing view feature extraction module and the feature completion module, including:

[0010] S31, feature extraction and completion:

[0011] A sample is randomly selected from the sample set, and the missing view feature extraction module is used to extract the missing view map of the current sample. init Perform feature extraction and processing to obtain the missing view feature map Z vis , the feature completion module is based on the missing view feature map Z vis To predict the missing view feature map Z vis The expanded feature map Z corresponding to the missing viewport pad At the same time, after the full-view feature extraction module extracts the features of the full-view image I of the current sample, the obtained full-view feature image Z is feature separated to separate the missing view image I init The real missing boundary feature map Z relative to the full view map I out ;

[0012] S32, Feature Merging:

[0013] For the missing view feature map Z vis And the expanded feature map Z pad Merge features to obtain the completed feature map;

[0014] S33, taking the completed feature map obtained in step S32 as the input of the box regression module, the box regression module predicts M groups of training stage prediction results according to the completed feature map, each group of training stage prediction results includes a position information c of a composition boundary box pred j And the confidence score p of the composition bounding box pred j ; Among them, c pred j is the position information of the bounding box in the j-th group of training stage prediction results, p pred j is the confidence score of the bounding box of the j-th group of training stage prediction results; j = 1, 2, ... M; M > N;

[0015] S4. Calculation of total loss value:

[0016] According to the real missing boundary feature map Z out And the predicted outward expansion feature map Z pad Perform comparative analysis to obtain the feature expansion loss function L extra The loss value is also used to compare and analyze the predicted results of the M groups of training phase with the N groups of calibration results to obtain the composition regression loss function L comp The loss value; and the feature expansion loss function L extra The loss value and the composition regression loss function L comp Sum the loss values ​​to get the total loss value;

[0017] Determine whether the total loss value is less than or equal to a preset threshold value. If yes, complete the training of the frame regression module, the missing view feature extraction module, and the feature completion module, and execute step S5. If no, update the parameters of the frame regression module, the missing view feature extraction module, and the feature completion module, assign the parameter value of the missing view feature extraction module to the full view feature extraction module, and retrain the frame regression module, and execute step S3.

[0018] S5, using the trained missing view feature extraction module, feature completion module and frame regression module to collect the missing view image I init The position information of the true composition bounding box is predicted to obtain M groups of true prediction results, and the position information of the composition bounding box corresponding to the maximum value of the confidence of the composition bounding box in the M groups of true prediction results is used as the position information of the optimal true composition bounding box.

[0019] Preferably, S31, obtaining the missing viewing angle feature map Z vis The implementation process includes:

[0020] S311, extracting the missing view image I of the current sample through the missing view feature extraction module init The image feature matrix h init ,h init =[C,H,W];

[0021] h init is a matrix composed of C, H and W, where H, W and C are image feature matrices h init The height, width and number of channels;

[0022] S312, the image feature matrix h init Flatten to the preprocessed matrix [L, d], encode the absolute position of the preprocessed matrix [L, d], and obtain the missing view feature map Z vis ; Wherein, L=HW, L represents the intermediate variable.

[0023] Preferably, the position information c of the cropping frame iand the position information c of the composition bounding box pred j There are four parameters in each, namely x, y, w and h: x and y represent the horizontal and vertical coordinates of the center point of the box, respectively, and w and h represent the width and height of the box, respectively.

[0024] Preferably, in step S4, the feature expansion loss function L is obtained extra The loss value is realized by:

[0025] L extra =smooth-l 1 (Z pad , sg(Z out ));

[0026] Among them, smooth-l 1 (·) indicates smooth-l 1 Loss function, sg(·) represents the operation of stopping gradient propagation.

[0027] Preferably, in step S4, the composition regression loss function L comp The loss value is realized by:

[0028] S41, the number of groups of calibration results in the current sample is expanded from N groups to M groups, and the position information of the cropping boxes in the N+1th to Mth groups is set to empty, and the confidence of the cropping boxes is set to 0;

[0029] S42, calculate the composition regression loss function L by the following formula comp The loss value is:

[0030]

[0031] Among them, L reg (·) represents the L1 loss function, λ IoU Represents the weighted coefficient of the L1 loss function, L IoU (·) represents the GIoU loss function, λ IoU Represents the weighted coefficient of the GIoU loss function, L focal (·) represents the focal loss function, λ focal Represents the weighting coefficient of the focal loss function.

[0032] Preferably, λ IoU =0.4,λ focal =0.1.

[0033] Preferably, the method of obtaining the cropping frame set of each sample in step S1 includes:

[0034] First, the full-view image I in each sample is cropped X times to obtain X cropping frames with different aspect ratios; the aspect ratio is the ratio of the width to the length of the cropping frame;

[0035] Secondly, the aesthetic quality score of the cropped image corresponding to each cropping frame is calculated using the aesthetic scoring guidance method, and the aesthetic quality score is used as the quality score of the cropping frame;

[0036] Finally, the quality scores of the X cropping frames are sorted from large to small, and the position information corresponding to the first N cropping frames in the sorting is used as the position information of the N cropping frames in the cropping frame collection of the current sample. At the same time, the confidence values ​​of the first N cropping frames in the sorting are all set to 1, and the confidences of the N cropping frames are used as the confidences of the N cropping frames in the cropping frame collection of the current sample.

[0037] Preferably, in S31, the feature completion module is based on the missing view feature map Z vis To predict the missing view feature map Z vis The expanded feature map Z corresponding to the missing viewport pad The implementation method is implemented using feature completion method.

[0038] Preferably, the composition bounding box recommendation method not restricted by the image boundary further includes step S6;

[0039] S6, for converting the aspect ratio of the optimal real composition bounding box into the aspect ratio of the camera according to the position information of the optimal real composition bounding box obtained in step S5.

[0040] Preferably, in step S6, the aspect ratio of the optimal real composition bounding box is converted into the aspect ratio of the camera by:

[0041] S61, the camera view v to be recommended pred The center position of shares the center position of the optimal true composition bounding box c1, specifically:

[0042] in, For the recommended camera view v pred The horizontal and vertical coordinates of the center position;

[0043] are the horizontal and vertical coordinates of the center position of the optimal true composition bounding box c1;

[0044] S62: Make the camera view v to be recommended pred has the same width and / or height as the optimal ground-truth bounding box c1, and or

[0045] Specifically:

[0046]

[0047] in, For the recommended camera view v pred The width of is the width of the optimal ground-truth composition bounding box c1;

[0048] For the recommended camera view v pred Height, is the height of the optimal ground-truth bounding box c1.

[0049] Principle analysis:

[0050] Each sample in the sample set of the present invention includes a full-view image I and any missing view image I obtained by cutting the full-view image I. init , randomly select a sample from the sample set to train the image feature extraction module, feature completion module and frame regression module. During the training process, the full view map I is used to predict the external feature map Z pad , through the missing perspective diagram I init Get the missing view feature map Z vis , the expanded feature map Z pad For the missing perspective I init The feature expansion of Z pad and Z vis After feature merging, it is sent to the box regression module for composition bounding box prediction and the feature expansion loss function L is calculated. extra And the composition regression loss function L comp The total loss value constituted is used to update the network parameters of the image feature extraction module, the feature completion module and the frame regression module, thereby completing the training of the image feature extraction module, the feature completion module and the frame regression module, and using the trained modules to obtain the missing view map of the real acquisition. init The position information of the real composition bounding box is predicted, and in specific applications, the missing view map I init As input to the box regression module, the box regression module outputs the location information of the predicted optimal ground-truth bounding box.

[0051] The beneficial effects brought by the present invention are:

[0052] The present invention provides a composition bounding box recommendation method that is not restricted by image boundaries, and for the first time realizes camera view angle and image composition recommendation that are not restricted by image boundaries. The present invention can help users easily take photos with higher aesthetic quality and provide shooting guidance. Users can adjust the current view based on the suggestions of the real composition bounding box predicted by the present invention to obtain pictures with higher aesthetic quality. Compared with the mainstream image cropping method in the prior art, the quantitative indicators of the present invention on the composition recommendation task that is not restricted by image boundaries have reached a better level. At the same time, the present invention also has good practicality and freedom, bringing new possibilities for the development of the field of image composition recommendation.

[0053] The present invention can freely go beyond the boundaries of the image, but prediction in invisible areas may lead to poor results. To solve this problem, the present invention chooses to expand in feature space and use the expanded content to predict camera movement and composition bounding boxes. Compared with expanding in image space, feature expansion avoids redundant information and computational burden, and can be well integrated into existing frameworks.

[0054] After the present invention predicts the real composition bounding box, the real composition bounding box provides the user with corresponding camera adjustment suggestions; however, adjusting the camera view alone is not enough to achieve the expected effect, because the aspect ratio of the predicted real composition bounding box is limited by the aspect ratio of the camera. Therefore, the present invention also converts the aspect ratio of the predicted real composition bounding box into the aspect ratio of the camera to obtain a better composition. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic diagram of the principle of the composition bounding box recommendation method not restricted by the image boundary of the present invention;

[0056] Figure 2 It is a schematic diagram of the principle of the cropping frame cropping the full-view image I in each sample. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0059] Embodiment 1:

[0060] See also Figure 1 In this embodiment of the specification, the composition bounding box recommendation method described in this embodiment is not restricted by the image boundary. The method is implemented based on the box regression module.

[0061] The method is implemented based on a frame regression module, a missing view feature extraction module, a full view feature extraction module and a feature completion module. The internal structures of the missing view feature extraction module and the full view feature extraction module are the same. The method is characterized in that it comprises the following steps:

[0062] S1. Construct sample set:

[0063] Each sample in the sample set includes a full-view image I and any missing view image I obtained by cropping the full-view image I. init , and missing perspective map I init The cropping frame collection corresponding to the full-view image I; the cropping frame collection includes N groups of calibration results, each group of calibration results includes the position information c of a cropping frame i And the confidence p of the crop box i ; Among them, c i is the position information of the cropping box in the i-th group calibration result in the cropping box collection, p i is the confidence of the cropping box in the i-th group of calibration results in the cropping box collection, i = 1, 2, ... N; p 1 to p N The value is 1, where each cropping frame is located in the area covered by the image content of the full view image I, and each cropping frame is regarded as the missing view image I. init The ideal composition bounding box;

[0064] S2, initializing the frame regression module, the missing view feature extraction module, the full view feature extraction module and the feature completion module, and the parameter values ​​of the missing view feature extraction module and the full view feature extraction module after initialization are exactly the same;

[0065] S3. Randomly extract a sample from the sample set and train the box regression module, the missing view feature extraction module and the feature completion module, including:

[0066] S31, feature extraction and completion:

[0067] A sample is randomly selected from the sample set, and the missing view feature extraction module is used to extract the missing view map of the current sample. init Perform feature extraction and processing to obtain the missing view feature map Z vis , the feature completion module is based on the missing view feature map Z vis To predict the missing view feature map Z vis The expanded feature map Z corresponding to the missing viewport padAt the same time, after the full-view feature extraction module extracts the features of the full-view image I of the current sample, the obtained full-view feature image Z is feature separated to separate the missing view image I init The real missing boundary feature map Z relative to the full view map I out ;

[0068] S32, Feature Merging:

[0069] For the missing view feature map Z vis And the expanded feature map Z pad Merge features to obtain the completed feature map;

[0070] S33, taking the completed feature map obtained in step S32 as the input of the box regression module, the box regression module predicts M groups of training stage prediction results according to the completed feature map, each group of training stage prediction results includes a position information c of a composition boundary box pred j And the confidence score p of the composition bounding box pred j ; Among them, c pred j is the position information of the bounding box in the j-th group of training stage prediction results, p pred j is the confidence score of the bounding box of the j-th group of training stage prediction results; j = 1, 2, ... M; M > N;

[0071] S4. Calculation of total loss value:

[0072] According to the real missing boundary feature map Z out And the predicted outward expansion feature map Z pad Perform comparative analysis to obtain the feature expansion loss function L extra The loss value is also used to compare and analyze the predicted results of the M groups of training phase with the N groups of calibration results to obtain the composition regression loss function L comp The loss value; and the feature expansion loss function L extra The loss value and the composition regression loss function L comp Sum the loss values ​​to get the total loss value;

[0073] Determine whether the total loss value is less than or equal to a preset threshold value. If yes, complete the training of the frame regression module, the missing view feature extraction module, and the feature completion module, and execute step S5. If no, update the parameters of the frame regression module, the missing view feature extraction module, and the feature completion module, assign the parameter value of the missing view feature extraction module to the full view feature extraction module, and retrain the frame regression module, and execute step S3.

[0074] S5, using the trained missing view feature extraction module, feature completion module and frame regression module to collect the missing view image I init The position information of the true composition bounding box is predicted to obtain M groups of true prediction results, and the position information of the composition bounding box corresponding to the maximum value of the confidence of the composition bounding box in the M groups of true prediction results is used as the position information of the optimal true composition bounding box.

[0075] In this embodiment, when it is specifically applied, in step S5, the missing view feature extraction module, feature completion module and frame regression module after training are used to extract the missing view image I of the real collection. init The process of predicting the position information of the real composition boundary box and obtaining M sets of real prediction results is the same. The present invention uses the box regression module to predict not only the missing viewpoint image I init In order to alleviate the difficulty of predicting the composition bounding box outside the image boundary, the feature space is expanded to predict the expanded feature map Z pad , where the external feature map Z pad It is supervised by the complete features extracted from the larger full view image I, and the view image I is missing in each sample init It is obtained by cropping from the corresponding full-view image I. Finally, the recommended composition bounding box can be obtained through the predicted position information of the composition bounding box, so the entire framework can achieve image composition recommendation without being restricted by the image boundary.

[0076] The full perspective image I (initial missing perspective image I init After feature extraction, the full-view feature map Z is obtained. Then the feature separation of the obtained full-view feature map Z is performed, and Z is divided into two parts, namely, init Z in range in and I init Z outside the range out ; Z out For the missing perspective I init The present invention separates the missing view image I from the full view image I. init The real missing boundary feature map Z relative to the full view map I out , providing supervision for feature expansion.

[0077] The present invention can freely go beyond the boundaries of the image. In the process of predicting in the invisible area, the present invention chooses to expand in the feature space and use the expanded content to predict the camera movement and the composition boundary box. Compared with expanding in the image space, feature expansion avoids redundant information and computational burden, avoids possible poor results, and can be well integrated into the existing framework.

[0078] The present invention uses the constructed samples to train an image feature extraction module, a feature completion module and a frame regression module. The frame regression module is used to predict the composition bounding box of the received image features. In the prior art, the frame regression module can be implemented by various networks as long as it can realize the prediction of the composition bounding box. Specifically, the frame regression module can be composed of six existing transformer blocks in cascade, each transformer block includes a multi-head self-attention network, a multi-head cross-attention network and a feedforward network, each transformer block uses the multi-head self-attention mechanism of the multi-head self-attention network to remove redundant frames, and uses the cross-attention mechanism of the multi-head cross-attention network to fuse image information to predict the boundary box, and then outputs it to the next transformer block through the feedforward network.

[0079] When applied, updating the network parameters of the image feature extraction module, the feature completion module and the frame regression module can be achieved through existing technical means.

[0080] Further, S31, obtain the missing view feature map Z vis The implementation process includes:

[0081] S311, extracting the missing view image I of the current sample through the missing view feature extraction module init The image feature matrix h init ,h init =[C,H,W];

[0082] h init is a matrix composed of C, H and W, where H, W and C are image feature matrices h init The height, width and number of channels;

[0083] S312, the image feature matrix h init Flatten to the preprocessed matrix [L, d], encode the absolute position of the preprocessed matrix [L, d], and obtain the missing view feature map Z vis ; Wherein, L=HW, L represents the intermediate variable.

[0084] In this preferred embodiment, by init Perform feature extraction and processing to obtain a more accurate missing view feature map Z vis , providing an accurate data basis for subsequent training.

[0085] Furthermore, the position information of the cropping frame c i and the position information c of the composition bounding box pred jThere are four parameters in each of them, namely x, y, w and h: x and y represent the horizontal and vertical coordinates of the center point of the frame, respectively, and w and h represent the width and height of the frame, respectively. pred j The four parameters in the figure show that the bounding box of the composition is in the missing perspective diagram I init relative position.

[0086] Furthermore, in step S4, the feature expansion loss function L is obtained extra The loss value is realized by:

[0087] L extra =smooth-l 1 (Z pad , sg(Z out ));

[0088] Among them, smooth-l 1 (·) indicates smooth-l 1 Loss function, sg(·) represents the operation of stopping gradient propagation.

[0089] In this preferred embodiment, the feature expansion loss function L extra For the real missing boundary feature map Z out And the expanded feature map Z pad To constrain.

[0090] Furthermore, in step S4, the composition regression loss function L comp The loss value is realized by:

[0091] S41, the number of groups of calibration results in the current sample is expanded from N groups to M groups, and the position information of the cropping boxes in the N+1th to Mth groups is set to empty, and the confidence of the cropping boxes is set to 0;

[0092] S42, calculate the composition regression loss function L by the following formula comp The loss value is:

[0093]

[0094] Among them, L reg (·) represents the L1 loss function, λ IoU Represents the weighted coefficient of the L1 loss function, L IoU (·) represents the GIoU loss function, λ IoU Represents the weighted coefficient of the GIoU loss function, L focal (·) represents the focal loss function, λ focal Represents the weighting coefficient of the focal loss function.

[0095] In this preferred embodiment, the L1 loss function and the GIoU loss function are specifically used to constrain the bounding box, and a focal loss function is used to constrain the confidence.

[0096] Furthermore, when applied, λ IoU =0.4,λ focal =0.1.

[0097] For further details, see Figure 2 , the method of obtaining the cropping frame set of each sample in step S1 includes:

[0098] First, use X cropping frames with different aspect ratios to crop the full view image I in each sample to obtain X missing view images I with different aspect ratios. init ; Wherein, the aspect ratio is the ratio of the width to the length of the cropping box;

[0099] Secondly, the aesthetic scoring guidance method is used to calculate each missing perspective map I init The quality score of the corresponding cropping frame, and the quality score of the cropping frame;

[0100] Finally, the quality scores of the X cropping frames are sorted from large to small, and the position information corresponding to the first N cropping frames in the sorting is used as the position information of the N cropping frames in the cropping frame collection of the current sample. At the same time, the confidence values ​​of the first N cropping frames in the sorting are all set to 1, and the confidences of the N cropping frames are used as the confidences of the N cropping frames in the cropping frame collection of the current sample.

[0101] In this preferred embodiment, based on the existing image cropping dataset, an unconstrained image composition dataset is recreated, that is, the sample set of the present invention. For a sample in the sample set, a full view image I and a missing view image I init Will be provided, missing perspective map I init The annotations corresponding to multiple ideal cropping boxes are used as the target domain, and the annotations include the position information and confidence of the cropping boxes. init , then the cropping box in the target domain may not be completely located in the currently selected missing view image I init within the image range.

[0102] Figure 2 The confidence of the cropping frame corresponding to the cropping of the full-view image I is given in , taking two cropping frames as an example. Figure 2 The quality scores of the two cropped boxes on the full view image I are 4.5 and 4.2 respectively, while the quality scores of the missing view image I are init The scores of the two crop boxes on are both 1, representing the confidence of each crop box.

[0103] Furthermore, in S31, the feature completion module is based on the missing view feature map Z vis To predict the missing view feature map Z vis The expanded feature map Z corresponding to the missing viewport pad The implementation method is implemented using feature completion method.

[0104] Furthermore, the composition bounding box recommendation method not restricted by image boundaries further includes step S6;

[0105] S6, for converting the aspect ratio of the optimal real composition bounding box into the aspect ratio of the camera according to the position information of the optimal real composition bounding box obtained in step S5.

[0106] Furthermore, in step S6, the aspect ratio of the optimal real composition bounding box is converted into the aspect ratio of the camera by:

[0107] S61, the camera view v to be recommended pred The center position of shares the center position of the optimal true composition bounding box c1, specifically:

[0108] in, For the recommended camera view v pred The horizontal and vertical coordinates of the center position;

[0109] are the horizontal and vertical coordinates of the center position of the optimal true composition bounding box c1;

[0110] S62: Make the camera view v to be recommended pred has the same width and / or height as the optimal ground-truth bounding box c1, and or

[0111] Specifically:

[0112]

[0113] in, For the recommended camera view v pred The width of is the width of the optimal ground-truth composition bounding box c1;

[0114] For the recommended camera view v pred Height, is the height of the optimal ground-truth bounding box c1.

[0115] In this preferred embodiment, the aspect ratio of the real composition bounding box predicted by the trained frame regression module of the present invention may be different from the aspect ratio of the camera. Therefore, the aspect ratio of the obtained real composition bounding box is converted into the aspect ratio of the camera. The adjustment process needs to satisfy the requirement that the camera view v to be recommended is pred The width and / or height are the same as those of the actual composition bounding box c1, and the aspect ratio of the camera is 3:4 or 4:3, so that the adjusted composition bounding box is obtained to provide camera adjustment suggestions.

[0116] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. It should therefore be understood that many modifications may be made to the exemplary embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in a manner different from that described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be used in other described embodiments.

Claims

1. A composition bounding box recommendation method that is not restricted by image boundaries. This method is based on a box regression module, a missing view feature extraction module, a full view feature extraction module, and a feature completion module. The internal structures of the missing view feature extraction module and the full view feature extraction module are the same. It is characterized in that The method comprises the following steps: S1. Construct sample set: Each sample in the sample set includes a full-view image I and any missing view image I obtained by cropping the full-view image I. init , and a cropping frame collection obtained by cropping the full-view image I; the cropping frame collection includes N groups of calibration results, each group of calibration results includes a cropping frame position information c i And the confidence p of the crop box i ; Among them, c i is the position information of the cropping box in the i-th group calibration result in the cropping box collection, p i is the confidence of the cropping box in the i-th group of calibration results in the cropping box collection, i = 1, 2, ... N; p 1 to p N The values ​​are all 1, where each cropping frame is located in the area covered by the image content of the full-view image I, and each cropping frame is a missing view image I init The ideal composition bounding box; S2, initializing the frame regression module, the missing view feature extraction module, the full view feature extraction module and the feature completion module, and the parameter values ​​of the missing view feature extraction module and the full view feature extraction module after initialization are exactly the same; S3. Randomly extract a sample from the sample set and train the box regression module, the missing view feature extraction module and the feature completion module, including: S31, feature extraction and completion: A sample is randomly selected from the sample set, and the missing view feature extraction module is used to extract the missing view map of the current sample. init Perform feature extraction and processing to obtain the missing view feature map Z vis , the feature completion module is based on the missing view feature map Z vis To predict the missing view feature map Z vis The expanded feature map Z corresponding to the missing viewport pad At the same time, after the full-view feature extraction module extracts the features of the full-view image I of the current sample, the obtained full-view feature image Z is feature separated to separate the missing view image I init The real missing boundary feature map Z relative to the full view map I out ; S32, Feature Merging: For the missing view feature map Z vis And the expanded feature map Z pad Merge features to obtain the completed feature map; S33, taking the completed feature map obtained in step S32 as the input of the box regression module, the box regression module predicts M groups of training stage prediction results according to the completed feature map, each group of training stage prediction results includes a position information c of a composition boundary box pred j And the confidence score p of the composition bounding box pred j ; Among them, c pred j is the position information of the bounding box of the composition in the prediction result of the jth training stage, p pred j is the confidence score of the bounding box of the j-th group of training stage prediction results; j = 1, 2, ... M; M > N; S4. Calculation of total loss value: According to the real missing boundary feature map Z out And the predicted outward expansion feature map Z pad Perform comparative analysis to obtain the feature expansion loss function L extra The loss value is also used to compare and analyze the predicted results of the M groups of training phase with the N groups of calibration results to obtain the composition regression loss function L comp The loss value; and the feature expansion loss function L extra The loss value and the composition regression loss function L comp Sum the loss values ​​to get the total loss value; Determine whether the total loss value is less than or equal to a preset threshold value. If yes, complete the training of the frame regression module, the missing view feature extraction module, and the feature completion module, and execute step S5. If no, update the parameters of the frame regression module, the missing view feature extraction module, and the feature completion module, assign the parameter value of the missing view feature extraction module to the full view feature extraction module, and retrain the frame regression module, and execute step S3. S5, using the trained missing view feature extraction module, feature completion module and frame regression module to collect the missing view image I init The position information of the true composition bounding box is predicted to obtain M groups of true prediction results, and the position information of the composition bounding box corresponding to the maximum value of the confidence of the composition bounding box in the M groups of true prediction results is used as the position information of the optimal true composition bounding box.

2. The composition bounding box recommendation method not restricted by image boundaries according to claim 1, It is characterized in that S31, obtain the missing view feature map Z vis The implementation process includes: S311, extracting the missing view image I of the current sample through the missing view feature extraction module init The image feature matrix h init ,h init =[C,H,W]; h init is a matrix composed of C, H and W, where H, W and C are image feature matrices h init The height, width and number of channels; S312, the image feature matrix h init Flatten to the preprocessed matrix [L, d], encode the absolute position of the preprocessed matrix [L, d], and obtain the missing view feature map Z vis ; Wherein, L=HW, L represents the intermediate variable.

3. The composition bounding box recommendation method not restricted by image boundaries according to claim 1, It is characterized in that The position information of the cropping frame c i and the position information c of the composition bounding box pred j There are four parameters in each, namely x, y, w and h: x and y represent the horizontal and vertical coordinates of the center point of the box, respectively, and w and h represent the width and height of the box, respectively.

4. The method for recommending a composition bounding box without being restricted by image boundaries according to claim 1, It is characterized in that In step S4, the feature expansion loss function L is obtained extra The loss value is realized by: in, express Loss function, sg(·) represents the operation of stopping gradient propagation.

5. The composition bounding box recommendation method not restricted by image boundaries according to claim 1, It is characterized in that In step S4, the composition regression loss function L comp The loss value is realized by: S41, the number of groups of calibration results in the current sample is expanded from N groups to M groups, and the position information of the cropping boxes in the N+1th to Mth groups is set to empty, and the confidence of the cropping boxes is set to 0; S42, calculate the composition regression loss function L by the following formula comp The loss value is: Among them, L reg (·) represents the L1 loss function, λ IoU Represents the weighted coefficient of the L1 loss function, L IoU (·) represents the GIoU loss function, λ IoU Represents the weighted coefficient of the GIoU loss function, L focal (·) represents the focal loss function, λ focal Represents the weighting coefficient of the focal loss function.

6. The composition bounding box recommendation method not restricted by image boundaries according to claim 5, It is characterized in that l IoU =0.4,λ focal = 0.

1.

7. The method for recommending a composition bounding box without being restricted by image boundaries according to claim 1, It is characterized in that The method of obtaining the cropping frame set of each sample in step S1 includes: First, the full-view image I in each sample is cropped X times to obtain X cropping frames with different aspect ratios; the aspect ratio is the ratio of the width to the length of the cropping frame; Secondly, the aesthetic quality score of the cropped image corresponding to each cropping frame is calculated using the aesthetic scoring guidance method, and the aesthetic quality score is used as the quality score of the cropping frame; Finally, the quality scores of the X cropping frames are sorted from large to small, and the position information corresponding to the first N cropping frames in the sorting is used as the position information of the N cropping frames in the cropping frame collection of the current sample. At the same time, the confidence values ​​of the first N cropping frames in the sorting are all set to 1, and the confidences of the N cropping frames are used as the confidences of the N cropping frames in the cropping frame collection of the current sample.

8. The method for recommending a composition bounding box without being restricted by image boundaries according to claim 1, It is characterized in that In S31, the feature completion module is based on the missing view feature map Z vis To predict the missing view feature map Z vis The expanded feature map Z corresponding to the missing viewport pad The implementation method is implemented using feature completion method.

9. The composition bounding box recommendation method not restricted by image boundaries according to claim 3, It is characterized in that Also includes step S6; S6, for converting the aspect ratio of the optimal real composition bounding box into the aspect ratio of the camera according to the position information of the optimal real composition bounding box obtained in step S5.

10. The method for recommending a composition bounding box without being restricted by image boundaries according to claim 1, It is characterized in that In step S6, the aspect ratio of the optimal real composition bounding box is converted into the aspect ratio of the camera by: S61, the camera view v to be recommended pred The center position of shares the center position of the optimal true composition bounding box c1, specifically: in, For the recommended camera view v pred The horizontal and vertical coordinates of the center position; are the horizontal and vertical coordinates of the center position of the optimal true composition bounding box c1; S62: Make the camera view v to be recommended pred has the same width and / or height as the optimal ground-truth bounding box c1, and or Specifically: in, For the recommended camera view v pred The width of is the width of the optimal ground-truth composition bounding box c1; For the recommended camera view v pred Height, is the height of the optimal ground-truth bounding box c1.

Citation Information

Patent Citations

  • Image completing method and device, computer equipment and storage medium

    CN108765315A

  • Image restoration method based on multi-scale discriminant generative adversarial network

    CN112270651A