Intelligent image transformation method and device and storage medium

Through object detection and automated cropping and adjustment technology, the problem of existing image transformation methods relying on manual operations is solved, and efficient and accurate image transformation is achieved to adapt to different application scenarios.

CN119941490APending Publication Date: 2025-05-06BEIJING THUNDERSTONE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510000965.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing image transformation methods rely on manual operations, which are prone to errors, are inefficient, and cannot adapt to different application scenarios and output requirements. It is difficult to accurately identify and process subject objects in the image when dealing with complex scenarios.

Method used

Through object detection, the main objects in the image are identified, bounding boxes are generated, and the images are cropped and adaptively adjusted to achieve automated image transformation.

Benefits of technology

Reduce manual intervention through automation, ensure consistency and accuracy of cropping and adjustment, improve processing efficiency, and optimize image composition to preserve more details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941490A_ABST
    Figure CN119941490A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an image intelligent transformation method, which comprises the steps of identifying at least one main object in a to-be-processed image through target detection; generating a bounding box of at least one main object in the to-be-processed image; according to the bounding box of the at least one main object, cutting the to-be-processed image to reserve the at least one main object; performing self-adaptive proportion adjustment on the image in which the at least one main object is reserved; and outputting the image after proportion adjustment. According to the technical scheme, the image can be efficiently and intelligently transformed, and more details can be reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image intelligent transformation method, device and storage medium. Background Art

[0002] After people use image acquisition devices (such as cameras) to capture images, in order to meet the needs of specific application scenarios (for example, ID photos), they usually need to perform some transformations on the images, such as cropping and scaling. However, most existing image transformation methods rely on manual operations, and users need to manually select the image area to be retained and adjust the size. This operation is prone to errors, is inefficient, and cannot adapt to different application scenarios and output requirements. In particular, it is difficult to accurately identify and process the main objects in the image when processing complex scenes. Summary of the invention

[0003] The present application provides an image intelligent transformation method, device and storage medium, which can efficiently perform intelligent transformation on images and retain more details.

[0004] On the one hand, the present application provides an image intelligent transformation method, the method comprising:

[0005] Identifying at least one main object in the image to be processed by object detection;

[0006] generating a bounding box of the at least one primary object;

[0007] According to the bounding box of the at least one main object, cropping the image to be processed to retain the at least one main object;

[0008] Adaptively scaling the image retaining the at least one main object;

[0009] The image after the scale adjustment is output.

[0010] On the other hand, the present application provides an image intelligent transformation device, the device comprising:

[0011] A detection module, used for identifying at least one main object in the image to be processed through object detection;

[0012] A generating module, configured to generate a bounding box of the at least one main object;

[0013] A cropping module, configured to crop the image to be processed according to a boundary box of the at least one main object to retain the at least one main object;

[0014] An adjustment module, used for adaptively adjusting the proportion of the image retaining the at least one main object;

[0015] An output module is used to output the image after the scale is adjusted.

[0016] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the technical solution of the above-mentioned image intelligent transformation method when executing the computer program.

[0017] In a fourth aspect, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the technical solution of the above-mentioned image intelligent transformation method.

[0018] From the technical solution provided by the present application, it can be seen that, on the one hand, the technical solution of the present application is automatically completed by a computer program according to preset rules or intelligent algorithms. Therefore, automation reduces manual intervention, ensures the consistency and accuracy of cropping and adjustment, and improves processing efficiency; on the other hand, after cropping the image to be processed to retain at least one main object, the image retaining at least one main object is adaptively proportionally adjusted. Since this technical means fully considers the content and structure of the image, it not only optimizes the composition of the image to make it more in line with aesthetic requirements, but also ensures that the main object is not over-stretched or compressed, thereby retaining more details. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 is a flow chart of the image intelligent transformation method provided by an embodiment of the present application;

[0021] Figure 2 is a structural schematic diagram of an image intelligent transformation device provided in an embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] In this specification, adjectives such as first and second may be used only to distinguish one element or action from another element or action, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be interpreted as being limited to only one of the elements, components, or steps, but may be one or more of the elements, components, or steps, etc.

[0025] In this specification, for the convenience of description, the sizes of various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0026] After people use image acquisition devices (such as cameras) to capture images, in order to meet the needs of specific application scenarios (for example, ID photos), they usually need to perform some transformations on the images, such as cropping and scaling. However, most existing image transformation methods rely on manual operations, and users need to manually select the image area to be retained and adjust the size. This operation is prone to errors, is inefficient, and cannot adapt to different application scenarios and output requirements. In particular, it is difficult to accurately identify and process the main objects in the image when processing complex scenes.

[0027] In view of the above problems in the prior art, this application proposes an image intelligent transformation method, the flow chart of which is shown in the attached figure. Figure 1 As shown, it mainly includes steps S101 to S105, which are described in detail as follows:

[0028] Step S101: identifying at least one main object in the image to be processed through object detection.

[0029] There may be more than one main object (for example, a face or an object) in the image to be processed, and it may contain multiple important targets, especially multiple objects or characters in complex scenes. In an embodiment of the present application, a target detection algorithm can be used to identify at least one main object in the image to be processed. As an embodiment of the present application, identifying at least one main object in the image to be processed through target detection can be based on an edge detection algorithm to identify at least one main object in the image to be processed. Specifically, an edge detection algorithm (such as a Sobel operator or a Canny edge detection) can be used to extract edge information in the image to be processed. These edge information constitute the boundary between the main object and the background. In this way, the main object in the image to be processed can be identified. In the above embodiment, if E(x, y) represents the edge strength of the image to be processed I(x, y) at the pixel (x, y), then The edge strength of each point of the image to be processed is calculated to obtain the boundary between the main object and the background in the image to be processed, where x and y represent the coordinates of the pixel.

[0030] Considering that the area with large color difference in the image usually represents the boundary between the main object and the background, therefore, as another embodiment of the present application, identifying at least one main object in the image to be processed through target detection can be based on color contrast to identify at least one main object in the image to be processed. In image segmentation, the boundary between the background and the main object is usually not the average value of the color difference between any two points, but the boundary is found based on the color difference of the local area. Generally speaking, the boundary (i.e., the boundary between the background and the main object) is often where the color or texture changes suddenly, and these changes occur between adjacent pixels. Therefore, it can be detected by Calculate the color difference ΔC in the image to be processed to measure the boundary between the main object and the background, so as to identify at least one main object in the image to be processed, where C i is the color value of a pixel in the local area marked as i in the image to be processed, C j is the color value of a pixel in the neighborhood of the local area marked as i in the image to be processed, and N is the number of sample points, i.e., the number of neighborhoods of a local area or the number of adjacent pixel pairs; if the value of ΔC exceeds the preset threshold, the line connecting these N adjacent pixel pairs constitutes the boundary between the main object and the background.

[0031] Step S102: Generate a bounding box of at least one main object in the image to be processed.

[0032] In the field of image processing, the bounding box of the main object is the smallest closed area surrounding these objects. If there is only one main object in the image to be processed, it is generally simple to process. If there is more than one main object in the image to be processed, an extreme value method needs to be used for processing. In other words, generating the bounding box of at least one main object in the image to be processed needs to be processed according to different situations. Specifically, as an embodiment of the present application, generating the bounding box of at least one main object in the image to be processed can be achieved by steps S1021 to S1023, as described in detail as follows:

[0033] Step S1021: If the image to be processed has only one main object, a minimum rectangular frame surrounding the main object is generated.

[0034] If (x min,i ,y min,i ) and (x max,i ,y max,i ) represent the coordinates of the upper left corner and lower right corner of the smallest rectangular box surrounding the i-th main object, which means that (x min,i ,y min,i ) and (x max,i ,y max,i ) are the minimum coordinates of the upper left corner and the lower right corner of all rectangular boxes surrounding the main object.

[0035] Step S1022: If the image to be processed contains multiple main objects, a minimum rectangular frame of each of the multiple main objects is generated.

[0036] Here, the method of generating the minimum rectangular frame of each of the plurality of main objects is the same as the method of generating the minimum rectangular frame surrounding one main object in the aforementioned step S1021.

[0037] Step S1023: generating a minimum rectangular frame surrounding the plurality of main objects according to the minimum rectangular frame of each of the plurality of main objects.

[0038] First, the geometric center of each main object is calculated. In the embodiment of the present application, Represents the coordinates of the geometric center of the ith principal object, where:

[0039]

[0040] Secondly, from all and Select the minimum and maximum values ​​from:

[0041]

[0042] Where n is the number of main objects, the obtained (xmin ,y min ,x max ,y max ) are the upper left corner coordinates and lower right corner coordinates of the smallest rectangular box that encloses multiple main objects.

[0043] Step S103: cropping the image to be processed according to the boundary box of at least one main object to retain at least one main object.

[0044] As an embodiment of the present application, according to the bounding box of at least one main object, cropping the image to be processed to retain at least one main object may be implemented by steps S1031 to S1033, which are described in detail as follows:

[0045] Step S1031: Generate a cropping box that encloses a bounding box of at least one main object.

[0046] If there is only one main object in the image to be processed (denoted as the i-th main object), the cropping box of the bounding box surrounding at least one main object is the one surrounding the coordinates (x min,i ,y min,i ) and (x max,i ,y max,i If there is more than one main object in the image to be processed, the cropping frame of the bounding box surrounding at least one main object is the rectangular frame of ... min ,y min ) and (x max ,y max )'s rectangular border.

[0047] Step S1032: Determine the golden section line of the image to be processed according to the width and height of the image to be processed.

[0048] In an embodiment of the present application, the strategy of cropping the image to be processed is not only about how to crop multiple main objects, but also includes how to optimize the cropping content so that the visual effect of the image to be processed after cropping is more natural and beautiful. This means that the cropping area should be adjusted according to the content of the image to be processed, the composition rules (such as the golden section, symmetry, etc.) and the spatial relationship between the main objects. For example, if the character in the image to be processed is located on one side of the image to be processed, a cropping frame can be selected to retain the integrity of the character as much as possible according to the relative position of the character, and the position and size of the cropping frame can be adjusted according to aesthetic rules (such as centering, golden section, etc.). For multiple main objects, when cropping, it is not only concerned with whether to cut off the main object, but also the overall composition of the image to be processed will be considered to ensure the visual balance of the image to be processed after cropping.

[0049] In the embodiment of the present application, the golden section line of the image to be processed may include a horizontal golden section line and a vertical golden section line. Specifically, assuming that the width of the image to be processed is W and the height is H, the position x of the horizontal golden section line is golden For x golden =W / φ, that is, the horizontal golden section line is located at φ of the width of the image to be processed and parallel to the horizontal axis of the image coordinate system, and the position of the vertical golden section line is y golden for y golden =H / φ, that is, the vertical golden section line is located at φ of the height of the image to be processed and is parallel to the longitudinal axis of the image coordinate system, where, In this way, the golden section line can divide the image to be processed into two areas, with a ratio of and

[0050] Step S1033: Align the cropping frame with the golden section line of the image to be processed or move at least one main object to a position whose distance from the golden section line of the image to be processed does not exceed a preset threshold, and then use the cropping frame to crop the image to be processed.

[0051] By adjusting the position of the cropping frame to align it with the golden section line, or making the main object (such as a face, object, etc.) as close to the golden section line as possible, that is, the distance from the golden section line of the image to be processed does not exceed the preset threshold, the optimization formula is as follows:

[0052] |C cx -x golden |,|C yx -y golden |

[0053] Among them, C cx and C cy is the center coordinate of one of the at least one main object. By optimizing and adjusting the position of the cropping frame to conform to the golden ratio, the image to be processed can be made more visually attractive.

[0054] It should be noted that it is a common operation to remove unnecessary background in the process of cropping the image to be processed. However, in some cases, simple background removal may cause the image content to be unnatural, for example, the edge between the object and the background appears stiff, the abrupt cropping effect, and the loss of space and depth, etc. It can be considered to slightly blur the background while retaining the main object, or to use the background blurring technology to achieve an aesthetic balance. Therefore, the above embodiment may also include:

[0055] Step S'1031: extracting the background of the image to be processed by using the segmentation mask of the image to be processed.

[0056] A segmentation mask is a binary image, usually used to indicate which areas in an image belong to a specific category (such as foreground or background). It distinguishes different areas by marking them at the pixel level. Specifically, each pixel value in the segmentation mask represents which category the pixel belongs to. The segmentation mask includes a foreground mask and a background mask. fg (x,y), if the pixel value M fg (x, y) = 1, it means that the pixel belongs to the foreground area (such as the main object in the image). fg (x, y) = 0, it means that the pixel belongs to the background area; the background mask M bg (x, y) is the complement of the foreground mask. If the pixel value M bg (x, y) = 1, it means that the pixel belongs to the background area. Conversely, if the pixel value M bg (x, y) = 0, it means that the pixel belongs to the foreground area. In the segmentation task in image processing, the segmentation mask, as a marking tool, can help distinguish different areas in the image. The foreground mask and background mask provide pixel-level classification information, which helps to accurately extract or separate the foreground and background. fg (x,y) and background mask M bg (x, y), the steps to extract the foreground of the image to be processed are as follows: for any pixel p in the image to be processed i (x,y), examine its foreground mask value M fg (x,y); if M fg (x,y)=1, then keep the pixel p i (x, y) (i.e. the pixel belongs to the foreground); if M fg (x, y) = 0, then set the pixel p i (x, y) is transparent or the pixel is removed (i.e. the pixel belongs to the background). The steps to extract the background of the image to be processed are as follows: j (x,y), examine its background mask value M bg (x,y); if M bg (x,y)=1, then keep the pixel p j (x, y) (i.e. the pixel belongs to the background); if M bg (x,y)=0, then the pixel p i (x,y) belongs to the foreground and should not be kept.

[0057] Step S'1032: Calculate the weighted average value of each pixel in the background and the pixels in its neighborhood through a convolution operation to obtain a background blurred image, wherein the weight of the weighted average value is determined by a Gaussian function.

[0058] In the embodiment of the present application, Gaussian blur can be used to obtain a background blurred image. Gaussian blur is a classic image blurring technique that calculates the value of each pixel in the image and the weighted average of the pixels in its neighborhood through a convolution operation, where the weight is given by a Gaussian function. The Gaussian function G(x, y, σ) can be expressed as:

[0059]

[0060] Where (x, y) is the coordinate of the image pixel and σ is the standard deviation of the Gaussian blur, which controls the intensity of the blur.

[0061] By bg (x,y) is Gaussian blurred to get a background blurred image

[0062]

[0063] The operator * indicates that I bg (x,y) and G(x,y,σ) perform convolution operation.

[0064] Step S'1033: Smoothing the background blurred image by bilateral filtering.

[0065] Step S104: adaptively adjust the scale of the image retaining at least one main object.

[0066] After the image to be processed is cropped, at least one main object is retained. The cropped image generally needs to be proportionally adjusted so that the main object is more coordinated in the entire image. If traditional proportional adjustment is used, for example, the entire image is scaled at the same ratio, it may cause image distortion, especially when the aspect ratio changes significantly. In order to prevent the problems caused by the above-mentioned traditional proportional adjustment, a more sophisticated adaptive algorithm can be introduced to automatically select the most appropriate scaling ratio and interpolation method to minimize distortion. Specifically, adaptive proportional adjustment of an image that retains at least one main object can be achieved through steps S1041 to S1043, as described in detail as follows:

[0067] Step S1041: Calculate the importance of each main object in at least one main object.

[0068] In one embodiment of the present application, one solution for calculating the importance of the main object is to use the area occupied by it in the image, the clarity of the outline, etc., specifically, if the AO i represents the area occupied by the ith main object, TIA represents the total area of ​​the image, and CF i represents the clarity evaluation index based on the edge or texture of the object, then the importance of the i-th main object Wi The calculation formula is as follows:

[0069]

[0070] In another embodiment of the present application, one solution for calculating the importance of the main object is to obtain the energy function E(x, y) of the main object at the position (x, y) of the image I(x, y), specifically:

[0071] E(x,y)=|▽I(x,y)| 2

[0072] Among them, ▽I(x,y) represents the gradient of the image I(x,y) at the position (x,y). Usually, the areas with obvious edges in the image have higher energy values, indicating that these areas are crucial to the image content.

[0073] Step S1042: assigning a corresponding zoom priority to each main object according to the importance of each main object.

[0074] In one embodiment, for the importance W i For main objects above a certain threshold, a higher scaling priority, i.e. a smaller scaling ratio, will be assigned. i Primary objects below a certain threshold are assigned a lower scaling priority, i.e. a larger scaling factor.

[0075] In another embodiment, once the energy function E(x, y) is calculated, the regions with lower energy (usually background regions) may be compressed preferentially during the scaling process, while the regions with higher energy (usually objects, edges, and regions with stronger textures) may be retained. The specific steps are as follows:

[0076] S1: Calculate the scaling ratio, that is, calculate the size of the main object, that is, the width and height.

[0077] S2: Energy-based scaling selection:

[0078] During the scaling process, the lower energy region is selected for compression according to the distribution of E(x,y), and the formula is:

[0079]

[0080] Where W and H are the width and height of the image, respectively, r x and r y are the local scaling ratios in the x and y directions, ∑ (x',y') E(x',y') is the sum of the energy of all pixels in the image. x and r yFrom the expression of , we can see that higher energy areas (e.g., objects, edges, etc.) will have a smaller scaling factor, while lower energy areas (such as flat backgrounds) will have a larger scaling factor.

[0081] Step S1043: resizing the image that has retained at least one main object using a scaling ratio corresponding to the scaling priority.

[0082] Although adaptive scaling of an image that retains at least one main object can intelligently select areas for scaling. However, when the background or low-importance areas are overstretched, blank or unnatural areas may appear, especially when there is less background information. Although the adaptive scaling scheme tries to retain important areas, on the one hand, some areas may still be distorted when the image content is complex or the scaling ratio is extremely large; on the other hand, when the size of the image changes greatly, content-aware scaling may not be handled perfectly, resulting in distortion of the image structure or artifacts. In order to overcome the defects of the above-mentioned adaptive scaling scheme, the above-mentioned embodiment also includes steps S'1041 to S'1043, which are described in detail as follows:

[0083] Step S'1041: Filling the background of the image retaining at least one main object based on the image texture information and color transition.

[0084] In the field of image processing, texture information generally refers to the local structural features of an area in an image, including surface details, patterns, repeating patterns, etc. Color transition refers to the smooth transition between different colors in an image. For example, in a natural scene, the color of the sky usually transitions from blue to white or yellow, and this change is manifested as a gradual change of color in the image. Color transition helps to merge different areas in the image, making the image look more natural. In an embodiment of the present application, the pixel values ​​of the background area are inferred by optimizing the loss function, so as to fill the background of the image retaining at least one main object. The loss function may include a data term (also referred to as data consistency, which is used to ensure that the pixels of the filled background area are consistent with the surrounding area in color and texture), a smoothing term (used to ensure that the color transition of the filled background area is smooth to avoid obvious color boundaries) and an edge retention term (used to ensure that the edge information in the image (such as the outline of an object) is maintained in the filled background area to avoid excessive blurring), which corresponds to the data consistency loss function, the smoothness loss function, and the edge retention loss function in turn. The process of filling the background is to continuously optimize these three loss functions until their values ​​are minimized.

[0085] Step S'1042: Perform bicubic interpolation on the filled background.

[0086] Step S'1043: fine-tune the scaling effect of the area corresponding to each main object according to the importance of each main object.

[0087] In order to fine-tune the scaling effect of each main object corresponding area, it can be done based on energy optimization. Specifically, the energy optimization objectives are as follows:

[0088]

[0089] Among them, W i is the importance of the i-th main object in the image before the zoom effect is fine-tuned, and its calculation method is as shown in the above embodiment, E repair (x', y') is the energy value of the image after fine-tuning the scaling effect, E(x, y) is the energy value of the image before fine-tuning the scaling effect, and the optimization goal is to minimize the energy difference between the images before and after fine-tuning.

[0090] Step S105: outputting the image after the ratio is adjusted.

[0091] From the above attached Figure 1 It can be seen from the example image intelligent transformation method that, on the one hand, the technical solution of the present application is automatically completed by a computer program according to preset rules or intelligent algorithms. Therefore, automation reduces manual intervention, ensures the consistency and accuracy of cropping and adjustment, and improves processing efficiency; on the other hand, after cropping the image to be processed to retain at least one main object, the image retaining at least one main object is adaptively proportionally adjusted. Since this technical means fully considers the content and structure of the image, it not only optimizes the composition of the image to make it more in line with aesthetic requirements, but also ensures that the main object is not over-stretched or compressed, thereby retaining more details.

[0092] Please see attached Figure 2 , is an image intelligent transformation device provided in an embodiment of the present application, the device may include a detection module 201, a generation module 202, a cropping module 203, an adjustment module 204 and an output module 205, which are described in detail as follows:

[0093] A detection module 201 is used to identify at least one main object in the image to be processed through object detection;

[0094] A generating module 202, configured to generate a bounding box of at least one main object in the image to be processed;

[0095] A cropping module 203, configured to crop the image to be processed according to a boundary box of at least one main object to retain at least one main object;

[0096] An adjustment module 204, configured to adaptively adjust the proportion of the image retaining at least one main object;

[0097] The output module 205 is used to output the image after the scale is adjusted.

[0098] From the above attached Figure 2 It can be seen from the example image intelligent transformation device that, on the one hand, the technical solution of the present application is automatically completed by a computer program according to preset rules or intelligent algorithms. Therefore, automation reduces manual intervention, ensures the consistency and accuracy of cropping and adjustment, and improves processing efficiency; on the other hand, after cropping the image to be processed to retain at least one main object, the image retaining at least one main object is adaptively proportionally adjusted. Since this technical means fully considers the content and structure of the image, it not only optimizes the composition of the image to make it more in line with aesthetic requirements, but also ensures that the main object is not over-stretched or compressed, thereby retaining more details.

[0099] Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program of the image intelligent transformation method. When the processor 30 executes the computer program 32, the steps in the above-mentioned image intelligent transformation method embodiment are implemented, such as Figure 1 Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned device embodiments are implemented, for example Figure 2 The functions of the detection module 201, the generation module 202, the cropping module 203, the adjustment module 204 and the output module 205 are shown.

[0100] Exemplarily, the computer program 32 of the image intelligent transformation method mainly includes: identifying at least one main object in the image to be processed through target detection; generating a bounding box of at least one main object in the image to be processed; cropping the image to be processed to retain at least one main object according to the bounding box of at least one main object; adaptively adjusting the scale of the image retaining at least one main object; and outputting the scaled image. The computer program 32 can be divided into one or more modules / units, one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program instruction segments that can perform specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of a detection module 201, a generation module 202, a cropping module 203, an adjustment module 204 and an output module 205 (modules in the virtual device), and the specific functions of each module are as follows: the detection module 201 is used to identify at least one main object in the image to be processed through target detection; the generation module 202 is used to generate a bounding box of at least one main object in the image to be processed; the cropping module 203 is used to crop the image to be processed according to the bounding box of at least one main object to retain at least one main object; the adjustment module 204 is used to adaptively scale the image that retains at least one main object; the output module 205 is used to output the image after scaling.

[0101] The electronic device 3 may include but is not limited to a processor 30 and a memory 31. Those skilled in the art will appreciate that Figure 3 It is only an example of the electronic device 3 and does not constitute a limitation of the electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0102] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0103] The memory 31 may be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 3. Further, the memory 31 may also include both an internal storage unit of the electronic device 3 and an external storage device. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or is to be output.

[0104] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0105] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0106] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0107] In the embodiments provided in the present application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0108] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0109] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0110] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program of the image intelligent transformation method can be stored in a storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments, that is, through target detection, identify at least one main object in the image to be processed; generate a bounding box of at least one main object in the image to be processed; according to the bounding box of at least one main object, crop the image to be processed to retain at least one main object; adaptively adjust the scale of the image retaining at least one main object; output the image after the scale adjustment. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. Storage media may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content contained in the storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, storage media do not include electric carrier signals and telecommunication signals.

[0111] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application is described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. The specific implementation methods described above further describe the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application, and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present invention.

Claims

1. An image intelligent transformation method, characterized in that: The method comprises: Identifying at least one main object in the image to be processed by object detection; generating a bounding box of the at least one primary object; According to the bounding box of the at least one main object, cropping the image to be processed to retain the at least one main object; Adaptively scaling the image retaining the at least one main object; The image after the scale adjustment is output.

2. The image intelligent transformation method according to claim 1, characterized in that: The step of identifying at least one main object in the image to be processed by target detection includes: Based on an edge detection algorithm, identifying at least one main object in the image to be processed; or At least one main object in the image to be processed is identified based on color contrast.

3. The image intelligent transformation method according to claim 1, characterized in that: The generating a bounding box of the at least one main object comprises: If the image to be processed has only one main object, generating a minimum rectangular frame surrounding the one main object; If the image to be processed includes multiple main objects, generating a minimum rectangular frame of each of the multiple main objects; A minimum rectangular frame surrounding the plurality of main objects is generated according to the minimum rectangular frame of each of the plurality of main objects.

4. The image intelligent transformation method according to claim 1, characterized in that: The step of cropping the image to be processed according to the bounding box of the at least one main object to retain the at least one main object comprises: generating a cropping box that encloses a bounding box of the at least one primary object; Determining the golden section line of the image to be processed according to the width and height of the image to be processed; The image to be processed is cropped using the cropping frame after the cropping frame is aligned with the golden section line of the image to be processed or the at least one main object is moved to a position where the distance from the golden section line of the image to be processed does not exceed a preset threshold.

5. The image intelligent transformation method according to claim 4, characterized in that: The method further comprises: Extracting the background of the image to be processed by using a segmentation mask of the background of the image to be processed; Calculating the value of each pixel in the background and the weighted average of the pixels in its neighborhood by a convolution operation to obtain a background blurred image, wherein the weight of the weighted average is determined by a Gaussian function; The background blurred image is smoothed by bilateral filtering.

6. The image intelligent transformation method according to claim 1, characterized in that: The step of adaptively adjusting the proportion of the image retaining the at least one main object comprises: calculating the importance of each of the at least one main object; According to the importance of each main object, a corresponding zoom priority is assigned to each main object; The image retaining the at least one main object is resized using a scaling ratio corresponding to the scaling priority.

7. The image intelligent transformation method according to claim 6, characterized in that: The method further comprises: Filling the background of the image retaining the at least one main object based on image texture information and color transition; Perform bicubic interpolation on the filled background; According to the importance of each main object, the scaling effect of the area corresponding to each main object is fine-tuned.

8. An image intelligent transformation device, characterized in that: The device comprises: A detection module, used for identifying at least one main object in the image to be processed through object detection; A generating module, configured to generate a bounding box of the at least one main object; A cropping module, configured to crop the image to be processed according to a boundary box of the at least one main object to retain the at least one main object; An adjustment module, used for adaptively adjusting the proportion of the image retaining the at least one main object; An output module is used to output the image after the scale is adjusted.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.