Image processing method and device, storage medium and computer device
By extracting visual features of the image and determining the cropping box, and using weight allocation and heatmap techniques to optimize the cropping box, the problem of inaccurate image cropping is solved, and high-quality image cropping effect is achieved.
Patent Information
- Application Number
- CN202011624292.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-30
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2040-12-30
AI Technical Summary
In the existing technology, inaccurate image cropping leads to distortion and incompleteness in the image content and composition.
By extracting visual features from the image, cropping boxes are determined based on these features, and the image is cropped according to the cropping boxes. Weighting and heatmap techniques are used to optimize the position and size of the cropping boxes, and padding is performed when necessary.
It enables reasonable cropping based on the visual content of the image, solving the problem of inaccurate cropping and improving the image quality and visual effect.
Smart Images

Figure CN114693535B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular, to an image processing method and device, a storage medium and a computer device. BACKGROUND
[0002] When processing an image, it is often necessary to adjust the frame. For example, the original image is stretched and compressed, or the original image is cut in the center cutting manner, and then a cut image is obtained. However, the visual content in the image is very complex and changes all the time. In the related art, after adjusting the frame, a high-quality image is often not obtained, and the content and composition of the image frame are distorted and missing.
[0003] In view of the above problems, an effective solution has not been proposed. SUMMARY
[0004] Embodiments of the present application provide an image processing method and device, a storage medium and a computer device to at least solve the technical problem of inaccurate cutting in the related art.
[0005] According to an aspect of an embodiment of the present application, an image processing method is provided, including: obtaining an image; extracting a visual feature of the image, wherein the visual feature includes one or more feature regions on the image, and the one or more feature regions are used to reflect the information content of the image; determining a cutting frame of the image according to the visual feature; and cutting the image according to the cutting frame to obtain a cut image.
[0006] Optionally, determining the cutting frame of the image according to the visual feature includes: in a case where the visual feature includes a plurality of feature regions, assigning different weights to the plurality of feature regions, wherein the size of the weight represents the importance of the feature region; and determining the cutting frame of the image according to the plurality of feature regions after assigning the weights, wherein the cutting frame includes a feature region with the largest weight in the plurality of feature regions.
[0007] Optionally, determining the cutting frame of the image according to the plurality of feature regions after assigning the weights includes: fusing the plurality of feature regions with different weights to form a heat map, wherein different colors in the heat map represent feature regions with different weights; and determining the size and position of the cutting frame according to the heat map.
[0008] Optionally, the cutting the image according to the cutting frame to obtain a cut image comprises: in a case that the cutting frame exceeds the image itself, determining an exceeding part of the cutting frame that exceeds the image; filling the exceeding part after cutting the image according to the cutting frame to obtain the cut image.
[0009] Optionally, in a case that the image comprises a plurality of shot images obtained by video cutting, the method further comprises: splicing the plurality of cut shot images to obtain a cut video.
[0010] Optionally, before the splicing the plurality of cut shot images to obtain a cut video, the method further comprises: performing a spline interpolation method on the cutting frame in each of the plurality of shot images to obtain a smoothed cutting frame, wherein the smoothed cutting frame is used for cutting to obtain the plurality of cut shot images.
[0011] Optionally, the extracting the visual feature of the image comprises: inputting the image into an image feature model to obtain the visual feature of the image, wherein the image feature model is obtained by machine learning using a first data set, and training data in the first data set comprises an image and a visual feature of the image.
[0012] Optionally, the image comprises a plurality of material images of a predetermined object, and after the cutting the image according to the cutting frame to obtain a cut image, the method further comprises: inputting each of the plurality of cut material images into an image special effect model to obtain a special effect image of each of the plurality of cut material images, wherein the image special effect model is obtained by machine learning using a second data set, and training data in the second data set comprises a material image and a special effect image of the material image; and splicing the special effect images corresponding to the plurality of material images to obtain a video of the predetermined object.
[0013] Optionally, the splicing the special effect images corresponding to the plurality of material images to obtain a video of the predetermined object comprises: displaying the special effect images corresponding to the plurality of material images on a display interface; receiving an adjustment operation of adjusting the special effect images corresponding to the plurality of material images; in response to the adjustment operation, adjusting the special effect images corresponding to the plurality of material images to obtain a plurality of adjusted special effect images; and splicing the plurality of adjusted special effect images to obtain a video of the predetermined object.
[0014] According to another aspect of the embodiments of the present application, there is also provided an image processing method, comprising: displaying an image on an interactive interface; displaying a visual feature of the image on the interactive interface, wherein the visual feature comprises one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image; displaying a cropping frame of the image on the interactive interface, wherein the cropping frame is determined according to the visual feature; and displaying a cropping result on the interactive interface, wherein the cropping result is obtained by cropping the image according to the cropping frame.
[0015] According to still another aspect of the embodiments of the present application, there is also provided an image processing apparatus, comprising: a first obtaining module configured to obtain an image; an extracting module configured to extract a visual feature of the image, wherein the visual feature comprises one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image; a determining module configured to determine a cropping frame of the image according to the visual feature; and a cropping module configured to crop the image according to the cropping frame to obtain a cropped image.
[0016] According to still another aspect of the embodiments of the present application, there is also provided an image processing apparatus, comprising: a first displaying module configured to display an image on an interactive interface; a second displaying module configured to display a visual feature of the image on the interactive interface, wherein the visual feature comprises one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image; a third displaying module configured to display a cropping frame of the image on the interactive interface, wherein the cropping frame is determined according to the visual feature; and a fourth displaying module configured to display a cropping result on the interactive interface, wherein the cropping result is obtained by cropping the image according to the cropping frame.
[0017] According to still another aspect of the embodiments of the present application, there is also provided a storage medium, comprising a stored program, wherein the program, when executed, controls a device in which the storage medium is located to perform any of the image processing methods described above.
[0018] According to still another aspect of the embodiments of the present application, there is also provided a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program stored in the memory, and the computer program, when executed, causes the processor to perform any of the image processing methods described above.
[0019] According to a further aspect of the embodiments of the present application, there is also provided an image processing method, comprising: obtaining material of an object; identifying information of the material by using an artificial intelligence recognition algorithm; and performing visual processing on the material by using an artificial intelligence visual algorithm according to the information of the material, to obtain a video of the object.
[0020] Optionally, in the case that the material is an image material, the identifying the information of the material by using the artificial intelligence recognition algorithm comprises at least one of the following: identifying a category to which the image material belongs by using a category algorithm in the artificial intelligence recognition algorithm; identifying a position of a target object in the image material by using a position algorithm in the artificial intelligence recognition algorithm; identifying a salient region in the image material by using a feature algorithm in the artificial intelligence recognition algorithm; and identifying an aesthetic score of the image material by using a scoring algorithm in the artificial intelligence recognition algorithm.
[0021] Optionally, the performing visual processing on the material by using the artificial intelligence visual algorithm comprises at least one of the following: performing screening on the material by using the artificial intelligence visual algorithm; performing cropping on a target object in an image material; performing arrangement on a plurality of types of material; and performing rendering on an image material.
[0022] Optionally, an intermediate result is output by an interactive device, wherein the intermediate result comprises at least one of the following: a result of the information of the material identified by using the artificial intelligence recognition algorithm; and a result of the visual processing on the material by using the artificial intelligence visual algorithm.
[0023] According to a further aspect of the embodiments of the present application, there is also provided an image processing device, comprising: a second obtaining module configured to obtain material of an object; an identifying module configured to identify information of the material by using an artificial intelligence recognition algorithm; and a processing module configured to perform visual processing on the material by using an artificial intelligence visual algorithm according to the information of the material, to obtain a video of the object.
[0024] In the embodiments of the present application, the image is obtained and the visual features of the image are extracted, the cropping frame of the image is determined according to the visual features, and the image is cropped according to the cropping frame, so as to obtain the image after cropping according to the requirement, thereby realizing the technical effect of reasonably cropping the image according to the visual content of the image, and further solving the technical problem of inaccurate cropping of the image in the related art. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and serve to explain the principles of the application. In the drawings:
[0026] Figure 1 is a computer terminal hardware structure block diagram for implementing the image processing method according to the embodiment of the present application;
[0027] Figure 2 is a flow chart of the image processing method one according to the embodiment 1 of the present application;
[0028] Figure 3 is a flow chart of the image processing method two according to the embodiment 1 of the present application;
[0029] Figure 4 is a structure schematic diagram for generating video by using the image processing according to the embodiment of the present application;
[0030] Figure 5 is a flow chart of the image processing method three according to the embodiment 1 of the present application;
[0031] Figure 6 is a flow chart of the image processing method according to the optional embodiment of the present application;
[0032] Figure 7 is a structure block diagram of the image processing device one according to the embodiment 2 of the present application;
[0033] Figure 8 is a structure block diagram of the image processing device two according to the embodiment 3 of the present application;
[0034] Figure 9 is a structure block diagram of the image processing device three according to the embodiment 4 of the present application;
[0035] Figure 10 is a structure block diagram of the computer terminal according to the embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by the person of ordinary skill in the art without creative labor should belong to the protection scope of the present application.
[0037] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present application and the above-described accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0038] Embodiment 1
[0039] According to the embodiments of the present application, an image processing method embodiment is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.
[0040] The method embodiment provided by Embodiment 1 of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method is shown. As Figure 1 shown, the computer terminal 10 (or mobile device) can include one or more processors 102 (the processor 102 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports in the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can also include more or fewer components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .
[0041] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be referred to herein generically as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. Furthermore, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any one of the other elements of the computer terminal 10 (or mobile device). As referred to in the embodiments of the present application, the data processing circuitry acts as a processor to control, for example, the selection of the variable resistance terminal path in connection with the interface.
[0042] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the image processing method of the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implements the image processing method of the application program described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory disposed remotely with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0043] The transmission device is used to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet in a wireless manner.
[0044] The display can be, for example, a touch screen type liquid crystal display (LCD), which can enable a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0045] In the above operating environment, the present application provides an image processing method as shown in Figure 2 Figure 2 is a flowchart of the image processing method one according to the embodiment 1 of the present application. As shown in Figure 2 the method includes the following steps:
[0046] Step S202, acquiring an image;
[0047] In step S204, a visual feature of the image is extracted, where the visual feature includes one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image.
[0048] In step S206, a cropping frame of the image is determined according to the visual feature.
[0049] In step S208, the image is cropped according to the cropping frame to obtain a cropped image.
[0050] In the above steps, the image is obtained and the visual feature of the image is extracted, the cropping frame of the image is determined according to the visual feature, and the image is cropped according to the cropping frame, so that the obtained image is the image cropped according to the information content of the image, thereby realizing the technical effect of reasonably cropping the image according to the visual content of the image, and further solving the technical problem of inaccurate cropping of the image in the related art.
[0051] As an optional embodiment, the image can be input into an image feature model to obtain the visual feature of the image, where the image feature model is obtained by machine learning using a first data set, and the training data in the first data set includes the image and the visual feature of the image. The image feature model can analyze the visual content in the image and extract the visual feature of the image therefrom. The image feature model is trained by machine learning using the training data including the image and the visual feature of the image, and thus can extract various types of image visual features. For example, the visual features such as a face, a text, and a subject can be extracted from the image. Using the above image feature model can improve the efficiency and accuracy of extracting the visual feature from the image, and facilitate subsequent adjustment of the image frame.
[0052] As an optional embodiment, the cropping frame of the image can be determined in the following manner: in a case where the visual feature includes a plurality of feature regions, different weights are assigned to the plurality of feature regions, where the size of the weight represents the importance of the feature region; and the cropping frame of the image is determined according to the plurality of feature regions after the weights are assigned, where the cropping frame includes the feature region with the largest weight among the plurality of feature regions. When the image includes a plurality of feature regions, using the cropping frame to crop the image sometimes needs to select the plurality of feature regions, because the cropping frame that meets the frame ratio requirement cannot frame all the feature regions. At this time, the above problem can be solved according to the present embodiment, which can assign weights representing the importance of the regions to the plurality of feature regions, and then ensure that the most important feature region is included in the cropping frame, thereby realizing the preservation of the most important information of the image in the process of cropping the image.
[0053] As an optional embodiment of the present application, the feature region can be assigned a weight value in various ways. For example, the feature region can be first identified for image type, area, relationship with surrounding feature regions, and other information, and then the weight value of each feature region can be determined according to the above information. For example, when the image type of a feature region is identified as a face region, a text region, or a human body region, and other types, different weight values can be assigned according to the type. When the area of a feature region is significantly large, a higher weight value can be assigned. When most feature regions are distributed around a particular feature region, a higher weight value can be assigned to the central feature region. By assigning weight values to multiple feature regions, the importance of the multiple feature regions can be well distinguished.
[0054] As an optional embodiment, according to the multiple feature regions assigned with weight values, the cropping frame of the image can be determined in the following way: the multiple feature regions are fused to form a heat map with different colors representing feature regions with different weight values; and the size and position of the cropping frame are determined according to the heat map. When there are many feature regions, it is complex to assign weight values to multiple isolated feature regions and perform overall analysis, and it is difficult to obtain a good effect of defining the cropping frame. According to the weight values of the multiple feature regions, the weight value distribution of the feature regions of the entire image is drawn as a heat map, which can simply and clearly determine the important regions in the entire image, and facilitate setting the size and position of the cropping frame. Different colors in the heat map represent the weight values of the regions, i.e., the importance of the regions. By this method, the discrete feature regions in the image can be fused into a continuous image weight value distribution function, and the basis for selecting the cropping frame is optimized.
[0055] As an optional embodiment, when the cropping frame exceeds the image itself, the exceeding part of the cropping frame that exceeds the image is determined; and the exceeding part is filled after the image is cropped according to the cropping frame, to obtain the cropped image. The cropping frame that meets the frame requirement can not be entirely located in the image, and there can be an exceeding part. In this embodiment, the exceeding part is filled, so that the exceeding part is not too conspicuous, and the overall style of the image regions in the frame is consistent, and the visual effect is similar.
[0056] As an optional embodiment, when the image includes multiple shot images obtained by cutting a video, the multiple cropped shot images can be spliced to obtain a cropped video. By splicing the cropped shot images belonging to the same shot, the technical effect of cropping the video can be achieved. Since the visual features in the video change rapidly, it is difficult to directly crop the video. The present embodiment provides a method for cropping a video according to the importance of visual features in multiple shot images, so that the process of cropping the video is simpler.
[0057] As an optional embodiment, before the plurality of cut images are spliced to obtain the cut video, the spline interpolation method can be used to perform smoothing processing on the cut frame in each of the plurality of cut images, to obtain a smoothed cut frame, wherein the smoothed cut frame is used for cutting to obtain the plurality of cut images. Because the positions of the cut frames selected from the plurality of images may not be completely consistent, when the images in the two cut frames are processed as a continuous video, the video may be jittered. For example, the same person's face is selected by using the cut frame in two images, but because the cut frame is created under the influence of other regions in the image, the respective cut frames cannot completely coincide. In this embodiment, the spline interpolation method is used to perform smoothing processing on the cut frame, so that the positions of the cut frames in the plurality of images are approximately the same, and when the images are processed as a video, the camera jitter of the video image can be avoided, and a better video playing effect can be obtained.
[0058] As an optional embodiment, in the case that the image includes a plurality of material images of a predetermined object, after the image is cut according to the cut frame and the cut image is obtained, each of the plurality of cut material images can be input into an image special effect model to obtain a special effect image of each of the plurality of cut material images, and then the special effect images corresponding to the plurality of material images are spliced to obtain a video of the predetermined object. The image special effect model is obtained by machine learning using a second data set, and the training data in the second data set includes material images and special effect images of the material images. By using the image special effect model to add special effects to the material images to generate special effect images, more rich effects can be added to the cut images. In addition, after the plurality of material images of the predetermined object are cut, the image special effect model is input to add special effects, the special effect style of the material images related to the predetermined object can be kept consistent, the style of the video about the predetermined object obtained by splicing is uniform, and the characteristics are distinct. The image special effect model obtained by machine learning can reduce the workload of adding special effects to the images, and the material images can obtain more stable and more targeted special effect effects.
[0059] As an optional embodiment, the plurality of special effect images corresponding to the plurality of material images can be spliced in the following manner: the plurality of special effect images corresponding to the plurality of material images are displayed on a display interface, an adjustment operation of adjusting the plurality of special effect images corresponding to the plurality of material images is received, the plurality of special effect images corresponding to the plurality of material images are adjusted in response to the adjustment operation to obtain adjusted plurality of special effect images, and the adjusted plurality of special effect images are spliced to obtain a video of the predetermined object. Through this embodiment, the user can adjust the added special effect type in real time or afterwards when adding special effects to the material images, so that the special effect effect of the video obtained by splicing can be adjusted according to the user's feedback, and the user's autonomy and the freedom of the video are increased.
[0060] Figure 3 is a flowchart of the image processing method two according to the embodiment 1 of the present application. As shown in the figure, the method comprises the following steps: Figure 3
[0061] Step S302, obtaining the material of the object;
[0062] Step S304, identifying the information of the material by using the artificial intelligence recognition algorithm;
[0063] Step S306, according to the information of the material, using the artificial intelligence visual algorithm to process the material visually, and obtaining the video of the object.
[0064] Through the above steps, the information of the material is identified by using the artificial intelligence recognition algorithm, and the material is processed visually by using the artificial intelligence visual algorithm, so as to achieve the purpose of using artificial intelligence to realize the production of video, thereby realizing the technical effect of efficient and accurate production of video.
[0065] As an optional embodiment, the object can be the target object for producing the video, which can be a specific real object or an abstract virtual object. For example, when the object is a commodity, it can be a specific real object or a virtual model of the commodity.
[0066] As an optional embodiment, when obtaining the material of the object, the type of the material can be various, such as image, text, video, audio, animation, etc. The above-mentioned various types of materials are the ways of describing the object in various ways. For various materials, the artificial intelligence algorithm used to identify the information of the material can be different.
[0067] As an optional embodiment, in the case of image material, the information of the material is identified by using an artificial intelligence recognition algorithm, including at least one of the following: using a category algorithm in the artificial intelligence recognition algorithm to identify the category to which the image material belongs; using a position algorithm in the artificial intelligence recognition algorithm to identify the position of the target object in the image material; using a feature algorithm in the artificial intelligence recognition algorithm to identify the salient region in the image material; using a scoring algorithm in the artificial intelligence recognition algorithm to identify the aesthetic score of the image material. Wherein the type of the image referred to above can be the category divided by the object, for example, when the object is a commodity, the category can be clothing, household goods, food and the like. The position of the target object in the image material can refer to the position of the person, the position of the animal, and the position of some markers, etc. The salient region in the image material can be the region for explanation or attention in the image material, for example, when the object is a bowl, whether the bowl can be used in the microwave oven, etc. The aesthetic score of the image material referred to above can be whether the image material can be directly used for subsequent video splicing or use.
[0068] As an optional embodiment, the visual processing of the material by the artificial intelligence visual algorithm can include multiple types, for example, can include at least one of the following: screening the material by the artificial intelligence visual algorithm, cutting the target object in the image material, arranging multiple types of materials, and rendering the image material. For example, some of the acquired materials are of poor quality, and although the subsequent processing effect can be better, there is still a possibility that they cannot be used. Therefore, unnecessary processing resources can be avoided by directly screening the collected materials and discarding the materials that do not meet the requirements. When processing the image material, sometimes the image material includes some unnecessary objects, for example, some more prominent backgrounds or reference objects. The existence of these objects cannot reflect the important part of the image itself (i.e., the target object that the image itself is supposed to reflect), so in order to make the target object display in a more prominent way in the image, the image material can be cut to make the target object prominent. When using multiple types of materials for video production, different types of materials have different effects and different emphases, so various types of materials can be effectively arranged to arrange them in the right place, not only to reflect the meaning of the material itself, but also to make the produced image more in line with the viewing habits and improve the viewing experience. When producing a video, in order to improve the viewing experience of the user, the image material can be beautified, for example, by using some colors and animations for beautification. Therefore, when using image materials to produce a video, the beautification degree of the image material can be determined first, for example, the beautification score of the image material is obtained first, and whether the image material needs to be beautified when producing a video is determined according to the beautification score.
[0069] As an optional embodiment, when artificial intelligence is used for video production, in order to realize the monitoring or adjustment of the production process, the intermediate results can be output through the interactive device, wherein the intermediate results include at least one of the following: the result of the information of the material identified by the artificial intelligence identification algorithm, and the result of the visual processing of the material by the artificial intelligence visual algorithm. The intermediate results of the artificial intelligence video production are adjusted or monitored by human adjustment or semi-human adjustment, so that the whole process is visible and controllable, avoiding the black box operation and unexplainability of artificial intelligence.
[0070] Figure 4 is a structure diagram for generating a video by image processing according to an embodiment of the application, as Figure 4 shown, in the process of generating a video by image processing, the above-mentioned cutting processing of the image can be part of the whole video generation, for example, it can be part of the material arrangement process in the video generation process. In the following description, the image is taken as a commodity image as an example.
[0071] In the related art, AI technology is also used to improve the efficiency of video clipping and editing process, but the whole production process is still dominated by people. In the video generation process provided by the embodiment of the application, although AI technology is also used, and an editor interface is provided, the whole video generation process is AI dominated, and the algorithm automatically processes a large number of underlying video processing processes, and the video producer only needs to fine-tune on the basis of the generated video, greatly reducing the workload of the video producer, and providing the ability of one-key video generation.
[0072] The video generated by the embodiment of the application mainly includes the following processing:
[0073] (1) Input a commodity link of an e-commerce platform, obtain the video corresponding to the commodity, the image, and the commodity information (category, price, comment, etc.).
[0074] (2) The input material is recognized and understood by calling a rich AI recognition algorithm, and various information such as the category of image material, target position, significant area, aesthetic score, etc. is given, which provides decision basis for subsequent intelligent arrangement and rendering.
[0075] (3) The material is filtered, cut, sorted, and rendered intelligent visual effect according to the material recognition information provided above, and the final e-commerce video is encoded and output.
[0076] It should be noted that the intermediate results of the AI algorithm decision process can be output to the intelligent editor while outputting the final synthesized video, so that the user can fine-tune the final result, such as replacing part of the display material and adjusting the display order. In addition, when the intelligent editor is used for adjustment, the material analysis results output by the intelligent editor are presented to the user, such as the category of the material.
[0077] Through the above optional embodiment, the video editor as the core of the video production process in the related art is abandoned, and a new AI-based video intelligent editing and generation process is proposed. The whole video generation process can be completed by AI, and then the video result generated by AI can be fine-tuned through a lightweight intelligent editor interaction interface. The new process can greatly improve the efficiency of video production.
[0078] The whole process is dominated by AI, but all intermediate results in the AI processing process can be presented through the intelligent editor, such as the category of the material, the video cutting position, the special effect category, etc. The whole process is visible and controllable, avoiding the black box operation and unexplainability of AI.
[0079] It should be noted that the intelligent editor is not an optional item of the video generation system. Because the editor does not undertake video rendering and synthesis and the like, the intelligent editor is not specifically a visual interactive interface with an interface, and can also be implemented through other interactive modes such as voice dialogue, as long as the user's demand can be transmitted to the AI-led process.
[0080] Figure 5 is a flowchart of the image processing method three according to the embodiment of the present application, as shown in Figure 5 the method comprises the following steps:
[0081] Step S502, displaying the image on the interactive interface;
[0082] Step S504, displaying the visual features of the image on the interactive interface, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used to embody the information content of the image;
[0083] Step S506, displaying the cropping frame of the image on the interactive interface, wherein the cropping frame is determined according to the visual features;
[0084] Step S508, displaying the cropping result on the interactive interface, wherein the cropping result is obtained by cropping the image according to the cropping frame.
[0085] In the above steps, by displaying the image, the visual features of the image, the cropping frame of the image and the cropping result on the interactive interface, the purpose of visualizing the process of cropping the image is achieved, so that the technical effect of reasonably cropping the image according to the visual content of the image and visualizing the image cropping process is realized, and the technical problem of being difficult to highlight the key information content of the image when adjusting the image frame is solved.
[0086] Figure 6 is a flowchart of the image processing method according to the optional embodiment of the present application, as shown in Figure 6 the present optional embodiment proposes an application scenario of the image processing method, that is, a process of cropping and frame adjusting the shot images in a video. The steps included in the method are briefly described below.
[0087] S1, shot segmentation of the video. Before processing the shot images in the video, the shot images are obtained by performing shot segmentation on the video, and a plurality of image sets belonging to different shots are obtained after segmentation.
[0088] S2, face / subject / text / saliency detection. By detecting each region in the image, the visual features of the image are extracted, such as the face, subject, text, etc. in the image. After detection, the above-mentioned visual features are fused with different weights to form a heat map to obtain the information heat map on the image. The greater the value, the more important the corresponding region. According to the heat map, each feature region can be labeled, and a cutting frame is generated to frame the feature region, so that the cutting frame contains the most important region.
[0089] S3, cutting frame estimation / smoothing. The cutting frame sequence of each shot is smoothed by spline interpolation to avoid cutting shot jitter.
[0090] S4, cutting / filling. According to the cutting frame position, frame-by-frame cutting or filling (when the important region is larger, to ensure that the cutting frame aspect ratio meets the requirements, the cutting frame is allowed to exceed the picture itself, and the exceeding part is filled) is performed.
[0091] S5, splicing and synthesizing all the cut shot images to output the adjusted video.
[0092] It should be noted that for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action sequence described, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0093] From the above description of the embodiments, those skilled in the art can clearly understand that the image processing method according to the above-mentioned embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the method of each embodiment of the present application.
[0094] Embodiment 2
[0095] According to the embodiments of the present application, a device for implementing the above-mentioned image processing method I is also provided, Figure 7 is a structural block diagram of the image processing device I according to the embodiment 2 of the present application, like Figure 7As shown in the figure, the device comprises a first acquisition module 72, an extraction module 74, a determination module 76 and a cutting module 78, which are specifically described as follows.
[0096] The first acquisition module 72 is configured to acquire an image; the extraction module 74 is connected to the first acquisition module 72 and configured to extract a visual feature of the image, wherein the visual feature comprises one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image; the determination module 76 is connected to the extraction module 74 and configured to determine a cutting frame of the image according to the visual feature; and the cutting module 78 is connected to the determination module 76 and configured to cut the image according to the cutting frame to obtain a cut image.
[0097] It should be noted that the first acquisition module 72, the extraction module 74, the determination module 76 and the cutting module 78 correspond to steps S202 to S208 in Embodiment 1, and the multiple modules have the same instances and application scenarios as the corresponding steps, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules as part of the device can run in the computer terminal 10 provided in Embodiment 1.
[0098] Embodiment 3
[0099] According to the embodiments of the present application, a device for implementing the above-mentioned image processing method two is further provided, Figure 8 is a structural block diagram of the image processing device two according to Embodiment 3 of the present application, as Figure 8 As shown in the figure, the device comprises a second acquisition module 82, an identification module 84 and a processing module 86, which are specifically described as follows.
[0100] The second acquisition module 82 is configured to acquire a material of an object; the identification module 84 is connected to the second acquisition module 82 and configured to identify information of the material by using an artificial intelligence identification algorithm; and the processing module 86 is connected to the identification module 84 and configured to perform visual processing on the material by using an artificial intelligence visual algorithm according to the information of the material to obtain a video of the object.
[0101] It should be noted that the second acquisition module 82, the identification module 84 and the processing module 86 correspond to steps S302 to S306 in Embodiment 1, and the multiple modules have the same instances and application scenarios as the corresponding steps, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules as part of the device can run in the computer terminal 10 provided in Embodiment 1.
[0102] Embodiment 4
[0103] According to the embodiments of the present application, a device for implementing the above-mentioned image processing method three is further provided,Figure 9 is a structural block diagram of the image processing device three according to the embodiment 4 of the present application, as shown, the device comprises: a first display module 92, a second display module 94, a third display module 96 and a fourth display module 98, and the device will be specifically described below. Figure 9
[0104] The first display module 92 is used for displaying the image on the interactive interface; the second display module 94 connected to the first display module 92 is used for displaying the visual features of the image on the interactive interface, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used for embodying the information content of the image; the third display module 96 connected to the second display module 94 is used for displaying the cropping frame of the image on the interactive interface, wherein the cropping frame is determined according to the visual features; and the fourth display module 98 connected to the third display module 96 is used for displaying the cropping result on the interactive interface, wherein the cropping result is obtained by cropping the image according to the cropping frame.
[0105] It should be noted that the first display module 92, the second display module 94, the third display module 96 and the fourth display module 98 correspond to the steps S502 to S508 in the embodiment 1, and the multiple modules have the same instances and application scenarios as the corresponding steps, but are not limited to the contents disclosed in the embodiment 1. It should be noted that the above modules as part of the device can run in the computer terminal 10 provided in the embodiment 1.
[0106] Embodiment 5
[0107] The embodiment of the present application can provide a computer terminal, which can be any one of the computer terminal devices in the computer terminal group. Alternatively, in the embodiment, the computer terminal can be replaced by a terminal device such as a mobile terminal.
[0108] Alternatively, in the embodiment, the computer terminal can be located in at least one of the multiple network devices of the computer network.
[0109] In the embodiment, the computer terminal can execute the program codes of the following steps in the image processing method of the application program: obtaining the image; extracting the visual features of the image, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used for embodying the information content of the image; determining the cropping frame of the image according to the visual features; and cropping the image according to the cropping frame to obtain the cropped image.
[0110] Alternatively, Figure 10 is a structural block diagram of a computer terminal according to the embodiment of the present application. As shown in Figure 10 As shown, the computer terminal can include one or more (only one is shown in the figure) processors 102, memory 104, etc.
[0111] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the image processing method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned image processing method. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the computer terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0112] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: acquiring an image; extracting visual features of the image, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used to reflect the information content of the image; determining a cropping frame of the image according to the visual features; and cropping the image according to the cropping frame to obtain a cropped image.
[0113] Optionally, the above-mentioned processor can further execute program codes of the following steps: determining the cropping frame of the image according to the visual features, including: in the case that the visual features include a plurality of feature regions, assigning different weights to the plurality of feature regions, wherein the size of the weight represents the importance of the feature region; and determining the cropping frame of the image according to the plurality of feature regions after assigning the weights, wherein the cropping frame includes the feature region with the largest weight in the plurality of feature regions.
[0114] Optionally, the above-mentioned processor can further execute program codes of the following steps: determining the cropping frame of the image according to the plurality of feature regions after assigning the weights, including: fusing the plurality of feature regions with different weights to form a heat map, wherein different colors in the heat map represent feature regions with different weights; and determining the size and position of the cropping frame according to the heat map.
[0115] Optionally, the above-mentioned processor can further execute program codes of the following steps: cropping the image according to the cropping frame to obtain a cropped image, including: in the case that the cropping frame exceeds the image itself, determining an exceeding part of the cropping frame that exceeds the image; and filling the exceeding part after cropping the image according to the cropping frame to obtain the cropped image.
[0116] Optionally, the processor can further execute program codes of the following steps: in the case that the image comprises a plurality of shot images obtained by video cutting, the method further comprises: splicing the plurality of cut shot images to obtain the cut video.
[0117] Optionally, the processor can further execute program codes of the following steps: before the step of splicing the plurality of cut shot images to obtain the cut video, the method further comprises: performing spline interpolation smoothing processing on the cut frame in each of the plurality of shot images to obtain a smoothed cut frame, wherein the smoothed cut frame is used for cutting to obtain the plurality of cut shot images.
[0118] Optionally, the processor can further execute program codes of the following steps: extracting the visual features of the image comprises: inputting the image into an image feature model to obtain the visual features of the image, wherein the image feature model is obtained by machine learning using a first data set, and the training data in the first data set comprises: the image and the visual features of the image.
[0119] Optionally, the processor can further execute program codes of the following steps: the image comprises a plurality of material images of a predetermined object, and after the step of cutting the image according to the cut frame to obtain the cut image, the method further comprises: inputting each of the plurality of cut material images into an image special effect model to obtain a special effect image of each of the plurality of cut material images, wherein the image special effect model is obtained by machine learning using a second data set, and the training data in the second data set comprises: the material image and the special effect image of the material image; and splicing the special effect images corresponding to the plurality of material images to obtain the video of the predetermined object.
[0120] Optionally, the processor can further execute program codes of the following steps: splicing the special effect images corresponding to the plurality of material images to obtain the video of the predetermined object comprises: displaying the special effect images corresponding to the plurality of material images on a display interface; receiving an adjustment operation for adjusting the special effect images corresponding to the plurality of material images; in response to the adjustment operation, adjusting the special effect images corresponding to the plurality of material images to obtain a plurality of adjusted special effect images; and splicing the plurality of adjusted special effect images to obtain the video of the predetermined object.
[0121] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: displaying the image on the interactive interface; displaying the visual features of the image on the interactive interface, wherein the visual features comprise one or more feature regions on the image, and the one or more feature regions are used to reflect the information content of the image; displaying the cut frame of the image on the interactive interface, wherein the cut frame is determined according to the visual features; and displaying the cutting result on the interactive interface, wherein the cutting result is obtained by cutting the image according to the cut frame.
[0122] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: obtaining the material of the object; identifying the information of the material by using an artificial intelligence recognition algorithm; and performing visual processing on the material by using an artificial intelligence visual algorithm according to the information of the material to obtain the video of the object.
[0123] Optionally, the processor can further execute the program code of the following steps: in the case that the material is an image material, identifying the information of the material by using the artificial intelligence recognition algorithm, including at least one of the following: identifying the category to which the image material belongs by using a category algorithm in the artificial intelligence recognition algorithm; identifying the position of the target object in the image material by using a position algorithm in the artificial intelligence recognition algorithm; identifying the salient region in the image material by using a feature algorithm in the artificial intelligence recognition algorithm; and identifying the aesthetic score of the image material by using a scoring algorithm in the artificial intelligence recognition algorithm.
[0124] Optionally, the processor can further execute the program code of the following steps: performing visual processing on the material by using the artificial intelligence visual algorithm, including at least one of the following: performing screening on the material by using the artificial intelligence visual algorithm, performing cutting on the target object in the image material, arranging the materials of multiple types, and rendering the image material.
[0125] Optionally, the processor can further execute the program code of the following steps: outputting the intermediate result through the interaction device, wherein the intermediate result includes at least one of the following: the result of identifying the information of the material by using the artificial intelligence recognition algorithm, and the result of performing visual processing on the material by using the artificial intelligence visual algorithm.
[0126] Those skilled in the art can understand that, Figure 10 The structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or other terminal devices. Figure 10 It does not limit the structure of the electronic device. For example, the computer terminal 10 can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 10 , or have a different configuration from that shown in Figure 10 .
[0127] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by instructing the terminal device related hardware through a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0128] Embodiment 4
[0129] The embodiment of the present application also provides a storage medium. Optionally, in the embodiment, the storage medium can be used to save the program code executed by the image processing method provided in the embodiment 1.
[0130] Optionally, in the embodiment, the storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0131] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: obtaining an image; extracting visual features of the image, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used to reflect the information content of the image; determining a cropping frame of the image according to the visual features; and cropping the image according to the cropping frame to obtain a cropped image.
[0132] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: determining the cropping frame of the image according to the visual features, including: in the case that the visual features include a plurality of feature regions, assigning different weights to the plurality of feature regions, wherein the size of the weight represents the importance of the feature region; and determining the cropping frame of the image according to the plurality of feature regions after assigning the weights, wherein the cropping frame includes the feature region with the largest weight in the plurality of feature regions.
[0133] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: determining the cropping frame of the image according to the plurality of feature regions after assigning the weights, including: fusing the plurality of feature regions with different weights to form a heat map, wherein different colors in the heat map represent feature regions with different weights; and determining the size and position of the cropping frame according to the heat map.
[0134] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: cutting the image according to the cutting frame to obtain the cut image, including: in the case that the cutting frame exceeds the image itself, determining an exceeding part of the cutting frame that exceeds the image; filling the exceeding part after cutting the image according to the cutting frame to obtain the cut image.
[0135] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: in the case that the image includes a plurality of shot images obtained by cutting a video, the method further includes: splicing the plurality of cut shot images to obtain the cut video.
[0136] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: before splicing the plurality of cut shot images to obtain the cut video, the method further includes: performing smoothing processing on the cutting frame in each of the plurality of shot images using a spline interpolation method to obtain a smoothed cutting frame, wherein the smoothed cutting frame is used for cutting to obtain the plurality of cut shot images.
[0137] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: extracting the visual feature of the image, including: inputting the image into an image feature model to obtain the visual feature of the image, wherein the image feature model is obtained by machine learning using a first data set, and the training data in the first data set includes: the image and the visual feature of the image.
[0138] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: the image includes a plurality of material images of a predetermined object, and after cutting the image according to the cutting frame to obtain the cut image, the method further includes: inputting each of the plurality of cut material images into an image special effect model to obtain a special effect image of each of the plurality of cut material images, wherein the image special effect model is obtained by machine learning using a second data set, and the training data in the second data set includes: the material image and the special effect image of the material image; and splicing the special effect images corresponding to the plurality of material images to obtain a video of the predetermined object.
[0139] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: splicing the special effect images corresponding to the plurality of material images to obtain the video of the predetermined object, comprising: displaying the special effect images corresponding to the plurality of material images on the display interface; receiving an adjustment operation of adjusting the special effect images corresponding to the plurality of material images; in response to the adjustment operation, adjusting the special effect images corresponding to the plurality of material images to obtain the plurality of adjusted special effect images; and splicing the plurality of adjusted special effect images to obtain the video of the predetermined object.
[0140] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: displaying the image on the interactive interface; displaying the visual features of the image on the interactive interface, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used to embody the information content of the image; displaying the cropping frame of the image on the interactive interface, wherein the cropping frame is determined according to the visual features; and displaying the cropping result on the interactive interface, wherein the cropping result is obtained by cropping the image according to the cropping frame.
[0141] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: obtaining the material of the object; identifying the information of the material by using the artificial intelligence recognition algorithm; and performing visual processing on the material by using the artificial intelligence visual algorithm according to the information of the material to obtain the video of the object.
[0142] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: in the case that the material is an image material, identifying the information of the material by using the artificial intelligence recognition algorithm, comprising at least one of the following: identifying the category to which the image material belongs by using a category algorithm in the artificial intelligence recognition algorithm; identifying the position of the target object in the image material by using a position algorithm in the artificial intelligence recognition algorithm; identifying the salient region in the image material by using a feature algorithm in the artificial intelligence recognition algorithm; and identifying the aesthetic score of the image material by using a scoring algorithm in the artificial intelligence recognition algorithm.
[0143] Optionally, in the embodiment, the storage medium is further configured to store program code for performing the following steps: performing visual processing on the material by using the artificial intelligence visual algorithm, comprising at least one of the following: performing screening on the material by using the artificial intelligence visual algorithm, performing cropping on the target object in the image material, arranging a plurality of types of materials, and rendering the image material.
[0144] Optionally, in the embodiment, the storage medium is further configured to store program code for executing the following step: outputting the intermediate result by the interaction device, wherein the intermediate result comprises at least one of the following: a result of identifying the information of the material by using the artificial intelligence identification algorithm, and a result of visually processing the material by using the artificial intelligence visual algorithm.
[0145] The above-mentioned sequence numbers of the embodiments of the application are only for description, and do not represent advantages or disadvantages of the embodiments.
[0146] In the above-mentioned embodiments of the application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0147] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other manners. Among them, the above-mentioned apparatus embodiment is only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.
[0148] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.
[0149] In addition, each functional unit in each embodiment of the application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be realized in the form of hardware or in the form of software functional unit.
[0150] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0151] The above is only the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. An image processing method, characterized by, The method comprises: acquiring an image; extracting visual features of the image, wherein the visual features comprise one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image; determining a cropping frame of the image according to the visual features; cropping the image according to the cropping frame to obtain a cropped image; wherein, according to the visual features, the cropping frame of the image is determined, including: in the case that the visual features comprise a plurality of feature regions, different weights are assigned to the plurality of feature regions, wherein the size of the weight represents the importance of the feature region, and the weight is determined based on at least one of the following information: image type, area, and relationship with the surrounding feature region of the feature region; and according to the plurality of feature regions after the weight assignment, the cropping frame of the image is determined, wherein the cropping frame comprises the feature region with the largest weight in the plurality of feature regions.
2. The method of claim 1, wherein, According to the plurality of feature regions after the weight assignment, the cropping frame of the image is determined, including: fusing the plurality of feature regions with different weights to form a heat map, wherein different colors in the heat map represent feature regions with different weights; determining the size and position of the cropping frame according to the heat map.
3. The method of claim 1, wherein, According to the cropping frame, the image is cropped to obtain a cropped image, including: in the case that the cropping frame exceeds the image itself, determining an exceeding part of the cropping frame that exceeds the image; after the image is cropped according to the cropping frame, the exceeding part is filled to obtain the cropped image.
4. The method of claim 1, wherein, In the case that the image comprises a plurality of shot images obtained by video cutting, the method further comprises: splicing the plurality of cropped shot images to obtain a cropped video.
5. The method of claim 4, wherein, Before splicing the plurality of cropped shot images to obtain a cropped video, the method further comprises: using a spline interpolation method to smooth the cropping frame in each shot image in the plurality of shot images to obtain a smoothed cropping frame, wherein the smoothed cropping frame is used for cropping to obtain the plurality of cropped shot images.
6. The method of claim 1, wherein, The visual features of the image are extracted, including: inputting the image into an image feature model to obtain the visual features of the image, wherein the image feature model is obtained by machine learning using a first data set, and the training data in the first data set comprises an image and visual features of the image.
7. The method of claim 6, wherein, The image comprises a plurality of material images of a predetermined object, and after the image is cropped according to the cropping frame to obtain a cropped image, the method further comprises: inputting each material image in the plurality of cropped material images into an image special effect model to obtain a special effect image of each material image, wherein the image special effect model is obtained by machine learning using a second data set, and the training data in the second data set comprises a material image and a special effect image of the material image; splicing the special effect images corresponding to the plurality of material images to obtain a video of the predetermined object.
8. The method of claim 7, wherein, Splicing the special effect images corresponding to the plurality of material images to obtain a video of the predetermined object, including: Display effect images corresponding to a plurality of material images on a display interface; Receive an adjustment operation for adjusting the effect images corresponding to the plurality of material images; In response to the adjustment operation, adjust the effect images corresponding to the plurality of material images to obtain a plurality of adjusted effect images; Splice the plurality of adjusted effect images to obtain a video of the predetermined object.
9. An image processing method characterized by, Comprise: Display an image on an interactive interface; Display visual features of the image on the interactive interface, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image; Display a cropping frame of the image on the interactive interface, wherein the cropping frame is determined according to the visual features, and in the case that the visual features include a plurality of feature regions, the cropping frame of the image is determined according to the plurality of feature regions after assigning weights, the cropping frame includes a feature region with the largest weight among the plurality of feature regions, different weights are assigned to the plurality of feature regions, the size of the weight represents the importance of the feature region, and the weight is determined based on at least one of the following information: image type, area, and relationship with surrounding feature regions of the feature region; Display a cropping result on the interactive interface, wherein the cropping result is obtained by cropping the image according to the cropping frame.
10. An image processing apparatus characterized by comprising: Comprise: A first acquisition module for acquiring an image; An extraction module for extracting visual features of the image, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image; A determination module for determining a cropping frame of the image according to the visual features; A cropping module for cropping the image according to the cropping frame to obtain a cropped image; The determination module is further configured to, in the case that the visual features include a plurality of feature regions, assign different weights to the plurality of feature regions, wherein the size of the weight represents the importance of the feature region, and the weight is determined based on at least one of the following information: image type, area, and relationship with surrounding feature regions of the feature region; and determine the cropping frame of the image according to the plurality of feature regions after assigning weights, wherein the cropping frame includes a feature region with the largest weight among the plurality of feature regions.
11. An image processing apparatus characterized by comprising: Comprise: A first display module for displaying an image on an interactive interface; A second display module for displaying visual features of the image on the interactive interface, wherein the visual features include one or more feature regions on the image, and the one or more feature regions are used to embody information content of the image; The third display module is configured to display a cutting frame of the image on the interactive interface, wherein the cutting frame is determined according to the visual features, in a case where the visual features include a plurality of feature regions, the cutting frame of the image is determined according to the plurality of feature regions after the weights are assigned, the cutting frame includes a feature region with the largest weight among the plurality of feature regions, different weights are assigned to the plurality of feature regions, the weight represents the importance of the feature region, and the weight is determined based on at least one of the following: an image type of the feature region, an area, and a relationship with a peripheral feature region. The fourth display module is configured to display a cutting result on the interactive interface, wherein the cutting result is obtained by cutting the image according to the cutting frame.
12. A storage medium, characterized by The storage medium includes a stored program, wherein the program controls a device where the storage medium is located to perform the image processing method in any one of claims 1 to 9 when the program is executed.
13. A computer device, comprising: Comprise: a memory and a processor, the memory stores a computer program; the processor is configured to execute the computer program stored in the memory, and the computer program causes the processor to execute the image processing method in any one of claims 1 to 9 when the computer program is executed.
14. An image processing method, characterized by, Comprise: obtain a material of an object; identify information of the material by using an artificial intelligence recognition algorithm; perform visual processing on the material by using an artificial intelligence visual algorithm to obtain a video of the object according to the information of the material, wherein the artificial intelligence visual algorithm is the image processing method in claim 1.
15. The method of claim 14, wherein, In a case where the material is an image material, identifying the information of the material by using the artificial intelligence recognition algorithm comprises at least one of the following: identifying a category to which the image material belongs by using a category algorithm in the artificial intelligence recognition algorithm; identifying a position of a target object in the image material by using a position algorithm in the artificial intelligence recognition algorithm; identifying a salient region in the image material by using a feature algorithm in the artificial intelligence recognition algorithm; identifying an aesthetic score of the image material by using a scoring algorithm in the artificial intelligence recognition algorithm.
16. The method of claim 14, wherein, Performing visual processing on the material by using an artificial intelligence visual algorithm comprises at least one of the following: performing screening on the material, cutting a target object in an image material, arranging a plurality of types of materials, and rendering an image material by using the artificial intelligence visual algorithm.
17. The method of any one of claims 14 to 16, characterized in that, outputting an intermediate result by using an interactive device, wherein the intermediate result comprises at least one of the following: a result of the information of the material identified by using the artificial intelligence recognition algorithm, and a result of the visual processing on the material by using the artificial intelligence visual algorithm.
18. An image processing apparatus characterized by comprising: Comprise: a second acquisition module configured to acquire a material of an object; an identification module configured to identify information of the material by using an artificial intelligence recognition algorithm; A processing module is configured to perform visual processing on the material by using an artificial intelligence visual algorithm according to information of the material, to obtain a video of the object, wherein the artificial intelligence visual algorithm is the image processing method in any one of claims 1 to 9.
Citation Information
Patent Citations
Intelligent image clipping method, system and equipment and storage medium
CN109712164A
Target video generation method and system
CN111739128A
Video processing method and device
CN111866585A