Image processing method, device and product based on artificial intelligence large model
Through the image processing method based on artificial intelligence large models, users can process images simply and conveniently, solving the problem of high professional skills of mobile image processing tools, and improving flexibility and convenience.
Patent Information
- Application Number
- CN202311713873.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2043-12-13
AI Technical Summary
Mobile image processing tools require high professional skills, making it difficult for users to process images simply and easily.
The image processing method based on the artificial intelligence big model is adopted, by obtaining the target image and user requests, the artificial intelligence big model is used to analyze the images and requests, and the images are processed according to the analysis results to meet user needs.
Improves the flexibility and convenience of image processing and improves the user experience.
Smart Images

Figure CN117710527B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the field of image processing technology, and more particularly to an image processing method, device, electronic device, storage medium, and computer program product based on an artificial intelligence large model, which can be applied in image processing scenarios. Background Art
[0002] Mobile devices are now widely used, and people have a strong demand for image processing on them. However, the barrier to entry for people to process images on mobile devices is high. For example, image processing tools like Photoshop require high professional skills, making it difficult for people to process images easily and conveniently. Summary of the Invention
[0003] The present disclosure provides an image processing method, device, electronic device, storage medium and computer program product based on an artificial intelligence large model.
[0004] According to a first aspect, an image processing method based on an artificial intelligence big model is provided, comprising: obtaining a target image and a user's image processing request for the target image; parsing the target image and the image processing request respectively through the artificial intelligence big model to obtain an image parsing result and a request parsing result; and processing the target image according to the image parsing result in accordance with the processing requirements represented by the request parsing result through the artificial intelligence big model.
[0005] According to the second aspect, an image processing device based on an artificial intelligence big model is provided, including: an acquisition unit, configured to acquire a target image and a user's image processing request for the target image; a parsing unit, configured to parse the target image and the image processing request respectively through the artificial intelligence big model to obtain an image parsing result and a request parsing result; a processing unit, configured to process the target image according to the image parsing result according to the processing requirements represented by the request parsing result through the artificial intelligence big model.
[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to execute the method described in any implementation manner of the first aspect.
[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method as described in any implementation of the first aspect.
[0008] According to a fifth aspect, a computer program product is provided, comprising: a computer program, which implements the method described in any implementation manner of the first aspect when executed by a processor.
[0009] According to the technology disclosed in the present invention, an image processing method and device based on an artificial intelligence big model are provided. The user only needs to send an image processing request to the artificial intelligence big model, and the artificial intelligence big model can process the target image according to the processing requirements represented by the request analysis results and the image analysis results of the target image, thereby achieving the image processing effect expected by the user, improving the flexibility and convenience of the image processing process, and enhancing the user experience in the image processing process.
[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0012] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;
[0013] Figure 2 is a flow chart of an embodiment of an image processing method based on an artificial intelligence large model according to the present disclosure;
[0014] Figure 3 is a schematic diagram of an application scenario of the image processing method based on the artificial intelligence large model according to this embodiment;
[0015] Figure 4 is a flowchart of another embodiment of the image processing method based on the artificial intelligence large model according to the present disclosure;
[0016] Figure 5 is a structural diagram of an embodiment of an image processing device based on an artificial intelligence large model according to the present disclosure;
[0017] Figure 6 It is a schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0019] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0020] Figure 1 An exemplary architecture 100 is shown to which the image processing method and apparatus based on the artificial intelligence big model disclosed herein can be applied.
[0021] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The communication connections between terminal devices 101, 102, 103 constitute a topological network, and network 104 is used to provide a medium for communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0022] Terminal devices 101, 102, and 103 can be hardware devices or software that support network connection for data interaction and data processing. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connection, information acquisition, interaction, display, processing, and other functions, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, for example, to provide distributed services, or they can be implemented as a single software or software module. No specific limitations are given here.
[0023] Server 105 can be a server that provides various services, such as a backend processing server that receives a target image and a user's image processing request sent by terminal devices 101, 102, and 103, and processes the target image based on the image analysis results of the target image according to the processing requirements represented by the request analysis results using a large artificial intelligence model. As an example, server 105 can be a cloud server.
[0024] It should be noted that the server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (e.g., software or software modules for providing distributed services), or as a single software or software module. No specific limitations are given here.
[0025] It should also be noted that the image processing method based on the artificial intelligence large model provided by the embodiments of the present disclosure can be executed by a server, can be executed by a terminal device, or can be executed by the server and the terminal device in cooperation with each other. Accordingly, the various parts (e.g., various units) included in the image processing device based on the artificial intelligence large model can be all set in the server, can be all set in the terminal device, or can be set separately in the server and the terminal device.
[0026] It should be understood that Figure 1 The number of terminal devices, networks, and servers is merely illustrative. Any number of terminal devices, networks, and servers may be provided as needed. When the electronic device on which the image processing method based on the artificial intelligence large model is running does not need to transmit data with other electronic devices, the system architecture may only include the electronic device (e.g., terminal device or server) on which the image processing method based on the artificial intelligence large model is running.
[0027] Please refer to Figure 2 , Figure 2 This is a flow chart of an image processing method based on an artificial intelligence large model provided in an embodiment of the present disclosure. In process 200, the following steps are included:
[0028] Step 201: Obtain a target image and a user's image processing request for the target image.
[0029] In this embodiment, the execution subject of the image processing method based on the artificial intelligence large model (for example, Figure 1 The terminal device or server in the system can obtain the target image and the user's image processing request for the target image remotely or locally through a wired network connection or a wireless network connection.
[0030] The target image is the image selected by the user to be processed. It can be a frame in a dynamic image sequence or a single static image. The target image can be of various types, classified based on various classification criteria. Under the color-based classification criteria, images include, but are not limited to, black and white images and color images; under the content-based classification criteria, images include, but are not limited to, images of people, landscapes, architecture, pets, products, sports, art, science fiction, and comics.
[0031] An image processing request represents a user's request to process a target image. It can be issued through various methods supported by existing or future technologies, including but not limited to voice, text, and actions (such as gestures). For example, an image processing request might be to delete the XX object in the target image or to merge multiple target images into a single image.
[0032] In step 202, the target image and the image processing request are analyzed respectively by the artificial intelligence large model to obtain the image analysis result and the request analysis result.
[0033] In this embodiment, the above-mentioned execution entity can respectively analyze the target image and the image processing request through the artificial intelligence large model to obtain the image analysis result and the request analysis result.
[0034] AI big models specifically refer to pre-trained AI models with large or ultra-large parameters. The term "AI big model" encompasses two concepts: "pre-training" and "big model." The combination of these two creates a new AI paradigm: once pre-trained on large datasets, the model can directly support a wide range of applications with minimal or no fine-tuning. These AI big models possess exceptional contextual understanding, language generation, learning capabilities, and transferability.
[0035] The above-mentioned execution entity can input the target image and image processing request into the artificial intelligence big model in sequence, parse the target image through the artificial intelligence big model to obtain the image parsing result; parse the image processing request through the artificial intelligence big model to obtain the request parsing result.
[0036] As an example, the artificial intelligence large model can extract features of the target image to obtain a feature map; then, based on the feature map, it can determine the semantic information, scene information, etc. represented by the target image to obtain the image analysis result.
[0037] The artificial intelligence large model can extract features from image processing requests to obtain feature data; then, based on the feature data, it can determine the processing requirements represented by the image processing request and obtain the request parsing results.
[0038] In some optional implementations of this embodiment, the above-mentioned execution entity can perform the target image parsing operation in the following manner: parse the target image through an artificial intelligence large model, determine the objects included in the target image and the positions of the objects, and obtain the image parsing results.
[0039] The objects may be all objects included in the target image, and the positions of the objects are generally represented by the area range in the image.
[0040] As an example, the artificial intelligence large model can extract features from the target image to obtain a feature map; then, based on the feature map, it can determine each object and the position of each object included in the target image to obtain an image analysis result.
[0041] In this implementation, the image analysis results include the objects in the target image and the positions of the objects, which provides data preparation for the subsequent artificial intelligence large model-based processing of the user's needs for the objects in the target image, and helps to improve the response speed and accuracy of the image processing process based on the artificial intelligence large model.
[0042] In some optional implementations of this embodiment, the above-mentioned execution entity can further perform the target image parsing operation in the following manner: parse the target image through an artificial intelligence large model, determine the objects included in the target image, the positions of the objects and the interaction relationships between the objects, and obtain the image parsing results.
[0043] When a target image includes multiple objects, there may be interactions between the multiple objects. The interaction relationships between the multiple objects include but are not limited to occlusion relationships, collision relationships, dependency relationships, spatial relationships, temporal relationships, combination relationships, and background relationships.
[0044] The artificial intelligence large model analyzes the target image and, based on determining each object and the position of each object included in the target image, can further determine the interaction relationship between each two objects to obtain the image analysis result.
[0045] When a user requests processing of a target object in a target image, adjustments to the target object generally affect other objects that interact with it. In this implementation, the image analysis results include the objects in the target image, their locations, and the interactions between them. This provides data preparation for subsequent AI-based large-scale model processing to meet the user's processing needs for the target object, helping to further improve the response speed and accuracy of the AI-based large-scale image processing process.
[0046] In some optional implementations of this embodiment, the above-mentioned execution entity can perform the parsing operation of the image processing request in the following manner: parse the image processing request through an artificial intelligence large model, determine the user's processing requirements for the target image and / or objects in the target image, and obtain the request parsing result.
[0047] In this implementation, the artificial intelligence big model can perform feature extraction on the image processing request to obtain feature data; and then determine the user's processing requirements for the target image and / or objects in the target image represented by the image processing request based on the feature data to obtain the request parsing result.
[0048] For example, the processing requirement may be a processing requirement for at least one object in the target image. Under this processing requirement, the artificial intelligence large model only needs to process at least one object. For another example, the processing requirement may be a processing requirement for the entire target image. Under this processing requirement, the artificial intelligence large model needs to process the entire target image, for example, perform super-resolution processing on the target image.
[0049] In this implementation, the artificial intelligence large model can clearly determine the object to be processed (target image and / or object in the target image) targeted by the image processing request, thereby improving the accuracy of the request parsing results and helping to improve the subsequent processing effect of the target image and the user experience.
[0050] Step 203: Process the target image according to the image analysis results through the artificial intelligence big model in accordance with the processing requirements represented by the request analysis results.
[0051] In this embodiment, the above-mentioned execution entity can use the artificial intelligence big model to process the target image according to the image analysis results in accordance with the processing requirements represented by the request analysis results.
[0052] As an example, the processing requirement represented by the request analysis result is to adjust the style of the target image to the style of a specific image, where the specific image can be an image with a certain style specified by the user; the artificial intelligence large model can analyze the style information of the specific image, and adjust the style of the target image based on the image analysis result to obtain the target image after style adjustment.
[0053] In some optional implementations of this embodiment, the processing requirement is a first attribute adjustment requirement for attribute information of the target image. The attribute information of the target image includes, but is not limited to, resolution, size, color, bit depth, transparency, grayscale, brightness, contrast, saturation, and other information.
[0054] For example, the first attribute adjustment requirement is a super-resolution requirement, a zoom-in or zoom-out requirement, or a contrast adjustment requirement for the target object.
[0055] In this implementation, the execution entity may perform step 203 as follows: by using an artificial intelligence big model, adjusting the attribute information of the target image according to the image analysis result in accordance with the first attribute adjustment requirement represented by the request analysis result.
[0056] As an example, the execution entity can use the AI model to analyze the request results and determine the attribute information and adjustment direction required for the first attribute adjustment request. Then, the attribute information of the target image can be adjusted in accordance with the adjustment direction. For example, for attributes such as resolution, size, and transparency, the adjustment direction includes increasing or decreasing; for attributes such as color, the adjustment direction includes lightening, darkening, or adjusting to another color.
[0057] In this implementation, an image processing method for the attribute information of a target image is provided under the requirement of attribute adjustment, which improves the convenience and flexibility of the adjustment process of the attribute information of the target image based on the artificial intelligence large model.
[0058] In some optional implementations of this embodiment, the processing requirement is a second attribute adjustment requirement for attribute information of a target object in a target image. The target object may be at least one object in the target image. Similarly, the attribute information of the target object includes, but is not limited to, information such as resolution, size, color, bit depth, transparency, grayscale, brightness, contrast, and saturation.
[0059] As an example, the above-mentioned execution entity can determine the target object, attribute information of the target object and the adjustment direction of the attribute information to be adjusted according to the request analysis results through the artificial intelligence big model; and then adjust the attribute information of the target object in the target image according to the adjustment direction.
[0060] When adjusting the target object's attribute information affects other objects with which it interacts, after adjusting the target object's attribute information in the target image, the other objects with which it interacts need to be adjusted based on the interaction relationship. For example, if you enlarge target object A in the target image, the occluded portion of object B, which is in an occlusion relationship with target object A, will become larger. This requires adjusting the occluded portion of object B, as well as the interaction between target objects A and B (e.g., the projection of target object A on object B).
[0061] In this implementation, an image processing method for the attribute information of the target object in the target image is provided under the requirement of attribute adjustment, which improves the convenience and flexibility of the adjustment process of the attribute information of the target object in the target image based on the artificial intelligence large model.
[0062] In some optional implementations of this embodiment, the processing request is an object adjustment request for a target object in the target image. The target object may be at least one object in the target image. The object adjustment request includes, but is not limited to, a request to move, delete, or replace the target object.
[0063] In the move requirement, the user expects to move the target object in the target image and change the structural relationship between the target objects in the target image without changing the number of objects; in the delete requirement, the user expects to delete the target object in the target image and change the structural relationship between the target objects in the target image while changing the number of objects; in the replace requirement, the user expects to replace the target object in the target image by specifying the object.
[0064] As an example, the above-mentioned execution entity can use the artificial intelligence big model to parse the request results to determine the target object to be adjusted and the adjustment method for the target object; then, adjust the target object in the target image according to the adjustment method.
[0065] When adjustments to the target object affect other objects that interact with it, after adjusting the target object in the target image, the other objects also need to be adjusted based on the interactive relationships. For example, if target object A in the target image is moved from a first position to a second position, target object A has an occlusion relationship with object B at the first position, and an occlusion relationship with object C at the second position after the move. In this case, the interactive portion between target object A and object B, as well as the interactive portion between target object A and object C, need to be adjusted.
[0066] In this implementation, an image processing method for a target object in a target image is provided under the requirement of object adjustment, which improves the convenience and flexibility of the adjustment process of the target object in the target image based on the artificial intelligence large model.
[0067] In some optional implementations of this embodiment, for a case where the object adjustment requirement is a requirement to delete a target object in a target image, the execution entity may perform the object adjustment process based on the object adjustment requirement in the following manner:
[0068] First, through the artificial intelligence large model, according to the object deletion requirements represented by the request analysis results, the target object in the target image is deleted according to the image analysis results.
[0069] As an example, the above-mentioned execution entity can determine the target object that is desired to be deleted in the object deletion requirement through an artificial intelligence large model, and determine whether the target object includes the target object from the image analysis result; in response to determining that the target object is included in the target image, the target object in the target image is deleted according to the position of the target object.
[0070] Second, the deleted area is filled with pixels based on the pixels in the surrounding area of the deleted area.
[0071] The deleted area is the area obtained after deleting the target object in the target image.
[0072] The deleted area is generally filled with blank pixels. Based on the pixels in the surrounding area of the deleted area, the artificial intelligence model predicts the background information of the original target object, and then uses pixel filling to make the deleted area present the background information blocked by the original target object.
[0073] Third, objects other than the target object in the target image are adjusted according to the interaction relationship.
[0074] Before the target object is deleted, there may be other objects that interact with it. In this case, after deleting the target object and filling the deleted area, you need to adjust the objects in the target image that interact with the target object. For example, before deleting target object A, object B has its shadow on it. After deleting target object A, you need to adjust the shadow on object B to make it disappear.
[0075] In this implementation, an image processing method for a target object in a target image is provided under the requirement of object deletion, which improves the convenience and flexibility of the deletion process of the target object in the target image based on the artificial intelligence large model.
[0076] In some optional implementations of this embodiment, the processing requirement is to add a specified object to the target image. The specified object can be any object, such as a cartoon image or a real person object. The number of the specified objects can be one or more.
[0077] The execution subject or an electronic device communicatively connected to the execution subject may be provided with a designated object library including a plurality of designated objects. When the processing requirement is determined to be an object addition requirement for adding a designated object to the target image, the artificial intelligence large model may determine the designated object to be added from the designated object library.
[0078] Users can add specific objects to the designated object library by adding them to the library. For example, users can use the AI model to segment the image or video frame they are browsing, and then add the selected target objects to the designated object library.
[0079] In this implementation, the execution subject can perform the image processing process based on the processing requirements in the following manner:
[0080] First, through the artificial intelligence large model, according to the object increase requirements represented by the request analysis result, the specified object is added to the target image.
[0081] As an example, the artificial intelligence big model is used to determine the designated object to be added in the object addition requirement and the location where the designated object is to be added; then, the designated object is added at the determined location where the object is to be added.
[0082] Second, the specified object and the object in the target image are adjusted according to the image parsing results.
[0083] As an example, the AI model adjusts the specified object and the object in the target image in the following way:
[0084] According to the attribute information of the object in the target image in the image analysis result, the attribute information of the added designated object is adjusted so that the attribute information of the added designated object is compatible with the attribute information of the object in the target image.
[0085] According to the interactive relationship between the objects in the target image in the image analysis result, the projection relationship, transition information, etc. of the interactive part between the added specified object and the object in the target image are adjusted.
[0086] In this implementation, an image processing method is provided for adding a specified object to a target image when an object is added. The convenience and flexibility of the process of adding a specified object to a target image are improved based on a large artificial intelligence model.
[0087] In some optional implementations of this embodiment, the processing requirement is an animation production requirement for a target image. In this implementation, the execution entity may perform step 203 as follows: using the artificial intelligence large model, in accordance with the animation production requirement represented by the request analysis result, and generating a dynamic image sequence corresponding to the target image based on the image analysis result.
[0088] As an example, the above-mentioned execution entity can use an artificial intelligence large model to predict the previous or future behavior trajectory or morphological information of each object in the target image based on the semantic information and scene information represented by the target image, and generate a predicted image of the target object at the predicted time based on the predicted behavior trajectory or morphological information at each preset time interval; then, according to the temporal relationship between multiple predicted images, multiple predicted images and the target image are combined to generate a dynamic image sequence corresponding to the target image.
[0089] In this implementation, an image processing method for target images is provided under the requirements of animation production, which improves the convenience and flexibility of the animation production process based on the artificial intelligence big model.
[0090] In some optional implementations of this embodiment, the execution subject may execute the process of generating the dynamic image sequence in the following manner:
[0091] First, through the artificial intelligence big model, according to the animation production requirements represented by the request analysis results, determine the animation type corresponding to each object in the image analysis results.
[0092] For example, the target image includes objects such as the sea, blue sky, white clouds, and flying birds. The animation production requirement expects the waves and white clouds to move with the wind and the birds to fly away.
[0093] The execution entity can determine each object indicated by the request parsing result and the animation type expected by each object through the artificial intelligence big model.
[0094] Second, for each object in the image parsing result, a dynamic object is generated according to the animation type corresponding to the object, a dynamic image sequence is generated, and the dynamic object is adjusted according to the interaction relationship.
[0095] As an example, for each object in the image analysis results, the execution entity can use the artificial intelligence model to predict the object's previous or future behavior trajectory and / or morphological information based on the object's corresponding animation type, and then sample the corresponding behavior trajectory and / or morphological information at multiple time points at preset time intervals to generate a dynamic object. After obtaining each dynamic object indicated by the animation production requirements, the multiple dynamic objects are merged to obtain a dynamic image sequence.
[0096] When multiple dynamic objects have interactive relationships, it is necessary to adjust the multiple dynamic objects based on the interactive relationships, for example, adjusting the projection relationship and collision relationship between the multiple dynamic objects.
[0097] In this implementation, an image processing method for target images is provided under the requirements of animation production, which further improves the convenience and flexibility of the animation production process based on the artificial intelligence big model.
[0098] In some optional implementations of this embodiment, the processing requirement is an image synthesis requirement for multiple target images. The image synthesis requirement includes, but is not limited to, splicing the multiple target images into a composite image in a certain manner, determining at least a partial area from each target image to assemble the composite image, and the like.
[0099] The execution subject may execute step 203 as follows: through a large artificial intelligence model, in accordance with the image synthesis requirements represented by the requested analysis results, the plurality of target images may be processed according to the image analysis results corresponding to each of the plurality of target images to obtain a composite image.
[0100] Taking the example of image synthesis requiring stitching multiple target images into a composite image in a certain way, the above-mentioned execution entity can determine the stitching method and stitching order of the multiple target images through the artificial intelligence big model, and then stitch the multiple target images according to the stitching method and stitching order to obtain a composite image.
[0101] Taking the example of an image synthesis requirement of determining at least a partial area from each target image to piece together a composite image, the above-mentioned execution entity can determine the pieced-to parts of each target image, as well as the pieced-to order and pieced-to method of the multiple pieced-to parts through an artificial intelligence large model based on the image analysis results and image synthesis requirements corresponding to each of the multiple target images, and then piece together the multiple target images according to the pieced-to order and pieced-to method to obtain a composite image.
[0102] In this implementation, an image processing method for the target image is provided under the image synthesis requirement, which further improves the convenience and flexibility of the image synthesis process based on the artificial intelligence large model.
[0103] In some optional implementations of this embodiment, the execution entity may perform the image synthesis process in the following manner:
[0104] First, through the artificial intelligence large model, according to the image synthesis requirements represented by the request analysis results, the object group composed of objects fused between multiple target images is determined.
[0105] For example, for each object in each target image, the execution entity can use the artificial intelligence model to randomly determine objects from other target images that are fused with the object to form an object group. For another example, based on a user's selection operation, the execution entity can determine objects that are fused with each other across multiple target images to form an object group.
[0106] Second, through the artificial intelligence large model, according to the image synthesis requirements, the objects in each object group in at least one object group are fused to obtain a composite image.
[0107] Fusion is, for example, fusing the features of each object in the object group to obtain a fused new object, and then forming the new object into a composite image.
[0108] In this implementation, another image processing method for the target image is provided under the image synthesis requirement, which enriches the image processing method and further improves the convenience and flexibility of the image synthesis process based on the artificial intelligence large model.
[0109] In some optional implementations of this embodiment, the processing requirement is a video synthesis requirement for multiple target images. In this implementation, the execution entity may perform step 203 as follows: using the artificial intelligence large model, in accordance with the video synthesis requirement represented by the request analysis result, and based on the image analysis results corresponding to the multiple target images, determine multiple to-be-processed images from the multiple target images, and generate a synthesized video based on the multiple to-be-processed images.
[0110] As an example, through the artificial intelligence large model, the selection method and synthesis order of the images to be processed can be determined according to the video synthesis requirements, and then multiple images to be processed can be determined from multiple target images according to the selection method, and a synthetic video can be generated based on the multiple images to be processed according to the synthesis order.
[0111] The selection method includes, but is not limited to, a selection method based on the shooting time corresponding to the target image and a selection method based on the object in the target image.
[0112] In this implementation, an image processing method for the target image is provided under the requirements of video synthesis, which enriches the image processing method and further improves the convenience and flexibility of the video synthesis process based on the artificial intelligence large model.
[0113] In some optional implementations of this embodiment, the processing requirement includes multiple processing sub-requirements for the target image. Each of the multiple processing sub-requirements can be the first attribute adjustment requirement, the second attribute adjustment requirement, the object adjustment requirement, the object addition requirement, the animation production requirement, the image synthesis requirement, or the video synthesis requirement.
[0114] In this implementation, the execution entity may perform step 203 as follows:
[0115] First, for each of the multiple processing sub-requirements, the target image is processed according to the processing sub-requirement through the artificial intelligence large model to obtain a processed image.
[0116] Second, through a large artificial intelligence model, the target video is generated based on multiple processed images.
[0117] As an example, the execution subject may combine multiple processed images obtained based on multiple processing sub-requirements in the order of the multiple processing sub-requirements to generate a target video.
[0118] As another example, after obtaining multiple processed images, the execution entity may receive a sequence-specifying operation from the user, and then combine the multiple processed images according to the sequence determined by the sequence-specifying operation to generate a target video.
[0119] In this implementation, an image processing method is provided for multiple processing sub-requirements of the target image, which enriches the image processing method and further improves the convenience and flexibility of image processing based on the artificial intelligence large model.
[0120] Continue to see Figure 3 , Figure 3 FIG3 is a schematic diagram 300 of an application scenario of the image processing method based on the artificial intelligence large model according to this embodiment. Figure 3In an application scenario, user 301 selects a target image from the photo album of mobile terminal 302 and enters an image processing request for the target image on mobile terminal 302. After obtaining the target image and user 301's image processing request for the target image, terminal device 302 sends the target image and image processing request to server 303. After determining the target image and the user's image processing request for the target image, the server first uses a large artificial intelligence model to parse the target image and image processing request, respectively, to obtain an image parsing result and a request parsing result. Then, the large artificial intelligence model processes the target image according to the image parsing result, based on the processing requirements represented by the request parsing result, and generates and feeds back a processed image or video to terminal device 302.
[0121] In this embodiment, an image processing method and device based on an artificial intelligence big model are provided. The user only needs to send an image processing request to the artificial intelligence big model. The artificial intelligence big model can process the target image according to the processing requirements represented by the request analysis results and the image analysis results of the target image, thereby achieving the image processing effect expected by the user, improving the flexibility and convenience of the image processing process, and enhancing the user experience in the image processing process.
[0122] Continue to refer Figure 4 , shows a schematic process 400 of another embodiment of the image processing method based on the artificial intelligence large model according to the present disclosure. In the process 400, the following steps are included:
[0123] Step 401: Obtain a target image and a user's image processing request for the target image.
[0124] Step 402: parse the target image using the artificial intelligence model to determine the objects included in the target image, the positions of the objects, and the interaction relationships between the objects to obtain an image parsing result.
[0125] Step 403: parse the image processing request through the artificial intelligence large model to determine the user's processing requirements for the target image and / or the object in the target image, and obtain the request parsing result.
[0126] Step 404: Process the target image according to the image analysis results through the artificial intelligence big model in accordance with the processing requirements represented by the request analysis results.
[0127] The processing requirement may be the first attribute adjustment requirement, the second attribute adjustment requirement, the object adjustment requirement, the object addition requirement, the animation production requirement, the image synthesis requirement, or the video synthesis requirement.
[0128] It can be seen from this embodiment that Figure 2Compared with the corresponding embodiments, the process 400 of the image processing method based on the artificial intelligence big model in this embodiment specifically illustrates the image parsing process based on the artificial intelligence big model and the request parsing process of the image processing request, accurately determining the image parsing results and the request parsing results, and helping to further improve the flexibility and convenience of the image processing process and enhance the user experience in the image processing process.
[0129] Continue to refer Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image processing device based on an artificial intelligence large model. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0130] like Figure 5 As shown, the image processing device 500 based on the artificial intelligence big model includes: an acquisition unit 501, configured to acquire a target image and a user's image processing request for the target image; a parsing unit 502, configured to parse the target image and the image processing request respectively through the artificial intelligence big model to obtain an image parsing result and a request parsing result; a processing unit 503, configured to process the target image according to the image parsing result according to the processing requirements represented by the request parsing result through the artificial intelligence big model.
[0131] In some optional implementations of this embodiment, the parsing unit 502 is further configured to: parse the target image through the artificial intelligence large model, determine the objects included in the target image and the positions of the objects, and obtain an image parsing result.
[0132] In some optional implementations of this embodiment, the parsing unit 502 is further configured to: parse the target image through an artificial intelligence large model, determine the objects included in the target image, the positions of the objects, and the interaction relationships between the objects, and obtain an image parsing result.
[0133] In some optional implementations of this embodiment, the parsing unit 502 is further configured to: parse the image processing request through the artificial intelligence large model, determine the user's processing requirements for the target image and / or the object in the target image, and obtain the request parsing result.
[0134] In some optional implementations of this embodiment, the processing requirement is a first attribute adjustment requirement for the attribute information of the target image, and the processing unit 503 is further configured to: adjust the attribute information of the target image according to the image analysis result according to the first attribute adjustment requirement represented by the request analysis result through the artificial intelligence big model.
[0135] In some optional implementations of this embodiment, the processing requirement is a second attribute adjustment requirement for the attribute information of the target object in the target image, and the processing unit 503 is further configured to: through the artificial intelligence big model, according to the second attribute adjustment requirement represented by the request parsing result, adjust the attribute information of the target object in the target image according to the image parsing result.
[0136] In some optional implementations of this embodiment, the processing requirement is an object adjustment requirement for a target object in a target image, and the processing unit 503 is further configured to: adjust the target object in the target image according to the image analysis result through the artificial intelligence big model, in accordance with the object adjustment requirement represented by the request analysis result.
[0137] In some optional implementations of this embodiment, the object adjustment requirement is an object deletion requirement for the target object in the target image, and the processing unit 503 is further configured to: delete the target object in the target image according to the image analysis result through the artificial intelligence big model in accordance with the object deletion requirement represented by the request analysis result; fill the deleted area with pixels according to the pixels in the surrounding area of the deleted area, wherein the deleted area is the area obtained after deleting the target object in the target image; and adjust objects other than the target object in the target image according to the interaction relationship.
[0138] In some optional implementations of this embodiment, the processing requirement is an object addition requirement for adding a specified object in the target image, and the processing unit 503 is further configured to: add the specified object to the target image according to the object addition requirement represented by the request parsing result through the artificial intelligence big model; and adjust the specified object and the object in the target image according to the image parsing result.
[0139] In some optional implementations of this embodiment, the processing requirement is an animation production requirement for the target image, and the processing unit 503 is further configured to: generate a dynamic image sequence corresponding to the target image according to the image analysis results through the artificial intelligence big model in accordance with the animation production requirement represented by the request analysis result.
[0140] In some optional implementations of this embodiment, the processing unit 503 is further configured to: determine the animation type corresponding to each object in the image analysis result according to the animation production requirements represented by the request analysis result through the artificial intelligence big model; for each object in the image analysis result, generate a dynamic object according to the animation type corresponding to the object, and adjust the dynamic object according to the interactive relationship.
[0141] In some optional implementations of this embodiment, the processing requirement is an image synthesis requirement for multiple target images, and the processing unit 503 is further configured to: through the artificial intelligence big model, according to the image synthesis requirement represented by the request analysis result, process the multiple target images according to the image analysis results corresponding to each of the multiple target images to obtain a synthesized image.
[0142] In some optional implementations of this embodiment, the processing unit 503 is further configured to: determine, through the artificial intelligence big model, an object group consisting of objects fused together between multiple target images according to the image synthesis requirements represented by the request parsing results; and fuse, through the artificial intelligence big model, the objects in each object group in at least one object group according to the image synthesis requirements to obtain a composite image.
[0143] In some optional implementations of this embodiment, the processing requirement is a video synthesis requirement for multiple target images, and the processing unit 503 is further configured to: through the artificial intelligence big model, according to the video synthesis requirement represented by the request analysis result, according to the image analysis results corresponding to each of the multiple target images, determine multiple images to be processed from the multiple target images, and generate a synthetic video based on the multiple images to be processed.
[0144] In some optional implementations of this embodiment, the processing requirements include multiple processing sub-requirements for the target image, and the processing unit 503 is further configured to: for each of the multiple processing sub-requirements, process the target image according to the processing sub-requirement through the artificial intelligence big model to obtain a processed image; and generate a target video based on the multiple processed images through the artificial intelligence big model.
[0145] In this embodiment, an image processing device based on an artificial intelligence big model is provided. The user only needs to send an image processing request to the artificial intelligence big model. The artificial intelligence big model can process the target image according to the processing requirements represented by the request analysis results and the image analysis results of the target image, thereby achieving the image processing effect expected by the user, improving the flexibility and convenience of the image processing process, and enhancing the user experience in the image processing process.
[0146] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the image processing method based on the artificial intelligence large model described in any of the above embodiments when executing.
[0147] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the image processing method based on the artificial intelligence large model described in any of the above embodiments when executed.
[0148] The embodiments of the present disclosure provide a computer program product, which, when executed by a processor, can implement the image processing method based on the artificial intelligence large model described in any of the above embodiments.
[0149] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0150] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0151] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0152] The computing unit 601 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as an image processing method based on an artificial intelligence large model. For example, in some embodiments, the image processing method based on an artificial intelligence large model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the image processing method based on the artificial intelligence large model described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute an image processing method based on an artificial intelligence large model in any other appropriate manner (e.g., by means of firmware).
[0153] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0154] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0155] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0157] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0158] A computer system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private servers (VPS). It may also be a server in a distributed system or a server integrated with blockchain.
[0159] According to the technical solution of the embodiments of the present disclosure, an image processing method and device based on an artificial intelligence big model are provided. The user only needs to send an image processing request to the artificial intelligence big model, and the artificial intelligence big model can process the target image according to the processing requirements represented by the request analysis results and the image analysis results of the target image, thereby achieving the image processing effect expected by the user, improving the flexibility and convenience of the image processing process, and enhancing the user experience in the image processing process.
[0160] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided by this disclosure can be achieved. This is not a limitation herein.
[0161] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. An image processing method based on an artificial intelligence large model, comprising: Acquire a target image and a user's image processing request for the target image; Parsing the target image and the image processing request respectively using an artificial intelligence large model to obtain an image parsing result and a request parsing result, wherein the image parsing result is used to characterize objects included in the target image, positions of objects, and interactions between objects, wherein the interactions include occlusion relationships, collision relationships, dependency relationships, spatial relationships, temporal relationships, combination relationships, and background relationships; Through the artificial intelligence big model, in accordance with the processing requirements represented by the request parsing result, the target image is processed according to the image parsing result to obtain a processing result in image form or video form, wherein, during the image processing process, when there is a processed object among multiple objects with an interactive relationship, other objects among the multiple objects are adaptively processed, and the processing requirements include at least one of the following: a first attribute adjustment requirement for the attribute information of the target image, a second attribute adjustment requirement for the attribute information of the target object in the target image, an object adjustment requirement for the target object in the target image, an animation production requirement for the target image, an image synthesis requirement for multiple target images, and a video synthesis requirement for multiple target images. The object adjustment requirement is an object addition requirement for adding a specified object to the target image, and / or an object deletion requirement for the target object in the target image.
2. The method according to claim 1, wherein The target image is parsed by the artificial intelligence model to obtain image parsing results, including: The target image is parsed by the artificial intelligence large model to determine the objects included in the target image and the positions of the objects, thereby obtaining the image parsing result.
3. The method according to claim 2, wherein: The step of parsing the target image by the artificial intelligence large model, determining the objects included in the target image and the positions of the objects, and obtaining the image parsing result includes: The target image is parsed by the artificial intelligence large model to determine the objects included in the target image, the positions of the objects and the interaction relationships between the objects, and obtain the image parsing result.
4. The method according to claim 1, wherein The image processing request is parsed by the artificial intelligence model to obtain the request parsing result, including: The image processing request is parsed by the artificial intelligence large model to determine the user's processing requirements for the target image and / or the object in the target image, and obtain the request parsing result.
5. The method according to claim 4, wherein The processing requirement is a first attribute adjustment requirement for the attribute information of the target image, and The step of processing the target image according to the image analysis result by the artificial intelligence large model and the processing requirements represented by the request analysis result includes: Through the artificial intelligence big model, according to the first attribute adjustment requirement represented by the request analysis result, the attribute information of the target image is adjusted according to the image analysis result.
6. The method according to claim 4, wherein: The processing requirement is a second attribute adjustment requirement for the attribute information of the target object in the target image, and The step of processing the target image according to the image analysis result by the artificial intelligence large model and the processing requirements represented by the request analysis result includes: Through the artificial intelligence big model, according to the second attribute adjustment requirement represented by the request analysis result, the attribute information of the target object in the target image is adjusted according to the image analysis result.
7. The method according to claim 4, wherein: The processing requirement is an object adjustment requirement for a target object in the target image, and The step of processing the target image according to the image analysis result by the artificial intelligence large model and the processing requirements represented by the request analysis result includes: Through the artificial intelligence big model, according to the object adjustment requirements represented by the request analysis result, the target object in the target image is adjusted according to the image analysis result.
8. The method according to claim 7, wherein: The object adjustment requirement is an object deletion requirement for the target object in the target image, and The step of adjusting the target object in the target image according to the image analysis result by using the artificial intelligence large model and the object adjustment requirement represented by the request analysis result includes: Deleting the target object in the target image according to the image analysis result by the artificial intelligence large model according to the object deletion requirement represented by the request analysis result; Filling the deleted area with pixels according to pixels in an area surrounding the deleted area, wherein the deleted area is an area obtained after deleting the target object from the target image; According to the interactive relationship, objects other than the target object in the target image are adjusted.
9. The method according to claim 4, wherein: The processing requirement is an object addition requirement for adding a specified object to the target image, and The step of processing the target image according to the image analysis result by the artificial intelligence large model and the processing requirements represented by the request analysis result includes: By using the artificial intelligence large model, according to the object addition requirement represented by the request parsing result, the specified object is added to the target image; The designated object and the object in the target image are adjusted according to the image analysis result.
10. The method according to claim 4, wherein: The processing requirement is an animation production requirement for the target image, and The step of processing the target image according to the image analysis result by the artificial intelligence large model and the processing requirements represented by the request analysis result includes: Through the artificial intelligence big model, according to the animation production requirements represented by the request analysis result, a dynamic image sequence corresponding to the target image is generated according to the image analysis result.
11. The method according to claim 10, wherein: The step of generating a dynamic image sequence corresponding to the target image according to the image analysis result by using the artificial intelligence large model and in accordance with the animation production requirements represented by the request analysis result includes: Determining, by means of the artificial intelligence large model, the animation type corresponding to each object in the image analysis result according to the animation production requirements represented by the request analysis result; For each object in the image analysis result, a dynamic object is generated according to the animation type corresponding to the object, the dynamic image sequence is generated, and the dynamic object is adjusted according to the interactive relationship.
12. The method according to claim 4, wherein: The processing requirement is an image synthesis requirement for multiple target images, and The step of processing the target image according to the image analysis result by the artificial intelligence large model and the processing requirements represented by the request analysis result includes: Through the artificial intelligence big model, according to the image synthesis requirements represented by the request analysis results, the multiple target images are processed according to the image analysis results corresponding to each of the multiple target images to obtain a synthesized image.
13. The method according to claim 12, wherein: The step of processing the plurality of target images by the artificial intelligence large model according to the image synthesis requirement represented by the request analysis result and according to the image analysis results corresponding to each of the plurality of target images to obtain a synthesized image includes: Determining, by the artificial intelligence large model, an object group consisting of objects fused between the multiple target images according to the image synthesis requirements represented by the request analysis result; Through the artificial intelligence large model, according to the image synthesis requirements, the objects in each object group in at least one object group are fused to obtain the synthesized image.
14. The method according to claim 4, wherein: The processing requirement is a video synthesis requirement for multiple target images, and The step of processing the target image according to the image analysis result by the artificial intelligence large model and the processing requirements represented by the request analysis result includes: Through the artificial intelligence big model, according to the video synthesis requirements represented by the request analysis results, and based on the image analysis results corresponding to each of the multiple target images, multiple images to be processed are determined from the multiple target images, and a synthetic video is generated based on the multiple images to be processed.
15. The method according to any one of claims 1 to 14, wherein The processing requirement includes a plurality of processing sub-requirements for the target image, and The step of processing the target image according to the image analysis result by the artificial intelligence large model and the processing requirements represented by the request analysis result includes: For each of the plurality of processing sub-requirements, the target image is processed according to the processing sub-requirement by the artificial intelligence macromodel to obtain a processed image; The target video is generated based on multiple processed images through the artificial intelligence large model.
16. An image processing device based on an artificial intelligence large model, comprising: an acquiring unit configured to acquire a target image and a user's image processing request for the target image; a parsing unit configured to parse the target image and the image processing request respectively using an artificial intelligence large model to obtain an image parsing result and a request parsing result, wherein the image parsing result is used to characterize objects included in the target image, positions of objects, and interactions between objects, wherein the interactions include occlusion relationships, collision relationships, dependency relationships, spatial relationships, temporal relationships, combination relationships, and background relationships; The processing unit is configured to process the target image according to the image analysis result through the artificial intelligence large model in accordance with the processing requirements represented by the request analysis result, and obtain a processing result in image form or video form, wherein, during the image processing process, when there is a processed object among multiple objects with an interactive relationship, other objects among the multiple objects are adaptively processed, and the processing requirements include at least one of the following: a first attribute adjustment requirement for the attribute information of the target image, a second attribute adjustment requirement for the attribute information of the target object in the target image, an object adjustment requirement for the target object in the target image, an animation production requirement for the target image, an image synthesis requirement for multiple target images, and a video synthesis requirement for multiple target images. The object adjustment requirement is an object addition requirement for adding a specified object to the target image, and / or an object deletion requirement for the target object in the target image.
17. The device according to claim 16, wherein The parsing unit is further configured to: The target image is parsed by the artificial intelligence large model to determine the objects included in the target image and the positions of the objects, thereby obtaining the image parsing result.
18. The device according to claim 17, wherein The parsing unit is further configured to: The target image is parsed by the artificial intelligence large model to determine the objects included in the target image, the positions of the objects and the interaction relationships between the objects, and obtain the image parsing result.
19. The device according to claim 16, wherein The parsing unit is further configured to: The image processing request is parsed by the artificial intelligence large model to determine the user's processing requirements for the target image and / or the object in the target image, and obtain the request parsing result.
20. The device according to claim 19, wherein The processing requirement is a first attribute adjustment requirement for the attribute information of the target image, and The processing unit is further configured to: Through the artificial intelligence big model, according to the first attribute adjustment requirement represented by the request analysis result, the attribute information of the target image is adjusted according to the image analysis result.
21. The apparatus according to claim 19, wherein The processing requirement is a second attribute adjustment requirement for the attribute information of the target object in the target image, and The processing unit is further configured to: Through the artificial intelligence big model, according to the second attribute adjustment requirement represented by the request analysis result, the attribute information of the target object in the target image is adjusted according to the image analysis result.
22. The apparatus according to claim 19, wherein The processing requirement is an object adjustment requirement for a target object in the target image, and The processing unit is further configured to: Through the artificial intelligence big model, according to the object adjustment requirements represented by the request analysis result, the target object in the target image is adjusted according to the image analysis result.
23. The device according to claim 22, wherein The object adjustment requirement is an object deletion requirement for the target object in the target image, and The processing unit is further configured to: Through the artificial intelligence large model, according to the object deletion requirement represented by the request parsing result, the target object in the target image is deleted according to the image parsing result; the deleted area is pixel-filled according to the pixels in the surrounding area of the deleted area, wherein the deleted area is the area obtained after deleting the target object in the target image; according to the interactive relationship, objects other than the target object in the target image are adjusted.
24. The apparatus according to claim 19, wherein The processing requirement is an object addition requirement for adding a specified object to the target image, and The processing unit is further configured to: Through the artificial intelligence big model, the specified object is added to the target image according to the object increase demand represented by the request analysis result; according to the image analysis result, the specified object and the objects in the target image are adjusted.
25. The apparatus according to claim 19, wherein The processing requirement is an animation production requirement for the target image, and The processing unit is further configured to: Through the artificial intelligence big model, according to the animation production requirements represented by the request analysis result, a dynamic image sequence corresponding to the target image is generated according to the image analysis result.
26. The device according to claim 25, wherein The processing unit is further configured to: Through the artificial intelligence big model, according to the animation production requirements represented by the request analysis result, the animation type corresponding to each object in the image analysis result is determined; for each object in the image analysis result, a dynamic object is generated according to the animation type corresponding to the object, and the dynamic object is adjusted according to the interactive relationship.
27. The apparatus according to claim 19, wherein The processing requirement is an image synthesis requirement for multiple target images, and The processing unit is further configured to: Through the artificial intelligence big model, according to the image synthesis requirements represented by the request analysis results, the multiple target images are processed according to the image analysis results corresponding to each of the multiple target images to obtain a synthesized image.
28. The apparatus according to claim 27, wherein The processing unit is further configured to: Determining, by the artificial intelligence large model, an object group consisting of objects fused between the multiple target images according to the image synthesis requirements represented by the request analysis result; Through the artificial intelligence large model, according to the image synthesis requirements, the objects in each object group in at least one object group are fused to obtain the synthesized image.
29. The apparatus according to claim 19, wherein The processing requirement is a video synthesis requirement for multiple target images, and The processing unit is further configured to: Through the artificial intelligence big model, according to the video synthesis requirements represented by the request analysis results, and based on the image analysis results corresponding to each of the multiple target images, multiple images to be processed are determined from the multiple target images, and a synthetic video is generated based on the multiple images to be processed.
30. The device according to any one of claims 16 to 29, wherein The processing requirement includes a plurality of processing sub-requirements for the target image, and The processing unit is further configured to: For each of the plurality of processing sub-requirements, the target image is processed according to the processing sub-requirement by the artificial intelligence macromodel to obtain a processed image; The target video is generated based on multiple processed images through the artificial intelligence large model.
31. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.
32. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 15.
33. A computer program product comprising: A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Methods for training generative large language models and for processing image tasks
CN117114063A
Image generation method and device, electronic equipment and readable storage medium
CN117115287A