Image processing method, device and equipment, computer readable storage medium and product
By obtaining the user's original image data, using the corpus matching library and image feature information generation model, the problem of single image generation method and poor quality is solved, and high-quality image generation is achieved that is more in line with user needs.
Patent Information
- Application Number
- CN202410122980.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2025-07-29
AI Technical Summary
The existing image generation method is relatively single, and the generated image quality is poor, which cannot meet the personalized needs of users.
By obtaining the user's original image data, the text description information is determined using the preset corpus matching library, and inputting the image generation model with the image feature information to generate the target image.
The accuracy of text description information and the quality of the target image are improved, so that the generated image is more consistent with the user's original image and improve the user experience.
Smart Images

Figure CN120388105A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to an image processing method, apparatus, device, computer-readable storage medium, and product. Background Art
[0002] With the improvement of the hardware performance of terminal devices and the continuous progress of artificial intelligence technologies, users can generate images on terminal devices based on custom text description content. However, the current image generation methods are often relatively single, and due to the inaccurate custom text description content of users, the quality of the generated images is poor. Summary of the Invention
[0003] Embodiments of the present disclosure provide an image processing method, apparatus, device, computer-readable storage medium, and product, which are used to solve the technical problems that the existing image generation methods are relatively single and the quality of the generated images is poor.
[0004] In a first aspect, an image processing method provided by an embodiment of the present disclosure includes:
[0005] Obtain original data determined by a user, where the original data includes at least an original image;
[0006] Determine text description information based on the original data and a preset corpus matching library, where the corpus matching library includes description information corresponding to a plurality of preset keywords;
[0007] Determine image feature information corresponding to the original image, where the image feature information includes image structure information and content main body information;
[0008] Input the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model.
[0009] In a second aspect, an image processing apparatus provided by an embodiment of the present disclosure includes:
[0010] An obtaining module, configured to obtain original data determined by a user, where the original data includes at least an original image;
[0011] A determining module, configured to determine text description information based on the original data and a preset corpus matching library, where the corpus matching library includes description information corresponding to a plurality of preset keywords;
[0012] A processing module, configured to determine image feature information corresponding to the original image, where the image feature information includes image structure information and content main body information;
[0013] A generation module, configured to input the text description information and the image feature information into a preset image generation model, and obtain a target image output by the image generation model.
[0014] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0015] The memory stores computer-executable instructions;
[0016] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the image processing method described in the first aspect and various possible designs of the first aspect above.
[0017] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the image processing method described in the first aspect and various possible designs of the first aspect above is implemented.
[0018] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which when executed by a processor, implements the image processing method described in the first aspect and various possible designs of the first aspect above.
[0019] The image processing method, apparatus, device, computer-readable storage medium and product provided in this embodiment determine text description information based on the original data and a preset corpus matching library by obtaining the original data determined by the user, so that the user does not need to perform a custom operation on the text description information. Determining the text description information associated with the original data through the preset description information in the corpus matching library can improve the accuracy of the text description information, and further improve the image quality of the target image generated based on the text description information. Further, by determining the image feature information corresponding to the original image and generating the target image based on the text description information and the image feature information, on the basis of improving the image quality of the target image, the generated target image can be made more conforming to the original image determined by the user, enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of the image processing method provided by the embodiment of the present disclosure;
[0022] Figure 2 Schematic flowchart of an image processing method provided by another embodiment of the present disclosure;
[0023] Figure 3 Schematic diagram of interface interaction provided by an embodiment of the present disclosure;
[0024] Figure 4 Schematic diagram of interface interaction provided by another embodiment of the present disclosure;
[0025] Figure 5 Schematic structural diagram of an image processing apparatus provided by an embodiment of the present disclosure;
[0026] Figure 6 Schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0027] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without making creative efforts shall fall within the protection scope of the present disclosure.
[0028] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure shall be informed to the user and the user's authorization shall be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0029] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0030] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0031] It can be understood that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0032] To solve the technical problems that the existing image generation methods are relatively single and the quality of the generated images is not good, the present disclosure provides an image processing method, apparatus, device, computer-readable storage medium, and product.
[0033] It should be noted that the image processing method, apparatus, device, computer-readable storage medium, and product provided by the present disclosure can be applied to any image processing scenario.
[0034] In the related art, users generally determine the text description content according to actual needs, and generate a target image based on the text description content through a preset image generation algorithm. However, since the text description content defined by users is often not accurate enough, the target image generated based on the user-defined text description content may not meet the personalized needs of users.
[0035] In the process of solving the above technical problems, the inventors found through research that in order to improve the image quality of the generated target image and make the generated target image more in line with the personalized needs of users, the target image can be generated based on the original image uploaded by the user.
[0036] Optionally, the original image uploaded by the user can be obtained, and the text description information matching the original image can be determined based on the analysis of the original image and a preset corpus matching library. Thus, there is no need for users to define the text description information by themselves, which improves the accuracy of the text description information. In addition, during the image generation process, the image feature information corresponding to the original image is determined, and the target image is generated jointly based on the text description information and the image feature information, thereby further improving the image quality of the target image.
[0037] Figure 1 For the flowchart of the image processing method provided by the embodiments of the present disclosure, as Figure 1 shown, the method includes:
[0038] Step 101, obtain the original data determined by the user, where the original data at least includes an original image.
[0039] The execution subject of this embodiment is an image processing apparatus. The image processing apparatus can be coupled to a terminal device, so as to be able to determine text description information based on the original data determined by the user on the terminal device, and generate a target image based on the text description information and the image feature information corresponding to the original data.
[0040] Alternatively, the image processing device can also be coupled to a server that can communicate with a terminal device, enabling it to obtain the original data determined by the user on the terminal device. Determine the original data to determine the text description information, and generate a target image based on the text description information and the image feature information corresponding to the original data.
[0041] In this embodiment, to generate a target image, the user can determine the original data. Among them, the original data at least includes the original image. In addition, the original data can also include at least one tag information determined by the user and / or the original description information input by the user, etc.
[0042] For example, when the user browses the media content stream within the media content display page, an image generation operation can be triggered based on a preset trigger control in the media content display page. Based on the image generation operation, a content collection page is displayed, and the original data is determined within the content collection page.
[0043] Step 102: Determine the text description information based on the original data and a preset corpus matching library, where the corpus matching library includes the description information corresponding to multiple preset keywords.
[0044] In this embodiment, since the text description information customized by the user is often not accurate enough, the target image generated based on the inaccurate information often fails to meet the actual needs of the user. Therefore, to improve the accuracy of the text description information, a corpus matching library can be preset, where the corpus matching library includes the description information corresponding to multiple preset keywords, and this description information can accurately describe the image effect.
[0045] Thus, after obtaining the original data determined by the user, the text description information can be determined based on the original data in the preset corpus matching library, without the need for the user to customize the generation of the text description information, and the accurate text description information can be automatically determined based on the original data. On the basis of ensuring the accuracy of the text description information, the efficiency of determining the text description information is improved.
[0046] Step 103: Determine the image feature information corresponding to the original image, where the image feature information includes image structure information and content main body information.
[0047] In this embodiment, due to the original image determined by the user. In the process of generating the target image based on the original image, in order to make the generated target image more conform to the original image and improve the generation efficiency of the target image, the image feature information corresponding to the original image can be determined, where the image feature information includes image structure information and content subject information. The image structure information includes, but is not limited to, the edge contour of the original image, depth information, pose information of the content subject, etc. The content subject information includes the feature information of the content subject in the original image. For example, face features, limb features, etc.
[0048] Step 104: Input the text description information and the image feature information into a preset image generation model to obtain the target image output by the image generation model.
[0049] In this embodiment, after obtaining the text description information and the image feature information respectively, the text description information and the image feature information can be input into a preset image generation model. Among them, the image generation model includes, but is not limited to, any model such as a diffusion model that can generate an image based on the text description information and the image feature information. Thus, the target image output by the image generation model can be obtained.
[0050] The image processing method provided in this embodiment determines the text description information based on the original data and the preset corpus matching library by obtaining the original data determined by the user, so that the user does not need to perform a custom operation on the text description information. Determining the text description information associated with the original data through the preset description information in the corpus matching library can improve the accuracy of the text description information, and further improve the image quality of the target image generated based on the text description information. Further, by determining the image feature information corresponding to the original image and generating the target image based on the text description information and the image feature information, the generated target image can be made more conform to the original image determined by the user on the basis of improving the image quality of the target image, thus enhancing the user experience.
[0051] Figure 2 It is a schematic flowchart of the image processing method provided in another embodiment of the present disclosure. On the basis of any of the above embodiments, as Figure 2 shown, step 102 includes:
[0052] Step 201: Identify the scene information corresponding to the original image and determine the first keyword corresponding to the scene information.
[0053] Step 202: Determine the scene description information matching the scene information in the corpus matching library based on the first keyword.
[0054] Step 203: Identify subject feature information corresponding to the content subject in the original image, and determine at least one second associated word corresponding to the subject feature information.
[0055] Step 204: Determine at least one subject description information matching the subject feature information in the corpus matching database based on the at least one second associated word.
[0056] Step 205: Determine text description information corresponding to the original image based on the scene description information and the at least one subject description information.
[0057] In this embodiment, after obtaining the original image, the scene information corresponding to the original image can be first identified to determine the first keyword corresponding to the scene information. This scene information includes, but is not limited to, scenes such as people, animals, and landscapes. The corpus matching library can be pre-populated with description information corresponding to different scenes. Therefore, after obtaining the first keyword corresponding to the scene information, scene description information that matches the scene information can be determined in the corpus matching library based on the first keyword.
[0058] Furthermore, to obtain more accurate text description information, subject feature information corresponding to the main content of the original image can be identified to determine at least one second associated word corresponding to the subject feature information. This subject feature information can be more detailed feature information. For example, this subject feature information includes, but is not limited to, hairstyle / color, clothing color, pet fur color, etc.
[0059] After determining the subject feature information, at least one subject description information matching the subject feature information may be determined in the corpus matching library based on the at least one second associated word. Text description information corresponding to the original image may be determined based on the scene description information and the at least one subject description information.
[0060] For example, the first keyword might be "Hanfu girl." Scene description information corresponding to the Hanfu girl can be obtained. Furthermore, it can be determined that the Hanfu girl has black hair, shoulder-length hair, and red Hanfu. Therefore, the resulting text description reads: "A girl with black shoulder-length hair, wearing red Hanfu." Based on this text description, the target image can be accurately generated.
[0061] The image processing method provided in this embodiment determines scene description information based on scene information corresponding to the original image, and determines at least one subject description information based on subject feature information corresponding to the content subject in the original image, thereby being able to determine more accurate text description information based on the scene description information and the at least one subject description information, eliminating the need for users to customize text description information, and improving the efficiency of image generation operations.
[0062] Further, based on any of the above embodiments, the original data further includes at least one tag information determined by the user. The method further includes:
[0063] Determining at least one scenario description information in the corpus matching library based on the at least one tag information.
[0064] And / or, the original data further includes original description information input by the user. The method further includes:
[0065] Identifying at least one keyword corresponding to the original description information.
[0066] Determining at least one scenario description information in the corpus matching library based on the at least one keyword.
[0067] In this embodiment, the original data may further include at least one tag information determined by the user and / or original description information input by the user, etc. Therefore, in order to more accurately determine the text description information, at least one scenario description information may also be determined in the corpus matching library based on at least one tag information. Among them, the tag information may be tags such as Hong Kong style, Hanfu, pixel style, realistic style, etc. And / or, identifying at least one keyword corresponding to the original description information. Determining at least one scenario description information in the corpus matching library based on at least one keyword. Among them, any keyword recognition method may be used to identify at least one keyword corresponding to the original description information, and the present disclosure does not limit this.
[0068] Figure 3 It is a schematic diagram of interface interaction provided by an embodiment of the present disclosure. As Figure 3 shown, the user can perform a trigger operation on a preset trigger control 32 within the media content playback page 31. In response to this trigger operation, a content collection page 33 can be displayed, so that the user can obtain an original image in the content collection page 33. After obtaining the original image 34 selected by the user, in response to the user's trigger operation on a preset generation control 35, a target image can be generated based on the original image.
[0069] Figure 4 It is a schematic diagram of interface interaction provided by another embodiment of the present disclosure. As Figure 4As shown, the user can perform a triggering operation on a preset triggering control 42 within the media content playback page 41. In response to this triggering operation, a tag determination page 43 can be displayed. Thus, the user can select at least one tag information 44 within the tag determination page 43. Further, a content collection page 45 can be displayed, so that the user can obtain an original image within the content collection page 45. After obtaining the original image 46 selected by the user, in response to the triggering operation of the user on the preset generation control 47, a target image can be generated based on the at least one tag information 44 and the original image 46.
[0070] The image processing method provided in this embodiment can determine more accurate text description information in the corpus matching library based on at least one tag information and / or original description information by obtaining at least one tag information and / or original description information determined by the user, and can jointly generate a target image based on the original image and at least one tag information and / or original description information, enriching the way of image generation.
[0071] Further, based on any of the above embodiments, step 103 includes:
[0072] Identify the edge contour information and / or image depth information and / or content body pose information in the original image to obtain the image structure information.
[0073] Identify at least one content body in the original image.
[0074] Obtain the attribute information of the target part corresponding to the content body to obtain the content body information.
[0075] In this embodiment, the image structure information includes edge contour information and / or image depth information and / or content body pose information. By determining the image structure information, the structure of the target image can be more accurately controlled to simulate the original image. The edge contour information and / or image depth information and / or content body pose information in the original image can be identified through a preset image recognition algorithm to obtain the image structure information.
[0076] Further, at least one content body in the original image can be identified through a preset content body recognition algorithm. The content body includes, but is not limited to, content bodies such as people, animals, and buildings. Obtain the attribute information of the target part corresponding to the content body to obtain the content body information. Taking the content body as a person as an example, the facial features of the person can be identified and the facial features can be determined as the content body information.
[0077] The image processing method provided in this embodiment can further improve the image quality of the target image by separately identifying the image structure information of the original image and the attribute information of the content subject, making the target image more conform to the original image determined by the user.
[0078] Optionally, based on any of the above embodiments, after step 101, the method further includes:
[0079] Performing a cropping operation on the original image according to a preset cropping size to obtain a cropped original image.
[0080] And / or,
[0081] Performing a scaling operation on the original image according to a preset long side threshold and the original ratio corresponding to the original image to obtain a scaled original image.
[0082] And / or,
[0083] Identifying the foreground information corresponding to the original image.
[0084] Performing a foreground segmentation operation on the original image based on the foreground information to obtain a segmented foreground region.
[0085] Performing background filling on the foreground region with a preset background to obtain a filled original image.
[0086] In this embodiment, after obtaining the original image, in order to improve the generation efficiency of the target image, a preprocessing operation can also be performed on the original image.
[0087] Optionally, a cropping operation can be performed on the original image according to a preset cropping size to obtain a cropped original image. Wherein, the cropping size can be set by the user according to actual needs, or can be determined based on the processing ability of the image generation model. Optionally, the user can also specify a cropping area or a resolution, etc. according to actual needs, and the present disclosure does not limit this. For example, the position of the face in the original image can be determined, the face position can be centered, and a cropping operation can be performed on other regions.
[0088] Optionally, a scaling operation can also be performed on the original image according to a preset long side threshold and the original ratio corresponding to the original image according to actual needs to obtain a scaled original image. For example, if the image generation model can process images with a preset resolution, a scaling operation can be performed on the original image based on the processing ability of the image generation model. For example, the resolution of the original image can be 720*1280, and it can be scaled proportionally to 540*960 according to the limit of the maximum long side of 960, so as to reduce the generation time consumed by the image generation model for image processing.
[0089] Optionally, in order to improve the processing accuracy of the image generation model, foreground segmentation can also be performed on the original image based on foreground information to obtain the segmented foreground region. The foreground region after segmentation is filled with a preset background to obtain the filled original image. For example, after performing foreground segmentation on the original image to obtain the segmented foreground region, the background of the original image can be filled with pure white to reduce the influence of the background on the generation effect of the algorithm.
[0090] It should be noted that the above image preprocessing methods can be implemented separately or in combination, and the present disclosure does not limit this.
[0091] The image processing method provided in this embodiment can facilitate subsequent image processing operations, reduce the computational amount in the image processing process, and improve the efficiency of image processing by cropping and / or scaling and / or background filling the original image after obtaining the original image.
[0092] Optionally, based on any of the above embodiments, after step 204, it further includes:
[0093] Performing super-resolution operation on the target image through a preset super-resolution algorithm to obtain the super-resolved target image.
[0094] In this embodiment, when generating the target image based on the image generation model, the image generation model may pre-specify the resolution upper limit of the long side of the target image, resulting in a low-resolution target image that cannot meet the personalized needs of users.
[0095] Therefore, in order to improve the clarity of the target image, after obtaining the target image, super-resolution operation can be performed on the target image through a preset super-resolution algorithm to obtain the super-resolved target image. During the super-resolution process, the user can specify the resolution of the super-resolved target image according to actual needs, and the present disclosure does not limit this.
[0096] The image processing method provided in this embodiment can improve the clarity of the target image and avoid the problem of unclear images caused by the resolution upper limit specified by the image generation model by performing super-resolution operation on the target image.
[0097] Optionally, based on any of the above embodiments, after step 204, it further includes:
[0098] Identifying at least one target subject in the target image that meets the preset conditions.
[0099] Performing image repair operation on the at least one target subject through a preset repair algorithm.
[0100] In this embodiment, since some of the main subjects in the target image, such as faces and buildings with relatively small sizes, may be unclear, in order to ensure the image quality of the target image, at least one target subject that meets the preset conditions in the target image can also be identified. Among them, the preset conditions can be a face, a body, a building, etc. with a size smaller than a preset threshold. An image restoration operation is performed on at least one target subject through a preset restoration algorithm.
[0101] For example, when the target subject is a face, an image restoration operation can be performed on the target subject through a preset face restoration algorithm to obtain a clearer face.
[0102] The image processing method provided in this embodiment can avoid the problem that smaller people or target subjects are blurred and invisible by performing image restoration on at least one target subject that meets the preset conditions in the target image, and further improve the image quality of the target image.
[0103] Optionally, based on any of the above embodiments, after step 204, the method further includes:
[0104] Determine the image style corresponding to the target image.
[0105] Perform a color adjustment operation on the target image using the color parameters corresponding to the image style.
[0106] In this embodiment, in order to make the content and color of the generated target image more coordinated, the image style corresponding to the target image can also be determined. For example, the style of the target image can be determined as Hong Kong style, pixel style, realistic style, anime style, etc. Determine the color parameters corresponding to the image style, and perform a color adjustment operation on the target image based on the color parameters. Among them, the color parameters corresponding to the image style can be preset empirical values, or can be set by designers based on different image segmentations. The present disclosure does not limit this.
[0107] The image processing method provided in this embodiment can make the color of the target image more matching with the image style of the target image by performing a color adjustment operation on the target image using the color parameters corresponding to different image styles, and further improve the image quality of the target image.
[0108] Figure 5 It is a schematic structural diagram of the image processing device provided in an embodiment of the present disclosure, as Figure 5As shown in the figure, the device includes: an acquisition module 51, a determination module 52, a processing module 53, and a generation module 54. The acquisition module 51 is configured to acquire the original data determined by the user, and the original data includes at least the original image. The determination module 52 is configured to determine the text description information based on the original data and a preset corpus matching library, where the corpus matching library includes description information corresponding to a plurality of preset keywords. The processing module 53 is configured to determine the image feature information corresponding to the original image, where the image feature information includes image structure information and content subject information. The generation module 54 is configured to input the text description information and the image feature information into a preset image generation model to obtain the target image output by the image generation model.
[0109] Further, based on any of the above embodiments, the determination module is configured to: identify the scene information corresponding to the original image, and determine the first keyword corresponding to the scene information. Determine the scene description information matching the scene information in the corpus matching library based on the first keyword. Identify the subject feature information corresponding to the content subject in the original image, and determine at least one second correlation word corresponding to the subject feature information. Determine at least one subject description information matching the subject feature information in the corpus matching library based on the at least one second correlation word. Determine the text description information corresponding to the original image based on the scene description information and the at least one subject description information.
[0110] Further, based on any of the above embodiments, the original data further includes at least one tag information determined by the user. The device further includes: a determination module, configured to determine at least one scene description information in the corpus matching library based on the at least one tag information. And / or, the original data further includes the original description information input by the user. The device further includes: an identification module, configured to identify at least one keyword corresponding to the original description information. A determination module, configured to determine at least one scene description information in the corpus matching library based on the at least one keyword.
[0111] Further, based on any of the above embodiments, the processing module is configured to: identify the edge contour information and / or the image depth information and / or the content subject pose information in the original image to obtain the image structure information. Identify at least one content subject in the original image. Obtain the attribute information of the target part corresponding to the content subject to obtain the content subject information.
[0112] Further, based on any of the above embodiments, the apparatus further includes: a cropping module, configured to crop the original image according to a preset cropping size to obtain a cropped original image. And / or, a scaling module, configured to scale the original image according to a preset long-side threshold and the original ratio corresponding to the original image to obtain a scaled original image. And / or, an identification module, configured to identify foreground information corresponding to the original image. A segmentation module, configured to perform foreground segmentation on the original image based on the foreground information to obtain a segmented foreground region. Fill the foreground region with a preset background to obtain a filled original image.
[0113] Further, based on any of the above embodiments, the apparatus further includes: a super-resolution module, configured to perform super-resolution on the target image through a preset super-resolution algorithm to obtain a super-resolved target image.
[0114] Further, based on any of the above embodiments, the apparatus further includes: an identification module, configured to identify at least one target object in the target image that meets a preset condition. A restoration module, configured to perform image restoration on the at least one target object through a preset restoration algorithm.
[0115] Further, based on any of the above embodiments, the apparatus further includes: a determination module, configured to determine the image style corresponding to the target image. An adjustment module, configured to perform color adjustment on the target image using color parameters corresponding to the image style.
[0116] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0117] To implement the above embodiments, the embodiments of the present disclosure also provide a computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the image processing method as described in any of the above embodiments is implemented.
[0118] To implement the above embodiments, the embodiments of the present disclosure also provide a computer program product, including a computer program, and when the computer program is executed by a processor, the image processing method as described in any of the above embodiments is implemented.
[0119] To implement the above embodiments, the embodiments of the present disclosure also provide an electronic device, including: a processor and a memory;
[0120] The memory stores computer-executable instructions;
[0121] The processor executes the computer-executable instructions stored in the memory, such that the processor executes the image processing method according to any of the foregoing embodiments.
[0122] Figure 6 FIG. 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. The electronic device 600 may be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0123] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may execute various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0124] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wireline to exchange data. Although Figure 6 the illustrated electronic device 600 has various devices, it should be understood that it is not required to implement or include all the illustrated devices. More or fewer devices may be alternatively implemented or included.
[0125] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.
[0126] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0127] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.
[0128] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the methods shown in the above embodiments.
[0129] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0131] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".
[0132] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0133] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0134] The above description is only a preferred embodiment of the present disclosure and an illustration of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0135] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0136] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An image processing method, characterized in that, Including: Obtain the original data determined by the user, where the original data at least includes an original image; Determine text description information based on the original data and a preset corpus matching library, where the corpus matching library includes description information corresponding to multiple preset keywords; Determine the image feature information corresponding to the original image, where the image feature information includes image structure information and content subject information; Input the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model.
2. The method according to claim 1, wherein The determining the text description information based on the original data and the preset corpus matching library includes: Identify the scene information corresponding to the original image and determine the first keyword corresponding to the scene information; Determine the scene description information matching the scene information in the corpus matching library based on the first keyword; Identify the subject feature information corresponding to the content subject in the original image and determine at least one second related word corresponding to the subject feature information; Determine at least one subject description information matching the subject feature information in the corpus matching library based on the at least one second related word; Determine the text description information corresponding to the original image based on the scene description information and the at least one subject description information.
3. The method according to claim 2, wherein The original data further includes at least one tag information determined by the user; the method further includes: Determine at least one scene description information in the corpus matching library based on the at least one tag information; And / or, the original data further includes the original description information input by the user; the method further includes: Identify at least one keyword corresponding to the original description information; Determine at least one scene description information in the corpus matching library based on the at least one keyword.
4. The method according to claim 1, wherein The determining the image feature information corresponding to the original image includes: Identify the edge contour information and / or image depth information and / or content subject pose information in the original image to obtain the image structure information; Identify at least one content subject in the original image; Obtain the attribute information of the target part corresponding to the content subject to obtain the content subject information.
5. The method according to any one of claims 1 to 4, characterized in that After obtaining the original data determined by the user, it further includes: Perform a cropping operation on the original image according to a preset cropping size to obtain a cropped original image; And / or, Perform a scaling operation on the original image according to a preset long side threshold and the original ratio corresponding to the original image to obtain a scaled original image; And / or, Identify the foreground information corresponding to the original image; Perform a foreground segmentation operation on the original image based on the foreground information to obtain a segmented foreground region; Fill the foreground region with a preset background to obtain a filled original image.
6. The method according to any one of claims 1-4, characterized in that, After inputting the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model, it further includes: Perform a super-resolution operation on the target image through a preset super-resolution algorithm to obtain a super-resolved target image.
7. The method according to any one of claims 1 to 4, characterized in that After inputting the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model, the method further includes: Identifying at least one target subject in the target image that meets a preset condition; Performing an image repair operation on the at least one target subject through a preset repair algorithm.
8. The method according to any one of claims 1-4, characterized in that, After inputting the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model, the method further includes: Determining an image style corresponding to the target image; Performing a color adjustment operation on the target image by using color parameters corresponding to the image style.
9. An image processing apparatus, characterized in that, Including: An acquisition module, configured to acquire original data determined by a user, where the original data includes at least an original image; A determination module, configured to determine text description information based on the original data and a preset corpus matching library, where the corpus matching library includes description information corresponding to a plurality of preset keywords; A processing module, configured to determine image feature information corresponding to the original image, where the image feature information includes image structure information and content subject information; A generation module, configured to input the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model.
10. An electronic device, characterized in that, Including: A processor and a memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the image processing method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the image processing method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the image processing method according to any one of claims 1 to 8 is implemented.