Image processing method and apparatus, device, computer-readable storage medium, and product
By obtaining the original data of users, using the corpus matching library and image feature information generation model, the existing image generation methods are solved and the problems of single and poor quality are achieved, and high-quality image generation that is more in line with user needs is achieved.
Patent Information
- Application Number
- PCT/CN2024/141413
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2024-12-23
- Publication Date
- 2025-08-07
AI Technical Summary
The existing image generation method is relatively simple, and the user-defined text description content is not accurate enough, resulting in poor quality of generated images.
By obtaining the original data of the user, the text description information is determined using the preset corpus matching library, and input it into the image generation model based on the image feature information of the original image to generate the target image.
The accuracy of text description information and the quality of generated images are improved, so that the generated target image is more in line with the user's personalized needs and improve the user experience.
Smart Images

Figure CN2024141413_07082025_PF_FP_ABST
Abstract
Description
Image processing method, device, equipment, computer-readable storage medium and product
[0001] This application claims priority to Chinese Patent Application No. 202410122980.6 filed on January 29, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] Embodiments of the present disclosure relate to an image processing method, apparatus, device, computer-readable storage medium, and product. Background Art
[0003] With the improvement of terminal device hardware performance and the continuous advancement of artificial intelligence technology, users can now generate images based on customized text descriptions on their devices. However, current image generation methods are often relatively simple and, due to the lack of precision of user-defined text descriptions, the generated images are of poor quality. Summary of the Invention
[0004] Embodiments of the present disclosure provide an image processing method, apparatus, device, computer-readable storage medium, and product.
[0005] The present disclosure provides an image processing method, including:
[0006] Acquiring original data determined by a user, wherein the original data at least includes an original image;
[0007] Determining text description information based on the original data and a preset corpus matching library, wherein the corpus matching library includes description information corresponding to a plurality of preset keywords;
[0008] Determining image feature information corresponding to the original image, wherein the image feature information includes image structure information and content main body information;
[0009] The text description information and the image feature information are input into a preset image generation model to obtain a target image output by the image generation model.
[0010] An embodiment of the present disclosure provides an image processing device, including:
[0011] an acquisition module configured to acquire original data determined by a user, wherein the original data at least includes an original image;
[0012] a determination module configured to determine text description information based on the original data and a preset corpus matching library, wherein the corpus matching library includes description information corresponding to a plurality of preset keywords;
[0013] A processing module is configured to determine image feature information corresponding to the original image, wherein the image feature information includes image structure information and content body information;
[0014] The generation module is configured to input the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model.
[0015] An embodiment of the present disclosure provides an electronic device, comprising: a processor and a memory;
[0016] The memory stores computer-executable instructions;
[0017] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the image processing method provided in any of the above embodiments.
[0018] An embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the image processing method provided in any of the above embodiments is implemented.
[0019] An embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the image processing method provided in any of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solution of the present disclosure, the following briefly introduces the drawings required for use in the description of the technical solution of the present disclosure. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] FIG1 is a schematic diagram of a flow chart of an image processing method provided by an embodiment of the present disclosure;
[0022] FIG2 is a schematic flow chart of an image processing method provided by another embodiment of the present disclosure;
[0023] FIG3 is a schematic diagram of an interface interaction provided by an embodiment of the present disclosure;
[0024] FIG4 is a schematic diagram of an interface interaction provided by another embodiment of the present disclosure;
[0025] FIG5 is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure; and
[0026] FIG6 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0028] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0029] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0030] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0031] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0032] In order to solve the technical problem that the image generation method is relatively single and the generated image quality is poor, the present disclosure provides an image processing method, device, equipment, computer-readable storage medium and product.
[0033] It should be noted that the image processing method, apparatus, device, computer-readable storage medium and product provided by the present disclosure can be applied in any image processing scenario.
[0034] For example, users typically define text descriptions based on their actual needs, and then generate target images based on these text descriptions using a pre-set image generation algorithm. However, since user-defined text descriptions are often inaccurate, the target images generated based on these customized text descriptions may not meet the user's personalized needs.
[0035] In the process of solving the above technical problems, the inventors of the present disclosure discovered through research that in order to improve the image quality of the generated target image and make the generated target image more in line with the user's personalized needs, the target image can be generated based on the original image uploaded by the user.
[0036] Optionally, the original image uploaded by the user can be obtained and, based on analysis of the original image and a pre-set corpus matching library, a text description matching the original image can be determined. This eliminates the need for the user to generate custom text descriptions, improving the accuracy of the text descriptions. Furthermore, by determining the image feature information corresponding to the original image during the image generation process and generating the target image based on both the text description and the image feature information, the target image's quality can be further improved.
[0037] FIG1 is a flow chart of an image processing method provided by an embodiment of the present disclosure. As shown in FIG1 , the method includes:
[0038] Step 101: Acquire original data determined by a user, where the original data at least includes an original image.
[0039] The execution subject of this embodiment is an image processing device. The image processing device can be coupled to a terminal device, so as to determine text description information based on the original data determined by the user on the terminal device, and generate a target image based on the text description information and image feature information corresponding to the original data.
[0040] Alternatively, the image processing apparatus may be coupled to a server that is communicatively connected to a terminal device, thereby acquiring raw data determined by a user on the terminal device, determining textual description information of the raw data, and generating a target image based on the textual description information and image feature information corresponding to the raw data.
[0041] In this embodiment, to achieve the generation of a target image, the user may determine the original data. The original data may include at least the original image. In addition, the original data may also include at least one tag information determined by the user and / or original description information input by the user.
[0042] For example, when a user browses a media content stream on a media content display page, an image generation operation can be triggered based on a trigger control preset in the media content display page. Based on the image generation operation, a content collection page is displayed, and raw data is determined on the content collection page.
[0043] Step 102: Determine text description information based on the original data and a preset corpus matching library, wherein the corpus matching library includes description information corresponding to a plurality of preset keywords.
[0044] In this embodiment, since user-defined text description information is often inaccurate, the target image generated based on this inaccuracy often fails to meet the user's actual needs. Therefore, to improve the accuracy of the text description information, a corpus matching library can be pre-set. The corpus matching library includes description information corresponding to multiple preset keywords, which can accurately describe the image effect.
[0045] After obtaining the original data specified by the user, the text description information can be determined based on the original data in a preset corpus matching library. The user does not need to generate the text description information by themselves, and the accurate text description information can be automatically determined based on the original data. This improves the efficiency of determining the text description information while ensuring the accuracy of the text description information.
[0046] Step 103: Determine image feature information corresponding to the original image, wherein the image feature information includes image structure information and content main body information.
[0047] In this embodiment, due to the original image determined by the user, in the process of generating a target image based on the original image, in order to make the generated target image more consistent with the original image and improve the generation efficiency of the target image, the image feature information corresponding to the original image can be determined, wherein the image feature information includes image structure information and content information. The image structure information includes but is not limited to the edge contour, depth information, and posture information of the content of the original image. The content information includes feature information of the content of the original image, such as facial features, body features, etc.
[0048] Step 104: Input the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model.
[0049] In this embodiment, after obtaining the text description information and image feature information, the text description information and image feature information can be input into a preset image generation model. The image generation model includes, but is not limited to, any model capable of generating an image based on the text description information and image feature information, such as a diffusion model. This allows the target image to be obtained as output by the image generation model.
[0050] The image processing method provided in this embodiment obtains original data determined by the user and determines text description information based on the original data and a preset corpus matching library, thereby eliminating the need for the user to customize the text description information. Determining the text description information associated with the original data using the description information preset in the corpus matching library can improve the accuracy of the text description information, thereby improving the image quality of the target image generated based on the text description information. Furthermore, by determining the image feature information corresponding to the original image and generating the target image based on the text description information and image feature information, the generated target image can be made to be more consistent with the original image determined by the user on the basis of improving the image quality of the target image, thereby enhancing the user experience.
[0051] FIG2 is a flowchart of an image processing method provided by another embodiment of the present disclosure. Based on any of the above embodiments, as shown in FIG2 , step 102 includes:
[0052] Step 201: Identify scene information corresponding to the original image and determine a first keyword corresponding to the scene information.
[0053] Step 202: Determine scene description information that matches the scene information in a corpus matching database based on the first keyword.
[0054] Step 203: Identify subject feature information corresponding to the content subject in the original image, and determine at least one second associated word corresponding to the subject feature information.
[0055] Step 204: Determine at least one subject description information that matches the subject feature information in the corpus matching database based on the at least one second associated word.
[0056] Step 205: Determine text description information corresponding to the original image based on the scene description information and at least one subject description information.
[0057] In this embodiment, after obtaining the original image, the scene information corresponding to the original image can be first identified to determine the first keyword corresponding to the scene information. This scene information includes, but is not limited to, scenes such as people, animals, and landscapes. The corpus matching library can be pre-populated with description information corresponding to different scenes. Therefore, after obtaining the first keyword corresponding to the scene information, scene description information that matches the scene information can be determined in the corpus matching library based on the first keyword.
[0058] Furthermore, to obtain more accurate text description information, subject feature information corresponding to the main content of the original image can be identified to determine at least one second associated word corresponding to the subject feature information. This subject feature information can be more detailed feature information. For example, this subject feature information includes, but is not limited to, hairstyle / color, clothing color, pet fur color, etc.
[0059] After determining the subject feature information, at least one subject description information matching the subject feature information may be determined in the corpus matching library based on the at least one second associated word. Text description information corresponding to the original image may be determined based on the scene description information and the at least one subject description information.
[0060] For example, the first keyword might be "Hanfu girl." Scene description information corresponding to the Hanfu girl can be obtained. Furthermore, it can be determined that the Hanfu girl has black hair, shoulder-length hair, and red Hanfu. Therefore, the resulting text description reads: "A girl with black shoulder-length hair, wearing red Hanfu." Based on this text description, the target image can be accurately generated.
[0061] The image processing method provided in this embodiment determines scene description information based on scene information corresponding to the original image, and determines at least one subject description information based on subject feature information corresponding to the content subject in the original image, thereby being able to determine more accurate text description information based on the scene description information and the at least one subject description information, eliminating the need for users to customize text description information, and improving the efficiency of image generation operations.
[0062] Furthermore, based on any of the above embodiments, the original data further includes at least one tag information determined by a user. The image processing method further includes:
[0063] At least one scene description information is determined in the corpus matching library based on the at least one tag information.
[0064] And / or, the original data also includes original description information input by the user. The image processing method also includes:
[0065] At least one keyword corresponding to the original description information is identified.
[0066] At least one scene description information is determined in a corpus matching library based on at least one keyword.
[0067] In this embodiment, the original data may also include at least one tag information determined by the user and / or original description information input by the user. Therefore, in order to more accurately determine the text description information, at least one scene description information may also be determined in the corpus matching library based on at least one tag information. The tag information may be a label such as Hong Kong style, Hanfu, pixel style, realistic style, etc. And / or, identify at least one keyword corresponding to the original description information. Determine at least one scene description information in the corpus matching library based on at least one keyword. Any keyword recognition method may be used to identify at least one keyword corresponding to the original description information, and the present disclosure does not limit this.
[0068] Figure 3 is a schematic diagram of an interface interaction provided by an embodiment of the present disclosure. As shown in Figure 3, a user can trigger a preset trigger control 32 within a media content playback page 31. In response to this triggering operation, a content acquisition page 33 is displayed, allowing the user to acquire an original image within the content acquisition page 33. After acquiring an original image 34 selected by the user, a target image can be generated based on the original image in response to the user triggering a preset generation control 35.
[0069] FIG4 is a schematic diagram of an interface interaction provided by another embodiment of the present disclosure. As shown in FIG4 , a user can trigger a preset trigger control 42 within a media content playback page 41. In response to the trigger operation, a tag determination page 43 can be displayed. The user can then select at least one tag information 44 within the tag determination page 43. Furthermore, a content acquisition page 45 can be displayed, allowing the user to acquire an original image within the content acquisition page 45. After acquiring the original image 46 selected by the user, a target image can be generated based on the at least one tag information 44 and the original image 46 in response to the user triggering the preset generation control 47.
[0070] The image processing method provided in this embodiment obtains at least one tag information and / or original description information determined by the user, thereby being able to determine more precise text description information in the corpus matching library based on the at least one tag information and / or original description information, and is able to achieve the generation of a target image based on the original image and at least one tag information and / or original description information, thereby enriching the image generation method.
[0071] Furthermore, based on any of the above embodiments, step 103 includes:
[0072] Identify edge contour information and / or image depth information and / or content main body posture information in the original image to obtain image structure information.
[0073] At least one content subject is identified in the original image.
[0074] Obtain the attribute information of the target part corresponding to the content body and obtain the content body information.
[0075] In this embodiment, the image structure information includes edge contour information and / or image depth information and / or content main body posture information. By determining the image structure information, the target image can be more accurately controlled to mimic the structure of the original image. The image structure information can be obtained by identifying the edge contour information and / or image depth information and / or content main body posture information in the original image using a preset image recognition algorithm.
[0076] Furthermore, at least one content subject in the original image can be identified using a preset content subject recognition algorithm. This content subject includes, but is not limited to, people, animals, buildings, and other content subjects. Attribute information of the target area corresponding to the content subject is obtained to obtain content subject information. For example, if the content subject is a person, the person's facial features can be identified and determined as the content subject information.
[0077] The image processing method provided in this embodiment can further improve the image quality of the target image by separately identifying the image structure information of the original image and the attribute information of the content body, so that the target image is more consistent with the original image determined by the user.
[0078] Optionally, based on any of the above embodiments, after step 101, the method further includes:
[0079] The original image is cropped according to a preset cropping size to obtain a cropped original image.
[0080] and / or,
[0081] The original image is scaled according to a preset long side threshold and an original ratio corresponding to the original image to obtain a scaled original image.
[0082] and / or,
[0083] Identify the foreground information corresponding to the original image.
[0084] A foreground segmentation operation is performed on the original image based on the foreground information to obtain the segmented foreground area.
[0085] The foreground area is filled with a preset background to obtain the original image after filling.
[0086] In this embodiment, after the original image is acquired, in order to improve the efficiency of generating the target image, a preprocessing operation may be performed on the original image.
[0087] Optionally, the original image can be cropped according to a preset cropping size to obtain a cropped original image. The cropping size can be set by the user based on actual needs, or can be determined based on the processing capabilities of the image generation model. Optionally, the user can also specify a cropping area or resolution based on actual needs, which is not limited by this disclosure. For example, the location of a face in the original image can be determined, the face position can be controlled to be centered, and the remaining area can be cropped.
[0088] Optionally, the original image can be scaled based on a preset long side threshold and the original ratio of the original image according to actual needs to obtain a scaled original image. For example, if the image generation model can process images with a preset resolution, the original image can be scaled based on the processing capabilities of the image generation model. For example, the resolution of the original image can be 720*1280, and the original image resolution can be proportionally scaled to 540*960 according to the maximum long side limit of 960, thereby reducing the generation time of the image generation model image processing.
[0089] Optionally, to improve the processing accuracy of the image generation model, a foreground segmentation operation can be performed on the original image based on the foreground information to obtain a segmented foreground region. The foreground region can then be filled with a preset background to obtain the filled original image. For example, after performing foreground segmentation on the original image to obtain the segmented foreground region, the original image background can be filled with pure white to reduce the impact of the background on the algorithm's generation results.
[0090] It should be noted that the above-mentioned image preprocessing methods can be implemented separately or in combination, and the present disclosure does not impose any limitation on this.
[0091] The image processing method provided in this embodiment can facilitate subsequent image processing operations, reduce the amount of calculation in the image processing process, and improve the efficiency of image processing by cropping and / or scaling and / or background filling the original image after acquiring the original image.
[0092] Optionally, based on any of the above embodiments, after step 204, the following steps are further included:
[0093] The target image is super-resolutioned by a preset super-resolution algorithm to obtain a super-resolved target image.
[0094] In this embodiment, when generating a target image based on an image generation model, the image generation model may pre-specify an upper limit on the resolution of the long side of the target image, resulting in a low resolution of the generated target image that cannot meet the personalized needs of the user.
[0095] Therefore, in order to improve the clarity of the target image, after obtaining the target image, a super-resolution operation can be performed on the target image using a preset super-resolution algorithm to obtain a super-resolved target image. During the super-resolution process, the user can specify the resolution of the super-resolved target image according to actual needs, and this disclosure does not impose any restrictions on this.
[0096] The image processing method provided in this embodiment can improve the clarity of the target image by performing a super-resolution operation on the target image, thereby avoiding the problem of unclear image caused by the upper limit of resolution specified by the image generation model.
[0097] Optionally, based on any of the above embodiments, after step 204, the following steps are further included:
[0098] Identify at least one target subject in the target image that meets preset conditions.
[0099] Perform an image restoration operation on at least one target subject using a preset restoration algorithm.
[0100] In this embodiment, since some smaller subjects such as faces and buildings in the target image may be unclear, to ensure the image quality of the target image, at least one target subject in the target image that meets a preset condition may be identified. The preset condition may be a face, body, or building whose size is smaller than a preset threshold. An image restoration operation is then performed on the at least one target subject using a preset restoration algorithm.
[0101] For example, when the target subject is a human face, an image restoration operation can be performed on the target subject using a preset face restoration algorithm to obtain a clearer face.
[0102] The image processing method provided in this embodiment performs image restoration on at least one target subject in the target image that meets preset conditions, thereby avoiding the problem of smaller people or target subjects being blurred and invisible, and further improving the image quality of the target image.
[0103] Optionally, based on any of the above embodiments, after step 204, the following steps are further included:
[0104] Determine the image style corresponding to the target image.
[0105] Use the color parameters corresponding to the image style to perform color adjustment on the target image.
[0106] In this embodiment, to achieve a more harmonious content and color of the generated target image, the target image's corresponding image style can also be determined. For example, the target image's style can be determined to be Hong Kong style, pixel style, realistic style, anime style, etc. Color parameters corresponding to the image style are determined, and color adjustment operations are performed on the target image based on the color parameters. The color parameters corresponding to the image style can be preset empirical values or set by the designer based on different image segmentation methods, and this disclosure does not impose any restrictions on this.
[0107] The image processing method provided in this embodiment performs color adjustment operations on the target image by using color parameters corresponding to the image style for different image styles, so that the color of the target image can better match the image style of the target image, thereby further improving the image quality of the target image.
[0108] Figure 5 is a structural diagram of an image processing device provided by an embodiment of the present disclosure. As shown in Figure 5, the device includes: an acquisition module 51, a determination module 52, a processing module 53 and a generation module 54. The acquisition module 51 is configured to acquire original data determined by a user, and the original data includes at least an original image. The determination module 52 is configured to determine text description information based on the original data and a preset corpus matching library, wherein the corpus matching library includes description information corresponding to a plurality of preset keywords. The processing module 53 is configured to determine image feature information corresponding to the original image, wherein the image feature information includes image structure information and content body information. The generation module 54 is configured to input the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model.
[0109] Furthermore, based on any of the above embodiments, the determination module is configured to: identify scene information corresponding to the original image, determine a first keyword corresponding to the scene information; determine scene description information matching the scene information in a corpus matching library based on the first keyword; identify subject feature information corresponding to the content body in the original image, and determine at least one second associated word corresponding to the subject feature information; determine at least one subject description information matching the subject feature information in the corpus matching library based on the at least one second associated word; and determine text description information corresponding to the original image based on the scene description information and the at least one subject description information.
[0110] Furthermore, based on any of the above embodiments, the raw data also includes at least one tag information determined by a user. The image processing apparatus further includes: a determination module configured to determine at least one scene description information in a corpus matching library based on the at least one tag information. And / or, the raw data also includes raw description information input by a user. The image processing apparatus further includes: an identification module configured to identify at least one keyword corresponding to the raw description information. The determination module is configured to determine at least one scene description information in the corpus matching library based on the at least one keyword.
[0111] Furthermore, based on any of the above embodiments, the processing module is configured to: identify edge contour information and / or image depth information and / or content body posture information in the original image to obtain image structure information; identify at least one content body in the original image; and obtain attribute information of a target portion corresponding to the content body to obtain content body information.
[0112] Furthermore, based on any of the above embodiments, the image processing device further includes: a cropping module configured to perform a cropping operation on the original image according to a preset cropping size to obtain a cropped original image. And / or, a scaling module configured to perform a scaling operation on the original image according to a preset long side threshold and the original ratio corresponding to the original image to obtain a scaled original image. And / or, an identification module configured to identify foreground information corresponding to the original image. A segmentation module configured to perform a foreground segmentation operation on the original image based on the foreground information to obtain a segmented foreground area. The foreground area is filled with a preset background to obtain a filled original image.
[0113] Furthermore, based on any of the above embodiments, the image processing apparatus further includes: a super-resolution module configured to perform a super-resolution operation on the target image through a preset super-resolution algorithm to obtain a super-resolved target image.
[0114] Furthermore, based on any of the above embodiments, the image processing apparatus further includes: a recognition module configured to recognize at least one target subject in the target image that meets preset conditions; and a restoration module configured to perform an image restoration operation on the at least one target subject using a preset restoration algorithm.
[0115] Furthermore, based on any of the above embodiments, the image processing apparatus further includes: a determination module configured to determine an image style corresponding to the target image; and an adjustment module configured to perform a color adjustment operation on the target image using color parameters corresponding to the image style.
[0116] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0117] In order to implement the above embodiments, the embodiments of the present disclosure further provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, an image processing method as described in any of the above embodiments is implemented.
[0118] In order to implement the above embodiments, the embodiments of the present disclosure further provide a computer program product, including a computer program. When the computer program is executed by a processor, the image processing method of any of the above embodiments is implemented.
[0119] In order to implement the above embodiment, the present disclosure further provides an electronic device, including: a processor and a memory;
[0120] Memory stores computer-executable instructions;
[0121] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the image processing method according to any one of the above embodiments.
[0122] FIG6 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. The electronic device 600 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG6 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0123] As shown in Figure 6, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0124] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 6 shows an electronic device 600 having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0125] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0126] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0127] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0128] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0129] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0131] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0132] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0133] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0134] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0135] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0136] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An image processing method, comprising: Acquiring original data determined by a user, wherein the original data at least includes an original image; Determining text description information based on the original data and a preset corpus matching library, wherein the corpus matching library includes description information corresponding to a plurality of preset keywords; Determining image feature information corresponding to the original image, wherein the image feature information includes image structure information and content main body information; The text description information and the image feature information are input into a preset image generation model to obtain a target image output by the image generation model.
2. The method according to claim 1, wherein The determining of text description information based on the original data and a preset corpus matching library includes: Identifying scene information corresponding to the original image, and determining a first keyword corresponding to the scene information; Determining, in the corpus matching library, scene description information that matches the scene information based on the first keyword; Identifying subject feature information corresponding to a content subject in the original image, and determining at least one second associated word corresponding to the subject feature information; Determining at least one subject description information matching the subject feature information in the corpus matching database based on the at least one second associated word; Text description information corresponding to the original image is determined based on the scene description information and the at least one subject description information.
3. The method according to claim 2, wherein: The original data further includes at least one tag information determined by the user; and the method further includes: Determining at least one scene description information in the corpus matching library based on the at least one tag information; And / or, the original data also includes original description information input by the user; the method further includes: Identifying at least one keyword corresponding to the original description information; At least one scene description information is determined in the corpus matching library based on the at least one keyword.
4. The method according to claim 1, wherein The determining of the image feature information corresponding to the original image includes: Identifying edge contour information and / or image depth information and / or content main body posture information in the original image to obtain the image structure information; identifying at least one content body in the original image; Acquire the attribute information of the target part corresponding to the content body to obtain the content body information.
5. The method according to any one of claims 1 to 4, wherein: After obtaining the original data determined by the user, the method further includes: Performing a cropping operation on the original image according to a preset cropping size to obtain a cropped original image; and / or, performing a scaling operation on the original image according to a preset long side threshold and an original ratio corresponding to the original image to obtain a scaled original image; and / or, Identifying foreground information corresponding to the original image; performing a foreground segmentation operation on the original image based on the foreground information to obtain a segmented foreground area; The foreground area is filled with a preset background to obtain the filled original image.
6. The method according to any one of claims 1 to 4, wherein: After inputting the text description information and the image feature information into a preset image generation model and obtaining a target image output by the image generation model, the method further includes: A super-resolution operation is performed on the target image using a preset super-resolution algorithm to obtain a super-resolved target image.
7. The method according to any one of claims 1 to 4, wherein: After inputting the text description information and the image feature information into a preset image generation model and obtaining a target image output by the image generation model, the method further includes: Identifying at least one target subject in the target image that meets preset conditions; An image restoration operation is performed on the at least one target subject using a preset restoration algorithm.
8. The method according to any one of claims 1 to 4, wherein: After inputting the text description information and the image feature information into a preset image generation model and obtaining a target image output by the image generation model, the method further includes: Determining an image style corresponding to the target image; A color adjustment operation is performed on the target image using color parameters corresponding to the image style.
9. An image processing device comprising: an acquisition module, configured to acquire original data determined by a user, wherein the original data at least includes an original image; a determination module configured to determine text description information based on the original data and a preset corpus matching library, wherein the corpus matching library includes description information corresponding to a plurality of preset keywords; a processing module configured to determine image feature information corresponding to the original image, wherein the image feature information includes image structure information and content body information; and The generation module is configured to input the text description information and the image feature information into a preset image generation model to obtain a target image output by the image generation model.
10. An electronic device comprising: processor and memory; wherein the memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the image processing method according to any one of claims 1 to 8.
11. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the image processing method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Virtual content generation method and device, electronic equipment and storage medium
CN114904270A
Image generation method and device, electronic equipment and storage medium
CN115115509A
Method and device for determining picture description information generation model, medium and equipment
CN116050496A
Generating Alternative Descriptions for Images
US20140146053A1
Cited By
Identification information identification method and device, equipment, readable storage medium and product
CN115222969A