IMAGE PROCESSING METHOD AND APPARATUS AND ELECTRONIC DEVICE

The image processing method and apparatus address the limitation of traditional photography by generating personalized photographs in multiple styles by merging objects from different images, improving user experience and efficiency.

DE102025103188A1Pending Publication Date: 2025-08-07LENOVO (BEIJING) LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102025103188
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-01
Filing Date
2025-01-29
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Current photography techniques do not allow for the direct capture of personalized photographs in various styles without post-processing.

Method used

An image processing method and apparatus that utilizes an image engine to determine feature information, generates a target image based on input contents, and merges objects from the first and second images to create a personalized image.

Benefits of technology

Enables the direct generation of personalized photographs in various styles, simplifying post-processing and enhancing user experience by utilizing existing elements in the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An image processing method comprises obtaining an image, processing the image based on an image engine to determine feature information contained in the image, obtaining input content, determining a target feature from the feature information contained in the image based on the input content, and generating a target image based at least on the target feature and the input content. The target image includes a first object corresponding to the target feature and a second object corresponding to a content description of the target image generated based on the input content.
Need to check novelty before this filing date? Find Prior Art

Description

REFERENCE TO RELATED APPLICATIONThis application claims priority to Chinese Patent Application No. 202410144292.X filed on Feb. 1, 2024, the entire contents of which are hereby incorporated by reference.TECHNICAL FIELDThe present disclosure relates generally to the field of image processing technology, and more particularly to an image processing method and apparatus and an electronic device.BACKGROUND ARTCurrently, the technique of photography can only capture real world images. If users wish to get personalized photos, they can only perform post-processing of the photos. It is impossible to directly obtain photographs of various styles by curative photography.SUMMARY OF THE INVENTIONAccording to the disclosure, there is provided an image processing method including obtaining an image, processing the image based on an image engine to determine feature information included in the image, obtaining input contents, determining a target feature from the feature information included in the image based on the input contents, and generating a target image based at least on the target feature and the input contents. The target image includes a first object corresponding to the target feature and a second object corresponding to a content description of the target image generated based on the input contents.Also according to the disclosure, there is provided an electronic apparatus including an image capturing device configured to acquire an image, and a processor configured to process the image based on an image engine to determine feature information included in the image, acquire input contents, determine a target feature from the feature information included in the image based on the input contents, and generate a target image based on at least the target feature and the input contents. The target image includes a first object corresponding to the target feature and a second object corresponding to a content description of the target image generated based on the input contents.Also according to the disclosure, there is provided an image processing method including inputting an image candidate to an image-to-text model to generate text information, determining, based on a semantic model, similarity between the text information of the image candidate and text information of input contents, and outputting, in response to the similarity satisfying a target threshold, the image candidate as a target image.BRIEF DESCRIPTION OF THE DRAWINGSIn order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the description of the embodiments will be briefly presented below. The drawings described below are some embodiments of the present disclosure. The skilled person can obtain other drawings based on these drawings without any creative work.In the drawings, like or corresponding parts bear like or corresponding reference numerals. FIG. 1 is a flowchart of an image processing method according to embodiments of the present disclosure. FIG. 2 is a schematic diagram illustrating generation of a target image according to embodiments of the present disclosure. FIG. 3 is a schematic structural diagram of an image processing apparatus according to embodiments of the present disclosure. FIG. 4 is a schematic structural diagram of an electronic device according to embodiments of the present disclosure.DETAILED DESCRIPTION OF THE EMBODIMENTSEmbodiments of the present disclosure will be described in conjunction with the drawings. Obviously, the described embodiments are only some embodiments of the present disclosure, not all embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without any creative work are within the scope of the present disclosure.The current technique of photography does not allow to directly obtain photographs of various styles by curative photography. In order to achieve conservative photography to obtain photographs of various styles, the present disclosure provides an image processing method, an image processing apparatus, and an electronic device. The electronic device provided by the present disclosure may be a mobile phone, a computer, a tablet computer, or another device.In this specification, various generative models are proposed that may be used to process natural language (NL) content and / or other input to generate outputs that reflect generative content that is a response to the input. For example, large language models (LLMs) have been developed that can be used to process NL content and / or other inputs to output LLM outputs that reflect NL generative content and / or other generative content that is a response to the inputs. These LLMs are typically trained using enormous amounts of a wide variety of data, including data from, but not limited to, web pages, electronic books, software code, electronic news articles, and machine translation data. Accordingly, these LLMs utilize the underlying data they were trained to perform in these various NLP tasks. For example, upon completion of a voice generation task, these LLMs may process a natural language (NL)-based input received from a client device and generate a response to the NL-based input to be displayed on the client device. In many cases, these LLMs may cause textual content to be included in the response. In some cases, these LLMs may additionally or alternatively cause multimedia content such as images to be included in the response (e.g., based on triggering an image fetch, based on causing image generation models to generate images, etc.).The present description describes a system implemented in the form of computer programs on one or more computers at one or more locations and implements a generative text-to-image model configured to generate an image based on a text prompt. In some cases, the text prompt specifies a particular style of the image, and the generative text-to-image model processes the text prompt to generate an image having the particular style specified by the text prompt. In some cases, the text prompt specifies a particular object (e.g., a particular person, animal, vehicle, boat, etc.) or a specific instance of an object to appear in the image, wherein the generative text-to-image model processes the text prompt to generate an image representing the particular object / instance specified by the text prompt. In some other cases, the text prompt specifies both the particular style of the image and the particular object / instance to appear in the image, wherein the generative text-to-image model processes the text prompt to generate an image that (i) has the particular style and (ii) represents the particular object / instance.The technical scheme according to embodiments of the present disclosure will be described below in conjunction with the drawings in the embodiments of the present disclosure.FIG. 1 shows a flowchart of an image processing method provided by the present disclosure. As illustrated in FIG. 1, the image processing method in one embodiment includes steps S 101 to S 105.In S 101, a first image is obtained.In one embodiment, an image capturing device may be used to immediately capture an image required by the user as the first image. Alternatively, in some other embodiments, the required image as the first image may also be obtained from the user's album. The user's album may be a local album or a cloud album. Alternatively, in other embodiments, the required image may be downloaded from the network as the first image. The image pickup device may be a camera or a mobile phone capable of capturing images.In S 102, the first image is processed based on an image engine to determine feature information included in the first image.The image engine may include an image feature extraction component. In one embodiment, the image feature extraction component may be used in the image engine to extract the feature information, such as texture, color, or shape, of each element in the first image. The elements in the first image may include foreground objects and background areas in the first image. Foreground objects may be people, animals, objects, etc. The feature extraction component may use a software program with a scale invariant feature transform (SIFT) algorithm, a histogram of oriented gradients (HOG) software program, or a deep learning model such as a convolutional neural network.In S 103, input contents are obtained, and the input contents are used to generate a content description of a target image.In one embodiment, the input content may be used to generate the content description of the target image, and the input content may include all of the elements needed for the target image. For example, when the input contents are "create a selfie of user A at a beach on the tenants", the items required to generate the target image may include "an image of user A" and "any beach landscape image on the tenants".The input contents may be voice contents or text contents. If the input contents are voice contents, the voice contents may be converted into corresponding text contents. Based on the natural language processing model, the text content may be understood and processed and the semantic features of the vocabulary included in the text content that may be used to generate the target image extracted. A feature set comprising the semantic features of the vocabulary that can be used to generate the target image can be obtained as the content description of the generated target image. The natural language processing model may include an RNN (Recurrent Neural Network) or a CNN (Convolutional Neural Network) capable of acquiring context information.In S 104, one or more target features are determined from the feature information included in the first image based on the input contents.In an embodiment, when the input contents include all the items required to generate the target image, all the items included in the first image may be determined by the feature information included in the first image. Then, the items included in the first image may be compared with the items included in the input contents, the items included in the first image that overlap with the items included in the input contents may be retained, and the features corresponding to the overlapping items may be determined as a target feature.By determining the target feature or features from the feature information included in the first image through the input contents, it may be possible to generate the target image using the elements present in the first image, thereby improving the generation efficiency of the target image.In S 105, the target image is generated based on the target feature or features and the input contents, wherein the target image includes the first object corresponding to the target feature or features and the second objects corresponding to the content description.As illustrated in FIG. 2, which is a schematic diagram showing the generation of the target image, in one embodiment, the input contents are "generate a photograph of user A in a full-sunny forest", and the input contents include the items "user A", "full-sunny forest", and "photograph" required for generating the target image. The first image 201 is obtained, and the figure included in the image 201 is the figure image of user A. The feature information included in the first image 201, i.e., the feature information of the half-body image 202 of user A and the feature information of the background area image, are extracted by the image engine. Then, the items included in the input contents are compared with the item or items corresponding to the feature information included in the first image 201, and the feature or features of the overlapping item "half-body image of user A" are obtained as a target feature. Then, based on the target feature and the other items "full sunny forest" and "photo" remaining in the input contents, the target image 203 is generated. The target image 203 includes the foreground area with the half-body image of user A and the background area with the full-sunny forest.In the image processing method provided by the present disclosure, the first image can be obtained. Then, the first image may be processed based on the image engine to determine the feature information included in the first image. The input contents may be obtained and the input contents may be used to generate the content description of the target image. The target feature or features may be determined based on the input contents from the feature information included in the first image. The target image may be generated based on the target feature or features and the input content, and the target image may include the first object corresponding to the target feature or features and the second object corresponding to the content description. The content description of the target image may be generated based on the input contents, and the target features required for generating the target image may be determined using the input contents from the first image, that is, the target features in the first image may be determined flexibly using the input contents. Then, the first object of the target image may be generated based on the target feature or features and the second objects of the target image may be generated using the content description, thereby achieving generation of multiple style, curative photos and improving the user experience.In an embodiment, determining the target feature or features from the feature information included in the first image based on the input contents may include:A1, based on a first feature set and a second feature set, determining one or more matching features in the first feature set and the second feature set, wherein the matching features are used as a target feature.The first feature set and the second feature set may be feature sets of the same type. The first feature set may be a feature set corresponding to the first image, and the second feature set may be a feature set corresponding to the input contents.In an embodiment, a set including the features of all elements in the extracted first image may be used as the first feature set, and a set including the features of each element included in the input contents may be used as the second feature set. The first feature set and the second feature set may be feature sets of the same type. For example, when the features included in the first feature set are a feature matrix corresponding to the elements included in the first image, the features included in the second feature set may also be a feature matrix corresponding to the features included in the input contents. When the features included in the first feature set are the elements included in the first image, the features included in the second feature set may also be the features included in the input contents.For each feature in the first feature set, the similarity between that feature and each feature in the second feature set may be calculated. When there is a feature whose similarity to the feature is greater than a preset similarity threshold in the second feature set, it may be determined that the feature in the second feature set whose similarity to the feature is greater than the preset similarity threshold is a matching feature. The determined matching features in the first feature set may be used as a target feature.The similarity between the feature in the first feature set and the feature in the second feature set may be calculated using a cosine similarity or root mean square algorithm.For example, when the first image is a landscape image with user A and the background scene in the landscape image is a road, the features included in the first feature set may be an image matrix X 1 representing users A and an image matrix X 2 representing the road. The input contents may be "Create a photograph of user A in a full sunny forest.". The features included in the second feature set may be an image matrix Y 1 corresponding to users A and an image matrix Y 2 of the full-sundly forest.For the image matrix X1representing users A, the similarity between the image matrix X1representing users A and the image matrix Y1corresponding to users A and the image matrix Y2of the full-sunny forest can be calculated. When the similarity between the image matrix X 1 representing users A and the image matrix Y 1 corresponding to users A is greater than the preset similarity threshold, it may be determined that the image matrix X 1 representing users A is a matching matrix. For the image matrix X2representing the road, the similarity between the image matrix X2representing the road and the image matrix Y1corresponding to users A and the image matrix Y2of the full-sunny forest can be calculated. When the similarity between the image matrix X 2 representing the road and the image matrix Y 1 corresponding to users A and the image matrix Y 2 of the full-sundock forest is not greater than the preset similarity threshold, the image matrix X 2 representing the road may be a non-matching matrix. Therefore, it can be determined that the image matrix X1 representing users A belongs to the target features.The preset similarity threshold may be set depending on the actual application scenario, for example, it may be set to 98% or 99%.In one embodiment, generating the target image based on the target feature or features and the input content may include steps B 1-B 3.In B 1, the first object corresponding to the target feature or features in the first image is determined.In an embodiment, the positions of the first object in the first image may be associated with the target features. The area corresponding to the target features in the first image may be extracted by the position and area information associated with the target features, and it may be determined that the objects in the area are the first object. For example, the position information of each pixel in the half-body image of user A may belong to the target features, and the region corresponding to the half-body image of user A in the first image may be extracted according to the position data of each pixel as the first object.In B 2, a second image is generated based on the non-matching features in the second feature set.Features in the second feature set that are capable of matching the first feature set may be matching features and the other features in the second feature set that are not matching features may be non-matching features. In the present disclosure, as a second image, a new image may be generated based on the non-matching features. For example, the second feature set may include a feature T 1 of the half-body image of user A, a feature T 2 of the didive beach image, and a feature T 3 of the image of blue sky and white clouds. The feature T 1 of the half-body image of user A may be a matching feature, and the feature T 2 of the didive beach image and the feature T 3 of the image of blue sky and white clouds may be non-matching features. An image including the didive beach and the blue sky and white clouds may be generated based on the feature T 2 of the didive beach image and the feature T 3 of the image of blue sky and white clouds as the second image.In B 3, the target image is obtained by fusion based on the first object and the second image.The first object may be merged into the second image and the generated merged image may be used as the target image. For example, the generated image with the didive beach and the blue sky and white clouds may be the second image, and the first object may include the area corresponding to the half-body image of user A in the first image. Therefore, the area corresponding to the half-body image of user A in the first image may be extracted and merged into the second image, the half-body image of user A may be used as the foreground area of the merged image, and the image area of the didive beach and the blue sky and the white cloud may be used as the background area of the merged image. The obtained merged image may be the target image. The target image may include all elements in the input contents.By applying this method, the new target image can be generated according to user needs by utilizing existing elements in the first image and the second image generated according to the non-matching features in the second feature set. This can not only meet various photography needs of the user, but also simplify or eliminate subsequent photo-editing operations of the user, thereby improving photography efficiency.In another embodiment, generating the target image based on the target feature and the input contents may include steps C 1-C 3.In C 1, the first object corresponding to the target feature or features in the first image is determined.For details of C1, reference may be made to the above description of B1, which is not repeated here.In C 2, the second image is determined by screening based on the non-matching features in the second feature set and the feature information of each image in an image set.In an embodiment, the image set may be a local album or a cloud album of an electronic device, for example, a mobile phone album or a cloud album of the user. In this step, the similarity between the image feature or features of each image in the image set and the non-matching features in the second feature set may be calculated based on the non-matching features in the second feature set.When an image having a similarity greater than a similarity threshold is present, it may be determined that the image is the second image. The similarity threshold may be set depending on the actual application scenario, for example, it may be set to 97% or 98%.In C 3, the target image is obtained by fusion based on the first object and the second image.The first object may be merged into the second image and the generated merged image may be used as the target image.By applying this method, the new target image can be generated according to the user needs by using the existing element or elements in the first image and the second image including the non-matching features in the second feature set, which not only satisfies various photography needs of the user but can also simplify or eliminate subsequent photo-processing operations of the user, thereby improving photography efficiency.In another embodiment, generating the target image based on the target feature and the input contents may include steps D 1-D 5:In D 1, the first object corresponding to the target feature in the first image is determined.For details of D1, reference may be made to the above description of B1, which is not repeated here.In D 2, a first feature and a second feature are determined based on the non-matching features of the second feature set, wherein the first feature is used to indicate the target object.In one embodiment, the target object may be a person, an animal, or a landscape element.In D 3, the target object is determined by screening based on the first feature and the feature information of each image in the image set.The target object may also appear in the images in the image set. Therefore, the first feature may be used to compare the feature information of each image in the image collection, calculate the feature similarity between the first feature and the feature information of each image in the image collection, perform filtering to obtain the feature information whose feature similarity is greater than the feature similarity threshold, and determine that the object of the image in the image collection corresponding to the feature information is the target object. The feature similarity threshold may be set depending on the actual application scenario, for example, it may be set to 97% or 98%.For example, the non-matching features of the second feature set may include the feature T 2 of the didive beach image, the feature T 3 of the image of blue sky and white clouds, and the feature T 4 of the half-body image of user B. Therefore, the feature T 4 of the half-body image of user B may be used as the first feature, and the feature T 2 of the didive beach image and the feature T 3 of the image of blue sky and white clouds may be used as the second feature. The target object indicated by the first feature may be the half-body image of user B. Then, the images in the local album can be obtained, and the feature similarity between the image feature or features in each image in the local album and the feature T 4 of the half-body image of user B can be calculated. It may be determined that the image having a feature similarity greater than the preset feature similarity threshold is the selected image, and the object corresponding to the feature T 4 of the half-body image of user B in the selected image is the target object.In D 4, the second image is generated based on the second feature.In an embodiment, as the second image, a new image according to each second feature may be generated. For example, the second feature may include the feature T 2 of the didive beach image and the feature T 3 of the image of blue sky and white clouds. Therefore, an image including the didive beach and the blue sky and the white clouds according to the feature T 2 of the didive beach image and the feature T 3 of the image of blue sky and white clouds can be generated as the second image.In D 5, the target image is obtained by fusion based on the first object, the second object, and the second image.The first object and the second object may be merged into the second image, and the generated merged image may be used as the target image. For example, the generated image with the didive beach and the blue sky and the white clouds may be the second image, the first object may be the area corresponding to the half-body image of user A in the first image, and the second image may be the area corresponding to the half-body image of user B in the image in the local album. Then, the area corresponding to the half-body image of user A in the first image and the area corresponding to the half-body image of user B in the image in the local album may be extracted, and the extracted area corresponding to the half-body image of user A and the area corresponding to the half-body image of user B may be merged into the second image. The half-body image of user A and the half-body image of user A may be used as foreground areas of the merged image, and the image area of the didive beach and the blue sky and the white cloud may be used as background areas of the merged image, and the obtained merged image may be the target image.By using this method, the new target image can be generated according to the user needs by using the existing items in the first image and the second image generated by the non-matching items in the second item set, which not only satisfies various photography needs of the user but also simplifies or eliminates subsequent photo-editing operations of the user, thereby improving photography efficiency. Additionally, objects that are not present in the first image may be augmented by the local album, and for objects that are not present in the local album, a corresponding image may be generated directly based on the second feature from the non-matching features in the input content, which significantly improves the user experience.In an embodiment, obtaining the first image may include:E1, obtaining the first image by an image capturing device, wherein the first image contains the geographical location at which the first image was obtained.Based on the first feature set and the second feature set, determining the matching features from the first feature set and the second feature set may include: determining one or more features in the first feature set corresponding to the geographic location as a matching feature.In an embodiment, when capturing the image, the image capturing device may simultaneously obtain the current geographic location information and time information, and store the geographic location information and time information as the image information of the captured image in the file corresponding to the image. For example, after the image is captured by the image capturing device, an image file may be generated, and the image file may include description information of the element or elements included in the image, as well as the location information and the time information of capturing the image. In one embodiment, a corresponding smart album may be generated based on the image description information included in the image file. For example, an album of images all captured at a particular location may be generated, and images captured at the same photographing location may be stored in the album. Or, an album of images all including the same object may be generated and images including the object may be stored in the album.The obtained first image may include the location information and time information corresponding to capturing the image. In determining a match of the features included in the first feature set and the second feature set, the location information of the image may also be one of the match conditions.In an embodiment, generating the target image based on the target feature and the input contents may include:F1 generating an image candidate based on the target feature and the input contents, and processing the image candidate based on an image-to-text model to generate text information; andF 2, determining the similarity between the text information of the image candidate and the text information of the input contents based on the semantic model, and when the similarity satisfies the target threshold, using the image candidate as the required target image.In one embodiment, the image-to-text model may be a model capable of generating descriptive text or speech for an image, and is an example of a generative model. A generative model is a machine learning model configured to generate new data that is similar to the training data. Artificial intelligence (AI) generative models learn the patterns and distributions of the training data and then apply these findings to generate novel content in response to new input data. Generative models have a wide range of applications including image generation, speech synthesis, and natural language generation. With the image-to-text model, an input image may be processed using, for example, convolutional neural networks (CNNs) to extract spatial features (e.g., shape, texture, and objects) from the image, and the extracted features may be processed using an RNN, a long-term memory (LTM) network, or a transformer to generate a sequence of words that may be output as descriptive text or speech.Other types of generative models may include text-to-image models and image-to-image models. A text-to-image model may generate images based on textual descriptions. With the text-to-image model, input text (e.g., a sentence or expression) is processed using a language model (for example, bidirectional encoder representation of transformers (BERT), generative pre-drawn transformer (GPT), or contrastive language-image pre-training (CLIP) to transform the input text into a vector representation, which can then be fed into, for example, a generative advanced network (GAN) or a diffusion model, to map textual features to visual features. The image-to-image model may generate new images by transforming an input image based on specific conditions or tasks. With the image-to-image model, an input image can be processed through multiple layers of a network to extract important features such as texture, shape, and spatial relationships within the image, which can then be used to generate a corresponding output image that reflects a specific transformation.For the generated target image, the description text corresponding to the target image, i.e., text information, may be generated from the image-to-text model. Then, the text similarity between the generated text information and the text information corresponding to the input contents can be calculated. When the similarity is not less than the target threshold, it may be determined that the generated target image is the required target image. A similarity less than the target threshold may indicate that there may be a large difference between the generated target image and the image described by the input contents, and it is necessary to re-generate the target image. The target threshold may be set to 99% or 99.5%, depending on the actual scenario.In applying this method, the generated target image may be checked to ensure that the generated target image is an image corresponding to the input contents, thereby improving the accuracy of the generated target image.The present disclosure also provides an image processing device. In an embodiment illustrated in FIG. 3, which is a schematic structural diagram of an image processing apparatus, the apparatus includes:an image acquisition module 301 configured to obtain a first image;an image engine 302 configured to process the first image and determine feature information included in the first image;a description content acquisition module 303 configured to obtain input content, wherein the input content is used to generate a content description of a target image;a feature extraction module 304 configured to determine a target feature from the feature information included in the first image based on the input contents; andan image generation module 305 configured to generate the target image based on the target feature and the input contents, wherein the target image includes a first object corresponding to the target feature and a second object corresponding to the content description.Using the image processing device provided by the present disclosure, the first image can be obtained. The first image may be processed based on the image engine to determine the feature information included in the first image. The input contents may be obtained and may be used to generate the content description of the target image. The target feature may be determined based on the input contents from the feature information included in the first image. The target image may be generated based on the target feature and the input contents. The target image may include a first object corresponding to the target feature and a second object corresponding to the content description. The content description of the target image may be generated by inputting the contents, and the target feature required for generating the target image may be determined using the input contents from the first image, that is, the target feature in the first image may be determined flexibly using the input contents. And then, the first object of the target image may be generated based on the target feature and the second object of the target image may be generated using the content description, thereby achieving generation of multiple style, curative photos and enhancing the user experience.In an embodiment, the feature extraction module 304 may be configured to determine the matching feature or features in the first feature set and the second feature set based on the first feature set and the second feature set. The matching feature or features may be used as a target feature. The first feature set and the second feature set may be feature sets of the same type. The first feature set may be the feature set corresponding to the first image, and the second feature set may be the feature set corresponding to the input contents.In one embodiment, the image generation module 305 may be configured to: determine the first object corresponding to the target feature in the first image; generate the second image based on the non-matching feature or features in the second feature set; and merge the target image based on the first object and the second image.In one embodiment, the image generation module 305 may be configured to: determine the first object corresponding to the target feature in the first image; determine the second image based on the non-matching feature or features in the second feature set and the feature information of each image in the image set; and merge the target image based on the first object and the second image.In one embodiment, the image generation module 305 may be configured to: determine the first object corresponding to the target feature in the first image; determine the first feature and the second feature based on the non-matching feature and features, respectively, in the second feature set, the first feature being used to indicate the target object; determine the target object based on the first feature and the feature information of each image in the image set; and generate the second image based on the second feature; and merge the target image based on the first object, the second object, and the second image.In an embodiment, the image acquisition module 301 may be configured to obtain the first image by an image acquisition device. The first image may include the geographic location of the first image. Based on the first feature set and the second feature set, determining the matched feature or features from the first feature set and the second feature set may include the first feature set including the geographic location to determine whether it belongs to the matched features.In an embodiment, the apparatus may further comprise a verification module (not shown in the drawings). The verification module may be configured to: process an image candidate generated based on the target feature and the input contents based on the image-to-text model to generate text information; determine, based on the semantic model, the similarity between the text information of the target image and the text information of the input contents; and use, when the similarity satisfies the target threshold, the target image as the required target image.The present disclosure also provides an electronic device. The electronic device may include an image capturing device configured to capture images and a processor.The processor may be configured to: obtain the first image by the image capturing device; process the first image based on the image engine to determine the feature information included in the first image; obtain the input contents, the input contents being used to generate a content description of a target image; determine the target feature from the feature information included in the first image based on the input contents; and generate the target image based on the target feature and the input contents, the target image including a first object corresponding to the target feature and a second object corresponding to the content description.As illustrated in FIG. 4, which is a schematic structural diagram of an example electronic device 400, the electronic device is intended to represent various forms of digital computers, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, or other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, or other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or needed herein.As illustrated in FIG. 4, the apparatus 400 includes: a computing unit 401 configured to perform various suitable actions and processes according to a computer program stored in a read only memory (ROM) 402, or a computer program loaded from a storage unit 408 to a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the device 400. The arithmetic unit 401, the ROM 402, and the RAM 403 may be connected to each other via a bus 404. An input / output (I / O) interface 405 may also be coupled to bus 404.Multiple components in device 400 may be connected to I / O interface 405 including: an input unit 406, for example, a keyboard, a mouse, etc.; an output unit 407, for example, various types of displays, speakers, etc.; a storage unit 408, for example, a hard disk, an optical disk, etc.; and a communication unit 409, for example, a network card, a modem, a wireless communication transceiver, etc. Communication unit 409 may allow device 400 to exchange information or data with other devices via a computer network such as the Internet and / or various telecommunications networks.The computing unit 401 may be various general and / or special purpose processing components with processing and computing capabilities. Examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), as well as any suitable processors, controllers, microcontrollers, etc. The computing unit 401 may perform the various methods and processes described above, for example, the image processing method. For example, in some embodiments, the image processing method may be implemented as computer program software that is physically embodied in a machine readable medium such as a storage unit 408. In some embodiments, the computer program may be fully or partially loaded onto and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the arithmetic unit 401, one or more steps of the above-described image processing method may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform the image processing method in any other suitable manner (e.g., using firmware).The electronic device 400 may also include an image capturing device.Various embodiments of the systems and approaches described above may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or a combination thereof. These various embodiments may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system that may include at least one programmable processor, which may be a special or general purpose programmable processor that may receive data and instructions from a storage system, at least one input device and at least one output device, and may transmit data and instructions to the storage system, the at least one input device and the at least one output device.The program code for implementing the methods disclosed herein may be embodied in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer or other programmable data processing device such that the program code, when executed by the processor or controller, causes the functions indicated in the flowchart and / or block diagram to be implemented. The program codes may execute entirely on the machine, partly on the machine, partly on the machine and partly on a remote machine as a stand-alone software package or entirely on a remote machine or server.In the context of the present disclosure, a machine-readable medium may be a physical medium that may include or have stored a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. Further specific examples of machine readable storage media would include electrical connections based on one or more wires, a portable computer drive, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or Flash memory), optical fibers, a portable CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination thereof.To provide for interaction with a user, the systems and approaches described herein may be implemented on a computer that includes: a display device (e.g., a cathode ray tube or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user may provide input to the computer. Other types of devices may also be used to provide the interaction with the user; for example, the feedback provided to the user may be sensory feedback of any form (e.g., visual feedback, audible feedback, or tactile feedback); and the input from the user may be received in any form (including as audible input, voice input, or tactile input).The systems and policies described herein may be implemented in a computing system that includes a back-end component (e.g., as a data server), or in a computing system that includes a middleware component (e.g., an application server), or in a computing system that includes a front-end component (e.g., a graphical user interface user computer or a web browser through which the user(s) may interact with an implementation of the systems and policies described herein), or in a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by digital data communication of any form or medium (e.g., through a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.A computer system may include a client and a server. The client and server may generally be remote to each other and normally interact over a communication network. The client-server relationship may be generated by computer programs running on respective computers and having a client-server relationship to each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to address the deficiencies of the difficult manageability and poor business scalability of traditional physical hosts and virtual private server (VPS) services. The server may also be a server of a distributed system or a server combined with a blockchain.It should be appreciated that the various forms of processes shown above may be used to rearrange, add, or omit steps. For example, the steps described in this disclosure may be performed in parallel, sequentially, or in another order as long as the desired results of the technical solutions disclosed in this disclosure can be achieved; this document does not impose any limitations on them herein.Moreover, the terms "first(r / s)" and "second(r / s)" are used for description purposes only and are not to be understood as indicating or impacting relative importance or as implicitly indicating the number of technical features. Thus, the features associated with "first(r / s)" and "second(r / s)" may explicitly or implicitly comprise at least one of the features. In the description of this disclosure, "multiple" means two or more unless clearly and concretely defined otherwise.Various embodiments have been described to illustrate the operating principles and implementation examples. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments described herein, and various other changes, rearrangements, and substitutions may be made. Thus, although the present disclosure has been described in detail with reference to the above-described embodiments, the present disclosure is not limited to the above-described embodiments, but may be embodied in other equivalent forms without departing from the spirit and scope of the present disclosure.References included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Patent Literature citedCN 202410144292

[0001]

Claims

An image processing method comprising: obtaining an image; processing the image based on an image engine to determine feature information included in the image; obtaining input contents; determining a target feature from the feature information included in the image based on the input contents; and generating a target image based at least on the target feature and the input contents, wherein the target image includes a first object corresponding to the target feature and a second object corresponding to a content description of the generated target image generated based on the input contents.The method of claim 1, wherein determining the target feature comprises: determining a matching feature from a first feature set and a second feature set as the target feature, wherein the first feature set and the second feature set are of the same type, the first feature set corresponds to the first image, and the second feature set corresponds to the input contents.The method of claim 2, wherein: the image is a first image; and generating the target image comprises: determining the first object; generating a second image based at least on a non-matching feature in the second feature set; and obtaining the target image by fusion based on the first object and the second image.The method of claim 2, wherein: the image is a first image; and generating the target image comprises: determining the first object; determining a second image by screening based at least on a non-matching feature in the second feature set and feature information of each image in an image set; and obtaining the target image by fusion based on the first object and the second image.The method of claim 2, wherein: the image is a first image; and generating the target image comprises: determining the first object; determining a first feature and a second feature based at least on a non-matching feature of the second feature set, the first feature indicating a target object; determining the target object by screening based on the first feature and feature information of each image in an image set; generating a second image based on the second feature; and obtaining the target image by fusion based on the first object, the second object, and the second image.The method of claim 2, wherein: obtaining the image comprises obtaining the image by an image capturing device, the image including a geographic location at which the image was obtained; and determining the matching feature comprises determining a feature in the first feature set corresponding to the geographic location as the matching feature.The method of claim 1, wherein generating the target image comprises: generating an image candidate based on the target feature and the input contents; processing the image candidate based on an image-to-text model to generate text information of the image candidate; determining a similarity between the text information of the image candidate and text information of the input contents based on a semantic model; and determining that the image candidate is the target image in response to the similarity satisfying a target threshold.An electronic device, comprising: an image capturing device configured to obtain an image; and a processor configured to: process the image based on an image engine to determine feature information included in the image; obtain input contents; determine a target feature from the feature information included in the image based on the input contents; and generate a target image based at least on the target feature and the input contents, wherein the target image includes a first object corresponding to the target feature and a second object corresponding to a content description of the target image generated based on the input contents.The electronic device of claim 8, wherein the processor is further configured to, when determining the target feature: determine a matching feature from a first feature set and a second feature set as the target feature, wherein the first feature set and the second feature set have the same type, the first feature set corresponds to the first image, and the second feature set corresponds to the input contents.The electronic device of claim 9, wherein: the image is a first image; and the processor is further configured to, when generating the target image: determine the first object; generate a second image based at least on a non-matching feature in the second feature set; and obtain the target image by fusion based on the first object and the second image.The electronic device of claim 9, wherein: the image is a first image; and the processor is further configured to, when generating the target image: determine the first object; determine a second image by screening at least based on a non-matching feature in the second feature set and feature information of each image in an image set; and obtain the target image by fusion based on the first object and the second image.The electronic device of claim 9, wherein: the image is a first image; and the processor is further configured to, when generating the target image: determine the first object; determine a first feature and a second feature based at least on a non-matching feature of the second feature set, the first feature indicating a target object; determine the target object by screening based on the first feature and feature information of each image in an image set; generate a second image based on the second feature; and obtain the target image by fusion based on the first object, the second object, and the second image.The electronic device of claim 9, wherein: the image includes a geographic location at which the image was obtained; and the processor is further configured to determine, upon determining the matching feature, a feature in the first feature set corresponding to the geographic location as the matching feature.The electronic device of claim 8, wherein the processor is further configured to, when generating the target image: generate an image candidate based on the target feature and the input contents; process the image candidate based on an image-to-text model to generate text information of the image candidate; determine a similarity between the text information of the image candidate and text information of the input contents based on a semantic model; and determine that the image candidate is the target image in response to the similarity satisfying a target threshold.An image processing method, comprising: inputting an image candidate into an image-to-text model to generate text information; determining, based on a semantic model, similarity between the text information of the image candidate and text information of input contents; and outputting the image candidate as a target image in response to the similarity satisfying a target threshold.The image processing method according to claim 15, further comprising: processing an input image based on an image engine to determine feature information included in the input image; determining, based on the input contents, a target feature from the feature information included in the input image; and generating the image candidate based on the target feature and the input contents.

Citation Information

Patent Citations

  • 202410144292