Image processing method and device and electronic equipment
By extracting and fusing image features using a pre-trained target neural network model, the problem of long generation time for group photos is solved, thus improving the efficiency of image processing and group photo generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2026-04-14
AI Technical Summary
The process of generating group photos in existing technologies is time-consuming, which affects the user experience.
By extracting image features from multiple first target images using a pre-trained target neural network model, outputting the corresponding second target image, and fusing it with the background image to generate a group photo, the step of uploading images for model training is avoided when generating a group photo.
It improves the efficiency of image visual processing and group photo generation, and reduces processing time.
Smart Images

Figure CN121860863A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to an image processing method, apparatus, and electronic device. Background Technology
[0002] With technological advancements, electronic devices can process images according to user needs, ensuring the processed image meets those needs. For example, an electronic device can process multiple object images (e.g., multiple face images) to obtain a group photo related to those object images. However, generating group photos still suffers from time-consuming issues. Summary of the Invention
[0003] In view of the above problems, this application proposes an image processing method, apparatus, and electronic device to improve the above problems.
[0004] In a first aspect, this application provides an image processing method, the method comprising: acquiring a plurality of first target object images; performing image visual processing on the plurality of first target object images respectively through a target neural network model to obtain a plurality of second target object images, wherein the plurality of second target object images include second target object images corresponding to each of the plurality of first target object images, wherein, during the image visual processing by the target neural network model, the target neural network model extracts image features of the obtained first target object images and outputs corresponding second target object images; and fusing the plurality of second target object images with a background image to obtain a target image.
[0005] Secondly, this application provides an image processing apparatus, the apparatus comprising: a target image acquisition unit for acquiring a plurality of first target object images; a target image processing unit for performing image visual processing on the plurality of first target object images respectively through a target neural network model to obtain a plurality of second target object images, the plurality of second target object images including second target object images corresponding to each of the plurality of first target object images, wherein, during the image visual processing by the target neural network model, the target neural network model extracts image features of the obtained first target object images and outputs corresponding second target object images; and an image fusion unit for fusing the plurality of second target object images with a background image to obtain a target image.
[0006] Thirdly, this application provides an electronic device, which includes at least a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above-described method.
[0007] Fourthly, this application provides a computer-readable storage medium storing program code, wherein the above-described method is executed when the program code is run by a processor.
[0008] This application provides an image processing method, apparatus, and electronic device. In this method, after obtaining multiple first target object images, image features of the first target object images are extracted using a target neural network model. Based on these image features, second target object images corresponding to the first target object images are output, thus obtaining multiple second target object images. Then, the multiple second target object images are fused with a background image to obtain a target image. This method, when the target neural network model is a pre-trained model for image visual processing, allows for the extraction of image features from the first target object images after obtaining them. This, combined with the target neural network model, yields the corresponding second target object images. The target image (group photo image) is then obtained by fusing the multiple second target images with the background image. This avoids the need for users to first upload images for model training before performing image visual processing based on the trained model when generating group photos, thereby improving the efficiency of image visual processing and group photo generation. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A schematic diagram illustrating an application scenario of the image processing method proposed in this application is shown.
[0011] Figure 2 A schematic diagram illustrating another application scenario of the image processing method proposed in the embodiments of this application is shown;
[0012] Figure 3 A flowchart of an image processing method according to an embodiment of this application is shown;
[0013] Figure 4 This illustration shows a background image in this application that includes multiple regions;
[0014] Figure 5 A schematic diagram of a layout position configuration interface according to this application is shown;
[0015] Figure 6A schematic diagram of an image fusion effect in this application is shown;
[0016] Figure 7 A flowchart of an image processing method according to another embodiment of this application is shown;
[0017] Figure 8 A structural block diagram of an image processing apparatus according to an embodiment of this application is shown;
[0018] Figure 9 A structural block diagram of another electronic device for performing an image processing method according to an embodiment of the present application is shown;
[0019] Figure 10 This is a storage unit in this application embodiment for storing or carrying program code that implements the image processing method according to this application embodiment. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0021] With advancements in technology, modern electronic devices can intelligently respond to user needs and process images. Taking group photo generation as an example, electronic devices can integrate images of multiple objects to create a combined group photo effect. However, despite the increasing maturity of the technology, the current process of generating such group photos still faces the challenge of being time-consuming and requires further optimization to improve the user experience.
[0022] Therefore, after discovering the above-mentioned problems in their research, the inventors proposed the image processing method, apparatus, and electronic device described in this application, which can improve upon these problems. In this method, after obtaining multiple first target object images, image features of the first target object images can be extracted using a target neural network model. Based on these image features, a second target object image corresponding to the first target object image is output, thereby obtaining multiple second target object images. Then, the multiple second target object images are fused with a background image to obtain a target image. Thus, by using the above method, when the target neural network model is a pre-trained model for image visual processing, after obtaining the first target object image, only the image features of the first target object image need to be extracted. Then, the corresponding second target object image can be obtained by combining the target neural network model. Finally, the target image (group photo image) is obtained by fusing the multiple second target images with a background image. This avoids the situation where users need to upload images for model training before performing image visual processing based on the trained model when generating group photos, thereby improving the efficiency of image visual processing and group photo generation.
[0023] Before providing a more detailed description of the embodiments of this application, an application environment related to the embodiments of this application will be introduced.
[0024] The application scenarios involved in the embodiments of this application will be introduced below.
[0025] In the embodiments of this application, the provided image processing method can be executed by an electronic device. In this manner, all steps of the image processing method provided in the embodiments of this application can be performed by the electronic device. For example, as Figure 1 As shown, all steps in the image processing method provided in this application embodiment can be executed by the processor of the electronic device 100.
[0026] Alternatively, the image processing method provided in this application embodiment can also be executed by a server. Correspondingly, in this method where the method is executed by a server, the server can begin executing the steps of the image processing method provided in this application embodiment in response to a triggering instruction. This triggering instruction can be sent by an electronic device used by a user, or it can be triggered locally by the server in response to some automated event.
[0027] In addition, such as Figure 2As shown, the image processing method provided in this application embodiment can also be executed collaboratively by an electronic device and a server. In this collaborative execution method, some steps of the image processing method provided in this application embodiment are executed by the electronic device, while other steps are executed by the server. For example, the electronic device 100 can execute the image processing method including: acquiring multiple first target object images; then, the electronic device 100 can transmit the multiple first target object images to the server 200; after receiving the input image, the server 200 can execute subsequent steps to obtain multiple second target object images and a target image, and then transmit the target image to the electronic device 100; after receiving the target image, the electronic device 100 can store, display, and share the target image.
[0028] It should be noted that in this method where electronic devices and servers work together, the steps performed by the electronic devices and servers are not limited to those described in the examples above. In practical applications, the steps performed by the electronic devices and servers can be dynamically adjusted according to the actual situation.
[0029] It should be noted that the electronic equipment 100, in addition to being for Figure 1 and Figure 2 In addition to smartphones, the device 200 can also be a tablet computer, smartwatch, smart voice assistant, or other device equipped with a camera. Server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud computing, cloud storage, network services, cloud communication, middleware services, and artificial intelligence platforms. In the case where the image processing method provided in this embodiment is executed by a server cluster or distributed system composed of multiple physical servers, different steps in the image processing method can be executed by different physical servers, or can be executed in a distributed manner by servers built on a distributed system.
[0030] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0031] Please see Figure 3 This application provides an image processing method, which includes:
[0032] S110: Acquire multiple images of the first target object.
[0033] In this embodiment, the first target image can be an image of a target object. The specific category of the target object is not specifically limited in this embodiment. For example, the target object can be a human face, or it can be an animal, plant, etc.
[0034] Furthermore, in the subsequent image fusion process, the first target image can be understood as the foreground image during fusion. Also, the first target image is the original target image before any visual processing has been performed. For example, if the first target image is a face image, then the first target image can be understood as the original face image, that is, a face image that has not yet undergone visual processing.
[0035] In this embodiment, the number of multiple first target images is not specifically limited. For example, the multiple first target images can be two first target images, three first target images, or even more first target images. Multiple first target images can be obtained in various ways.
[0036] One approach is to recognize multiple input images separately to obtain multiple first target object images from them. It should be noted that in this approach, a user can input multiple images into the electronic device, in which case the electronic device can obtain multiple input images. For each of these multiple input images, the electronic device can recognize it separately to obtain a first target object image from each input image, thus obtaining multiple first target object images.
[0037] As another approach, the user can input only one image, which includes multiple first target objects. In this case, the electronic device receives only one input image, but because the single image contains multiple first target object images, multiple first target objects can still be obtained by recognizing the image.
[0038] As a recognition process, after obtaining the input image, the number of input images can be detected first.
[0039] If only one first target object is detected in the input image, the number of first target objects included in the input image can be detected. If multiple first target objects are detected in the input image, the input image can be directly identified to obtain the multiple first target objects included in the input image. If only one first target object is detected in the input image or no first target object is detected, the user can be prompted to continue inputting images until multiple first target objects are obtained from the user's input image.
[0040] If multiple input images are detected, the first target object can be identified in each of the multiple input images. If only one first target object is identified in the multiple input images, or if no first target object is included in any of the multiple input images, the user can be prompted to continue inputting images until multiple first target objects are obtained from the images input by the user.
[0041] In this embodiment, the source of the input image can be various. For example, the input image can be an image captured by a camera of an electronic device, or it can be transmitted from another electronic device. Another example is that the input image can be obtained from a network. In one approach, when there are multiple input images, the multiple input images can simultaneously include images from at least two sources. For example, the multiple input images can include images captured by a camera of an electronic device, or they can simultaneously include images transmitted from other devices. Another example is that the multiple input images can include images captured by a camera of an electronic device, or they can simultaneously include images obtained from a network.
[0042] Optionally, in this embodiment of the application, by recognizing the input image, the region where the first target object is located in the input image can be obtained, and then the first target object can be separated from the input image based on the obtained region to obtain the image of the first target object.
[0043] In the embodiments of this application, the first target image can be obtained from the input image through various image recognition methods.
[0044] One approach is to obtain the first target image from the input image using the SegmentAnything model. The SegmentAnything model is an image segmentation model. It can perform zero-shot generalization on unfamiliar input images without additional training, accurately segmenting various objects within the input image. This means it can be directly applied to new image domains, effectively segmenting underwater photographs, cell microscope images, etc. The SegmentAnything model's interface supports various types of cues, such as foreground / background points, rough bounding boxes, masks, and free text. Through these cues, the SegmentAnything model can understand the object or region the user wants to segment, thus achieving flexible segmentation tasks. For example, in this embodiment, when acquiring the first target image using the SegmentAnything model, the category information of the first target image can be simultaneously input into the SegmentAnything model. For instance, if the user expects to extract a face from the input image to obtain the first target image, the input category information could be a face.
[0045] One approach is to obtain the first target image from the input image using the InSPyReNet model. The InSPyReNet model combines multi-scale feature extraction and thinning operations to achieve high-precision image segmentation. The core of the InSPyReNet model is a saliency map with a rigorous image pyramid structure. This structure allows for image analysis and understanding at different scales, acquiring multi-scale feature information. This multi-scale feature set is crucial for handling complex scenes and targets of varying sizes, enabling better capture of details and global information within the image.
[0046] The aforementioned SegmentAnything model and InSPyReNet model are both based on CNN (Convolutional Neural Networks) and multi-scale feature extraction to process the first target object in the image.
[0047] In the embodiments of this application, when an input image can be recognized by multiple image recognition methods, there are multiple ways to determine the current method used to recognize the input image.
[0048] One approach is to have it pre-configured by the developers.
[0049] Alternatively, the electronic device can determine the method based on the actual situation. Optionally, it can determine the method based on the complexity of the input image. It should be noted that different image recognition methods may have different characteristics. For example, some image recognition methods may be more suitable for recognizing relatively complex images, while others are only suitable for recognizing relatively simple images. In this approach, after determining the corresponding image recognition method based on the complexity of the input image, the input image can be recognized using the determined method to obtain the region where the first target object is located. This method improves the flexibility of image recognition and allows for more targeted selection of recognition methods.
[0050] In one sense, image complexity can be understood as the degree of difficulty a user faces in comprehending or describing an image from the perspectives of global abstraction and local details. In another sense, image complexity can be viewed as the number of details and content variations within an image, which may include the complexity of color distribution, shape distribution, texture distribution, and structural distribution.
[0051] In the embodiments of this application, there are multiple ways to determine the complexity of an image.
[0052] One approach is to determine the complexity of an image by calculating its information entropy. This involves calculating the probability of each gray level occurring to obtain the image's gray-level information entropy value, which reflects the inherent complexity of the image's gray-level domain.
[0053] One approach is to transform the image to the frequency domain and use mathematical methods to extract its frequency distribution characteristics and the contrast of each frequency component, which can then be used as the basis for measuring the image's complexity and thus determining the image's complexity.
[0054] One approach is to determine the complexity of an image by analyzing its color distribution. Images with a greater variety of colors and uneven color distribution are generally considered more complex. In this case, the more colors an image contains, the higher its complexity will be; conversely, the more uneven the color distribution, the higher the complexity will be. Optionally, the color complexity can be assessed by calculating the image's color histogram and analyzing its distribution and entropy values.
[0055] One approach is to use image segmentation algorithms to identify different regions or objects in an image, and then assess the structural complexity by analyzing the number, size, shape, and interrelationships of these regions or objects, thereby using the structural complexity as the image complexity.
[0056] S120: The target neural network model performs image visual processing on multiple first target images to obtain multiple second target images. The multiple second target images include the second target images corresponding to each of the multiple first target images. During the image visual processing by the target neural network model, the target neural network model extracts the image features of the obtained first target images and outputs the corresponding second target images.
[0057] The number of second target object images obtained can be the same as the number of first target object images. For example, if there are two first target object images, the number of second target object images obtained after performing image visual processing on each of the first target object images will also be two. Similarly, if there are three first target object images, the number of second target object images obtained after performing image visual processing on each of the first target object images will also be three.
[0058] In this embodiment of the application, the target image obtained by subsequent fusion can be an image with a specific visual effect. If the target neural network model is capable of processing the input image into an image with the specific visual effect, then the target neural network model can first perform image visual processing on multiple first target images, so that the obtained second target image is an image with the specific visual effect.
[0059] In this embodiment of the application, no specific limitation is made to the specific visual effect. For example, the specific visual effect can be a photo style, a cartoon style, a sketch style, or other visual effects.
[0060] In this embodiment, a target neural network model can be obtained by training a neural network model to be trained using training data. The training data may include multiple training images and corresponding labels for each image. When the images included in the training data are images with a specific visual effect, training the neural network model to be trained using the training data enables the trained target neural network model to possess the ability to process that specific visual effect. This ability to process specific visual effects can be understood as the target neural network model being able to perform image visual processing on the input image, thereby enabling the processed image to possess that specific visual effect.
[0061] For example, if the images included in the training data are photorealistic, then the trained target neural network model can convert the input image into a photorealistic image. For example, if the images included in the training data are cartoon-style images, then the trained target neural network model can convert the input image into a cartoon-style image.
[0062] In this embodiment, the target neural network model can generalize the knowledge acquired during training to process the input image in practical applications. Specifically, the target neural network model extracts features from the input image, converting the extracted image features into a vector, and then compares it with the vectors of images already learned in the training data. Since there is a certain correlation between the images already learned in the training data and the input image, the target neural network model can utilize this correlation to perform image visual processing on the input image.
[0063] It should be noted that in the embodiments of this application, there can be multiple target objects. Correspondingly, the image features corresponding to the first target object image will also differ depending on the first target object. For example, if the first target object image is a face image, the extracted image features can be the facial features of the face image. For example, if the first target object image is an animal image, the extracted image features can be the animal features of that animal image.
[0064] Optionally, in this embodiment, the neural network model to be trained can be a deep learning model, and correspondingly, the trained target neural network model is also a deep learning model. The target neural network model in this embodiment can be obtained based on zero-shot learning technology.
[0065] Furthermore, as mentioned above, the image visual processing performance of the trained target neural network model is related to the training data used. When the training data includes images with multiple visual effects, the target neural network model trained on that data can process the input image into images with multiple visual effects. In this case, to ensure the target neural network model clearly understands the current conversion target (i.e., the visual effect of the output image), target visual information can be simultaneously input when the first target image is input into the target neural network model. This target visual information represents the conversion target that the user expects the target neural network model to achieve.
[0066] It should be noted that when target visual information is input into the target neural network model, the visual effect of the final output second target image corresponds to that target visual information, and the visual effect of the subsequently fused target image also corresponds to that target visual information. The correspondence between the target image and target visual information can be understood as the visual effect of the target image corresponding to the image visual processing target represented by the target visual information. Similarly, the correspondence between the second target image and target visual information can be understood as the visual effect of the second target image corresponding to the image visual processing target represented by the target visual information. For example, the target visual information can be "photographic style." In this case, the visual effect of the second target image can be photographic style, and correspondingly, the visual effect of the final target image will also be photographic style. For example, the target visual information can be "cartoon style." In this case, the visual effect of the second target image can be cartoon style, and correspondingly, the visual effect of the final target image will also be cartoon style.
[0067] One approach is to first detect the number of visual effects that the target neural network model can achieve after acquiring multiple first target object images. If the target neural network model can achieve multiple visual effects, a prompt message can be triggered to encourage the user to select one. In this case, the target visual information represents the selected visual effect. Therefore, the final second target object image visual effect output by the target neural network model will be the selected visual effect. If the target neural network model can achieve only one visual effect, the multiple input first target object images can be processed directly.
[0068] S130: Fuse multiple images of the second target object with the background image to obtain the target image.
[0069] After obtaining multiple images of the second target objects, these images can be fused with the background image to obtain the target image. Fusing multiple second target objects with the background image can be understood as adding the multiple second target objects to the background image, using the background image with the added target object images as the target image.
[0070] In the embodiments of this application, the background image can be obtained in a variety of ways.
[0071] One approach is to obtain the background image based on target visual information. It should be noted that target visual information refers to information that characterizes the visual effect of the subsequently obtained target image. In this case, the background image can be generated based on the target visual information. Optionally, the target visual information can be in text form, thus allowing the background image to be generated based on a text-to-image approach.
[0072] In the process of generating background images based on text-based image generation, a pre-trained text encoder (such as BERT) can first be used to convert the visual information of the target in text form into vector representations. These vector representations can capture the semantic and contextual information in the text. After obtaining the vector representations, they can be input into the text-based image model to output the background image. Optionally, the text-based image model can be a diffusion model, generative adversarial networks (GANs), or an auto-regressive model. Among them, the diffusion model adds random noise to the data by defining a Markov chain of diffusion steps, and then learns the inverse diffusion process to generate images.
[0073] As one approach, when multiple first target object images are identified from multiple input images, the background of one of the input images can be used as the background image. It should be noted that when the input image includes the first target object image, the image content other than the first target object image in the input image can be considered as the background of the input image.
[0074] Optionally, when there are multiple input images, one input image can be randomly selected to obtain the background image. Optionally, when the user inputs multiple input images sequentially, the background content of the first input image can be used as the background image. Optionally, after obtaining multiple input images, the user can be prompted to select one input image to determine the subsequent background image.
[0075] In the embodiments of this application, multiple second target images can be fused with a background image in various ways.
[0076] One method is to fuse multiple second target object images with a background image using alpha blending technology. Alpha blending is an important technique in computer graphics used to achieve transparency effects for images or objects. It fuses two or more images based on the transparency information (i.e., alpha value) of each pixel. Alpha values are typically between 0 and 1, where 0 represents complete transparency, 1 represents complete opacity, and intermediate values represent varying degrees of semi-transparency. The blending formula is: destinationcolor.rgb = (sourcecolor.rgb * sourcecolor.a) + (destinationcolor.rgb * (1 - sourcecolor.a)). Here, sourcecolor is the color and transparency value of the source images to be blended (i.e., the second target object image and the background image), and destinationcolor is the color value of the target image. This formula means that the final pixel color is a weighted sum of the source image color proportionally to its transparency and the target image color proportionally to its remaining transparency.
[0077] As another approach, ImageBlend technology can be used to fuse multiple images of a second target object with a background image. This technology is mainly used to blend two or more images using specific algorithms and parameters to generate a new image (i.e., the target image). In this method, the images to be fused (multiple images of the second target object and the background image) are first preprocessed, including grayscale conversion, brightness adjustment, contrast enhancement, and histogram equalization, to ensure consistency in brightness, contrast, etc., facilitating subsequent fusion processing. Then, fusion parameters can be set according to actual needs, such as transparency (alpha value) and blending mode (e.g., weighted average, Laplacian pyramid method, wavelet transform method, etc.). These parameters directly affect the final fused image. Based on the set fusion parameters, pixel-level calculations and fusion are performed on the images to be fused to obtain the target image.
[0078] It should be noted that when multiple second target images are blended with a background image, the positions of the multiple second target objects in the fused target image can be determined in various ways.
[0079] One approach is to determine the positions randomly. In this approach, the positions of multiple second target images appearing in the target image are not fixed and can be randomly determined by the device performing the fusion.
[0080] For example, the obtained multiple second target images include second target image P1, second target image P2, and second target image P3, such as... Figure 4 As shown, the background image includes regions Q1, Q2, and Q3. In this case, a second target image can be randomly selected from the second target images P1, P2, and P3 and placed in region Q1. Then, a second target image can be randomly selected from the remaining second target images and placed in region Q2. Finally, the last remaining second target image can be placed in region Q3.
[0081] As another approach, based on a preset arrangement strategy, multiple second target images are fused with a background image to obtain a target image, wherein the arrangement strategy is used to determine the arrangement position of the multiple second target images in the fused target image.
[0082] Optionally, the preset layout strategy can be: to arrange multiple second target images according to the order determined by the user.
[0083] In this method, the user can determine the arrangement position of multiple second target object images through the displayed arrangement position configuration interface. After acquiring multiple first target object images, the electronic device can display the arrangement position configuration interface, which shows thumbnails of the multiple first target object images and prompts the user to manipulate these thumbnails to edit the arrangement position of the multiple second target object images. The thumbnails corresponding to the first target object images can be obtained by cropping a portion of the first target object image, or by reducing the size of the first target object image.
[0084] For example, such as Figure 5 As shown, in Figure 5 The layout configuration interface 10 shown includes an image candidate region 11 and an image layout region 12. The image candidate region 11 displays thumbnails of multiple first target object images. For example, taking a face image as an example, in... Figure 5 The image candidate region 11 shown contains three face images. In this case, the user can drag the face images from the image candidate region 11 to the image arrangement region 12 for arrangement, so as to determine the arrangement position of the multiple second target images corresponding to the multiple first target images in the target image.
[0085] Users can determine the arrangement of multiple second target images using various layout options. These options can characterize the placement of the core target, such as placing the core target in the center, on the far left, or on the far right. For example, if the first target image is a face, and the core target is a user of an electronic device, and the user selects that the core target should always be in the center, the core target will be positioned in the center of the final target image. The core target can also be determined by the user. For instance, if the core target is a person, it can be either the user of the electronic device or other persons selected by the user.
[0086] Optionally, the preset layout strategy can be: determining the layout positions of multiple second target images based on the degree of adaptation between the second target image and the content in the background image.
[0087] It should be noted that the background image can contain image content, and the image content of some areas within the background image may differ. Furthermore, the image content of multiple second target images may also differ. In such cases, multiple second target images can be adapted to the image content of the background image, and the arrangement position of the second target images within the target image can be determined based on the degree of adaptation.
[0088] Optionally, in this embodiment, the second target image can be adapted to the image content in the background image using a feature matching method. In the feature matching method, key points (such as SIFT, SURF, or ORB feature points) and their feature descriptors can be extracted from the two images. The similarity between feature descriptors (such as Euclidean distance, Hamming distance, etc.) is used to match feature point pairs. The degree of adaptation between the two images is evaluated based on the number, distribution, and consistency of the matched feature points (e.g., using the RANSAC algorithm to remove mismatches). The two images include the second target image and the image content of the region in the background image currently being used to assess the adaptation.
[0089] Optionally, in this embodiment, the second target image can be adapted to the image content in the background image using template matching. In template matching, the second target image can be used as a template, which is slid across the background image, and the normalized cross-correlation coefficient (NCC) is calculated at each position reached. This quickly detects the similarity between the template and sub-regions of the background image by comparing the accumulated pixel differences to reduce computation. The position with the largest NCC (Normalized Cross-Correlation) value or the smallest SSDA (Sequential Similarity Detection Algorithm) value is selected as the optimal adaptation position. Here, the NCC value refers to the similarity index between two image regions calculated using the normalized cross-correlation algorithm. These are two techniques used to evaluate the similarity or adaptation degree between images.
[0090] Optionally, the preset arrangement strategy can be: determining the arrangement position of multiple second target images based on the degree of correlation between two second target images among multiple second target images.
[0091] The degree of correlation between two second target images can be determined in several ways.
[0092] As a method for determining the degree of association, when the second target image is a face image, the degree of association can be determined based on the social relationship and / or the degree of facial similarity between the two second target images. The social relationship between the two second target images can be understood as the social relationship between the people corresponding to the two second target images. The closer the social relationship between the two second target images, the closer their positions will be in the target image. Conversely, the higher the degree of facial similarity between the two second target images, the closer their positions will be in the target image.
[0093] The closeness of the social relationship between two second target images can be determined based on the communication frequency and / or number of times between the people corresponding to the two second target images.
[0094] For example, taking a face image as the first target image, such as... Figure 6 As shown, in this case, the resulting multiple second target object images are also face images. Finally, the target image obtained by fusing the multiple second target object images (multiple face images) with the background image can be as follows: Figure 6 As shown, when the target neural network model processes the first target image and the resulting second target image has a photorealistic visual effect, multiple second target images are photorealistic face images, and the final fused target image is also a photorealistic image.
[0095] This embodiment provides an image processing method that, through the above-described manner, when the target neural network model is a pre-trained model for image visual processing, after obtaining the first target image, only the image features of the first target image need to be extracted. Then, the corresponding second target image can be obtained by combining the target neural network model. By fusing multiple second target images with a background image, a target image (group photo image) can be obtained. This avoids the situation where users need to upload images for model training before performing image visual processing based on the trained model when generating group photos, thereby improving the efficiency of image visual processing and the efficiency of generating group photos.
[0096] Please see Figure 7 This application provides an image processing method, which includes:
[0097] S210: Acquire multiple images of the first target object.
[0098] In this application embodiment, there can be multiple methods for determining the target object.
[0099] One approach is to determine the target image based on the current image processing mode. Optionally, if the electronic device is currently in a portrait group photo processing mode, it will use a human face image as the target. Therefore, in this mode, all resulting first target image images will also be human face images. Alternatively, if the electronic device is currently in an animal group photo processing mode, it will use an animal as the target. Therefore, in this mode, all resulting first target image images will also be animal images.
[0100] In one approach, the electronic device can identify all foreground objects in the input image and provide them to the user for selection, using the selected object as the target object. For example, the multiple foreground objects identified by the electronic device from the input image include: foreground object W1, foreground object W2, foreground object W3, and foreground object W4. In this case, the electronic device can display foreground objects W1, W2, W3, and W4. If the user selects one of these foreground objects, W2, W3, or W4, then W2, W3, and W4 can be used as multiple first target objects, and correspondingly, the images of each of the foreground objects W2, W3, and W4 can be used as multiple first target object images.
[0101] S220: The target neural network model performs image visual processing on multiple first target images to obtain multiple second target images. The multiple second target images include the second target images corresponding to each of the multiple first target images. During the image visual processing by the target neural network model, the target neural network model extracts the image features of the obtained first target images and outputs the corresponding second target images.
[0102] S230: Acquire the differences between multiple second target object images and background images in preset image parameters.
[0103] S240: Adjust the preset image parameters of multiple second target object images and / or background images based on the differences to obtain the adjusted background image and the adjusted multiple second target object images.
[0104] S250: The adjusted images of multiple second target objects are fused with the adjusted background image to obtain the target image.
[0105] For example, in realistic group photo scenarios, during the fusion process, there are differences in color and brightness between the foreground image (multiple images of second target objects) and the background image. In this case, a lighting model is used to light the foreground image and / or the background image, thereby adjusting the image lighting and color and improving the naturalness of the fusion.
[0106] This embodiment provides an image processing method that improves the efficiency of image visual processing and the efficiency of generating group photos. Furthermore, in this embodiment, before fusing the second target image with the background image, differences in preset image parameters between the second target image and the background image are detected. Based on these differences, the preset image parameters of the second target image and / or the background image are adjusted before image fusion, resulting in a more natural fused target image with better visual effects.
[0107] It should be noted that, in this embodiment of the application, the more first target object images that need to be processed, the greater the resource consumption of the processing device. Therefore, in this case, the processing device can determine the number or type of first target object images based on current practical considerations.
[0108] In one approach, when the device is currently under high load, the acquired images of multiple first target objects can be images of some target objects identified from multiple input images (or multiple images). For example, if the identified target objects include foreground object W1, foreground object W2, foreground object W3, and foreground object W4, then when the device is currently under high load, the acquired images of multiple first target objects can include images of foreground object W1 and foreground object W2.
[0109] In one approach, when the device is currently under low load, the acquired images of multiple first target objects can be images of all target objects identified from the input image (or multiple images). For example, if the identified target objects include foreground object W1, foreground object W2, foreground object W3, and foreground object W4, then when the device is currently under high load, the acquired images of multiple first target objects can include images of foreground object W1, foreground object W2, foreground object W3, and foreground object W4.
[0110] Optionally, the current state can be determined directly based on the remaining processing resources. Specifically, if the remaining processing resources are greater than the resource threshold, the current state is determined to be low-load; if the remaining processing resources are less than or equal to the resource threshold, the current state is determined to be high-load.
[0111] Optionally, the current state can be determined based on the current time period. Specifically, if a high-load time period is detected, the current state is determined to be high-load. If a low-load time period is detected, the current state is determined to be low-load.
[0112] In this embodiment of the application, taking the image processing method performed by an electronic device as an example, the low-load time period and the high-load time period can be divided according to the historical operation of the electronic device.
[0113] One approach is to divide low-load and high-load periods based on the number of programs running on the electronic device.
[0114] Optionally, multiple time periods can be pre-defined, and the number of programs running in each time period can be counted. Based on the number of programs running in each time period, each time period can be classified as a low-load or high-load period. For example, a time period where the number of running programs is less than a first threshold can be defined as a low-load period, and a time period where the number of running programs is not less than the first threshold can be defined as a high-load period. As another example, a time period where processor utilization is less than a first utilization threshold can be defined as a low-load period, and a time period where processor utilization is not less than the first utilization threshold can be defined as a high-load period.
[0115] The length of each pre-defined time segment can vary. For example, each time segment can be 1 hour, 2 hours, or 3 hours long.
[0116] Alternatively, time periods can be divided by statistically analyzing the running status of programs in the electronic device, and the resulting time periods can be simultaneously determined as either low-load or high-load periods. In this method, the moment when the number of running programs changes compared to a first threshold can be used as the boundary between low-load and high-load time periods. Specifically, after the electronic device starts, it can begin detecting the number of currently running programs, and the start time can be used as the start time of the low-load period. If the detected number of running programs exceeds the first threshold, the moment when the detected number of running programs exceeds the first threshold can be used as the start time of the high-load period, which is also the end time of the preceding low-load period. If, after detecting that the number of running programs exceeds the first threshold, the detected number of running programs does not exceed the first threshold, the moment when the detected number of running programs does not exceed the first threshold is used as the end time of the current high-load period, which is also the start time of a new low-load period. This method allows the electronic device to avoid pre-dividing multiple time periods, thus providing greater flexibility.
[0117] Optionally, time periods can be divided by statistically analyzing processor usage in electronic devices, and the resulting time periods can be simultaneously determined as either low-load or high-load periods. In this approach, the moment when the current processor usage changes compared to a first usage threshold can be used as the boundary between low-load and high-load periods. Specifically, after the electronic device starts up, it can begin detecting the current processor usage, and the startup time can be used as the start time of the low-load period. If the detected processor usage exceeds the first usage threshold, the moment when the detected usage exceeds the first threshold can be used as the start time of the high-load period, which also marks the end time of the preceding low-load period. If, after detecting usage exceeding the first threshold, the detected usage does not exceed the first threshold, the moment when the detected usage does not exceed the first threshold is used as the end time of the current high-load period, which also marks the start time of a new low-load period.
[0118] Given the availability of multiple ways to divide low-load and high-load time periods, electronic equipment can flexibly determine which division method to use at any given time.
[0119] One approach is for the user of the electronic device to determine how to divide the low-load and high-load time periods. Another approach is for the electronic device itself to determine the specific method for dividing the low-load and high-load time periods based on the actual situation.
[0120] Optionally, the method for dividing low-load and high-load time periods can be determined based on the total operating time of the electronic device after its first startup. Furthermore, different division methods will result in different low-load and high-load time periods.
[0121] It's important to note that in the initial stages of using electronic devices, users' program usage habits are relatively unpredictable. Therefore, directly dividing time periods based on the execution of programs on the device and simultaneously determining whether these periods are low-load or high-load periods may result in inaccurate classifications. Conversely, after users have used the electronic device for a period of time, the program execution patterns become relatively fixed. In this case, dividing low-load or high-load periods in real-time based on execution data will be more accurate.
[0122] Therefore, if the total running time of the electronic device after its first startup is less than the first specified time, multiple time periods can be pre-divided, and the number of programs running in each time period can be counted. Based on the number of programs running in each time period, each time period can be classified as a low-load or high-load period. Conversely, if the total running time of the electronic device after its first startup is not less than the first specified time, the time periods can be divided by counting the running programs in the electronic device, and simultaneously determining whether the divided time periods are low-load or high-load periods.
[0123] One approach is to determine the number and types of first target object images based on user preference data. For example, if the user preference data indicates a preference for images of people, then all the multiple first target object images could be images of people (one type of target object). Similarly, if the user preference data indicates a preference for images of animals, then all the multiple first target object images could be images of animals (one type of target object).
[0124] It should be noted that preference data can be understood as data representing a user's habits or preferences. Alternatively, it can be understood as data representing a user profile. The target user can be understood as the user referenced when acquiring preference data, thus ensuring that the acquired preference data corresponds to the target user. The correspondence between preference data and the target user can be understood as the acquired preference data representing the target user's habits or preferences. In this embodiment, the target user can be understood as a user of an electronic device, or a user currently using an electronic device. The user of an electronic device can be understood as the user corresponding to the identification information recorded in the electronic device. This identification information may include a user account or a current phone number.
[0125] One approach is to store user preference data on the electronic device itself. This locally stored preference data can cover multiple aspects, such as dietary, travel, sports, and health preferences. Alternatively, the electronic device can store the user's identification information locally, allowing the user identified by this information to be targeted. Optionally, this identification information can be an account number or phone number. Alternatively, all user preference data can be centrally stored in the cloud. In this approach, when a specified event is triggered, preference data can be retrieved from the cloud based on the stored identification information and the triggered event.
[0126] As another approach, the current user of an electronic device may not be the owner of the device. For example, in some cases, the owner of an electronic device may lend it to another user. In this situation, directly obtaining preference data based on the owner's preferences might result in the obtained preference data not meeting the current user's needs. For instance, electronic device C1 belongs to user P1, and user P1 lends it to user P2. If user P2 triggers an image processing method while using electronic device C1, but electronic device C1 still uses user P1's preference data as the basis for obtaining the first target image, it will not only lead to a privacy breach for user P1 but also prevent user P2 from obtaining the required feedback. In this approach, when triggering an image processing method, the electronic device can first perform user identification to identify the identified user as the target user.
[0127] Please see Figure 8 This application provides an image processing apparatus 300, which includes:
[0128] The target image acquisition unit 310 is used to acquire multiple images of the first target object.
[0129] In one manner, the target image acquisition unit 310 is specifically used to identify multiple input images respectively in order to acquire multiple first target object images from the multiple input images; or, to identify input images in order to acquire multiple first target object images from the input images.
[0130] Optionally, the target image acquisition unit 310 is specifically used to identify the input image to obtain the region where the first target object is located in the input image; and to separate the first target object from the input image based on the region to obtain the first target object image.
[0131] Optionally, the target image acquisition unit 310 is specifically used to determine the complexity of the input image, so as to determine the corresponding image recognition method based on the complexity; and to recognize the input image based on the determined image recognition method to obtain the region where the first target object is located in the input image.
[0132] The target image processing unit 320 is used to perform image visual processing on multiple first target images through a target neural network model to obtain multiple second target images. The multiple second target images include the second target images corresponding to each of the multiple first target images. During the image visual processing of the target neural network model, the target neural network model extracts the image features of the obtained first target images and outputs the corresponding second target images.
[0133] The image fusion unit 330 is used to fuse multiple second target object images with the background image to obtain a target image.
[0134] In one approach, the image fusion unit 330 is specifically used to fuse multiple second target object images with a background image based on a preset arrangement strategy to obtain a target image, wherein the arrangement strategy is used to determine the arrangement positions of the multiple second target object images in the fused target image.
[0135] In one approach, the image fusion unit 330 is specifically used to acquire the differences between multiple second target images and background images in preset image parameters; adjust the preset image parameters of the multiple second target images and / or background images based on the differences to obtain an adjusted background image and multiple adjusted second target images; and fuse the adjusted multiple second target images with the adjusted background image to obtain a target image.
[0136] This embodiment provides an image processing apparatus that improves the efficiency of image visual processing and the efficiency of generating group photos.
[0137] It should be noted that the device embodiments in this application correspond to the aforementioned method embodiments. The specific principles in the device embodiments can be found in the content of the aforementioned method embodiments, and will not be repeated here.
[0138] The following will combine Figure 9 This application describes an electronic device.
[0139] Please see Figure 9 Based on the aforementioned image processing methods and apparatus, this application also provides another electronic device 100 capable of executing the aforementioned image processing methods. The electronic device 100 includes one or more (only one shown in the figure) processors 102, a memory 104, a network module 106, and a camera 108 coupled together. The memory 104 stores programs capable of executing the contents of the aforementioned embodiments, and the processor 102 can execute the programs stored in the memory 104.
[0140] The processor 102 may include one or more processing cores. The processor 102 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by calling data stored in the memory 104. Optionally, the processor 102 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 102 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 102 and may be implemented separately using a communication chip.
[0141] The memory 104 may include random access memory (RAM) or read-only memory (ROM). The memory 104 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the terminal 100 during use (such as phonebook data, audio and video data, chat log data, etc.).
[0142] Network module 106 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, thereby communicating with communication networks or other devices, such as audio playback devices. Network module 106 may include various existing circuit elements used to perform these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, user identity module (SIM) cards, memory, etc. Network module 106 can communicate with various networks such as the Internet, corporate intranets, and wireless networks, or communicate with other devices via wireless networks. The aforementioned wireless networks may include cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs). For example, network module 106 can exchange information with base stations.
[0143] Please refer to Figure 10 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 800 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0144] The computer-readable storage medium 800 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 800 has storage space for program code 810 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 810 may be compressed, for example, in a suitable form.
[0145] In summary, the image processing method, apparatus, and electronic device provided in this application, after obtaining multiple first target object images, can extract image features from the first target object images using a target neural network model, and output second target object images corresponding to the first target object images based on the image features, thereby obtaining multiple second target object images. Then, the multiple second target object images are fused with a background image to obtain a target image. Thus, by the above method, when the target neural network model is a pre-trained model for image visual processing, after obtaining the first target object images, only the image features of the first target object images need to be extracted, and then the corresponding second target object images can be obtained by combining them with the target neural network model. The target image (group photo image) is then obtained by fusing the multiple second target object images with the background image, avoiding the need for users to first upload images for model training before performing image visual processing based on the trained model when generating group photos. This improves the efficiency of image visual processing and the efficiency of generating group photos.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An image processing method, characterized in that, The method includes: Acquire multiple images of the first target object; The multiple first target images are processed by a target neural network model to obtain multiple second target images. The multiple second target images include the second target images corresponding to each of the multiple first target images. During the image visual processing by the target neural network model, the target neural network model extracts the image features of the obtained first target images and outputs the corresponding second target images. The multiple second target images are fused with the background image to obtain the target image.
2. The method according to claim 1, characterized in that, The step of fusing the plurality of second target object images with the background image to obtain a target image includes: Based on a preset arrangement strategy, the multiple second target images are fused with the background image to obtain a target image, wherein the arrangement strategy is used to determine the arrangement position of the multiple second target images in the fused target image.
3. The method according to claim 1, characterized in that, The target neural network model outputs a corresponding second target image by inputting target visual information and extracting image features from the first target image. The background image is obtained based on the target visual information.
4. The method according to claim 1, characterized in that, The acquisition of multiple first target object images includes: Multiple input images are identified separately to obtain multiple first target object images from the multiple input images; Alternatively, the input image can be identified to obtain multiple first target object images from the input image.
5. The method according to claim 4, characterized in that, The method further includes: The input image is identified to obtain the region where the first target object is located in the input image; The first target object is separated from the input image based on the region to obtain an image of the first target object.
6. The method according to claim 5, characterized in that, Before performing recognition on the input image to obtain the region where the first target object is located in the input image, the process further includes: Determine the complexity of the input image, and determine the corresponding image recognition method based on the complexity; The step of recognizing the input image to obtain the region where the target object is located in the input image includes: The input image is identified based on the determined image recognition method to obtain the region where the first target object is located in the input image.
7. The method according to claim 1, characterized in that, The step of fusing the plurality of second target object images with the background image to obtain the target image includes: Acquire the differences between multiple images of the second target object and the background image in preset image parameters; Based on the differences, the preset image parameters of the plurality of second target object images and / or the background image are adjusted to obtain the adjusted background image and the plurality of second target object images; The adjusted images of the multiple second target objects are fused with the adjusted background image to obtain the target image.
8. The method according to any one of claims 1-7, characterized in that, The first target image is a face image.
9. An image processing apparatus, characterized in that, The device includes: The target image acquisition unit is used to acquire multiple images of the first target object; The target image processing unit is used to perform image visual processing on the plurality of first target images respectively through a target neural network model to obtain a plurality of second target images. The plurality of second target images include the second target images corresponding to each of the plurality of first target images. In the process of image visual processing by the target neural network model, the target neural network model extracts the image features of the obtained first target images and outputs the corresponding second target images. The image fusion unit is used to fuse the plurality of second target object images with the background image to obtain a target image.
10. An electronic device, characterized in that, It includes a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, wherein the program code, when executed by a processor, performs the method according to any one of claims 1-8.