Image processing method and device, electronic equipment and storage medium
By using a target facial annotation model generated from training sample fusion images, the facial features are accurately annotated, solving the problem of inaccurate facial feature recognition in existing technologies and improving the effect of special effects processing and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2022-03-29
- Publication Date
- 2026-07-21
AI Technical Summary
The existing facial feature recognition is inaccurate, resulting in poor special effects and a poor user experience.
At least two facial features of the target object in the image to be processed are annotated using a target facial annotation model. The target facial annotation model is trained based on a first training sample, which includes the original sample image and the corresponding sample fusion image. The sample fusion image is obtained by fusing at least two sample local annotation images.
It improves the accuracy and effectiveness of facial feature annotation, enhances the visualization effect of special effects processing, and improves the user experience.
Smart Images

Figure CN116935455B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to an image processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of internet technology, more and more applications have entered users' lives, especially a series of short video shooting apps, which are very popular. When shooting short videos, users have certain demands for the richness and entertainment value of the video content.
[0003] In short video shooting, special effects are often applied to the user's facial features. However, the current facial feature recognition is inaccurate, which makes it impossible to effectively apply special effects, resulting in a poor user experience. Summary of the Invention
[0004] This disclosure provides an image processing method, apparatus, electronic device, and storage medium to accurately and effectively label at least two facial features in a facial image, thereby facilitating subsequent special effects processing.
[0005] In a first aspect, embodiments of this disclosure provide an image processing method, the method comprising:
[0006] In response to the detection of a trigger condition, acquire the image to be processed, including the target object;
[0007] Based on the target facial annotation model, at least two facial features of the target object in the image to be processed are annotated to obtain a target image corresponding to the image to be processed.
[0008] The target face annotation model is trained based on a first training sample, which includes original sample images and corresponding sample fusion images. The sample fusion images are obtained by fusing at least two sample local annotation images.
[0009] Secondly, embodiments of this disclosure also provide an image processing apparatus, the apparatus comprising:
[0010] The image acquisition module is used to acquire an image to be processed, including the target object, in response to the detection of a trigger condition;
[0011] The image processing module is used to annotate at least two facial features of the target object in the image to be processed based on the target facial annotation model, so as to obtain a target image corresponding to the image to be processed.
[0012] The target face annotation model is trained based on a first training sample, which includes original sample images and corresponding sample fusion images. The sample fusion images are obtained by fusing at least two sample local annotation images.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in any of the embodiments of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image processing method as described in any of the embodiments of this disclosure.
[0018] The technical solution of this disclosure involves acquiring an image to be processed, including a target object; annotating at least two facial features of the target object in the image to be processed based on a target facial annotation model to obtain a target image corresponding to the image to be processed; wherein, the target facial annotation model is trained based on a first training sample, the first training sample including an original sample image and a corresponding sample fusion image, the original sample image including the object to be annotated, the sample fusion image being obtained by fusing at least two sample local annotation images, and the at least two sample local annotation images being generated by annotating the corresponding parts of the at least two facial features of the object to be annotated by the corresponding target local annotation model. The technical solution of this disclosure solves the technical problem in the prior art where inaccurate facial feature annotation leads to poor special effects processing, resulting in a poor user experience. It achieves the determination of the annotation image corresponding to the same original image based on each target local annotation model, and after fusing the annotation images to obtain a sample fusion image, the target facial annotation model is trained based on the original image and the corresponding sample fusion image, improving the accuracy and effectiveness of facial feature annotation. Furthermore, when performing special effects processing based on the annotated facial features, the technical effect of the special effects processing can be further improved. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure;
[0021] Figure 2 This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure;
[0022] Figure 3 This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure;
[0023] Figure 4 This is a structural block diagram of an image processing apparatus provided in an embodiment of the present disclosure;
[0024] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0030] Before introducing this technical solution, we can first provide an example of the application scenario.
[0031] Existing facial image annotation methods often focus on annotating only one facial feature, resulting in scattered data; for example, only the mouth and eyes might be annotated. To obtain a model capable of annotating multiple facial features, each feature needs to be annotated, significantly increasing costs. Furthermore, even with these methods, the richness of the annotated samples is limited. For instance, the obtained samples might not cover facial features under exaggerated facial expressions, leading to poor general applicability. Therefore, existing annotation methods suffer from high costs and poor general applicability.
[0032] It should also be noted that the technical solution disclosed herein can be applied to any scene requiring facial annotation, such as when facial features are annotated in an image corresponding to a photographed user, as in short video shooting scenarios. It can also be applied to still image shooting scenarios, for example, after capturing an image using a camera built into a terminal device, at least two facial features can be annotated in the image; it can also be applied in instant messaging software to annotate at least two facial features of a user making a video call.
[0033] Figure 1 This is a flowchart illustrating an image processing method provided in an embodiment of this disclosure. This embodiment is applicable to situations in any special effects video where facial features need to be annotated or special effects processing is applied to facial features based on the annotation results. The method can be executed by an image processing device, which can be implemented in software and / or hardware. The hardware can be an electronic device, such as a mobile terminal, PC, or server. This technical solution can be implemented by a client or server, or by a combination of both.
[0034] like Figure 1 The method in this embodiment includes:
[0035] S110. Obtain the image to be processed, which includes the target object.
[0036] It should be noted that the apparatus for executing the image processing method provided in the embodiments of this disclosure can be integrated into application software that supports image processing functions, and this software can be installed on an electronic device, optionally a mobile terminal or a PC. The application software can be a type of software for image / video processing; specific application software will not be described in detail here, as long as it can achieve image / video processing. It can also be a specially developed application program, or integrated into a corresponding page, allowing users to perform image processing through the integrated page on the PC.
[0037] The target object can be a user or any pet with facial features. The image to be processed includes the target object. The user can shoot a short video using a short video shooting app. The short video consists of multiple video frames, and the video frame including the target object is used as the image to be processed. The image to be processed can also be a static image, for example, an image taken by the user using a camera app on a mobile device, and if facial feature annotation processing is required, then that image can be used as the image to be processed.
[0038] Specifically, images with facial features in a real-world scene can be captured using a mobile device's camera, and the resulting image can be used as the image to be processed. Alternatively, in special effects video processing scenarios, video frames that include the target object can be used as the image to be processed during real-time video capture. Or, during post-processing of captured video, video frames that include the target object can be used as the image to be processed.
[0039] It should be noted that the image to be processed can contain one or more target objects. If there is only one, at least two facial features need to be labeled. If the image contains multiple target objects, facial features can be labeled for all target objects, as well as for user-defined target objects. This means that even if the image contains multiple target objects, only the user-defined target objects will have their facial features labeled. For example, if the user-defined target object is object A, and the image to be processed includes both object A and object B, only the facial features of object A need to be labeled.
[0040] In this embodiment, obtaining the image to be processed, including the target object, can be done as follows: when the triggering of the image annotation control is detected, the page is redirected to the image upload page, and the image corresponding to the image upload completion instruction is used as the image to be processed; or when the triggering of the image annotation control is detected, the image displayed on the display interface is used as the image to be processed; or when the triggering of the image annotation control is detected, the page is redirected to the image library, and the image triggered in the image library is used as the image to be processed; or when the target object is detected in the on-screen image, the image to be processed, including the target object, is captured; or when the wake word is detected, the image to be processed, including the target object, is captured.
[0041] The image annotation control can be a control on the client side. By triggering the control, corresponding image annotation instructions can be generated so that the client can annotate the image to be processed.
[0042] As can be seen from the above, there are multiple ways to obtain the image to be processed:
[0043] The first method is as follows: When an image needs to be retrieved, the user can touch a virtual button on the client's display interface. Touching this virtual button triggers an image annotation control. When the client detects that the control has been triggered, it will redirect the display interface to the image upload page, which contains a corresponding upload control. For example, this control could be an upward arrow representing the upload of an image. After the user clicks the control, the display interface can redirect to any image library, which could be the mobile device's internal image library or an online image library. The user can then select an image from the image library and click the confirmation button on the display interface. The image will then be displayed on the upload page, i.e., the image to be processed.
[0044] The second method is: when the current display interface shows an image, the user can touch a virtual button on the client display interface. Touching the virtual button can trigger the image annotation control. When the client detects that the control has been triggered, the client will use the image displayed on the current page as the image to be processed.
[0045] The third method is as follows: Users pre-capture and save many images containing the target object on their mobile devices. When users need to annotate the facial features in the images, they can trigger the image annotation control by touching a virtual button on the client's display interface. After the image annotation control is triggered, the client detects the trigger and generates a control command to switch the client's display interface to the image library. The image library displays the pre-captured facial images in the form of thumbnails, which users can preview. Users can then select the facial image to be annotated by clicking on one or more thumbnails. The client uses the image selected by the user as the image to be processed.
[0046] The fourth method is: when the client detects that the image annotation control is triggered, it can generate a control command to turn on the mobile terminal's camera and display the image captured by the camera in real time on the client's display interface. When the camera captures a clear facial image, it uses that facial image as the image to be processed.
[0047] The fifth method is: the client can be set with a wake word, such as "start taking pictures". When the user starts the client, he can say "start taking pictures" to the mobile terminal. The mobile terminal can record the audio and the client can analyze it. If it matches the wake word, the camera will start to take pictures and the captured pictures will be used as images to be processed.
[0048] It should be noted that this technical solution can be applied to special effects video processing scenarios. When a user triggers a facial annotation effect prop, facial features can be annotated on the target object in the frame. Alternatively, in special effects video processing scenarios, such as virtual makeup scenarios, facial features can be determined based on this technical solution, and makeup rendering can be applied to those features.
[0049] S120. Based on the target face annotation model, at least two facial features of the target object in the image to be processed are annotated to obtain the target image corresponding to the image to be processed.
[0050] The target facial annotation model is trained based on at least one first training sample. The first training sample includes an original sample image and a corresponding fused sample image. The original image includes the object to be annotated. The fused sample image is obtained by fusing at least two local annotated sample images. The annotated facial features of the object to be annotated differ in each local annotated sample image. The set of annotated features in each local annotated sample image corresponds to at least two facial feature sets. The local annotated sample images are generated by the corresponding target local annotation model. The target image can be the image output by the target facial annotation model after annotating the facial features in the image to be processed.
[0051] Specifically, the image to be processed can be input into the target face annotation model. The target face model can annotate at least two facial features in the face region of the image to be processed, and obtain the annotated image as the target image.
[0052] For example, the image to be processed is a user's selfie, and the target facial annotation model has at least two facial features: the left eyebrow and the right eyebrow. Accordingly, the target facial annotation model is a model that annotates the user's left and right eyebrows. When the client detects special effects props associated with facial annotation or audio information related to facial annotation, it can trigger an image annotation control within the client and, in response to this control, acquire the user's selfie image as the image to be processed. After inputting the image to be processed into the model, the target facial annotation model can output an image annotated with the user's left and right eyebrows, i.e., the target image. Specific annotation methods could include adding a special color above the left and right eyebrow areas in the image, or adding special effects. The image to be processed can also be an image containing a kitten, which can be input into a target facial annotation model that can annotate the eyes and mouth. After inputting the image into the target facial annotation model, the model can annotate the kitten's eyes and mouth, for example, by adding flower effects to the eye and mouth positions. Furthermore, the model outputs the annotated image as the target image.
[0053] In this embodiment, at least two facial features include at least two of the following: left eyebrow, right eyebrow, left eye, right eye, nose, upper lip, lower lip, left ear, and right ear. Each facial feature corresponds to a target local annotation model. To clearly annotate each facial feature in the image to be processed, a target local annotation model corresponding to each feature can be trained.
[0054] The technical solution of this disclosure involves acquiring an image to be processed, including a target object; annotating at least two facial features of the target object in the image to be processed based on a target facial annotation model to obtain a target image corresponding to the image to be processed; wherein, the target facial annotation model is trained based on a first training sample, the first training sample including an original sample image and a corresponding sample fusion image, the original sample image including the object to be annotated, the sample fusion image being obtained by fusing at least two sample local annotation images, and the at least two sample local annotation images being generated by annotating the corresponding parts of the at least two facial features of the object to be annotated by the corresponding target local annotation model. The technical solution of this disclosure solves the technical problem in the prior art where inaccurate facial feature annotation leads to poor special effects processing, resulting in a poor user experience. It achieves the determination of the annotation image corresponding to the same original image based on each target local annotation model, and after fusing the annotation images to obtain a sample fusion image, the target facial annotation model is trained based on the original image and the corresponding sample fusion image, improving the accuracy and effectiveness of facial feature annotation. Furthermore, when performing special effects processing based on the annotated facial features, the technical effect of the special effects processing can be further improved.
[0055] Figure 2 This is a schematic flowchart of an image processing method provided in an embodiment of this disclosure. Based on the foregoing embodiments, before labeling the image to be processed, a first training sample can be determined, and the face labeling model to be trained can be trained to obtain a target face labeling model. Specific implementation details can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0056] like Figure 2 As shown, the method specifically includes the following steps:
[0057] S210. Determine at least one first training sample for training to obtain the target facial annotation model.
[0058] The first training sample includes original sample images and corresponding fused sample images. The number of first training samples can be one or more, as long as it is sufficient to train the target facial annotation model. The original sample images are captured images or can be fake facial images generated based on random noise. The fused sample images are determined by fusing multiple locally annotated images. For example, there are three locally annotated images: one annotating the left eyebrow, one annotating the right eyebrow, and one annotating the upper lip. The schematic diagram obtained by fusing these three locally annotated images is the fused sample image.
[0059] It should also be noted that the fused parts in the sample fusion image correspond to the actual situation, and users can set them according to their actual needs; there are no further restrictions here.
[0060] In this technical solution, determining at least one first training sample for training a target facial annotation model includes: acquiring at least one original sample image; for each original sample image, inputting the current original sample image into each target local annotation model to obtain at least one sample local annotation image corresponding to the original sample image; for each original sample image, fusing the at least one sample local annotation image corresponding to the current original sample image into one image to obtain a sample fusion image corresponding to the current original sample image; and determining at least one first training sample for training a target facial annotation based on each original sample image and the corresponding sample fusion image.
[0061] The original sample images can be images of different objects pre-collected by the developers. These objects could be users, pets with facial features, etc., and users can customize them according to their needs. The target local annotation model is pre-trained and used to annotate specific facial features in the face image. There are at least two target local annotation models. The set of annotated parts for each target local annotation model corresponds to at least two facial features of the target annotated area. The original image containing the object to be annotated can be input into a local annotation model, and then sample local annotation images annotated for a specific facial feature can be obtained. There are at least two sample local annotation images corresponding to each original image, and their specific number corresponds to the number of at least two facial features. That is, there are three samples with at least two facial features, and three sample local annotation images corresponding to the original image, with their parts corresponding to at least two facial features.
[0062] In practical applications, to improve the accuracy of target local annotation models and target facial annotation models, the richness of training samples can be increased. Richness can be reflected in: pre-collecting facial images in different poses as raw sample images. For example, different poses of facial features could include facial images of raised eyebrows, smiling faces, and open mouths. Images of the user's head in different poses could also be included, such as images corresponding to faces tilted upwards at 45 degrees or downwards at 30 degrees.
[0063] It should be noted that in this embodiment, there can be multiple target local annotation models, and different target local annotation models can annotate different regions of the image. For example, the left eye local annotation model can annotate the left eye region of a facial image and output an image corresponding to that image that only annotates the left eye region. It should also be noted that the output local annotated image can only include the outline of the annotated part, and the pixel values of the parts other than the outline are set to 255 or 0.
[0064] Specifically, one or more images can be selected from a large number of pre-collected images as original sample images, and these original sample images can be input into various target local annotation models. For example, the original sample images can be input into the left eye local annotation model, mouth local annotation model, ear local annotation model, and nose local annotation model. Each target local annotation model can output a sample local annotation image corresponding to the original sample image.
[0065] Optionally, at least one sample local annotation image corresponding to the current original sample image is fused into one image to obtain a sample fusion image corresponding to the current original sample image, including: extracting the annotation parts in each sample local annotation image and stitching the annotation parts onto one image to obtain a sample fusion image corresponding to the current original sample image.
[0066] The labeled area can be a part that has already been labeled in the local labeled image of the sample. For example, if the labeled area in local labeled image A is the left eye, then the left eye is the labeled area. If the labeled area in local labeled image B is the nose, then the nose is the labeled area.
[0067] Specifically, by extracting and stitching the labeled parts from the local labeled image of the sample, a sample fusion image of the current original sample image can be obtained.
[0068] For example, the original sample image is A, and its corresponding locally annotated images are: image 1 (annotated for the nose region), image 2 (annotated for the left eye), and image 3 (annotated for the upper lip). The nose region in image 1, the left eye region in image 2, and the upper lip region in image 3 can be extracted. These extracted images are then fused using image processing techniques to obtain a fused image, i.e., the sample fused image. The eyes, nose, and mouth in this sample fused image retain their original annotations.
[0069] S220. For each first training sample, the original sample image in the current training sample is used as the input parameter of the face annotation model to be trained, and the actual output image corresponding to the original sample image is obtained.
[0070] The face annotation model to be trained is a model whose parameters are either initial or default. The actual output image is the image output after the original sample images from the current training samples are input into the face annotation model to be trained.
[0071] Specifically, the original training sample images in the current training samples can be used as input parameters for the face annotation model to be trained. The face annotation model to be trained can output the corresponding image, and the image output by the model can be used as the actual output image.
[0072] In this embodiment, the facial annotation model to be trained can be a ResNet network model. It should be noted that it is sufficient to annotate the original sample images; the specific model type is not specifically limited.
[0073] S230. Based on the sample fusion image of the actual output image and the current training sample, determine the loss value, and correct the model parameters of the face annotation model to be trained based on the loss value.
[0074] It should be noted that the model parameters in the facial annotation model being trained do not meet the expected requirements. Therefore, the actual output image based on the current model parameters differs from the theoretical output image, and the model cannot accurately annotate the facial features in the image. Therefore, the sample fusion image corresponding to the current original sample image can be used as the theoretical output image. The error loss value between the actual output image and the sample fusion image can be determined, and the model parameters can be corrected based on this loss value.
[0075] Specifically, after inputting the original sample image into the face annotation model to be trained, the face annotation model can obtain the actual output image corresponding to the original sample image. Based on the actual output image and the sample fusion image, the corresponding loss value can be determined, and the model parameters in the face annotation model to be trained can be corrected using the backpropagation method.
[0076] S240. Take the convergence of the loss function in the face annotation model to be trained as the training objective to obtain the target face annotation model.
[0077] Specifically, the training error of the loss function, i.e., the loss parameter, can be used as a condition to detect whether the loss function has reached convergence. For example, whether the training error is less than a preset error, whether the error trend is stable, or whether the current number of iterations equals a preset number. If convergence is detected, such as the training error of the loss function being less than the preset error or the error trend being stable, it indicates that the training of the facial annotation model is complete, and iterative training can be stopped. If convergence is not detected, the first training sample can be obtained to train the facial annotation model until the training error of the loss function is within a preset range. When the training error of the loss function converges, the facial annotation model to be trained can be used as the target facial annotation model.
[0078] S250. Obtain the image to be processed, which includes the target object.
[0079] S260. Based on the target facial annotation model, at least two facial features of the target object in the image to be processed are annotated to obtain the target image corresponding to the image to be processed.
[0080] The target facial annotation model is trained based on at least one first training sample. The first training sample includes an original sample image and a corresponding sample fusion image. The original image includes the object to be annotated. The sample fusion image is obtained by fusing at least two sample local annotation images. The facial features annotated in each sample local annotation image are different. The set of annotated parts in each sample local annotation image corresponds to at least two facial features. The sample local annotation image is generated by the corresponding target local annotation model.
[0081] The technical solution of this disclosure, by determining the original image and the sample fusion image in the first training sample, and training the target annotation model based on the training sample, obtains the target annotation model. The image can be processed based on the target facial annotation model, thereby improving the technical effect of subsequent facial image annotation effectiveness and accuracy.
[0082] Figure 3 This is a schematic flowchart of an image processing method provided in an embodiment of this disclosure. Based on the foregoing embodiments, a second training sample can be obtained to train the local annotation model to be trained, thereby obtaining the target local annotation model. Specific implementation methods can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0083] like Figure 3 As shown, the method specifically includes the following steps:
[0084] S310. For each local annotation model to be trained, obtain the second training sample that is the same as the current local annotation model to be trained.
[0085] The second training sample includes the original sample image and the training labeled image that corresponds to the labeled parts of the local labeling model to be trained. The labeled parts correspond to one of at least two facial features. The training labeled image is an image in which the user has pre-labeled a certain facial feature, and the labeled facial feature is consistent with the facial feature labeled by the corresponding local labeling model to be trained.
[0086] The local annotation model of the training sample is a model whose model parameters are the initial parameters or default parameters. It needs to be trained on a second training sample to obtain the target local annotation model. The annotation image to be trained can be an image that accurately annotates a facial feature region corresponding to the original sample image, an image that has been accurately annotated by developers in advance using other annotation methods, or an image that has been manually annotated. The annotation of this image is accurate.
[0087] It should be noted that the local annotation model to be trained corresponds to a certain facial feature to be annotated, and the set of annotated parts of each local annotation model to be trained corresponds to the at least two facial features.
[0088] S320. For each second training sample, input the original sample image in the current second training sample into the local annotation model to be trained to obtain the output annotation image with the annotation of the annotated parts.
[0089] The output labeled image is the image that the local labeling model actually outputs after the original labeled image is input into the local labeling model to be trained. However, since the parameters of the local labeling model to be trained are default parameters or initial parameters, the image differs from the theoretically output image.
[0090] Specifically, the original sample image is used as input to the local annotation model to be trained. The local annotation model performs local annotation on the image and actually outputs an image with local annotation, which is used as the output annotation image.
[0091] S330. Based on the output labeled image and the labeled image to be trained in the second training sample, determine the loss value, and correct the model parameters in the local labeled model to be trained based on the loss value.
[0092] Specifically, the model parameters in the local annotation model to be trained do not meet the expected requirements. Therefore, the output annotation image based on the current model parameters differs from the theoretically expected training annotation image. Thus, the corresponding error loss value can be determined based on the training annotation image corresponding to the original sample image and the output annotation image. Furthermore, based on this loss value, the backpropagation method is used to correct the model parameters in the local annotation model to be trained.
[0093] S340. Take the convergence of the loss function in the local annotation model to be trained as the training objective to obtain the target local annotation model.
[0094] Specifically, the training error of the loss function, i.e., the loss parameter, can be used as a condition to detect whether the loss function has reached convergence. For example, whether the training error is less than a preset error, whether the error trend is stable, or whether the current number of iterations equals a preset number. If the convergence condition is met, such as the training error of the loss function being less than the preset error or the error trend being stable, it indicates that the illumination estimation model to be trained has completed training, and iterative training can be stopped. If the convergence condition is not met, a second training sample can be obtained to train the local annotation model to be trained until the training error of the loss function is within a preset range. When the training error of the loss function converges, the local annotation model to be trained can be used as the target local annotation model.
[0095] The target facial annotation model is trained based on at least one first training sample. The first training sample includes an original sample image and a corresponding sample fusion image. The original image includes the object to be annotated. The sample fusion image is obtained by fusing at least two sample local annotation images. The facial features annotated in each sample local annotation image are different. The set of annotated parts in each sample local annotation image corresponds to at least two facial features. The sample local annotation image is generated by the corresponding target local annotation model.
[0096] In this embodiment, after obtaining the target local annotation model, the original sample image can be processed based on each target local annotation model to obtain a sample local annotation image. By fusing the sample local annotation images, a sample fusion image in the first training sample is obtained. Further, the face annotation model to be trained is trained based on the original sample image and the sample fusion image to obtain a target face annotation model. At least two facial features are annotated on the acquired image to be processed based on the target face annotation model. The technical solution of this embodiment, after training each local annotation model, can obtain a local annotation image corresponding to the original image based on each local annotation model, and determine a sample fusion image corresponding to the original image based on each local annotation image, thereby improving the technical effect of subsequent facial feature annotation accuracy.
[0097] Figure 4 This is a structural block diagram of an image processing apparatus provided in an embodiment of the present disclosure. It can execute the image processing method provided in any embodiment of the present disclosure and has corresponding functional modules and beneficial effects for executing the method. For example... Figure 4 As shown, the device specifically includes an image acquisition module 410 and an image processing module 420.
[0098] Image acquisition module 410 is used to acquire an image to be processed, including the target object, in response to the detection of a trigger condition;
[0099] Image processing module 420 is used to annotate at least two facial features of a target object in the image to be processed based on a target facial annotation model, so as to obtain a target image corresponding to the image to be processed;
[0100] The target face annotation model is trained based on at least one first training sample. The first training sample includes an original sample image and a sample fusion image corresponding to the original sample image. The sample fusion image is obtained by fusing at least two sample local annotation images.
[0101] Based on the above technical solutions, the image acquisition module 410 includes:
[0102] The trigger module is used to jump to the image upload page when the trigger image annotation control is detected, and to use the image corresponding to the image upload completion instruction received as the image to be processed;
[0103] The display module is used to take the image displayed on the display interface as the image to be processed when the image annotation control is triggered.
[0104] The image library module is used to jump to the image library when the trigger image annotation control is detected, and to use the triggered image in the image library as the image to be processed;
[0105] The camera detection module is used to capture an image of the target object that is included in the camera view when the camera detects that the target object is included in the camera view.
[0106] The wake word detection module is used to capture an image of the target object when a wake word is detected.
[0107] Based on the above technical solutions, the image processing device further includes:
[0108] The first training sample determination module is used to determine at least one first training sample for training the target facial annotation model; wherein, the first training sample includes the original sample image and the corresponding sample fusion image;
[0109] The actual output image acquisition module takes the original sample image in the current training sample as the input parameter of the face annotation model to be trained for each first training sample, and obtains the actual output image corresponding to the original sample image.
[0110] The model parameter correction module is used to determine the loss value based on the sample fusion image of the actual output image and the current training sample, and to correct the model parameters of the face annotation model to be trained based on the loss value.
[0111] The target facial annotation model determination module uses the convergence of the loss function in the facial annotation model to be trained as the training objective to obtain the target facial annotation model.
[0112] Based on the above technical solutions, the first training sample determination module includes:
[0113] The original sample image acquisition module is used to acquire at least one original sample image;
[0114] The sample local annotation module is used to input the current original sample image into each target local annotation model for each original sample image to obtain at least one sample local annotation image corresponding to the original sample image; wherein, the sample local annotation image corresponds to the annotation part of the target local annotation model;
[0115] The sample fusion module is used to fuse at least one locally annotated sample image corresponding to the current original sample image into one image for each original sample image, thereby obtaining a sample fusion image corresponding to the current original sample image;
[0116] The sample determination module is used to determine at least one first training sample for training the target face annotation based on each original sample image and the corresponding sample fusion image.
[0117] Based on the above technical solutions, the sample fusion module includes:
[0118] The sample fusion unit is used to extract the labeled parts in the local labeled images of each sample and stitch the labeled parts onto an image to obtain a sample fusion image corresponding to the current original sample image.
[0119] Based on the above technical solutions, the at least two facial features include at least two of the following: left eyebrow, right eyebrow, left eye, right eye, nose, upper lip, lower lip, left ear, and right ear, with each facial feature corresponding to a target local annotation model.
[0120] Based on the above technical solutions, the image processing device also includes:
[0121] The second training sample acquisition module is used to acquire a second training sample for each local annotation model to be trained, which is consistent with the current local annotation model to be trained; wherein, the second training sample includes the original sample image and the annotation image to be trained that is consistent with the annotation part of the local annotation model to be trained, and the annotation part corresponds to one of the at least two facial features.
[0122] The output labeled image acquisition module is used to input the original sample image in the current second training sample into the local labeling model to be trained for each second training sample, so as to obtain the output labeled image of the labeled part;
[0123] The loss value determination module determines the loss value based on the output labeled image and the training labeled image in the second training sample, and corrects the model parameters in the training local labeled model based on the loss value.
[0124] The target local annotation model determination module is used to converge the loss function in the local annotation model to be trained as the training target to obtain the target local annotation model.
[0125] The technical solution of this disclosure involves acquiring an image to be processed, including a target object; annotating at least two facial features of the target object in the image to be processed based on a target facial annotation model to obtain a target image corresponding to the image to be processed; wherein, the target facial annotation model is trained based on a first training sample, the first training sample including an original sample image and a corresponding sample fusion image, the original sample image including the object to be annotated, the sample fusion image being obtained by fusing at least two sample local annotation images, and the at least two sample local annotation images being generated by annotating the corresponding parts of the at least two facial features of the object to be annotated by the corresponding target local annotation model. The technical solution of this disclosure solves the technical problem in the prior art where inaccurate facial feature annotation leads to poor special effects processing, resulting in a poor user experience. It achieves the determination of the annotation image corresponding to the same original image based on each target local annotation model, and after fusing the annotation images to obtain a sample fusion image, the target facial annotation model is trained based on the original image and the corresponding sample fusion image, improving the accuracy and effectiveness of facial feature annotation. Furthermore, when performing special effects processing based on the annotated facial features, the technical effect of the special effects processing can be further improved. The image processing apparatus provided in this disclosure can execute the image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0126] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0127] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 5 of this disclosure. Refer to the following... Figure 5 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0128] like Figure 5As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 506 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.
[0129] Typically, the following devices can be connected to I / O interface 505: editing devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0130] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 506, or installed from ROM 502. When the computer program is executed by processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0131] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0132] The electronic device provided in this embodiment and the image processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0133] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the image processing method provided in the above embodiments.
[0134] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0135] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0136] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0137] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0138] Acquire the image to be processed, which includes the target object;
[0139] Based on the target facial annotation model, at least two facial features of the target object in the image to be processed are annotated to obtain a target image corresponding to the image to be processed.
[0140] The target facial annotation model is trained based on at least one first training sample. The first training sample includes an original sample image and a corresponding sample fusion image. The original image includes the object to be annotated. The sample fusion image is obtained by fusing at least two sample local annotation images. The facial features annotated in each sample local annotation image are different. The set of annotated parts in each sample local annotation image corresponds to the at least two facial features. The sample local annotation image is generated by the corresponding target local annotation model.
[0141] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0143] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0144] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0146] According to one or more embodiments of this disclosure, [Example 1] provides an image processing method, the method comprising:
[0147] In response to the detection of a trigger condition, acquire the image to be processed, including the target object;
[0148] Based on the target facial annotation model, at least two facial features of the target object in the image to be processed are annotated to obtain a target image corresponding to the image to be processed.
[0149] The target face annotation model is trained based on a first training sample, which includes original sample images and corresponding sample fusion images. The sample fusion images are obtained by fusing at least two sample local annotation images.
[0150] According to one or more embodiments of this disclosure, [Example 2] provides an image processing method, further comprising:
[0151] Optionally, the original sample image includes an object to be labeled, and the at least two sample local labeled images are generated by corresponding target local labeling models, with each target local labeling model labeling different facial features of the object to be labeled.
[0152] According to one or more embodiments of this disclosure, [Example 3] provides an image processing method, further comprising:
[0153] Optionally, the triggering condition includes at least one of the following:
[0154] Special effects props associated with facial annotations;
[0155] Audio information associated with facial annotations.
[0156] According to one or more embodiments of this disclosure, [Example 4] provides an image processing method, further comprising:
[0157] Optionally, the at least one first training sample is determined by the following steps:
[0158] Obtain at least one original sample image; and
[0159] For each original sample image,
[0160] The current original sample image is input into each target local annotation model to obtain at least two sample local annotation images corresponding to the current original sample image;
[0161] The at least two locally annotated sample images are fused into one image to obtain a sample fused image corresponding to the current original sample image;
[0162] Based on each original sample image and the corresponding sample fusion image, at least one first training sample is determined.
[0163] According to one or more embodiments of this disclosure, [Example 5] provides an image processing method, further comprising:
[0164] Optionally, the at least two locally annotated sample images are fused into one image to obtain a fused sample image corresponding to the original sample image, including:
[0165] Extract the labeled areas from the local labeled images of each sample;
[0166] The labeled areas are stitched onto an image to obtain a sample fusion image corresponding to the current original sample image.
[0167] According to one or more embodiments of this disclosure, [Example Six] provides an image processing method, further comprising:
[0168] Optionally, the at least two locally annotated sample images are fused into one image to obtain a fused sample image corresponding to the original sample image, including:
[0169] Extract the labeled areas from the local labeled images of each sample;
[0170] The labeled areas are stitched onto an image to obtain a sample fusion image corresponding to the current original sample image.
[0171] According to one or more embodiments of this disclosure, [Example Seven] provides an image processing method, further comprising:
[0172] Optionally, the target facial annotation model is trained through the following steps:
[0173] For each first training sample, the original sample image in the current first training sample is used as the input parameter of the face annotation model to be trained, and the actual output image corresponding to the original sample image is obtained.
[0174] Based on the actual output image and the sample fusion image of the first training sample, a loss value is determined, and the model parameters of the face annotation model to be trained are corrected based on the loss value.
[0175] The convergence of the loss function in the face annotation model to be trained is used as the training objective to obtain the target face annotation model.
[0176] According to one or more embodiments of this disclosure, [Example Eight] provides an image processing method, further comprising:
[0177] Optionally, the target local annotation model is trained through the following steps:
[0178] For each local annotation model to be trained, a second training sample is obtained that corresponds to the current local annotation model to be trained; wherein, the second training sample includes the original sample image and the annotation image to be trained that is consistent with the annotation part of the local annotation model to be trained, and the annotation part corresponds to one of the at least two facial features.
[0179] For each second training sample, the original sample image in the current second training sample is input into the local annotation model to be trained to obtain the output annotation image with the annotation of the annotation part;
[0180] Based on the output labeled image and the labeled image to be trained in the second training sample, a loss value is determined, and the model parameters in the local labeled model to be trained are corrected based on the loss value;
[0181] The convergence of the loss function in the local annotation model to be trained is used as the training objective to obtain the target local annotation model.
[0182] According to one or more embodiments of this disclosure, [Example Nine] provides an image processing apparatus, including:
[0183] The image acquisition module is used to acquire the image to be processed, including the target object;
[0184] The image processing module is used to annotate at least two facial features of a target object in the image to be processed based on a target facial annotation model to obtain a target image corresponding to the image to be processed; wherein, the target facial annotation model is trained based on a first training sample, the first training sample includes an original sample image and a corresponding sample fusion image, and the sample fusion image is obtained by fusing at least two sample local annotation images.
[0185] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0186] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0187] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image processing method, characterized in that, include: In response to the detection of a trigger condition, acquire the image to be processed, including the target object; Based on the target facial annotation model, at least two facial features of the target object in the image to be processed are annotated to obtain a target image corresponding to the image to be processed. The target face annotation model is trained based on a first training sample, which includes an original sample image and a corresponding sample fusion image. The sample fusion image is obtained by fusing at least two sample local annotation images. The original sample image includes an object to be labeled. The at least two sample local labeled images are generated by corresponding target local labeling models. Each target local labeling model labels different facial features of the object to be labeled, and each facial feature corresponds to one target local labeling model.
2. The method according to claim 1, characterized in that, The triggering condition includes at least one of the following: Special effects props associated with facial annotations; Audio information associated with facial annotations.
3. The method according to claim 1, characterized in that, The at least one first training sample is determined by the following steps: Obtain at least one original sample image; and For each original sample image, The current original sample image is input into each target local annotation model to obtain at least two sample local annotation images corresponding to the current original sample image; The at least two locally annotated sample images are fused into one image to obtain a sample fused image corresponding to the current original sample image; Based on each original sample image and the corresponding sample fusion image, at least one first training sample is determined.
4. The method according to claim 3, characterized in that, The at least two locally annotated sample images are fused into one image to obtain a fused sample image corresponding to the original sample image, including: Extract the labeled areas from the local labeled images of each sample; The labeled areas are stitched onto an image to obtain a sample fusion image corresponding to the current original sample image.
5. The method according to claim 1, characterized in that, The at least two facial features include at least two of the following: left eyebrow; right eyebrow; left eye; right eye; nose; upper lip; lower lip; left ear; and right ear.
6. The method according to claim 3, characterized in that, The target facial annotation model is trained through the following steps: For each first training sample, the original sample image in the current first training sample is used as the input parameter of the face annotation model to be trained, and the actual output image corresponding to the original sample image is obtained. Based on the actual output image and the sample fusion image of the first training sample, a loss value is determined, and the model parameters of the face annotation model to be trained are corrected based on the loss value. The convergence of the loss function in the face annotation model to be trained is used as the training objective to obtain the target face annotation model.
7. The method according to claim 1, characterized in that, The target local annotation model is trained through the following steps: For each local annotation model to be trained, a second training sample is obtained that corresponds to the current local annotation model to be trained; wherein, the second training sample includes the original sample image and the annotation image to be trained that is consistent with the annotation part of the local annotation model to be trained, and the annotation part corresponds to one of the at least two facial features. For each second training sample, the original sample image in the current second training sample is input into the local annotation model to be trained to obtain the output annotation image with the annotation of the annotation part; Based on the output labeled image and the labeled image to be trained in the second training sample, a loss value is determined, and the model parameters in the local labeled model to be trained are corrected based on the loss value; The convergence of the loss function in the local annotation model to be trained is used as the training objective to obtain the target local annotation model.
8. An image processing apparatus, characterized in that, include: The image acquisition module is used to acquire the image to be processed, including the target object; An image processing module is used to annotate at least two facial features of a target object in the image to be processed based on a target facial annotation model, thereby obtaining a target image corresponding to the image to be processed; wherein, the target facial annotation model is trained based on a first training sample, the first training sample including an original sample image and a corresponding sample fusion image, the sample fusion image being obtained by fusing at least two sample local annotation images; The original sample image includes an object to be labeled. The at least two sample local labeled images are generated by corresponding target local labeling models. Each target local labeling model labels different facial features of the object to be labeled, and each facial feature corresponds to one target local labeling model.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in any one of claims 1-7.
10. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the image processing method as described in any one of claims 1-7.