Portrait sketch generation method, device, electronic device and storage medium

By combining the saliency map and the segmentation contour map to optimize the portrait sketch generation method, the problem of low accuracy of portrait sketches in the existing technology is solved, and portrait sketch generation with higher precision is achieved.

CN120451321BActive Publication Date: 2025-09-16BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510942041.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-16
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

In existing portrait sketch generation methods, the accuracy of generated portrait sketches is low, and the initialization point positions obtained from the saliency map are not accurate enough, resulting in the loss of key features and unreasonable stroke distribution.

Method used

By obtaining the original portrait and the number of strokes, the saliency map is combined with the arbitrary segmentation module to extract the initial stroke points and optimize their positions. The differentiable rasterization module and the contrastive language image pre-training module are combined to optimize the stroke parameter information and finally generate a portrait sketch.

Benefits of technology

The accuracy of portrait sketch generation is improved. By combining the saliency map and the segmentation contour map to optimize the stroke position, the accuracy of the initialization point position is enhanced, and the accuracy of the generated results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451321B_ABST
    Figure CN120451321B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, electronic device, and storage medium for generating a portrait sketch. The method comprises: obtaining an original portrait for generating a portrait sketch, and the number of strokes in the portrait sketch; inputting the original portrait and the number of strokes into a portrait sketch generation model to obtain a saliency map corresponding to the original portrait, and obtaining a number of first initial stroke points that matches the number of strokes based on the saliency map; inputting the original portrait into a segmentation arbitrary module in the portrait sketch generation model to obtain a segmentation contour map of the original portrait, and obtaining a number of second initial stroke points that matches the number of strokes based on the segmentation contour map; optimizing the first initial stroke position of each first initial stroke point using the second initial stroke position of each second initial stroke point to obtain a number of third initial stroke positions that matches the number of strokes, and generating a portrait sketch based on the third initial stroke position. The present disclosure can improve the accuracy of the generated portrait sketch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a method, device, electronic device and storage medium for generating a portrait sketch. Background Art

[0002] With the development of image processing technology, a technique for generating portrait sketches using artificial intelligence has emerged. For example, the CLIP model (Contrastive Language Image Pre-training) can be used to generate portrait sketches. This model can extract semantic concepts from sketches and images. This model defines the sketch as a set of Bezier curves and uses a differentiable rasterizer to directly optimize the curve parameters, targeting a perceptual loss based on the CLIP model. The number of strokes is used to control the degree of abstraction. Using the CLIP model to generate portrait sketches eliminates the need for training the model, improving the efficiency of portrait sketch generation.

[0003] The above-mentioned technology for generating portrait sketches through the CLIP model relies on the semantic perception ability of CLIP, namely the saliency map, to obtain the initialization point positions of the sketch strokes. However, the initialization point positions obtained in this way are often not accurate enough, which may lead to the inability to accurately depict the key features of the portrait. The generated portrait sketch will have problems such as missing key features and unreasonable stroke distribution. Therefore, the accuracy of generated portrait sketches in existing portrait sketch generation methods is low. Summary of the Invention

[0004] The present disclosure provides a method, device, electronic device, and storage medium for generating a portrait sketch, to at least address the problem of low accuracy in generating portrait sketches in related technologies. The technical solutions of the present disclosure are as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a method for generating a portrait sketch is provided, comprising:

[0006] Obtaining an original portrait for generating a portrait sketch and the number of strokes of the portrait sketch;

[0007] Inputting the original portrait and the number of strokes into a portrait sketch generation model to obtain a saliency map corresponding to the original portrait, and acquiring a number of first initial stroke points that matches the number of strokes based on the saliency map;

[0008] Inputting the original portrait into a segmentation module in the portrait sketch generation model, obtaining a segmentation contour map of the original portrait through the segmentation module, and obtaining a number of second initial stroke points that matches the number of strokes based on the segmentation contour map;

[0009] The first initial stroke position of each of the second initial stroke points is optimized by using the second initial stroke position of each of the first initial stroke points to obtain a third initial stroke position matching the number of strokes, and the portrait sketch is generated based on the third initial stroke position.

[0010] In an exemplary embodiment, the segmentation contour map of the original portrait is obtained by the arbitrary segmentation module, and based on the segmentation contour map, a second initial stroke point matching the number of strokes is obtained, including: extracting the portrait features of the original portrait by the arbitrary segmentation module, and obtaining the segmentation contour map of the original portrait based on the portrait features; sampling the pixel points of the segmentation contour map according to the number of strokes, and using each sampling point as the second initial stroke point; wherein the arc length difference between each second initial stroke point is the same.

[0011] In an exemplary embodiment, the use of the second initial stroke position of each second initial stroke point to optimize the first initial stroke position of each first initial stroke point to obtain a number of third initial stroke positions that matches the number of strokes includes: for each current first initial stroke point, obtaining a current second initial stroke point corresponding to the current first initial stroke point from each second initial stroke point; the current first initial stroke point is any first initial stroke point; when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is less than or equal to a preset difference threshold, the first initial stroke position of the current first initial stroke point is used as the third initial stroke position; when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is greater than the preset difference threshold, the second initial stroke position of the current second initial stroke point is used as the third initial stroke position.

[0012] In an exemplary embodiment, the generating of the portrait sketch based on the third initial stroke position includes: constructing stroke parameter information of each stroke used to generate the portrait sketch according to each of the third initial stroke positions; inputting the stroke parameter information into the differentiable rasterization module in the portrait sketch generation model, and obtaining a candidate portrait sketch through the differentiable rasterization module; inputting the candidate portrait sketch and the original portrait into the comparative language image pre-training module in the portrait sketch generation model, obtaining the difference between the candidate portrait sketch and the original portrait through the comparative language image pre-training module, optimizing the stroke parameter information using the difference, and obtaining optimized stroke parameter information; when the optimized stroke parameter information meets a preset condition, inputting the optimized stroke parameter information into the differentiable rasterization module, and obtaining the portrait sketch through the differentiable rasterization module.

[0013] In an exemplary embodiment, each of the stroke parameter information includes: any one of the third initial stroke positions, and the position information to be optimized that matches the third initial stroke position; optimizing the stroke parameter information using the difference to obtain the optimized stroke parameter information includes: updating the position information to be optimized that matches any one of the third initial stroke positions using the difference to obtain the position information to be optimized after the third initial stroke position is updated; obtaining any one of the optimized stroke parameter information based on the third initial stroke position and the position information to be optimized after the third initial stroke position is updated.

[0014] In an exemplary embodiment, obtaining the difference between the candidate portrait sketch and the original portrait through the comparative language image pre-training module includes: obtaining the geometric difference between the candidate portrait sketch and the original portrait, and the semantic difference between the candidate portrait sketch and the original portrait through the comparative language image pre-training module; and obtaining the difference between the candidate portrait sketch and the original portrait based on the geometric difference and the semantic difference.

[0015] In an exemplary embodiment, the comparative language image pre-training module includes multiple encoding layers; obtaining the geometric difference between the candidate portrait sketch and the original portrait, as well as the semantic difference between the candidate portrait sketch and the original portrait through the comparative language image pre-training module, includes: obtaining the first encoding information corresponding to the candidate portrait sketch and the first encoding information corresponding to the original portrait through the first encoding layer of the comparative language image pre-training module; the first encoding layer is the last of the multiple encoding layers; taking the difference between the first encoding information corresponding to the candidate portrait sketch and the first encoding information corresponding to the original portrait as the semantic difference; obtaining the second encoding information corresponding to the candidate portrait sketch and the second encoding information corresponding to the original portrait through the second encoding layer of the comparative language image pre-training module; the second encoding layer is a preset intermediate encoding layer among the multiple encoding layers; taking the difference between the second encoding information corresponding to the candidate portrait sketch and the second encoding information corresponding to the original portrait as the geometric difference.

[0016] In an exemplary embodiment, after obtaining the optimized stroke parameter information, it also includes: when the optimized stroke parameter information does not meet the preset conditions, using the optimized stroke parameter information as new stroke parameter information, and returning to execute the step of inputting the stroke parameter information into the differentiable rasterization module in the portrait sketch generation model.

[0017] According to a second aspect of an embodiment of the present disclosure, there is provided a device for generating a portrait sketch, comprising:

[0018] an original portrait obtaining unit, configured to obtain an original portrait for generating a portrait sketch, and the number of strokes of the portrait sketch;

[0019] A first stroke point acquisition unit is configured to input the original portrait and the number of strokes into a portrait sketch generation model to obtain a saliency map corresponding to the original portrait, and acquire a number of first initial stroke points that matches the number of strokes based on the saliency map;

[0020] a second stroke point acquisition unit configured to execute a segmentation arbitrary module in the portrait sketch generation model by inputting the original portrait into the segmentation arbitrary module, obtain a segmentation contour map of the original portrait through the segmentation arbitrary module, and obtain a number of second initial stroke points that matches the number of strokes based on the segmentation contour map;

[0021] The portrait sketch generation unit is configured to utilize the second initial stroke position of each second initial stroke point to optimize the first initial stroke position of each first initial stroke point, obtain a third initial stroke position matching the number of strokes, and generate the portrait sketch based on the third initial stroke position.

[0022] In an exemplary embodiment, the second stroke point acquisition unit is further configured to extract the portrait features of the original portrait through the arbitrary segmentation module, and obtain a segmentation contour map of the original portrait based on the portrait features; sample the pixel points of the segmentation contour map according to the number of strokes, and use each sampling point as the second initial stroke point; wherein the arc length difference between each of the second initial stroke points is the same.

[0023] In an exemplary embodiment, the portrait sketch generation unit is further configured to execute, for each current first initial stroke point, obtaining a current second initial stroke point corresponding to the current first initial stroke point from each second initial stroke point; the current first initial stroke point is any first initial stroke point; when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is less than or equal to a preset difference threshold, the first initial stroke position of the current first initial stroke point is used as the third initial stroke position; when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is greater than the preset difference threshold, the second initial stroke position of the current second initial stroke point is used as the third initial stroke position.

[0024] In an exemplary embodiment, the portrait sketch generation unit is further configured to execute, based on each of the third initial stroke positions, construct stroke parameter information for generating the portrait sketch; input the stroke parameter information into the differentiable rasterization module in the portrait sketch generation model, and obtain a candidate portrait sketch through the differentiable rasterization module; input the candidate portrait sketch and the original portrait into the comparative language image pre-training module in the portrait sketch generation model, obtain the difference between the candidate portrait sketch and the original portrait through the comparative language image pre-training module, optimize the stroke parameter information using the difference, and obtain optimized stroke parameter information; when the optimized stroke parameter information meets a preset condition, input the optimized stroke parameter information into the differentiable rasterization module, and obtain the portrait sketch through the differentiable rasterization module.

[0025] In an exemplary embodiment, each of the stroke parameter information includes: any one of the third initial stroke positions, and the position information to be optimized that matches the third initial stroke position; the portrait sketch generation unit is further configured to execute the update of the position information to be optimized that matches any one of the third initial stroke positions using the difference to obtain the updated position information to be optimized of the third initial stroke position; and obtain any one of the optimized stroke parameter information based on the third initial stroke position and the updated position information to be optimized of the third initial stroke position.

[0026] In an exemplary embodiment, the portrait sketch generation unit is further configured to execute the comparative language image pre-training module to obtain the geometric difference between the candidate portrait sketch and the original portrait, as well as the semantic difference between the candidate portrait sketch and the original portrait; and obtain the difference between the candidate portrait sketch and the original portrait based on the geometric difference and the semantic difference.

[0027] In an exemplary embodiment, the comparative language image pre-training module includes multiple encoding layers; the portrait sketch generation unit is further configured to execute the first encoding layer of the comparative language image pre-training module to obtain the first encoding information corresponding to the candidate portrait sketch and the first encoding information corresponding to the original portrait; the first encoding layer is the last of the multiple encoding layers; the difference between the first encoding information corresponding to the candidate portrait sketch and the first encoding information corresponding to the original portrait is used as the semantic difference; the second encoding information corresponding to the candidate portrait sketch and the second encoding information corresponding to the original portrait are obtained through the second encoding layer of the comparative language image pre-training module; the second encoding layer is a preset intermediate encoding layer among the multiple encoding layers; the difference between the second encoding information corresponding to the candidate portrait sketch and the second encoding information corresponding to the original portrait is used as the geometric difference.

[0028] In an exemplary embodiment, the portrait sketch generation unit is further configured to execute the step of using the optimized stroke parameter information as new stroke parameter information when the optimized stroke parameter information does not meet the preset conditions, and return to execute the step of inputting the stroke parameter information into the differentiable rasterization module in the portrait sketch generation model.

[0029] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the portrait sketch generation method as described in any one of the embodiments of the first aspect.

[0030] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the portrait sketch generation method as described in any one of the embodiments in the first aspect.

[0031] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes instructions. When the instructions are executed by a processor of an electronic device, the electronic device is able to execute the portrait sketch generation method as described in any embodiment of the first aspect.

[0032] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:

[0033] The method comprises the following steps: obtaining an original portrait for generating a portrait sketch and the number of strokes of the portrait sketch; inputting the original portrait and the number of strokes into a portrait sketch generation model to obtain a saliency map corresponding to the original portrait, and obtaining first initial stroke points whose number matches the number of strokes based on the saliency map; inputting the original portrait into an arbitrary segmentation module in the portrait sketch generation model to obtain a segmentation contour map of the original portrait through the arbitrary segmentation module, and obtaining second initial stroke points whose number matches the number of strokes based on the segmentation contour map; optimizing the first initial stroke position of each first initial stroke point using the second initial stroke position of each second initial stroke point to obtain a third initial stroke position whose number matches the number of strokes, and generating a portrait sketch based on the third initial stroke position. After obtaining the original portrait used to generate a portrait sketch and the number of strokes in the portrait sketch, the present invention can not only obtain a saliency map through the portrait sketch generation model to extract the first initial stroke point, but also use the segmentation arbitrary module in the portrait sketch generation model to obtain a segmentation contour map to extract the second initial stroke point, and then use the second initial stroke position of the second initial stroke point to optimize the first initial stroke position of each first initial stroke point to obtain the third initial stroke position to generate the portrait sketch. Compared with the existing technology that only obtains the initialization point position of the sketch stroke through the saliency map, the present application can combine the segmentation arbitrary module to correct the initialization point position of the sketch stroke, thereby improving the accuracy of the initialization point position acquisition result, thereby improving the accuracy of the generated portrait sketch.

[0034] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0036] Figure 1 The figure is a flowchart of a method for generating a portrait sketch according to an exemplary embodiment.

[0037] Figure 2 FIG. 4 is a flowchart of obtaining a third initial stroke position according to an exemplary embodiment.

[0038] Figure 3 The figure is a flowchart of generating a portrait sketch according to an exemplary embodiment.

[0039] Figure 4 The figure is a flowchart of obtaining geometric difference and semantic difference according to an exemplary embodiment.

[0040] Figure 5 The figure is a flowchart of a method for generating a portrait sketch based on SAM segmentation guidance according to an exemplary embodiment.

[0041] Figure 6 The figure is a block diagram of a device for generating a portrait sketch according to an exemplary embodiment.

[0042] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0043] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0044] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0045] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0046] Figure 1 FIG. 1 is a flow chart of a method for generating a portrait sketch according to an exemplary embodiment. Figure 1As shown, the portrait sketch generation method is used in the server and includes the following steps.

[0047] In step S101 , an original portrait for generating a portrait sketch and the number of strokes of the portrait sketch are obtained.

[0048] The original portrait refers to the original person image used to generate the portrait sketch, such as the original face image. The portrait sketch refers to the line drawing sketch image generated based on the original face image. The number of strokes refers to the number of strokes contained in the generated portrait sketch. This number of strokes can be used to control the abstraction level of the portrait sketch and can be set by the user. Specifically, when a user needs to generate a portrait sketch, they can enter the original portrait used to generate the portrait sketch and the number of strokes of the portrait sketch. The server can then obtain the original portrait and the number of strokes.

[0049] In step S102, the original portrait and the number of strokes are input into a portrait sketch generation model to obtain a saliency map corresponding to the original portrait, and first initial stroke points matching the number of strokes are obtained based on the saliency map.

[0050] A saliency map is an image that visualizes the importance of input features (such as pixels) to model decisions in the form of a heat map. It can be used to reveal the focus areas of deep learning models. A portrait sketch generation model is a neural network model used to generate portrait sketches. This model takes an original portrait and the number of strokes as input and generates a portrait sketch of the original portrait with the specified number of strokes. The first initial stroke point refers to the initial position of each stroke in the portrait sketch, extracted from the saliency map.

[0051] Specifically, after obtaining the original portrait and the number of strokes, the original portrait and the number of strokes can be input into the portrait sketch generation model, and the portrait sketch generation model can be used to obtain the saliency map corresponding to the original portrait. Afterwards, the saliency map can be used as a distribution to sample the initial stroke points that match the number of strokes, thereby obtaining the first initial stroke point.

[0052] In step S103, the original portrait is input into the arbitrary segmentation module in the portrait sketch generation model, and a segmentation contour map of the original portrait is obtained by the arbitrary segmentation module. According to the segmentation contour map, a number of second initial stroke points matching the number of strokes is obtained.

[0053] The Segment Any Module can be the SAM module in the portrait sketch generation model. This module can achieve zero-shot segmentation of objects in any image using user prompts (such as points, boxes, and text). The segmentation contour map is the contour map of the original portrait segmented by the Segment Any Module. The second initial stroke points refer to the initial positions of each stroke in the portrait sketch, extracted from the segmentation contour map.

[0054] In this embodiment, a segmentation module is also provided in the portrait sketch generation model. By inputting the original portrait into the segmentation module, the segmentation module can be used to extract a segmentation contour map of the original portrait. Then, the initial stroke points matching the number of strokes can be extracted from the segmentation contour map as the second initial stroke points.

[0055] In step S104, the first initial stroke position of each first initial stroke point is optimized using the second initial stroke position of each second initial stroke point to obtain a third initial stroke position matching the number of strokes, and a portrait sketch is generated based on the third initial stroke position.

[0056] The first initial stroke position refers to the coordinate position of each first initial stroke point, the second initial stroke position refers to the coordinate position of each second initial stroke point, and the third initial stroke position refers to the coordinate position of each initial stroke point after optimization using the second initial stroke position. In this embodiment, in order to improve the accuracy of obtaining the initial stroke points, after obtaining the first initial stroke points through the saliency map and extracting the second initial stroke points by segmenting any module, the second initial stroke positions of each second initial stroke point can be used to optimize the first initial stroke position of each first initial stroke point, thereby obtaining the third initial stroke positions that match the number of strokes, and then a portrait sketch can be generated based on the above third initial stroke positions.

[0057] In the above-mentioned portrait sketch generation method, the original portrait used to generate the portrait sketch and the number of strokes of the portrait sketch are obtained; the original portrait and the number of strokes are input into the portrait sketch generation model to obtain a saliency map corresponding to the original portrait, and based on the saliency map, the first initial stroke points whose number matches the number of strokes are obtained; the original portrait is input into the arbitrary segmentation module in the portrait sketch generation model, and the segmentation contour map of the original portrait is obtained through the arbitrary segmentation module, and based on the segmentation contour map, the second initial stroke points whose number matches the number of strokes are obtained; the second initial stroke position of each second initial stroke point is used to optimize the first initial stroke position of each first initial stroke point to obtain a third initial stroke position whose number matches the number of strokes, and the portrait sketch is generated based on the third initial stroke position. After obtaining the original portrait used to generate a portrait sketch and the number of strokes in the portrait sketch, the present invention can not only obtain a saliency map through the portrait sketch generation model to extract the first initial stroke point, but also use the segmentation arbitrary module in the portrait sketch generation model to obtain a segmentation contour map to extract the second initial stroke point, and then use the second initial stroke position of the second initial stroke point to optimize the first initial stroke position of each first initial stroke point to obtain the third initial stroke position to generate the portrait sketch. Compared with the existing technology that only obtains the initialization point position of the sketch stroke through the saliency map, the present application can combine the segmentation arbitrary module to correct the initialization point position of the sketch stroke, thereby improving the accuracy of the initialization point position acquisition result, thereby improving the accuracy of the generated portrait sketch.

[0058] In an exemplary embodiment, step S103 may further include: extracting portrait features of the original portrait by segmenting any module, and obtaining a segmentation contour map of the original portrait based on the portrait features; sampling pixel points of the segmentation contour map according to the number of strokes, and using each sampling point as a second initial stroke point; wherein the arc length differences between each second initial stroke point are the same.

[0059] Among them, portrait features can refer to the key features of the portrait corresponding to the original portrait, for example, they can include anatomical boundary features, such as head contour features, and can also include facial features and significant area features, such as eye structure, nose shape, lip contour, eyebrow direction, etc., and can also include texture and color contrast features, such as skin color and background color difference and hair texture, etc.

[0060] Specifically, the process of obtaining the second initial stroke point can be that the server first inputs the original portrait into the segmentation arbitrary module in the portrait sketch generation model, extracts the key portrait features of the original portrait through the segmentation arbitrary module, and further outputs the segmentation contour map of the original portrait based on the portrait features. After that, the pixel points of the segmentation contour map can be equidistantly sampled according to the number of strokes to ensure that the arc length difference between each sampling point is the same, and each of the above sampling points is used as the second initial stroke point. For example, the total length of each arc length in the segmentation contour map can be first obtained, and the arc length difference between the two sampling points is calculated based on the total arc length and the number of strokes, so as to implement pixel sampling according to the above arc length difference to obtain multiple second initial stroke points.

[0061] In this embodiment, the portrait features of the original portrait can also be extracted by segmenting any module, and the segmentation contour map of the original portrait can be output in combination with the portrait features. After that, the pixel points can be equidistantly sampled according to the number of strokes to obtain the second initial stroke points. This can ensure that the initialization stroke points are set on the key features of the portrait, and the accurate generation of the portrait features can be guaranteed. The equal difference distribution can be conducive to the rationalization of the stroke distribution of the final sketch. In this way, the accuracy of the initial stroke point setting can be improved.

[0062] Further, if Figure 2 As shown, step S104 may further include:

[0063] In step S201, for each current first initial stroke point, a current second initial stroke point corresponding to the current first initial stroke point is obtained from each second initial stroke point; the current first initial stroke point is any first initial stroke point.

[0064] The current first initial stroke point refers to any first initial stroke point extracted from the saliency map, and the current second initial stroke point refers to the second initial stroke point corresponding to the current first initial stroke point. In this embodiment, the number of first initial stroke points and second initial stroke points is the same, which is equal to the number of strokes. Therefore, each first initial stroke point is matched with a corresponding second initial stroke point. After obtaining each current first initial stroke point, the server can obtain the current second initial stroke point corresponding to each current first initial stroke point based on the above matching relationship.

[0065] In step S202, when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is less than or equal to a preset difference threshold, the first initial stroke position of the current first initial stroke point is used as the third initial stroke position;

[0066] In step S203, when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is greater than a preset difference threshold, the second initial stroke position of the current second initial stroke point is used as the third initial stroke position.

[0067] Afterwards, it can be determined whether the difference between the first initial stroke position of each current first initial stroke point and the second initial stroke position of the corresponding current second initial stroke point is less than or equal to a preset difference threshold, that is, whether the difference between the position of each current first initial stroke point and the position of the current second initial stroke point is large. If the difference is small, that is, the distance between the current first initial stroke point and the current first initial stroke point is shorter, then the first initial stroke position of the current first initial stroke point can be used as one of the third initial stroke positions. If the difference is large, that is, the distance between the current first initial stroke point and the current first initial stroke point is longer, then the second initial stroke position of the current second initial stroke point can be used as one of the third initial stroke positions.

[0068] Taking the number of strokes 3 as an example, the first initial stroke points obtained through the saliency map may include stroke point A1, stroke point B1 and stroke point C1, and the second initial stroke points obtained through the segmentation contour map may include stroke point A2, stroke point B2 and stroke point C2. Then, the position difference between stroke point A1 and stroke point A2, the position difference between stroke point B1 and stroke point B2, and the position difference between stroke point C1 and stroke point C2 can be calculated respectively. If the position difference between stroke point A1 and stroke point A2 is less than the preset difference threshold, the position difference between stroke point B1 and stroke point B2 is greater than the preset difference threshold, and the position difference between stroke point C1 and stroke point C2 is less than the preset difference threshold, then the position of the first initial stroke point of stroke point A1, the position of the second initial stroke point of stroke point B2, and the position of the first initial stroke point of stroke point C1 can be used as the third initial stroke position.

[0069] In this embodiment, the position difference between each current second initial stroke point and the current first initial stroke point can also be used to obtain the third initial stroke position. In this way, the initial stroke position can be optimized and the accuracy of the initial stroke position setting can be improved.

[0070] In an exemplary embodiment, Figure 3 As shown, step S104 may further include:

[0071] In step S301, stroke parameter information of each stroke for generating a portrait sketch is constructed according to each third initial stroke position.

[0072] Stroke parameter information refers to parameter information used to control the depiction of a stroke. For example, each stroke can be represented as a parameterized Bezier curve. A Bezier curve typically includes four control points, and the stroke parameter information can be the coordinate information of the four control points in the Bezier curve. In this embodiment, each third initial stroke position can correspond to a stroke used to generate a portrait sketch. Therefore, stroke parameter information for each stroke used to generate the portrait sketch can be constructed based on each third initial stroke position.

[0073] In step S302, the stroke parameter information is input into the differentiable rasterization module in the portrait sketch generation model, and a candidate portrait sketch is obtained through the differentiable rasterization module;

[0074] In step S303, the candidate portrait sketch and the original portrait are input into the comparative language image pre-training module in the portrait sketch generation model, the difference between the candidate portrait sketch and the original portrait is obtained through the comparative language image pre-training module, and the stroke parameter information is optimized using the difference to obtain the optimized stroke parameter information.

[0075] The differentiable rasterizer module is the differentiable rasterizer included in the portrait sketch generation model. This rasterizer can be used to output a sketch image. The candidate portrait sketch is the portrait sketch directly output by the differentiable rasterizer module based on stroke parameter information. This candidate portrait sketch can be optimized and iterated to ultimately obtain the portrait sketch corresponding to the original portrait. This optimization and iteration process is achieved through the comparative language image pre-training module in the portrait sketch generation model, namely the CLIP module in the portrait sketch generation model.

[0076] Specifically, after the server constructs the stroke parameter information used in the first round of iteration based on the third initial stroke positions, the stroke parameter information can be input into the differentiable rasterization module in the portrait sketch generation model to obtain a candidate portrait sketch. Then, the candidate portrait sketch and the original portrait can be input into the comparative language image pre-training module in the portrait sketch generation model, that is, the CLIP module in the portrait sketch generation model. The CLIP module calculates the difference between the candidate portrait sketch and the original portrait, and uses the difference to perform reverse transmission optimization on the stroke parameter information of each stroke to obtain the optimized stroke parameter information.

[0077] In step S304, when the optimized stroke parameter information meets the preset conditions, the optimized stroke parameter information is input into the differentiable rasterization module, and the portrait sketch is obtained through the differentiable rasterization module.

[0078] Afterwards, it can be determined whether the optimized stroke parameter information meets the preset conditions. For example, the preset condition can be whether the number of optimization iterations of the stroke parameter information meets the preset number. If the optimized stroke parameter information meets the preset conditions, the server can input the optimized stroke parameter information into the differentiable rasterization module, and obtain the final portrait sketch through the output of the differentiable rasterization module.

[0079] In this embodiment, the stroke parameter information of each stroke can also be constructed through the third initial stroke position, so that the stroke parameter information is input into the differentiable rasterization module, and the candidate portrait sketch is obtained by the output of the differentiable rasterization module. The original portrait input is combined with the comparison language image pre-training module to obtain the difference between the candidate portrait sketch and the original portrait, thereby completing the optimization of the stroke parameter information to output the portrait sketch. In this way, the fineness of the portrait sketch output can be improved.

[0080] Furthermore, each stroke parameter information includes: any third initial stroke position, and position information to be optimized that matches the third initial stroke position; step S303 may further include: using the difference to update the position information to be optimized that matches any third initial stroke position, to obtain the updated position information to be optimized of the third initial stroke position; according to the third initial stroke position, and the updated position information to be optimized of the third initial stroke position, to obtain any optimized stroke parameter information.

[0081] In this embodiment, each stroke parameter information can be composed of two parts, namely, any third initial stroke position, and the position information to be optimized that matches the third initial stroke position. For example, the stroke parameter information can be the coordinate information of four control points in the Bezier curve, wherein the coordinate information of one control point is the third initial stroke position, and the coordinate information of the remaining three control points is the position information to be optimized corresponding to the third initial stroke position, and the process of updating each stroke parameter information using the difference can be to update each position information to be optimized using the difference.

[0082] Specifically, after obtaining the difference between the candidate portrait sketch and the original portrait by comparing the language image pre-training module, the position information to be optimized that matches any third initial stroke position can be updated based on the difference. After obtaining the updated position information to be optimized of the third initial stroke position, the third initial stroke position and the updated position information to be optimized can be used to obtain any optimized stroke parameter information.

[0083] Taking stroke parameter information 1 as an example, the stroke parameter information is used to control stroke 1 in portrait sketches. The stroke parameter 1 may include coordinate information of 4 control points. The coordinate information of the 4 control points includes 1 third initial stroke position and 3 position information to be optimized. When the stroke parameter 1 is updated using the difference, the 3 position information to be optimized can be optimized and updated to obtain 3 optimized position information to be optimized, thereby combining 1 third initial stroke position and 3 optimized position information to be optimized as the optimized stroke parameter 1.

[0084] In this embodiment, the stroke parameter information may include two parts, a third initial stroke position, and the position information to be optimized corresponding to the third initial stroke position. When the stroke parameter information is updated, the position information to be optimized may be updated, thereby combining the third initial stroke position and the updated position information to be optimized to obtain the optimized stroke parameter information. In this way, the accuracy of the stroke parameter information update can be improved.

[0085] In an exemplary embodiment, step S303 may further include: obtaining the geometric difference between the candidate portrait sketch and the original portrait, and the semantic difference between the candidate portrait sketch and the original portrait by comparing the language image pre-training module; and obtaining the difference between the candidate portrait sketch and the original portrait based on the geometric difference and the semantic difference.

[0086] In this embodiment, the difference between the candidate portrait sketch and the original portrait output by the comparative language image pre-training module is mainly composed of two parts: a geometric difference for measuring the geometric similarity between the candidate portrait sketch and the original portrait, and a semantic difference for measuring the degree of semantic similarity between the candidate portrait sketch and the original portrait. Specifically, after the candidate portrait sketch and the original portrait are input into the comparative language image pre-training module, the comparative language image pre-training module can first output the geometric difference and the semantic difference between the candidate portrait sketch and the original portrait, respectively, so as to obtain the difference between the candidate portrait sketch and the original portrait based on the geometric difference and the semantic difference. For example, the geometric difference and the semantic difference can be weighted by weighting the geometric difference and the semantic difference to obtain the difference between the candidate portrait sketch and the original portrait.

[0087] For example, the difference between a candidate portrait sketch and the original portrait can be calculated using the following formula:

[0088]

[0089] in, Characterize geometric differences, Representing semantic differences, while The characterization weight can be set to 0.1. It represents the parameter information of each stroke, including n.

[0090] In this embodiment, the geometric differences and semantic differences between the candidate portrait sketch and the original portrait can be output by comparing the language image pre-training module, so as to obtain the difference between the candidate portrait sketch and the original portrait by combining the geometric differences and semantic differences. In this way, the accuracy of obtaining the difference between the candidate portrait sketch and the original portrait can be improved.

[0091] Furthermore, the contrastive language image pre-training module includes multiple encoding layers; e.g. Figure 4 As shown, by comparing the language image pre-training module, obtaining the geometric difference between the candidate portrait sketch and the original portrait, as well as the semantic difference between the candidate portrait sketch and the original portrait, can further include:

[0092] In step S401, first coding information corresponding to the candidate portrait sketch and first coding information corresponding to the original portrait are obtained by comparing the first coding layer of the language image pre-training module; the first coding layer is the last one of the multiple coding layers;

[0093] In step S402, the difference between the first coding information corresponding to the candidate portrait sketch and the first coding information corresponding to the original portrait is used as a semantic difference.

[0094] In this embodiment, the comparative language image pre-training module may be a CLIP image encoder model comprising multiple encoding layers, wherein the first encoding layer is the last of the multiple encoding layers of the comparative language image pre-training module, and the first encoding information refers to the encoding information output by the first encoding layer. After the candidate portrait sketch and the original portrait are input into the comparative language image pre-training module, the first encoding information corresponding to the candidate portrait sketch and the first encoding information corresponding to the original portrait output by the first encoding layer of the comparative language image pre-training module may also be obtained. Thereafter, the difference between the first encoding information corresponding to the candidate portrait sketch and the first encoding information corresponding to the original portrait may be used as a semantic difference, and the distance between the first encoding information corresponding to the candidate portrait sketch and the first encoding information corresponding to the original portrait may be used as the semantic difference.

[0095] For example, semantic difference can be calculated by the following formula:

[0096]

[0097]

[0098] in, Representing semantic differences, Characterizes the first coding information corresponding to the original portrait, and The first encoding information corresponding to the candidate portrait sketch is represented, is the cosine distance.

[0099] In step S403, second coding information corresponding to the candidate portrait sketch and second coding information corresponding to the original portrait are obtained by comparing the second coding layer of the language image pre-training module; the second coding layer is a preset intermediate coding layer among the multiple coding layers;

[0100] In step S404, the difference between the second coded information corresponding to the candidate portrait sketch and the second coded information corresponding to the original portrait is used as a geometric difference.

[0101] The second coding layer is an intermediate coding layer of a predetermined number of layers among the multiple coding layers of the comparative language image pre-training module. For example, the third and fourth layers of the comparative language image pre-training module can be used as the second coding layer, and the second coding information is the coding information output by the second coding layer. Similarly, after the candidate portrait sketch and the original portrait are input into the comparative language image pre-training module, the second coding information corresponding to the candidate portrait sketch output by the second coding layer of the comparative language image pre-training module and the second coding information corresponding to the original portrait can also be obtained. Thereafter, the difference between the second coding information corresponding to the candidate portrait sketch and the second coding information corresponding to the original portrait can be used as a geometric difference, and the L2 distance between the second coding information corresponding to the candidate portrait sketch and the second coding information corresponding to the original portrait can be used as the geometric difference.

[0102] For example, the geometric difference can be calculated using the following formula:

[0103]

[0104] in, Representing semantic differences, Characterizing the original portrait corresponds to the coding layer l, that is, the second coding information of the second coding layer, and The sketch representing the candidate portrait corresponds to the coding layer 1, ie, the second coding information of the second coding layer. The coding layer 1 may be the third layer and the fourth layer.

[0105] In this embodiment, the difference between the first coding information output by the last coding layer in the comparison language image pre-training module can also be used as the semantic difference, and the difference between the second coding information output by the preset intermediate coding layer in the comparison language image pre-training module can be used as the geometric difference. In this way, the accuracy of obtaining semantic differences and geometric differences can be improved.

[0106] In addition, after step S303, the method may further include: if the optimized stroke parameter information does not meet the preset conditions, using the optimized stroke parameter information as new stroke parameter information, and returning to step S302.

[0107] If the optimized stroke parameter information does not meet the preset conditions, for example, the number of optimization times of the stroke parameter information does not reach the set number, the optimized stroke parameter information can be input as new stroke parameter information into the differentiable rasterization module in the portrait sketch generation model to obtain a new candidate portrait sketch, and the candidate portrait sketch and the original portrait are again input into the comparative language image pre-training module to calculate the difference and optimize the stroke parameter information. In this way, iterative optimization of the stroke parameter information is achieved, and the accuracy of portrait sketch generation is further improved.

[0108] In this embodiment, when the optimized stroke parameter information does not meet the preset conditions, the stroke parameter information can be iteratively optimized. In this way, iterative optimization of the stroke parameter information is achieved, further improving the accuracy of portrait sketch generation.

[0109] In an exemplary embodiment, a portrait sketch generation method based on SAM segmentation guidance is also provided. On the basis of the sketch generation method based on semantic guidance, SAM is introduced as an additional guidance, thereby solving the problems of missing key features and unreasonable stroke distribution when generating portrait sketches in existing methods. Figure 5 As shown, the method specifically includes the following steps:

[0110] Given the target person image I and the number of strokes n, we first use the saliency map as a distribution to sample the initial stroke positions Then input the portrait into the Segment Anything module, that is, the segmentation module, to obtain the segmentation contour map based on the key features of the portrait. Take n points at equal distances on the contour map and use the saliency map as the distribution to sample the initial stroke positions. Optimize and ensure that initial stroke points are set on the key features of the portrait, so as to ensure the accurate generation of the portrait features, and the equal distance distribution can be conducive to the rationalization of the stroke distribution of the final sketch.

[0111] Next, a differentiable rasterizer R is used to create a rasterized sketch S. Both the sketch and the image are fed into a pre-trained CLIP model to evaluate the geometric distance Lg and semantic distance Ls between them. The loss is backpropagated through R to optimize the stroke parameters until convergence.

[0112] Loss function: We utilize the pre-trained CLIP image encoder model, which is trained on various image modalities so that it can encode information from natural images and sketches without further training. CLIP encodes high-level semantic attributes in the last layer because it is trained on images and text. Therefore, we convert the sketch and images The distance between the embeddings of is defined as:

[0113]

[0114] in, is the cosine distance. However, the network’s final encoding is agnostic to low-level spatial features such as pose and structure. To measure the geometric similarity between the image and the sketch, and thus allow some control over the output appearance, the L2 distance between the activations of the intermediate CLIP layers is calculated:

[0115]

[0116] in, is the activation of the CLIP encoder at layer l. Specifically, layers 3 and 4 of the ResNet101 CLIP model are used. The final optimization objective is defined as:

[0117]

[0118] in Can be set to 0.1. It can be expressed as The greater the weight, the greater the effect of semantic loss. It can be any number 0 or above, but is generally not set to 0 to ensure that the semantic loss can play a certain role in the loss function.

[0119] Through this embodiment, a training-free framework is provided, which mainly relies on the multimodal guided optimization capability of CLIP. It can output accurate portrait sketches for any portrait, laying a solid data generation foundation for application scenarios such as multimedia interaction. In addition, by introducing the segmentation guidance of SAM, the distribution of the initialization stroke points of the portrait sketch is reasonably planned, which rationalizes the stroke distribution of the portrait sketch while retaining the key features of the portrait.

[0120] It should be understood that, although the various steps in the flowchart of the present disclosure are shown in sequence as indicated by the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0121] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can be referred to each other, and each embodiment focuses on the differences from other embodiments. For related parts, please refer to the description of other method embodiments.

[0122] Figure 6 1 is a block diagram of a device for generating a portrait sketch according to an exemplary embodiment. Figure 6 The device includes an original portrait acquisition unit 601, a first stroke point acquisition unit 602, a second stroke point acquisition unit 603 and a portrait sketch generation unit 604.

[0123] The original portrait acquisition unit 601 is configured to acquire an original portrait for generating a portrait sketch and the number of strokes of the portrait sketch;

[0124] A first stroke point acquisition unit 602 is configured to input the original portrait and the number of strokes into a portrait sketch generation model, obtain a saliency map corresponding to the original portrait, and obtain a number of first initial stroke points that matches the number of strokes based on the saliency map;

[0125] The second stroke point acquisition unit 603 is configured to execute the arbitrary segmentation module of the portrait sketch generation model by inputting the original portrait into the arbitrary segmentation module, obtain a segmentation contour map of the original portrait through the arbitrary segmentation module, and obtain a number of second initial stroke points that matches the number of strokes based on the segmentation contour map;

[0126] The portrait sketch generation unit 604 is configured to utilize the second initial stroke position of each second initial stroke point to optimize the first initial stroke position of each first initial stroke point, obtain a third initial stroke position that matches the number of strokes, and generate a portrait sketch based on the third initial stroke position.

[0127] In an exemplary embodiment, the second stroke point acquisition unit 603 is further configured to extract the portrait features of the original portrait by segmenting any module, and obtain a segmentation contour map of the original portrait based on the portrait features; sample the pixel points of the segmentation contour map according to the number of strokes, and use each sampling point as the second initial stroke point; wherein the arc length difference between each second initial stroke point is the same.

[0128] In an exemplary embodiment, the portrait sketch generation unit 604 is further configured to execute, for each current first initial stroke point, obtaining a current second initial stroke point corresponding to the current first initial stroke point from each second initial stroke point; the current first initial stroke point is any first initial stroke point; when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is less than or equal to a preset difference threshold, the first initial stroke position of the current first initial stroke point is used as the third initial stroke position; when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is greater than the preset difference threshold, the second initial stroke position of the current second initial stroke point is used as the third initial stroke position.

[0129] In an exemplary embodiment, the portrait sketch generation unit 604 is further configured to execute the construction of stroke parameter information of each stroke for generating the portrait sketch based on each third initial stroke position; input the stroke parameter information into the differentiable rasterization module in the portrait sketch generation model, and obtain the candidate portrait sketch through the differentiable rasterization module; input the candidate portrait sketch and the original portrait into the comparative language image pre-training module in the portrait sketch generation model, obtain the difference between the candidate portrait sketch and the original portrait through the comparative language image pre-training module, optimize the stroke parameter information using the difference, and obtain the optimized stroke parameter information; when the optimized stroke parameter information meets the preset conditions, input the optimized stroke parameter information into the differentiable rasterization module, and obtain the portrait sketch through the differentiable rasterization module.

[0130] In an exemplary embodiment, each stroke parameter information includes: any third initial stroke position, and position information to be optimized that matches the third initial stroke position; the portrait sketch generation unit 604 is further configured to execute the update of the position information to be optimized that matches any third initial stroke position using the difference to obtain the updated position information to be optimized after the third initial stroke position is updated; and any optimized stroke parameter information is obtained based on the third initial stroke position and the position information to be optimized after the third initial stroke position is updated.

[0131] In an exemplary embodiment, the portrait sketch generation unit 604 is further configured to execute a comparison of the language image pre-training module to obtain the geometric difference between the candidate portrait sketch and the original portrait, as well as the semantic difference between the candidate portrait sketch and the original portrait; and obtain the difference between the candidate portrait sketch and the original portrait based on the geometric difference and the semantic difference.

[0132] In an exemplary embodiment, the comparative language image pre-training module includes multiple coding layers; the portrait sketch generation unit 604 is further configured to execute the first coding layer of the comparative language image pre-training module to obtain the first coding information corresponding to the candidate portrait sketch and the first coding information corresponding to the original portrait; the first coding layer is the last of the multiple coding layers; the difference between the first coding information corresponding to the candidate portrait sketch and the first coding information corresponding to the original portrait is used as a semantic difference; the second coding information corresponding to the candidate portrait sketch and the second coding information corresponding to the original portrait are obtained through the second coding layer of the comparative language image pre-training module; the second coding layer is a preset intermediate coding layer among the multiple coding layers; the difference between the second coding information corresponding to the candidate portrait sketch and the second coding information corresponding to the original portrait is used as a geometric difference.

[0133] In an exemplary embodiment, the portrait sketch generation unit 604 is further configured to execute the step of using the optimized stroke parameter information as new stroke parameter information when the optimized stroke parameter information does not meet the preset conditions, and return to execute the step of inputting the stroke parameter information into the differentiable rasterization module in the portrait sketch generation model.

[0134] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0135] Figure 7 FIG. 7 is a block diagram of an electronic device 700 for generating a portrait sketch according to an exemplary embodiment. For example, the electronic device 700 may be a server. Figure 7 The electronic device 700 includes a processing component 720, which further includes one or more processors, and a memory resource represented by a memory 722 for storing instructions executable by the processing component 720, such as an application. The application stored in the memory 722 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 720 is configured to execute the instructions to perform the above method.

[0136] The electronic device 700 may further include a power supply component 724 configured to perform power management of the electronic device 700, a wired or wireless network interface 726 configured to connect the electronic device 700 to a network, and an input / output (I / O) interface 728. The electronic device 700 may operate based on an operating system stored in the memory 722, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, or the like.

[0137] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 722 including instructions. The instructions may be executed by a processor of the electronic device 700 to perform the above method. The storage medium may be a computer-readable storage medium, such as a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0138] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions, and the instructions can be executed by a processor of the electronic device 700 to implement the above method.

[0139] It should be noted that the above-mentioned devices, electronic devices, computer-readable storage media, computer program products, etc. can also include other implementation methods according to the description of the method embodiments. The specific implementation methods can refer to the description of the relevant method embodiments and will not be described one by one here.

[0140] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0141] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for generating a portrait sketch, characterized in that: include: Obtaining an original portrait for generating a portrait sketch and the number of strokes of the portrait sketch; Inputting the original portrait and the number of strokes into a portrait sketch generation model to obtain a saliency map corresponding to the original portrait, and acquiring a number of first initial stroke points that matches the number of strokes based on the saliency map; Inputting the original portrait into a segmentation module in the portrait sketch generation model, obtaining a segmentation contour map of the original portrait through the segmentation module, and obtaining a number of second initial stroke points that matches the number of strokes based on the segmentation contour map; Utilizing the second initial stroke position of each second initial stroke point, optimizing the first initial stroke position of each first initial stroke point, obtaining a third initial stroke position matching the number of strokes, and generating the portrait sketch based on the third initial stroke position; including: obtaining the current second initial stroke point corresponding to each current first initial stroke point; when the difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is less than or equal to a preset difference threshold, using the first initial stroke position of the current first initial stroke point as the third initial stroke position; when the difference is greater than the preset difference threshold, using the second initial stroke position of the current second initial stroke point as the third initial stroke position.

2. The method according to claim 1, characterized in that The step of obtaining a segmentation contour map of the original portrait by the arbitrary segmentation module, and obtaining a number of second initial stroke points matching the number of strokes according to the segmentation contour map, includes: Extracting portrait features of the original portrait through the arbitrary segmentation module, and obtaining a segmentation contour map of the original portrait according to the portrait features; Pixel sampling is performed on the segmentation contour image according to the number of strokes, and each sampling point is used as the second initial stroke point; wherein the arc length differences between each second initial stroke point are the same.

3. The method according to claim 1, characterized in that Generating the portrait sketch based on the third initial stroke position includes: constructing stroke parameter information of each stroke for generating the portrait sketch according to each of the third initial stroke positions; Inputting the stroke parameter information into a differentiable rasterization module in the portrait sketch generation model, and obtaining a candidate portrait sketch through the differentiable rasterization module; inputting the candidate portrait sketch and the original portrait into a comparative language image pre-training module in the portrait sketch generation model, obtaining a difference between the candidate portrait sketch and the original portrait through the comparative language image pre-training module, and optimizing the stroke parameter information using the difference to obtain optimized stroke parameter information; In a case where the optimized stroke parameter information meets a preset condition, the optimized stroke parameter information is input into the differentiable rasterization module, and the portrait sketch is obtained through the differentiable rasterization module.

4. The method according to claim 3, characterized in that Each of the stroke parameter information includes: any of the third initial stroke positions, and position information to be optimized that matches the third initial stroke position; optimizing the stroke parameter information using the difference to obtain optimized stroke parameter information includes: Using the difference, updating the position information to be optimized matched by any of the third initial stroke positions, to obtain the updated position information to be optimized of the third initial stroke position; According to the third initial stroke position and the updated position information to be optimized of the third initial stroke position, any optimized stroke parameter information is obtained.

5. The method according to claim 3, characterized in that The obtaining of the difference between the candidate portrait sketch and the original portrait by the comparative language image pre-training module includes: Obtaining, through the comparative language image pre-training module, geometric differences between the candidate portrait sketch and the original portrait, and semantic differences between the candidate portrait sketch and the original portrait; The difference between the candidate portrait sketch and the original portrait is obtained according to the geometric difference and the semantic difference.

6. The method according to claim 5, characterized in that The comparative language image pre-training module includes a plurality of encoding layers; obtaining the geometric difference between the candidate portrait sketch and the original portrait, and the semantic difference between the candidate portrait sketch and the original portrait through the comparative language image pre-training module includes: Obtaining first encoding information corresponding to the candidate portrait sketch and first encoding information corresponding to the original portrait through the first encoding layer of the comparative language image pre-training module, wherein the first encoding layer is the last of the plurality of encoding layers; taking the difference between the first coded information corresponding to the candidate portrait sketch and the first coded information corresponding to the original portrait as the semantic difference; Obtaining second encoding information corresponding to the candidate portrait sketch and second encoding information corresponding to the original portrait through the second encoding layer of the comparative language image pre-training module; the second encoding layer is a preset intermediate encoding layer among the plurality of encoding layers; The difference between the second coded information corresponding to the candidate portrait sketch and the second coded information corresponding to the original portrait is used as the geometric difference.

7. The method according to any one of claims 3 to 6, characterized in that After obtaining the optimized stroke parameter information, the method further includes: In the case that the optimized stroke parameter information does not meet the preset conditions, the optimized stroke parameter information is used as new stroke parameter information, and the step of inputting the stroke parameter information into the differentiable rasterization module in the portrait sketch generation model is returned to be executed.

8. A portrait sketch generating device, characterized in that: include: an original portrait obtaining unit, configured to obtain an original portrait for generating a portrait sketch, and the number of strokes of the portrait sketch; A first stroke point acquisition unit is configured to input the original portrait and the number of strokes into a portrait sketch generation model to obtain a saliency map corresponding to the original portrait, and acquire a number of first initial stroke points that matches the number of strokes based on the saliency map; a second stroke point acquisition unit configured to execute a segmentation arbitrary module in the portrait sketch generation model by inputting the original portrait into the segmentation arbitrary module, obtain a segmentation contour map of the original portrait through the segmentation arbitrary module, and obtain a number of second initial stroke points that matches the number of strokes based on the segmentation contour map; The portrait sketch generating unit is configured to optimize the first initial stroke positions of each of the first initial stroke points using the second initial stroke positions of each of the second initial stroke points to obtain a number of third initial stroke positions that matches the number of strokes, and generate the portrait sketch based on the third initial stroke positions; further configured to obtain a current second initial stroke point corresponding to each current first initial stroke point; and, if a difference between the first initial stroke position of the current first initial stroke point and the second initial stroke position of the current second initial stroke point is less than or equal to a preset difference threshold, use the first initial stroke position of the current first initial stroke point as the third initial stroke position; In a case where the difference is greater than the preset difference threshold, the second initial stroke position of the current second initial stroke point is used as the third initial stroke position.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the portrait sketch generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the portrait sketch generation method according to any one of claims 1 to 7.

11. A computer program product comprising instructions, characterized in that: When the instruction is executed by a processor of an electronic device, the electronic device is enabled to execute the portrait sketch generation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Face art image generation method and device, computer equipment and storage medium

    CN117197283A

  • Image video semantic coding and decoding method and device based on sketch

    CN117392247A