Image processing method and apparatus, and device, computer-readable storage medium and product
By providing an image determination area and a pose selection area in the image processing page, users can quickly adjust the target subject posture in the image, solving the problems of low processing efficiency and inconsistency of posture in the prior art, and achieving a more efficient and personalized image processing effect.
Patent Information
- Application Number
- PCT/CN2024/129539
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-11-02
- Publication Date
- 2025-05-08
AI Technical Summary
The existing posture adjustment methods are less efficient in processing, cumbersome in operation, and the adjusted target subject is inconsistent with the background, which cannot meet the actual needs of users.
An image processing method is provided. By displaying an image determination area and a posture selection area in an image processing page, a user can determine the image to be processed in the image determination area, and select a target posture template in the posture selection area, thereby generating a target image and displaying it in the result display area.
The efficiency of posture transformation is improved, allowing users to quickly realize the transformation of the target subject posture in the image to be processed, and the generated target images are more coordinated to meet the user's personalized needs.
Smart Images

Figure CN2024129539_08052025_PF_FP_ABST
Abstract
Description
Image processing method, device, equipment, computer-readable storage medium and product
[0001] This application claims priority to Chinese Patent Application No. 202311446675.4 filed on November 2, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby cited in their entirety as part of this application. Technical Field
[0002] Embodiments of the present disclosure relate to an image processing method, apparatus, device, computer-readable storage medium, and product. Background Art
[0003] With the gradual advancement of image processing technology, users can now adjust the pose of the subject in an image to better suit their needs. However, current pose adjustment methods are often cumbersome, inefficient, and the resulting pose is inconsistent with the background, failing to meet users' actual needs.
[0004] Summary of the Invention
[0005] The embodiments of the present disclosure provide an image processing method, apparatus, device, computer-readable storage medium, and product, which are used to solve the problem of low processing efficiency of existing posture adjustment methods.
[0006] The present disclosure provides an image processing method, including:
[0007] In response to an image processing operation triggered by a user, displaying an image processing page, wherein the image processing page includes an image determination area and a gesture selection area;
[0008] Acquire the image to be processed determined by the user in the image determination area, and acquire the target posture template selected by the user in the posture selection area;
[0009] generating a target image based on the image to be processed and the target posture template, wherein a posture of a main content in the target image is consistent with a posture in the target posture template;
[0010] The target image is displayed in a preset result display area.
[0011] An embodiment of the present disclosure provides an image processing device, including:
[0012] A display module, configured to display an image processing page in response to an image processing operation triggered by a user, wherein the image processing page includes an image determination area and a posture selection area;
[0013] An acquisition module, configured to acquire the image to be processed determined by the user in the image determination area, and acquire the target posture template selected by the user in the posture selection area;
[0014] a generating module, configured to generate a target image based on the image to be processed and the target posture template, wherein the posture of the main content in the target image is consistent with the posture in the target posture template;
[0015] The display module is used to display the target image in a preset result display area.
[0016] An embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0017] The memory stores computer-executable instructions;
[0018] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the various possible image processing methods described above.
[0019] An embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, various possible image processing methods as described above are implemented.
[0020] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements the above various possible image processing methods when executed by a processor.
[0021] An embodiment of the present disclosure provides a computer program product, wherein the computer program product includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for various possible image processing methods as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] FIG1 is a schematic diagram of a flow chart of an image processing method provided by an embodiment of the present disclosure;
[0024] FIG2 is a schematic diagram of a display interface provided by an embodiment of the present disclosure;
[0025] FIG3 is a schematic diagram of an interface interaction provided by an embodiment of the present disclosure;
[0026] FIG4 is another schematic diagram of an interface interaction provided by an embodiment of the present disclosure;
[0027] FIG5 is a schematic diagram of another display interface provided by an embodiment of the present disclosure;
[0028] FIG6 is a schematic flow chart of an image processing method provided by yet another embodiment of the present disclosure;
[0029] FIG7 is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure; and
[0030] FIG8 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0032] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0033] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0034] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0035] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0036] In order to solve the problem of low processing efficiency of existing posture adjustment methods, the present disclosure provides an image processing method, device, equipment, computer-readable storage medium and product.
[0037] It should be noted that the image processing method, apparatus, device, computer-readable storage medium and product provided by the present disclosure can be applied in any scenario of image posture transformation.
[0038] Current pose transformation methods require users to manually specify an editable area after obtaining the image to be transformed, and then drag key points within the editable area to achieve the pose transformation. However, this operation is often cumbersome, resulting in low image processing efficiency.
[0039] In solving the above-mentioned technical problem, the inventors discovered through research that, to improve the efficiency of posture transformation, an image determination area and a posture selection area can be displayed on the image processing page. This allows the user to determine the image to be processed in the image determination area and select a target posture template based on actual needs in the posture selection area. This allows the posture of the target subject in the image to be processed to be transformed into the posture in the target posture template, thereby obtaining the target image. To enable the user to more intuitively view the target image, the image processing page can also include a result display area, where the target image can be displayed.
[0040] The target subject can be a human object or an animal object. Therefore, based on the posture transformation method, the transformation of the target subject's posture can be achieved, and the transformation of the animal's posture can be achieved. For example, the image to be processed can be a human image, and the posture of the human body area in the image to be processed can be a standing posture with both hands making a heart shape, and the target posture template selected by the user can be a standing posture with one hand making a scissors hand. Based on the target posture template, the posture of the human body area in the image to be processed can be adjusted to a standing posture with one hand making a scissors hand, to obtain the target image. Alternatively, the image to be processed can be an image including a cat, and the posture of the cat in the image to be processed can be a standing posture, and the target posture template can be a lying posture. Based on the target posture template, the posture of the cat in the image to be processed can be adjusted to a lying posture, to obtain the target image.
[0041] Furthermore, to accurately implement pose transformation operations, a diffusion model for pose transformation can be pre-trained. After obtaining the image to be processed and the target pose template, the pose of the target subject in the image to be processed can be adjusted to be consistent with the pose in the target pose template, thereby obtaining a target skeleton image. The image to be processed and the target skeleton image are used as reference conditions and pose conditions, respectively. The reference conditions and pose conditions are used to jointly constrain the diffusion model to generate a transformed image. The reference conditions constrain the diffusion model to generate a transformed image that retains the features of the original image, while the pose conditions enable pose control during the generation of the transformed image.
[0042] In addition, in order to retain the facial features in the image to be processed in the target image, after obtaining the transformed image, the facial area in the image to be processed may be replaced with the transformed image to obtain the target image.
[0043] FIG1 is a flow chart of an image processing method provided by an embodiment of the present disclosure. As shown in FIG1 , the method includes:
[0044] Step 101: In response to an image processing operation triggered by a user, an image processing page is displayed, wherein the image processing page includes an image determination area and a posture selection area.
[0045] The execution subject of this embodiment is an image processing device. The image processing device can be coupled to a terminal device, so that, based on a trigger operation of the user on the terminal device, it can obtain a user-determined image to be processed and a target posture template, and generate and display a target image consistent with the posture in the target posture template based on the image to be processed and the target posture template.
[0046] In this embodiment, the user can trigger an image processing operation to change the pose of the subject in the image being processed. For example, after taking a photo, if the user feels that the pose of the subject in the photo is not aesthetically pleasing, the user can trigger an image processing operation to adjust the pose of the subject in the photo to improve the photo's aesthetics.
[0047] The target subject includes but is not limited to people, animals, objects with adjustable joint positions, etc.
[0048] Optionally, a gesture change control may be pre-set in the image processing application, and the user may trigger the image processing operation by triggering the gesture change control.
[0049] Furthermore, in response to the image processing operation triggered by the user, an image processing page may be displayed, wherein the image processing page includes an image determination area and a posture selection area. The image determination area is used to determine the image to be processed. The posture selection area is used to select a target posture template.
[0050] In addition, a result display area can be pre-set to display the target image generated based on the image to be processed and the target posture template. The result display area can be displayed on the image processing page. Alternatively, after obtaining the image to be processed and the target posture template determined by the user, the user can jump to a preset display page and display the result display area on the preset display page. The present disclosure does not limit the display location of the result display area.
[0051] As an operative method, the result display area can be displayed within the image processing page. The image determination area and the posture selection area can be displayed on the left side of the image processing page, and can be displayed vertically on the left side of the image processing page. The result display area can be displayed on the right side of the image processing page. Users can adjust the display parameters such as the display position and display size of the image determination area, the posture selection area, and the result display area according to actual needs, and this disclosure does not impose any restrictions on this.
[0052] Step 102: Acquire the image to be processed determined by the user in the image determination area, and acquire the target posture template selected by the user in the posture selection area.
[0053] In this embodiment, after the image determination area and the posture selection area are displayed respectively, the user can determine the image to be processed in the image determination area, wherein the image to be processed may include the target subject area.
[0054] Optionally, to implement a posture transformation operation, a plurality of candidate postures may be pre-set for the user to select. Within the posture selection area, the user may select a target posture template based on the plurality of pre-set candidate postures. This allows the posture of the target subject area in the image to be processed to be transformed into the target posture corresponding to the currently selected target posture template.
[0055] Step 103: Generate a target image based on the image to be processed and the target posture template, wherein the posture of the main content in the target image is consistent with the posture in the target posture template.
[0056] In this embodiment, after respectively acquiring the image to be processed and the target posture template, a target image may be generated based on the image to be processed and the target posture template, wherein the posture of the main content in the target image is consistent with the posture in the target posture template.
[0057] Optionally, a preset diffusion model may be used to generate a target image based on the image to be processed and the target posture template.
[0058] As an operative approach, after obtaining the image to be processed and the target posture template, the target image can be generated directly based on the image to be processed and the target posture template. Alternatively, after obtaining the image to be processed and the target posture template, the target image can be generated in response to a user-triggered generation instruction.
[0059] For example, the image to be processed may be an image of a person, and the target subject in the image to be processed may be a human body region. The posture of the person in the image to be processed may be a sitting posture with hands hanging down. The target posture template may be a sitting posture with hands raised. After obtaining the image to be processed and the target posture template, the posture of the person in the image to be processed may be adjusted to a sitting posture with hands raised, thereby obtaining the target image.
[0060] Step 104: Display the target image in a preset result display area.
[0061] In this embodiment, in order to enable the user to view the target image more intuitively, after the target image is generated based on the image to be processed and the target posture template, the target image can be displayed in the result display area.
[0062] Optionally, the result display area can be displayed within the image processing page, allowing the user to simultaneously view the image to be processed, the target posture template, and the target image within the image processing page. Alternatively, the result display area can be displayed within a preset display page. Thus, after generating the target image based on the image to be processed and the target posture template, the user can jump to the preset display page, allowing the user to view the target image within the preset display page.
[0063] The image processing method provided in this embodiment displays an image determination area and a posture selection area on the image processing page. This allows the user to determine the image to be processed in the image determination area and select a target posture template based on actual needs in the posture selection area. This allows the posture of the target subject in the image to be processed to be transformed into the posture in the target posture template, thereby obtaining the target image. To enable the user to more intuitively view the target image, the target image can be displayed in the result display area. This allows the posture of the target subject in the image to be processed to be transformed quickly, improving image processing efficiency.
[0064] FIG2 is a schematic diagram of a display interface provided by an embodiment of the present disclosure. As shown in FIG2 , the image processing page 21 includes an image determination area 22, a posture selection area 23, and a result display area 24. The image determination area 22 is used to determine the image to be processed. The image to be processed may be a human image, which includes a human body area. The posture selection area 23 is used to select a target posture template. When the image to be processed is a human image, the user can select a target posture template that matches the human posture in the posture selection area. The result display area 24 is used to display a target image generated based on the image to be processed and the target posture template.
[0065] The image processing method provided in this embodiment displays an image determination area and a posture selection area on the image processing page. This allows the user to determine the image to be processed in the image determination area and select a target posture template based on actual needs in the posture selection area. The posture of the target subject in the image to be processed can then be transformed into the posture in the target posture template to obtain the target image. To enable the user to more intuitively view the target image, the target image can be displayed in a preset result display area. This allows the target subject in the image to be processed to be quickly transformed in posture, improving image processing efficiency.
[0066] Optionally, based on any of the above embodiments, the image processing page includes a preset generation control.
[0067] Step 103 includes:
[0068] In response to a user triggering operation on the generation control, a target image is generated based on the image to be processed and the target posture template.
[0069] In this embodiment, after the image to be processed and the target posture template are acquired respectively, the target image generation operation may be performed in response to a generation instruction triggered by the user.
[0070] Optionally, a generation control may be pre-set in the image processing page. After the image to be processed and the target posture template are acquired respectively, in response to a user triggering operation on the generation control, a target image may be generated based on the image to be processed and the target posture template.
[0071] The image processing method provided in this embodiment pre-sets a generation control in the image processing page, thereby enabling rapid generation of a target image based on a triggering operation of the generation control, thereby improving the efficiency of image posture transformation.
[0072] Furthermore, based on any of the above embodiments, the image determination area includes an image upload control.
[0073] Step 102 includes:
[0074] In response to the user triggering the image upload control, an image selection page is displayed, wherein the image selection page includes a plurality of images to be selected.
[0075] In response to the user's selection operation on the image to be selected in the image selection page, the target candidate image selected by the user is determined as the image to be processed.
[0076] Alternatively, the image determination area includes an image acquisition control.
[0077] Step 102 includes:
[0078] In response to the user's triggering operation on the image acquisition control, the image to be processed is acquired by a preset image acquisition device.
[0079] In this embodiment, an image upload control and / or an image acquisition control may be pre-set in the image determination area. The user may select a corresponding image determination method according to actual needs to perform a determination operation on the image to be processed.
[0080] Optionally, the image determination area includes an image upload control. In response to the user triggering the image upload control, an image selection page may be displayed, displaying multiple images to be selected from a preset storage path. The user may select from the multiple images to be selected based on actual needs. In response to the user selecting an image to be selected on the image selection page, the target candidate image selected by the user is determined as the image to be processed.
[0081] In one practicable embodiment, the image determination area includes an image acquisition control. In response to a user triggering the image acquisition control, the image to be processed is acquired by a preset image acquisition device. For example, in response to the user triggering the image acquisition control, a preset camera in the terminal device may be invoked to capture an image. The captured image is determined as the image to be processed.
[0082] FIG3 is a schematic diagram of the interface interaction provided by an embodiment of the present disclosure. As shown in FIG3 , the image processing page 31 includes an image determination area 32, a posture selection area 33, and a result display area 34. The image determination area 32 includes an image upload control 35. Optionally, a prompt message 36 may also be displayed in the image determination area 32, and the prompt message 36 is used to prompt the user on the method of determining the image to be processed. For example, the prompt message 36 may be: Click to upload the image to be processed. In response to the user's triggering operation on the image upload control 35, an image selection page 37 may be displayed, wherein the image selection page 37 includes multiple images to be selected 38. In response to the user's selection operation on the image to be selected 38, the target image to be selected 38 currently selected by the user may be determined as the image to be processed 39, and the image to be processed 39 may be displayed in the image determination area 32.
[0083] The image processing method provided in this embodiment provides an image upload control and / or an image acquisition control within the image determination area, allowing users to select a corresponding image determination method to determine the image to be processed based on their actual needs. This makes the determination method for the image to be processed more tailored to the user's personalized needs. Furthermore, it can improve the efficiency of determining the image to be processed.
[0084] Further, based on any of the above embodiments, step 102 includes:
[0085] A plurality of preset postures to be selected are displayed in the posture selection area.
[0086] In response to the user's selection operation on the plurality of candidate postures, a target candidate posture selected by the user is determined as the target posture template.
[0087] In this embodiment, to allow the user to more intuitively view the postures to be selected, a plurality of preset postures to be selected may be displayed in the posture selection area. The user may select from the plurality of postures to be selected based on actual needs. In response to the user's selection operation for the plurality of postures to be selected, the target posture selected by the user may be determined as the target posture template.
[0088] As an implementable method, when the image to be processed is a person image, multiple postures associated with the person can be selected in the posture selection area. For example, the postures to be selected include but are not limited to standing, waving, making a heart shape, making a thumbs-up gesture, and the like.
[0089] The image processing method provided in this embodiment displays a plurality of preset postures to be selected in the posture selection area, so that the user can view the postures to be selected more intuitively, and can then quickly select the target posture template according to actual needs, thereby improving the efficiency of image processing.
[0090] Optionally, based on any of the above embodiments, the posture selection area includes a category selection sub-area and a posture display sub-area. Step 102 includes:
[0091] In response to a target category selected by the user in the category selection subarea, a plurality of gestures to be selected associated with the target category are displayed in the gesture display subarea.
[0092] In response to the user's selection operation on the plurality of candidate postures, a target candidate posture selected by the user is determined as the target posture template.
[0093] In this embodiment, in order to facilitate the user to more accurately select the target posture template, multiple postures to be selected can be classified in advance to obtain multiple posture categories, such as a sitting posture category and a standing posture category.
[0094] As an practicable approach, when the image to be processed is a person image, the posture categories may include a sitting posture category and a standing posture category corresponding to the person, etc. When the image to be processed is an animal image, the posture categories may include a standing posture category and a lying posture category corresponding to the animal, etc.
[0095] Optionally, the gesture selection area includes a category selection subarea and a gesture display subarea. A user can select a target category in the category selection subarea. In response to a user triggering the category selection subarea, a category list can be displayed, wherein the category list includes multiple gesture categories. In response to the user selecting a gesture category, the target category selected by the user in the category selection subarea can be determined.
[0096] After determining the target category, multiple candidate poses associated with the target category may be displayed in the pose display sub-area. For example, if the user currently selects a target category of standing poses, multiple candidate poses of standing poses may be displayed in the pose display sub-area, including, for example, a standing hand-raising pose, a standing scissors hand pose, a standing heart gesture, and other different candidate poses.
[0097] After multiple standing postures are displayed in the posture display sub-area, the user can select the posture to be selected according to actual needs. In response to the selection operation triggered by the user, the posture currently selected by the user can be determined as the target posture template.
[0098] FIG4 is another interface interaction diagram provided by an embodiment of the present disclosure. As shown in FIG4 , the image processing page 41 includes an image determination area 42, a posture selection area 43, and a result display area 44. Among them, the posture selection area 43 includes a category selection sub-area 45 and a posture display sub-area 46. The user can display a category list 47 by triggering the category selection sub-area 45, wherein the category list 47 includes multiple preset categories 48. The user can select a target category based on the category list 47. Multiple postures to be selected 49 associated with the target category are displayed in the posture display sub-area 46. In this way, the user can perform a selection operation on the multiple postures to be selected 49 to determine the target posture template.
[0099] The image processing method provided in this embodiment provides a category selection sub-area within the posture selection area, allowing the user to select a target category within the category selection sub-area. Multiple postures associated with the target category are then displayed in the posture display sub-area for the user to select. This allows the displayed content within the image processing page to better meet the user's personalized needs, improving the user experience. Furthermore, by dividing the selected postures into different posture categories, a larger number of postures can be displayed within the image processing page.
[0100] Further, based on any of the above embodiments, the displaying of multiple candidate postures associated with the target category in the posture display sub-area includes:
[0101] The display area in the posture display sub-area displays preset postures to be selected according to a preset first display size, and the selection area in the posture display sub-area displays all postures to be selected associated with the target category according to a preset second display size.
[0102] The first display size is larger than the second display size.
[0103] In this embodiment, the gesture display sub-area may include a display area and a selection area. The display area may display preset gestures to be selected according to a preset first display size, and the selection area may display all gestures to be selected associated with the target category according to a preset second display size.
[0104] The first display size is larger than the second display size. Thus, the user can more intuitively view the currently selected postures in the display area. In addition, a selection operation can also be performed based on all the postures to be selected in the selection area.
[0105] FIG5 is a schematic diagram of another display interface provided by an embodiment of the present disclosure. As shown in FIG5 , an image processing page 51 includes an image determination area 52, a posture selection area 53, and a result display area 54. The posture selection area 53 includes a category selection sub-area 55 and a posture display sub-area 56. A display area 57 within the posture display sub-area 56 displays a preset candidate posture 58 in a preset first display size, and a selection area 59 within the posture display sub-area 56 displays all candidate postures 510 associated with the target category in a preset second display size.
[0106] Furthermore, based on any of the above embodiments, the method further includes:
[0107] In response to the user's selection operation on all the to-be-selected postures associated with the target category within the selection area, the posture to be displayed selected by the user is determined, and the posture to be displayed is switched to a selected state.
[0108] The gesture to be displayed is switched and displayed in the display area according to the first display size.
[0109] In this embodiment, after all available postures are displayed in the selection area, the user can select a posture to be selected. "Want to Love You" senses the user's selection of any of the available postures and determines the posture to be displayed, switching the posture to be displayed to a selected state. In this selected state, the posture to be displayed can be highlighted or enlarged, distinguishing it from other available postures.
[0110] Furthermore, after determining the gesture to be displayed currently selected by the user, the gesture to be displayed may be switched and displayed in the display area according to the first display size, so that the user can view the gesture to be displayed more intuitively in the display area.
[0111] The image processing method provided in this embodiment displays a preset candidate pose in a first preset display size in the display area of the pose display sub-area, and displays all candidate poses associated with a target category in a second preset display size in the selection area of the pose display sub-area. This allows the user to more intuitively view the currently selected candidate pose. Furthermore, the user can switch between multiple candidate poses, making the display content within the image processing page more tailored to the user's personalized needs and enhancing the user experience.
[0112] FIG6 is a flowchart of an image processing method provided by another embodiment of the present disclosure. Based on any of the above embodiments, as shown in FIG6 , step 103 includes:
[0113] Step 601: Identify a skeleton graph to be processed corresponding to the image to be processed, wherein the skeleton graph to be processed includes a plurality of key points corresponding to a target subject.
[0114] Step 602: Adjust the display position of each key point in the skeleton graph to be processed according to the target posture template to obtain a target skeleton graph that matches the target posture template.
[0115] Step 603: Determine the image to be processed as the first constraint condition, and determine the target skeleton graph as the second constraint condition.
[0116] Step 604: Generate the transformed image by jointly constraining the preset diffusion model through the first constraint condition and the second constraint condition.
[0117] Step 605: Generate the target image based on the transformed image and the image to be processed.
[0118] In this embodiment, to implement a posture transformation operation, after acquiring an image to be processed, a skeleton graph corresponding to the image to be processed can be first identified. The skeleton graph includes multiple key points corresponding to the target subject and may also include position information corresponding to each key point. Any algorithm capable of identifying key points can be used to identify the skeleton graph corresponding to the image to be processed, and this disclosure does not impose any restrictions on this.
[0119] Furthermore, in order to adjust the posture of the target subject in the image to be processed to the posture corresponding to the target posture template, the display position of each key point in the skeleton image to be processed can be adjusted according to the target posture template to obtain a target skeleton image that matches the target posture template.
[0120] Before performing a pose transformation operation, a diffusion model can be trained. The image to be processed is defined as a first constraint, and the target skeleton image is defined as a second constraint. The first and second constraints are combined to constrain the preset diffusion model to generate a transformed image.
[0121] The diffusion model generates a transformed image, which not only transforms the pose of the target subject in the target image but also adjusts the facial region in the target image, resulting in differences between the face in the transformed image and the target image. Therefore, to preserve the facial features in the target image, a target image can be generated based on the transformed image and the target image.
[0122] Optionally, the control weight corresponding to the second constraint is greater than the control weight corresponding to the first constraint. When the control weight corresponding to the second constraint is greater than the control weight corresponding to the first constraint, the corresponding posture can be successfully adjusted while preserving the original image features as much as possible, thereby improving the image quality of the transformed image.
[0123] As an operative approach, when the image to be processed includes an animal region, a corresponding skeleton image to be processed can be identified. Key points in the skeleton image to be processed can be adjusted based on a target pose template to obtain a target skeleton image. This allows the image to be processed and the target skeleton image to serve as a first constraint and a second constraint, respectively. A constrained diffusion model is then used to generate a transformed image based on the first and second constraints. The pose of the animal in the transformed image matches the pose in the target pose template. Furthermore, the facial region of the animal in the image to be processed can be replaced with the facial region of the animal in the transformed image to obtain the target image.
[0124] The image processing method provided in this embodiment uses the image to be processed and the target skeleton graph as the first and second constraints, respectively, and then uses the first and second constraints to jointly constrain a diffusion model to generate a transformed image. The first constraint constrains the diffusion model to generate a transformed image that retains the features of the original image, while the second constraint allows for posture control during the generation of the transformed image. This allows the transformed image output by the diffusion model to generate a more realistic target image that retains the features of the original image, thereby improving the quality of the target image.
[0125] Further, based on any of the above embodiments, step 602 includes:
[0126] At least one first key point is determined in the skeleton image to be processed, and at least one second key point that matches the at least one first key point is determined in the target posture template.
[0127] Adjusting the first key point to a display position that matches the second key point;
[0128] An adjustment operation is performed on at least one other key point associated with the first key point based on the adjustment parameter corresponding to the first key point to obtain the target skeleton graph.
[0129] In this embodiment, the skeleton image to be processed includes information about multiple key points of the target subject in the image to be processed. The target posture template includes information about multiple key points corresponding to the target posture. For example, taking the target subject as a human body, key points at each corresponding position can be pre-numbered. For example, the key points of the cervical spine can be numbered 1, the key points of the shoulder can be numbered 2, the key points of the elbow joint can be numbered 3, and the key points of the hand can be numbered 4.
[0130] It's important to note that when a keypoint moves, it can also cause other keypoints to move. For example, when shoulder joint keypoint 2 moves, it can also cause elbow joint keypoint 3 and hand keypoint 4 to move. Therefore, when performing posture transfer, the first step is to identify the keypoint at the corresponding position in the target skeleton image and the target posture template. Then, the posture adjustment operation is performed based on the relative position of the keypoint and other keypoints with pre-defined relationships.
[0131] For example, if the shoulder joint key point in the image to be processed is in the first position and the shoulder joint key point in the target posture template is in the second position, the shoulder joint key point in the image to be processed can be adjusted to the second position, and the elbow joint key point and the hand key point can be adjusted accordingly based on the adjustment parameters generated during the adjustment process.
[0132] To achieve a posture transfer operation, at least one first key point can be first determined in the skeleton image to be processed, and at least one second key point that matches the at least one first key point can be determined in the target posture template. The first key point is adjusted to a display position that matches the second key point. Adjustment parameters during the adjustment of the first key point are determined. The adjustment parameters include, but are not limited to, adjustment parameters such as distance, direction, and angle. To make the posture of the target subject in the image more coordinated, at least one other key point associated with the first key point can be adjusted based on the adjustment parameters corresponding to the first key point to obtain a target skeleton image.
[0133] It should be noted that since many postures require the coordination of multiple joints to be achieved, during the posture migration process, at least one key point can be selected, and the position adjustment operation can be performed on at least one key point and at least one other key point associated with it in turn to obtain the target skeleton diagram.
[0134] The image processing method provided in this embodiment can accurately realize the migration of the target posture based on the first key point in the skeleton image to be processed and the second key point in the target posture template, thereby obtaining a target skeleton image that matches the target posture template.
[0135] Further, based on any of the above embodiments, step 605 includes:
[0136] The facial area in the image to be processed is identified and extracted using a preset image processing algorithm.
[0137] The target image is obtained by replacing the facial region in the image to be processed with the facial region in the transformed image.
[0138] In this embodiment, the diffusion model generates a transformed image, which not only transforms the pose of the target subject in the target image but also adjusts the facial region in the target image, resulting in differences between the face in the transformed image and the target image. Therefore, to preserve the facial features of the target image in the target image, the facial region in the target image can be replaced with the facial region in the transformed image.
[0139] Optionally, to replace the facial region, the facial region in the image to be processed may be first identified and extracted using a preset image processing algorithm. The facial region extraction operation may be performed using any image processing algorithm capable of identifying and cropping the facial region, and this disclosure does not impose any limitation thereto.
[0140] After obtaining the facial region in the image to be processed, the facial region in the transformed image can be replaced with the facial region in the image to be processed to obtain the target image. The display position of the facial region in the transformed image can be determined, and the facial region replacement operation is performed based on the display position.
[0141] As an implementable approach, facial detection can be performed on the target image and the transformed image using an open-source image processing library such as OpenCV. Facial key points are obtained from both images, and the facial region is extracted based on these key points. The facial region is then triangulated, and a triangular mesh is drawn within the facial region. Each triangle in the facial region corresponding to the target image is mapped to the corresponding triangle in the transformed image using an affine transformation, resulting in the transformed facial region. The transformed facial region is then applied to the corresponding position in the transformed image to produce the target image.
[0142] The image processing method provided in this embodiment replaces the facial area in the transformed image with the facial area in the image to be processed, thereby retaining the facial features in the image to be processed in the target image and improving the image quality of the target image.
[0143] FIG7 is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure. As shown in FIG7 , the device includes: a display module 71, an acquisition module 72, a generation module 73, and a presentation module 74. The display module 71 is used to display an image processing page in response to an image processing operation triggered by a user, wherein the image processing page includes an image determination area and a posture selection area. The acquisition module 72 is used to acquire the image to be processed determined by the user in the image determination area, and acquire the target posture template selected by the user in the posture selection area. The generation module 73 is used to generate a target image based on the image to be processed and the target posture template, wherein the posture of the main content in the target image is consistent with the posture in the target posture template. The presentation module 74 is used to display the target image in a preset result presentation area.
[0144] Furthermore, based on any of the above embodiments, the image determination area includes an image upload control. The acquisition module is configured to: in response to the user triggering the image upload control, display an image selection page, wherein the image selection page includes multiple images to be selected. In response to the user selecting an image to be selected on the image selection page, the target candidate image selected by the user is determined as the image to be processed. Alternatively, the image determination area includes an image acquisition control. The acquisition module is configured to: in response to the user triggering the image acquisition control, acquire the image to be processed through a preset image acquisition device.
[0145] Furthermore, based on any of the above embodiments, the acquisition module is configured to: display a plurality of preset candidate postures in the posture selection area, and in response to the user selecting the plurality of candidate postures, determine the target candidate posture selected by the user as the target posture template.
[0146] Furthermore, based on any of the above embodiments, the posture selection area includes a category selection subarea and a posture display subarea. The acquisition module is configured to: in response to a target category selected by the user in the category selection subarea, display a plurality of candidate postures associated with the target category in the posture display subarea. In response to the user selecting the plurality of candidate postures, determine the target candidate posture selected by the user as the target posture template.
[0147] Furthermore, based on any of the above embodiments, the acquisition module is configured to: display a preset candidate posture in a display area within the posture display sub-area according to a preset first display size, and display all candidate postures associated with the target category in a selection area within the posture display sub-area according to a preset second display size. The first display size is larger than the second display size.
[0148] Furthermore, based on any of the above embodiments, the device further includes: a determination module configured to, in response to the user selecting all the candidate postures associated with the target category within the selection area, determine the posture to be displayed selected by the user and switch the posture to be displayed to a selected state; and a display module configured to switch the display of the posture to be displayed within the display area according to the first display size.
[0149] Furthermore, based on any of the above embodiments, the image processing page includes a preset generation control, wherein the generation module is configured to generate a target image based on the image to be processed and the target posture template in response to a user triggering operation on the generation control.
[0150] Furthermore, based on any of the above embodiments, the generation module is configured to: identify a skeleton graph to be processed corresponding to the image to be processed, wherein the skeleton graph to be processed includes multiple key points corresponding to the target subject; adjust the display position of each key point in the skeleton graph to be processed according to the target posture template to obtain a target skeleton graph that matches the target posture template; determine the image to be processed as a first constraint condition, and determine the target skeleton graph as a second constraint condition; generate the transformed image by jointly constraining a preset diffusion model with the first and second constraints; and generate the target image based on the transformed image and the image to be processed.
[0151] Furthermore, based on any of the above embodiments, the generation module is configured to: determine at least one first key point in the skeleton image to be processed, and determine at least one second key point in the target posture template that matches the at least one first key point; adjust the first key point to a display position that matches the second key point; and adjust at least one other key point associated with the first key point based on an adjustment parameter corresponding to the first key point, thereby obtaining the target skeleton image.
[0152] Furthermore, based on any of the above embodiments, the generating module is configured to: identify and extract a facial region in the image to be processed using a preset image processing algorithm, and obtain the target image by replacing the facial region in the image to be processed with the facial region in the transformed image.
[0153] Furthermore, based on any of the above embodiments, the control weight corresponding to the second constraint condition is greater than the control weight corresponding to the first constraint condition.
[0154] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0155] In order to implement the above embodiments, the embodiments of the present disclosure further provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the image processing method described in any of the above embodiments is implemented.
[0156] In order to implement the above embodiments, the embodiments of the present disclosure further provide a computer program product, including a computer program, which implements the image processing method as described in any of the above embodiments when executed by a processor.
[0157] In order to implement the above embodiment, the present disclosure further provides an electronic device, including: a processor and a memory;
[0158] The memory stores computer-executable instructions;
[0159] The processor executes the computer-executable instructions stored in the memory, so that the processor performs the image processing method as described in any of the above embodiments.
[0160] FIG8 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. The electronic device 800 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG8 is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0161] As shown in Figure 8, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0162] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 shows an electronic device 800 having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0163] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0164] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0165] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0166] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0167] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0169] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0170] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0171] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0172] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0173] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0174] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An image processing method, comprising: In response to an image processing operation triggered by a user, displaying an image processing page, wherein the image processing page includes an image determination area and a posture selection area; Acquire the image to be processed determined by the user in the image determination area, and acquire the target posture template selected by the user in the posture selection area; generating a target image based on the image to be processed and the target posture template, wherein the posture of the main content in the target image is consistent with the posture in the target posture template; The target image is displayed in a preset result display area.
2. The method according to claim 1, wherein: The image determination area includes an image upload control, and obtaining the image to be processed determined by the user in the image determination area includes: In response to the user triggering the image upload control, displaying an image selection page, the image selection page including a plurality of images to be selected, and In response to a selection operation of the user on the image selection page for the plurality of images to be selected, determining the target candidate image selected by the user as the image to be processed; Alternatively, the image determination area includes an image acquisition control, and acquiring the image to be processed determined by the user in the image determination area includes: In response to the user's triggering operation on the image acquisition control, the image to be processed is acquired through a preset image acquisition device.
3. The method according to claim 1 or 2, wherein: The acquiring of the target posture template selected by the user in the posture selection area comprises: Displaying a plurality of preset postures to be selected in the posture selection area; In response to the user's selection operation on the plurality of candidate postures, a target candidate posture selected by the user is determined as the target posture template.
4. The method according to claim 1 or 2, wherein: The posture selection area includes a category selection sub-area and a posture display sub-area; The acquiring of the target posture template selected by the user in the posture selection area comprises: In response to a target category selected by the user in the category selection subarea, displaying a plurality of to-be-selected gestures associated with the target category in the gesture display subarea; In response to the user's selection operation on the plurality of candidate postures, a target candidate posture selected by the user is determined as the target posture template.
5. The method according to claim 4, wherein: The step of displaying a plurality of postures to be selected and associated with the target category in the posture display sub-area includes: The display area in the posture display sub-area displays a preset posture to be selected according to a preset first display size, and the selection area in the posture display sub-area displays the multiple postures to be selected associated with the target category according to a preset second display size; The first display size is larger than the second display size.
6. The method according to claim 5, further comprising: In response to the user's selection operation on the plurality of to-be-selected postures associated with the target category in the selection area, determining the to-be-displayed posture selected by the user, and switching the to-be-displayed posture to a selected state; The gesture to be displayed is displayed in the display area in a switching manner according to the first display size.
7. The method according to any one of claims 1 to 6, wherein: The image processing page includes a preset generation control; The generating a target image based on the image to be processed and the target posture template comprises: In response to the user's triggering operation on the generating control, the target image is generated based on the image to be processed and the target posture template.
8. The method according to any one of claims 1 to 7, wherein: The generating a target image based on the image to be processed and the target posture template comprises: Identifying a skeleton graph to be processed corresponding to the image to be processed, wherein the skeleton graph to be processed includes a plurality of key points corresponding to a target subject in the image to be processed; Adjusting the display position of each key point in the skeleton image to be processed according to the target posture template to obtain a target skeleton image that matches the target posture template; Determine the image to be processed as a first constraint condition, and determine the target skeleton graph as a second constraint condition; Generate a transformation image by jointly constraining a preset diffusion model through the first constraint condition and the second constraint condition; The target image is generated based on the transformed image and the image to be processed.
9. The method according to claim 8, wherein: The step of adjusting the display position of each key point in the skeleton image to be processed according to the target posture template to obtain a target skeleton image matching the target posture template includes: Determining at least one first key point in the skeleton image to be processed, and determining at least one second key point matching the at least one first key point in the target posture template; Adjusting the first key point to a display position matching the second key point; An adjustment operation is performed on at least one other key point associated with the first key point based on the adjustment parameter corresponding to the first key point to obtain the target skeleton graph.
10. The method according to claim 8 or 9, wherein: The control weight corresponding to the second constraint condition is greater than the control weight corresponding to the first constraint condition.
11. An image processing device, wherein: include: A display module, configured to display an image processing page in response to an image processing operation triggered by a user, wherein the image processing page includes an image determination area and a posture selection area; An acquisition module is configured to acquire the image to be processed determined by the user in the image determination area, and acquire the target posture template selected by the user in the posture selection area; a generating module configured to generate a target image based on the image to be processed and the target posture template, wherein the posture of the main content in the target image is consistent with the posture in the target posture template; and The display module is configured to display the target image in a preset result display area.
12. An electronic device comprising: Processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the image processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the image processing method according to any one of claims 1 to 10 is implemented.
14. A computer program product, wherein: The computer program product comprises a computer program carried on a non-transitory computer-readable medium, wherein the computer program comprises a program code for executing the image processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
An image editing method and device
CN109472795A
Image processing method and device, equipment and storage medium
CN111246113A
Imaging method and device and electronic equipment
CN111800574A
Image processing method, image processing device, electronic equipment and storage medium
CN114913104A
Image processing method, electronic equipment and readable storage medium
CN115423752A