Image processing method and device, equipment and storage medium

By obtaining the local image of the target object and determining its bone information, the problem of inaccurate image expansion in the prior art is solved, and accurate image expansion from local to whole body is achieved, improving image quality and user experience.

CN120107376APending Publication Date: 2025-06-06BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311666394.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing image processing technologies are difficult to effectively expand from local images to global images, especially from local to whole body expansion. There are often problems with errors in the number or position of limbs and unreasonable limb proportions, resulting in poor quality of extended images.

Method used

By obtaining a local image of the target object, the bone information of the area is determined, and based on the bone information, it expands to a larger area, and a full-body image is generated. The specific steps include: acquiring the first image, determining the bone information of the first region, determining the bone information of the second region based on the information, and generating the second image using the second bone information and the first image.

Benefits of technology

Improve the accuracy of image expansion, ensuring the rationality of the number, position and proportion of extended image limbs, thereby improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107376A_ABST
    Figure CN120107376A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring a first image associated with a target object, wherein the first image presents a first area of the target object; determining first skeleton information corresponding to the first area based on the first image; based on the first skeleton information, second skeleton information corresponding to a second area of the target object is determined, and the second area has a larger range than the first area; and generating a second image associated with the target object based on the second skeleton information and the first image, the second image presenting the second region of the target object. In this way, the problem that the number of limbs and the size proportion of the expanded human body are improper can be avoided, and the quality of the expanded image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to image processing methods, devices, apparatuses, and computer-readable storage media. Background Art

[0002] With the development of image technology, users' demand for image processing is increasing. In some application scenarios, it is expected that the acquired or captured images can be better presented through image processing technology. For example, users expect to achieve the expansion from local images to global images, especially from local body parts to whole body images through image processing technology. This is a challenge for current image processing technology. Summary of the invention

[0003] In a first aspect of the present disclosure, an image processing method is provided. The method includes acquiring a first image associated with a target object, the first image presenting a first region of the target object; determining first bone information corresponding to the first region based on the first image; determining second bone information corresponding to a second region of the target object based on the first bone information, the second region having a larger range than the first region; and generating a second image associated with the target object based on the second bone information and the first image, the second image presenting the second region of the target object.

[0004] In a second aspect of the present disclosure, an image processing device is provided. The device includes an image acquisition module configured to acquire a first image associated with a target object, wherein the first image presents a first region of the target object; a first skeleton information determination module configured to determine first skeleton information corresponding to the first region based on the first image; a second skeleton information determination module configured to determine second skeleton information corresponding to a second region of the target object based on the first skeleton information, wherein the second region has a larger range than the first region; and an image generation module configured to generate a second image associated with the target object based on the second skeleton information and the first image, wherein the second image presents the second region of the target object.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the program is executed by a processor, the method of the first aspect is implemented.

[0007] It should be understood that the contents described in the summary of the present invention are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0010] Figure 2A A block diagram showing a key point extraction process according to some embodiments of the present disclosure is shown;

[0011] Figure 2B A schematic diagram showing key point extraction according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A block diagram of a model training process for key point extraction according to some embodiments of the present disclosure is shown;

[0013] Figure 4A A block diagram of an image expansion process according to some embodiments of the present disclosure is shown;

[0014] Figure 4B A schematic diagram showing an image expansion process according to some embodiments of the present disclosure is shown;

[0015] Figure 5 A flowchart showing a process of image processing according to some embodiments of the present disclosure;

[0016] Figure 6 A block diagram showing an image processing apparatus according to some embodiments of the present disclosure; and

[0017] Figure 7 A block diagram of a device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0020] It is understandable that the image data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0021] It is understandable that before using the technical solutions disclosed in the various embodiments of the present disclosure, the types, scopes of use, usage scenarios, etc. of the data involved in the present disclosure should be informed to relevant users and their authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. The relevant users may include any type of right holders, such as individuals, enterprises, and groups.

[0022] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multi-layer processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0023] As mentioned above, users expect to expand a larger image from a partial image. In particular, they expect to expand a larger or even full-body portrait from a partial portrait. However, based on current image processing solutions, the expanded image often has defects such as the wrong number or position of limbs and unreasonable limb proportions, resulting in poor quality of the expanded portrait and low user satisfaction.

[0024] The technical solution of the present disclosure provides an image processing solution. According to the image processing solution of the embodiment of the present disclosure, the first skeleton information corresponding to the first area can be determined based on the first image acquired and presenting the first part of the target object, and the second skeleton information corresponding to the second area of ​​the target object can be determined based on the first skeleton information. Thus, a second image presenting the second area of ​​the target object is generated based on the second skeleton information and the first image. In this way, a second image corresponding to the second area expanded from the first image corresponding to the first area can be accurately generated based on the generated second skeleton information, thereby improving the accuracy of the expanded image and improving the user experience.

[0025] Example Environment

[0026] See first Figure 1 , which schematically illustrates a diagram of an example environment 100 in which example implementations according to the present disclosure may be implemented.

[0027] In the example environment 100, the electronic device 110 may receive an original image 101. The original image 101 may be, for example, an image input by a user showing a portion of a target object 103. It is understood that the target object 103 may generally include a human body, but may also include any animal body with bones.

[0028] In some other embodiments, the original image 101 may also be an image captured by the electronic device 110 through a capture device. In some embodiments, the capture device may be configured to be connected to the electronic device 110. In some other embodiments, the capture device may also be integrated inside the electronic device 110.

[0029] After acquiring the original image 101 that presents a portion of the target object 103 , the electronic device 110 may generate an expanded image 102 that presents a larger portion or even the entire region of the target object 103 by using image processing technology.

[0030] The electronic device 110 can generate the expanded image 102 based on the original image 101 through the server 120. For example, it can be generated by an image processing model deployed on the server. The original image 101 can be the input of the image processing model, and the expanded image 102 is the output of the image processing model. In some embodiments, the server 120 can be arranged independently of the electronic device 110, for example, it can be deployed as a remote server. In other embodiments, the server 120 can be integrated in the electronic device 110 or implemented in the electronic device 110.

[0031] The electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication systems (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the user (such as "wearable" circuits, etc.). The server 120 can be various types of electronic systems / servers that can provide computing power, including but not limited to mainframes, edge computing nodes, servers in cloud environments, and the like.

[0032] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and does not imply any limitation on the scope of the present disclosure.

[0033] Skeleton information extraction process

[0034] Figure 2A A block diagram of a key point extraction process according to some embodiments of the present disclosure is shown. Figure 2A The process shown can be implemented in the electronic device 110, or in the server 120. For ease of discussion, reference will be made to Figure 1 The process shown in FIG. 2 is described with reference to the environment 100 of FIG.

[0035] The electronic device 110 obtains an original image 101 (also referred to as a "first image" in this disclosure). The original image 101 presents a first area of ​​a target object 103. Figure 1 In the example shown in FIG. 1 , the first region may include at least a portion of the head and upper body of the target object 103. Based on the original image 101, the electronic device may determine first skeleton information of the first region. The first skeleton information may be determined, for example, by determining a first group of key points of the first region of the target object 103 presented in the original image 101.

[0036] like Figure 2A As shown, after the electronic device 110 acquires the original image 101, the original image 101 may be input into the first model 201. In the first model 201, the feature extraction module 210 extracts the features of the original image 101. The feature extraction module 210 may be regarded as a backbone network. For example, the feature extraction module 210 may include a convolutional neural network (CNN) or a residual neural network (ResNet).

[0037] The first model 201 further includes a classification module 220 . The extracted features of the original image 101 are provided from the feature extraction module 210 to the classification module 220 , so as to obtain a classification feature representation of the original image 101 through the classification module 220 .

[0038] The first model 201 also includes a codebook 230. The codebook 230 may include a set of candidate feature representations. The classification feature representation of the original image 101 obtained by the classification module 220 will be compared with a set of candidate feature representations in the codebook 230 to determine a target feature representation that matches the classification feature representation obtained by the classification module 220 from the set of candidate feature representations. For example, the candidate feature representation in the codebook 230 that is closest to the classification feature representation obtained by the classification module 220 is determined as the target feature representation.

[0039] The determined target feature representation is decoded by the decoder 240 in the first model 201, and the output of the decoder 240, that is, the output of the first model 201, is a first set of key points 250 of the first region of the target object 103 presented in the original image 101. Based on the first set of key points 250, the first skeleton information corresponding to the first region of the target object 103 can be determined.

[0040] like Figure 2B As shown, after the original image 101 is input into the first model 201 , a first group of key points 250 may be output. The first group of key points 250 may correspond to the skeleton information of the first region of the target object 103 .

[0041] As described above, in the first model 201 , the codebook 230 includes a set of candidate feature representations. The candidate feature representations included in the codebook 230 can be determined by learning the relationship information between the key points of the skeleton of the target object 103 . Figure 3 A block diagram illustrating a model training process for keypoint extraction according to some embodiments of the present disclosure is shown.

[0042] Specifically, Figure 3 As shown, skeleton key points can be extracted from multiple training images, and the extracted training key points 301 are input into the encoder 310 for key point information structure encoding. It should be understood that the training key points 301 can include multiple groups of key points extracted from multiple training images. The training key points 301 can be two-dimensional coordinates of two-dimensional key points obtained from multiple training images.

[0043] The training key point encoded by the encoder 310 is input into the codebook 230 to find the candidate feature representation closest to the encoded training key point in the codebook 230 and update the corresponding candidate feature representation in the codebook 230. As a result, after the feature representation output from the encoder 230 is decoded by the decoder 320, the obtained reference key point 302 approaches the input training key point 301.

[0044] In addition, the feature extraction module 210 and the classification module 220 in the first model 201 can also be obtained by learning the training data. For example, the training data of the feature extraction module 210 and the classification module 220 can be a human body image or picture, such as an image containing the whole body, half of the body, and a part of the human body area. Based on the training data, the feature extraction module 210 and the classification module 220 can output a feature representation of the human body image or picture based on the image or picture.

[0045] After determining the first skeleton information of the first region of the target object 103 in the original image 101, the electronic device 110 can determine the second skeleton information corresponding to the second region of the target object 103 based on the first skeleton information. It should be understood that the second region should have a larger range than the first region. In other words, the second region can be understood as an extension of the first region. Figure 4A A block diagram of an image expansion process according to some embodiments of the present disclosure is shown.

[0046] As described above, the first group of key points 250 corresponding to the first region of the target object 103 in the original image 101 can be obtained through the first model 201 and the first skeleton information of the first region of the target object 103 can be determined based on the first group of key points 250. Figure 4A As shown, the first skeleton information (ie, the “posture” 401 of the first region) is input into the second model 410 .

[0047] The second model 410 may determine, based on the first skeleton information, second skeleton information corresponding to a second region of the target object 103. The second region has a larger body range of the target object 103 than the first region.

[0048] like Figure 4B As shown, after the first group of key points 250 corresponding to the first region of the target object 103 obtained by the first model 201 is input into the second model 410, a second group of key points 404 of the second region can be obtained. Based on the second group of key points 404, the second skeleton information corresponding to the second region of the target object 103 can be determined. Figure 4B As shown, the second region may be the whole body of the target object 103 , which has a larger range than the first region of the target object 103 presented in the original image 101 .

[0049] For example, the second model 410 may include but is not limited to a controlNet model. The second model 410 may be trained based on a plurality of sample pairs. Each sample pair includes first training skeleton information and corresponding second training skeleton information, wherein the second training skeleton information corresponds to a body part with a larger range than the first training skeleton information.

[0050] During the training process, the second training skeleton information may be determined based on the first training image corresponding to the entire body part. The first skeleton information may be determined based on the second training image generated by intercepting a portion of the first training image.

[0051] In some embodiments, the first training skeleton information and the second skeleton information can be Figure 2A That is, the first training image is input into the first model 201 to obtain the second training skeleton information, and the second training image generated by intercepting a part of the first training image is input into the first model 201 to obtain the first training skeleton information.

[0052] In other embodiments, the first training skeleton information and the second skeleton information may also be obtained by manually labeling the second training image and the first training image respectively.

[0053] Re-reference Figure 4A , the second group of key points 404 output by the second model 410 (ie, second skeleton information corresponding to the second region of the target object 103 ) is input into the third model 420 .

[0054] In some embodiments, the third model 420 can generate an expanded image 102 (also referred to as a "second image" or a "third image" in the present disclosure) based on the second skeleton information and the original image 101. For example, the third model 420 can be implemented as a stable-diffusion (SD) image restoration (Inpainting) model. However, it should be understood that the third model 420 should not be limited to this.

[0055] exist Figure 4A In the embodiment, the image information 402 input to the third model 420 may include information obtained from the original image 101 and / or information associated with the original image and / or information associated with the expanded image 102 to be generated.

[0056] For example, the image information 402 may include a noise image. The noise image has a size corresponding to the expanded image 102 to be generated. For another example, the image information 402 may also include mask information. The mask information indicates an area in the noise image corresponding to the original image. For example, the mask information may be binary, the area of ​​the expanded image to be generated is, for example, 255, and the area corresponding to the original image is, for example, 0. In practice, it can be calculated by the ratio of expansion up, down, left, and right, or it can be obtained by placing the original image on a blank canvas.

[0057] The image information 402 is encoded by the encoding module 420 and then input to the third model 430 . The third model 430 processes the encoded image information 402 and the second skeleton information, and inputs the processed data to the decoding module 440 for decoding to generate the expanded image 102 .

[0058] In some other embodiments, the third model 430 may also receive a prompt word 403. The prompt word may, for example, describe details such as clothing and accessories of the target object 103. Based on the prompt word, the expanded image 102 may present a more accurate and vivid effect.

[0059] The solution disclosed in the present invention can avoid the problem of inappropriate limb number and size ratio in the expanded human body. By extracting the first skeleton information of the human body part of the original image to determine the second skeleton information of the larger human body part and generating the expanded image based on this, the accuracy of the portrait expansion is significantly improved, so that it can truly reflect the proportions of the human body and the corresponding human body parts, thereby significantly improving the user experience.

[0060] Example Process

[0061] Figure 5 FIG. 5 is a flowchart of a process 600 of image processing according to some embodiments of the present disclosure. The process 500 may be implemented at the electronic device 110. It should be understood that the process 500 may also be implemented at the server 120.

[0062] In block 510 , the electronic device 110 acquires a first image associated with a target object, where the first image presents a first area of ​​the target object.

[0063] In block 520 , the electronic device 120 determines first skeleton information corresponding to the first region based on the first image.

[0064] In some embodiments, the electronic device 110 may process the first image using a first model to determine a first group of key points corresponding to the first area; and determine the first skeleton information based on the first group of key points.

[0065] In some embodiments, the first model includes a classification module, a coding book and a decoding module, and the electronic device 110 can use the classification module to determine the classification feature representation of the first image; based on the classification feature representation, determine a target feature representation that matches the classification feature representation from a set of candidate feature representations in the coding book; and use the decoding module to process the target feature representation to determine the first group of key points corresponding to the first area.

[0066] In box 530, the electronic device 110 determines second skeleton information corresponding to a second area of ​​the target object based on the first skeleton information, where the second area has a larger range than the first area.

[0067] In some embodiments, the electronic device 110 may process the first skeleton information using a second model to determine a second group of key points corresponding to the second area; and determine the second skeleton information based on the second group of key points.

[0068] In some embodiments, the second model is trained based on multiple samples, and the sample pairs include first training skeleton information and corresponding second training skeleton information, and the second training skeleton information corresponds to a larger range of body parts.

[0069] In some embodiments, the second training skeleton information is determined based on a first training image corresponding to all body parts, the first skeleton information is determined based on a second training image, and the second training image is generated by intercepting a portion of the first training image.

[0070] In block 540 , the electronic device 110 generates a second image associated with the target object based on the second skeleton information and the first image, wherein the second image presents the second region of the target object.

[0071] In some embodiments, the electronic device 110 may provide the second skeleton information and the first image to a third model; and determine the second image based on a third image generated by the third model.

[0072] In some embodiments, the electronic device 110 may provide a noise image and mask information to the third model, wherein the noise image has a size corresponding to the second image to be generated, and the mask information indicates an area in the noise image corresponding to the first image.

[0073] In some embodiments, the electronic device 110 may provide a prompt word to the third model so that the third model generates the third image based on the prompt word.

[0074] In some embodiments, the electronic device 110 may fuse the first image and the third image to determine the second image.

[0075] In some embodiments, the target object includes a human object, and the first image corresponds to a half-body image of the human object, and the second image corresponds to a full-body image of the human object.

[0076] Example devices and equipment

[0077] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 6 A schematic structural block diagram of an apparatus 600 for data processing according to some embodiments of the present disclosure is shown.

[0078] like Figure 6 As shown, the device 600 may include an image acquisition module 610, which is configured to acquire a first image associated with a target object, and the first image presents a first area of ​​the target object. The device 600 may also include a first skeleton information determination module 620, which is configured to determine first skeleton information corresponding to the first area based on the first image. The device 600 may also include a second skeleton information determination module 630, which is configured to determine second skeleton information corresponding to a second area of ​​the target object based on the first skeleton information, and the second area has a larger range than the first area. The device 600 may also include an image generation module 640, which is configured to generate a second image associated with the target object based on the second skeleton information and the first image, and the second image presents the second area of ​​the target object.

[0079] In some embodiments, the first skeleton information determination module 620 may also include a first image processing module, which is configured to: process the first image using a first model to determine a first group of key points corresponding to the first area; and determine the first skeleton information based on the first group of key points.

[0080] In some embodiments, the first model includes a classification module, a coding book and a decoding module, and the first image processing module can also be configured to: use the classification module to determine the classification feature representation of the first image; based on the classification feature representation, determine the target feature representation that matches the classification feature representation from a set of candidate feature representations in the coding book; and use the decoding module to process the target feature representation to determine the first group of key points corresponding to the first area.

[0081] In some embodiments, the second skeleton information determination module 630 can also be configured to process the first skeleton information using a second model to determine a second group of key points corresponding to the second area; and determine the second skeleton information based on the second group of key points.

[0082] In some embodiments, the second model is trained based on multiple samples, and the sample pairs include first training skeleton information and corresponding second training skeleton information, and the second training skeleton information corresponds to a larger range of body parts.

[0083] In some embodiments, the second training skeleton information is determined based on a first training image corresponding to the entire body part, the first skeleton information is determined based on a second training image, and the second training image is generated by intercepting a portion of the first training image.

[0084] In some embodiments, the image generation module 640 may also be configured to provide the second bone information and the first image to a third model; and determine the second image based on a third image generated by the third model.

[0085] In some embodiments, the image generation module 640 may also be configured to provide a noise image and mask information to the third model, wherein the noise image has a size corresponding to the second image to be generated, and the mask information indicates an area in the noise image corresponding to the first image.

[0086] In some embodiments, the image generation module 640 may be further configured to provide a prompt word to the third model, so that the third model generates the third image based on the prompt word.

[0087] In some embodiments, the apparatus 600 may further include a fusion module, which may be configured to fuse the first image and the third image to determine the second image.

[0088] In some embodiments, the target object includes a human object, and the first image corresponds to a half-body image of the human object, and the second image corresponds to a full-body image of the human object.

[0089] The units included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 600 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0090] Figure 7 1 shows a block diagram of a computing device / server 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 7 The computing device / server 700 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein.

[0091] like Figure 7 As shown, computing device / server 700 is in the form of a general computing device. Components of computing device / server 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 760, and one or more output devices 760. Processing unit 710 may be an actual or virtual processor and is capable of performing various processes according to a program stored in memory 720. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to increase the parallel processing capabilities of computing device / server 700.

[0092] The computing device / server 700 typically includes a plurality of computer storage media. Such media may be any available media accessible to the computing device / server 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 may be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 may be a removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which may be capable of being used to store information and / or data (e.g., training data for training) and may be accessed within the computing device / server 700.

[0093] The computing device / server 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 7 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.

[0094] The communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device / server 700 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device / server 700 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0095] Input device 750 may be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 may be one or more output devices, such as a display, speaker, printer, etc. Computing device / server 700 may also communicate with one or more external devices (not shown) as needed, such as storage devices, display devices, etc., with one or more devices that enable users to interact with computing device / server 700, or with any device that enables computing device / server 700 to communicate with one or more other computing devices (e.g., network card, modem, etc.) through communication unit 740. Such communication may be performed via an input / output (I / O) interface (not shown).

[0096] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, wherein the one or more computer instructions are executed by a processor to implement the method described above.

[0097] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0098] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0099] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0100] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0101] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the marketplace, or to enable other persons of ordinary skill in the art to understand the implementations disclosed herein.

Claims

1. An image processing method, include: Acquire a first image associated with a target object, wherein the first image presents a first area of ​​the target object; Based on the first image, determining first bone information corresponding to the first region; Based on the first skeleton information, determining second skeleton information corresponding to a second area of ​​the target object, where the second area has a larger range than the first area; as well as A second image associated with the target object is generated based on the second skeleton information and the first image, wherein the second image presents the second region of the target object.

2. The method according to claim 1, wherein first bone information corresponding to the first region is determined based on the first image. include: Processing the first image using a first model to determine a first set of key points corresponding to the first region; as well as Based on the first group of key points, the first skeleton information is determined.

3. The method according to claim 2, wherein the first model comprises a classification module, a codebook and a decoding module, and the first image is processed using the first model include: Determining, using the classification module, a classification feature representation of the first image; Based on the classification feature representation, determining a target feature representation matching the classification feature representation from a set of candidate feature representations in the codebook; as well as The target feature representation is processed using the decoding module to determine the first group of key points corresponding to the first area.

4. The method according to claim 1, wherein second skeleton information corresponding to a second area of ​​the target object is determined based on the first skeleton information include: Processing the first skeleton information using a second model to determine a second set of key points corresponding to the second region; as well as Based on the second group of key points, the second skeleton information is determined.

5. The method according to claim 4, wherein the second model is trained based on multiple samples, the sample pairs include first training skeleton information and corresponding second training skeleton information, and the second training skeleton information corresponds to a larger range of body parts.

6. The method according to claim 5, wherein the second training skeleton information is determined based on a first training image corresponding to the entire body part, the first skeleton information is determined based on a second training image, and the second training image is generated by intercepting a portion of the first training image.

7. The method according to claim 1, wherein a second image associated with the target object is generated based on the second skeleton information and the first image. include: Providing the second bone information and the first image to a third model; as well as The second image is determined based on a third image generated by the third model.

8. The method according to claim 7, further comprising: include: A noise image and mask information are provided to the third model, wherein the noise image has a size corresponding to the second image to be generated, and the mask information indicates a region in the noise image corresponding to the first image.

9. The method according to claim 7, further comprising: include: The third model is provided with a cue word so that the third model generates the third image based on the cue word.

10. The method according to claim 7, wherein the second image is determined based on a third image generated by the third model include: The first image and the third image are fused to determine the second image.

11. An image processing device, include: An image acquisition module is configured to acquire a first image associated with a target object, wherein the first image presents a first area of ​​the target object; A first skeleton information determination module is configured to determine first skeleton information corresponding to the first region based on the first image; A second skeleton information determination module is configured to determine second skeleton information corresponding to a second area of ​​the target object based on the first skeleton information, the second area having a larger range than the first area; as well as The image generation module is configured to generate a second image associated with the target object based on the second skeleton information and the first image, wherein the second image presents the second area of ​​the target object.

12. An electronic device, include: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.