Image processing method and apparatus, and device and storage medium
By obtaining the local image of the target object and determining its bone information, the problem of inaccurate image expansion in the prior art is solved, high-quality global image generation is achieved, and user experience is improved.
Patent Information
- Application Number
- PCT/CN2024/136879
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-12-04
- Publication Date
- 2025-06-12
AI Technical Summary
Existing image processing technologies are difficult to effectively expand from local images to global images, especially portrait images, which often have problems such as errors in the number or position of limbs and unreasonable limb proportions, resulting in poor quality of extended images.
By obtaining a local image of the target object, the bone information of the area is determined, and based on the bone information, it expands to a larger area, and generates a global image. The method includes acquiring the first image, determining the bone information of the first region, determining the bone information of the second region based on the information, and generating a second image presenting the second region.
Improve the accuracy of image expansion, ensuring the rationality of the number, position and proportion of extended image limbs, thereby improving the user experience.
Smart Images

Figure CN2024136879_12062025_PF_FP_ABST
Abstract
Description
Image processing method, device, equipment and storage medium
[0001] This application claims priority to the Chinese invention patent application entitled “Image processing method, device, equipment and storage medium” filed on December 6, 2023, with application number 202311666394.X. The entire contents of that application are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to image processing methods, apparatuses, devices, and computer-readable storage media. Background Art
[0003] With the advancement of imaging technology, users' demands for image processing are increasing. In some application scenarios, image processing technology is expected to enhance the presentation of captured or captured images. For example, users expect image processing to enable expansion from a local image to a global image, particularly from a part of the body to the entire body. This presents a challenge for current image processing technologies. Summary of the Invention
[0004] In a first aspect of the present disclosure, an image processing method is provided. The method includes acquiring a first image associated with a target object, the first image representing a first region of the target object; determining first skeletal information corresponding to the first region based on the first image; determining second skeletal information corresponding to a second region of the target object based on the first skeletal information, the second region having a larger range than the first region; and generating a second image associated with the target object based on the second skeletal information and the first image, the second image representing the second region of the target object.
[0005] In a second aspect of the present disclosure, an image processing device is provided. The device includes an image acquisition module configured to acquire a first image associated with a target object, the first image presenting a first region of the target object; a first skeletal information determination module configured to determine first skeletal information corresponding to the first region based on the first image; a second skeletal information determination module configured to determine second skeletal information corresponding to a second region of the target object based on the first skeletal information, the second region having a larger range than the first region; and an image generation module configured to generate a second image associated with the target object based on the second skeletal information and the first image, the second image presenting the second region of the target object.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the medium, and when the program is executed by a processor, the method of the first aspect is implemented.
[0008] In a fifth aspect of the present disclosure, there is provided a computer program product, which is tangibly stored in a computer storage medium and comprises computer executable instructions, which when executed by a device cause the device to perform the method of the first aspect.
[0009] It should be understood that the contents described in the summary of the present invention are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] FIG2A shows a block diagram of a key point extraction process according to some embodiments of the present disclosure;
[0013] FIG2B shows a schematic diagram of key point extraction according to some embodiments of the present disclosure;
[0014] FIG3 shows a block diagram of a model training process for key point extraction according to some embodiments of the present disclosure;
[0015] FIG4A shows a block diagram of an image expansion process according to some embodiments of the present disclosure;
[0016] FIG4B shows a schematic diagram of an image expansion process according to some embodiments of the present disclosure;
[0017] FIG5 shows a flowchart of an image processing process according to some embodiments of the present disclosure;
[0018] FIG6 shows a block diagram of an image processing apparatus according to some embodiments of the present disclosure; and
[0019] FIG7 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.
[0022] It is understandable that the image data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0023] It is understandable that before using the technical solutions disclosed in the various embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the data involved in the present disclosure should be informed to relevant users and authorization should be obtained from relevant users in an appropriate manner in accordance with relevant laws and regulations. The relevant users may include any type of right holders, such as individuals, enterprises, and groups.
[0024] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.
[0025] As mentioned above, users expect to expand a partial image into a larger image. In particular, they expect to expand a partial portrait into a larger or even full-body portrait. However, current image processing solutions often produce expanded images with defects such as incorrect number or position of limbs, as well as illogical limb proportions. This results in poor quality of the expanded portrait and low user satisfaction.
[0026] The technical solution disclosed in the present invention provides an image processing solution. According to the image processing solution of the embodiment of the present invention, the first skeleton information corresponding to the first area can be determined based on the first image acquired and presenting the first part of the target object, and the second skeleton information corresponding to the second area of the target object can be determined based on the first skeleton information. A second image presenting the second area of the target object is thereby generated based on the second skeleton information and the first image. In this way, a second image corresponding to the second area expanded from the first image corresponding to the first area can be accurately generated based on the generated second skeleton information, thereby improving the accuracy of the expanded image and improving the user experience.
[0027] Sample Environment
[0028] First, reference is made to FIG1 , which schematically illustrates a diagram of an example environment 100 in which example implementations according to the present disclosure may be implemented.
[0029] In the example environment 100, the electronic device 110 may receive an original image 101. The original image 101 may be, for example, an image input by a user, showing a portion of a target object 103. It is understood that the target object 103 may generally include a human body, but may also include any animal body with bones.
[0030] In some other embodiments, the original image 101 may also be an image captured by the electronic device 110 through a capture device. In some embodiments, the capture device may be configured to be connected to the electronic device 110. In some other embodiments, the capture device may also be integrated into the electronic device 110.
[0031] After acquiring the original image 101 that presents a portion of the target object 103 , the electronic device 110 may generate an expanded image 102 that presents a larger portion or even the entire region of the target object 103 through image processing technology.
[0032] The electronic device 110 can generate the expanded image 102 based on the original image 101 via the server 120. For example, the expanded image 102 can be generated by an image processing model deployed on the server. The original image 101 can be the input of the image processing model, while the expanded image 102 is the output of the image processing model. In some embodiments, the server 120 can be arranged independently of the electronic device 110, for example, as a remote server. In other embodiments, the server 120 can be integrated into or implemented in the electronic device 110.
[0033] The electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.). The server 120 can be various types of electronic systems / servers that can provide computing power, including but not limited to mainframes, edge computing nodes, servers in cloud environments, and the like.
[0034] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0035] Skeleton information extraction process
[0036] FIG2A shows a block diagram of a key point extraction process according to some embodiments of the present disclosure. The process shown in FIG2A can be implemented in electronic device 110 or in server 120. For ease of discussion, the process shown in FIG2 will be described with reference to environment 100 of FIG1.
[0037] Electronic device 110 obtains an original image 101 (also referred to as a "first image" in this disclosure). Original image 101 presents a first region of a target object 103. As shown in the example of FIG1 , the first region may include at least a portion of the head and upper torso of target object 103. Based on original image 101, the electronic device may determine first skeletal information of the first region. The first skeletal information may be determined, for example, by determining a first set of key points of the first region of target object 103 presented in original image 101.
[0038] As shown in FIG2A , after electronic device 110 acquires original image 101, the original image 101 may be input into first model 201. In first model 201, feature extraction module 210 extracts features of original image 101. Feature extraction module 210 may be considered a backbone network. For example, feature extraction module 210 may include a convolutional neural network (CNN) or a residual neural network (ResNet).
[0039] The first model 201 further includes a classification module 220 . The extracted features of the original image 101 are provided from the feature extraction module 210 to the classification module 220 , so as to obtain a classification feature representation of the original image 101 through the classification module 220 .
[0040] The first model 201 also includes a codebook 230. The codebook 230 may include a set of candidate feature representations. The classification feature representation of the original image 101 obtained by the classification module 220 is compared with the set of candidate feature representations in the codebook 230 to determine a target feature representation from the set of candidate feature representations that matches the classification feature representation obtained by the classification module 220. For example, the candidate feature representation in the codebook 230 that is closest to the classification feature representation obtained by the classification module 220 is determined as the target feature representation.
[0041] The determined target feature representation is decoded by the decoder 240 in the first model 201, and the output of the decoder 240, that is, the output of the first model 201, is a first set of key points 250 of the first region of the target object 103 presented in the original image 101. Based on the first set of key points 250, first skeleton information corresponding to the first region of the target object 103 can be determined.
[0042] As shown in FIG2B , after the original image 101 is input into the first model 201 , a first set of key points 250 may be output. The first set of key points 250 may correspond to the skeleton information of the first region of the target object 103 .
[0043] As described above, in the first model 201, the codebook 230 includes a set of candidate feature representations. The candidate feature representations included in the codebook 230 can be determined by learning the relationship information between the key points of the skeleton of the target object 103. FIG3 shows a block diagram illustrating the model training process for key point extraction according to some embodiments of the present disclosure.
[0044] Specifically, as shown in FIG3 , skeletal key points can be extracted from multiple training images, and the extracted training key points 301 are input into an encoder 310 for key point information structure encoding. It should be understood that the training key points 301 can include multiple groups of key points extracted from multiple training images. The training key points 301 can be two-dimensional coordinates of two-dimensional key points obtained from the multiple training images.
[0045] The training keypoints encoded by the encoder 310 are input into the codebook 230 to find the candidate feature representation in the codebook 230 that is closest to the encoded training keypoints and update the corresponding candidate feature representations in the codebook 230. This ensures that after the feature representation output from the encoder 230 is decoded by the decoder 320, the resulting reference keypoint 302 approaches the input training keypoint 301.
[0046] Furthermore, the feature extraction module 210 and classification module 220 in the first model 201 can also be acquired by learning from training data. For example, the training data for the feature extraction module 210 and classification module 220 can be human images or pictures, such as images of the entire body, half of the body, or a portion of a human body region. Based on the training data, the feature extraction module 210 and classification module 220 can output feature representations of the human image or picture based on the image or picture.
[0047] After determining the first skeleton information of the first region of the target object 103 in the original image 101, the electronic device 110 can determine the second skeleton information corresponding to the second region of the target object 103 based on the first skeleton information. It should be understood that the second region should have a larger range than the first region. In other words, the second region can be understood as an extension of the first region. Figure 4A shows a block diagram of the image expansion process according to some embodiments of the present disclosure.
[0048] As described above, the first model 201 can be used to obtain a first set of key points 250 corresponding to the first region of the target object 103 in the original image 101, and based on this, first skeletal information of the first region of the target object 103 can be determined. As shown in FIG4A , this first skeletal information (i.e., the "pose" 401 of the first region) is input into the second model 410.
[0049] The second model 410 may determine, based on the first skeleton information, second skeleton information corresponding to a second region of the target object 103. The second region has a larger body area of the target object 103 than the first region.
[0050] As shown in FIG4B , after inputting the first set of key points 250 corresponding to the first region of the target object 103 obtained by the first model 201 into the second model 410, a second set of key points 404 for the second region can be obtained. Based on the second set of key points 404, second skeletal information corresponding to the second region of the target object 103 can be determined. As shown in FIG4B , the second region can be the entire body of the target object 103, which has a larger area than the first region of the target object 103 represented in the original image 101.
[0051] For example, the second model 410 may include but is not limited to a controlNet model. The second model 410 may be trained based on multiple sample pairs. Each sample pair includes first training skeleton information and corresponding second training skeleton information, where the second training skeleton information corresponds to a larger range of body parts than the first training skeleton information.
[0052] During the training process, the second training skeleton information may be determined based on the first training image corresponding to the entire body part. The first skeleton information may be determined based on the second training image generated by intercepting a portion of the first training image.
[0053] In some embodiments, the first training skeleton information and the second training skeleton information can be determined by the first model 201 shown in FIG2A . That is, the first training image is input into the first model 201 to obtain the second training skeleton information, and the second training image generated by intercepting a portion of the first training image is input into the first model 201 to obtain the first training skeleton information.
[0054] In other embodiments, the first training skeleton information and the second skeleton information may also be obtained by manually labeling the second training image and the first training image respectively.
[0055] Referring back to FIG. 4A , the second set of key points 404 output by the second model 410 (ie, second skeleton information corresponding to the second region of the target object 103 ) is input to the third model 420 .
[0056] In some embodiments, the third model 420 can generate an expanded image 102 (also referred to as a "second image" or a "third image" in this disclosure) based on the second skeleton information and the original image 101. For example, the third model 420 can be implemented as a stable-diffusion (SD) image restoration (Inpainting) model. However, it should be understood that the third model 420 should not be limited to this.
[0057] In FIG. 4A , the image information 402 input to the third model 420 may include information acquired from the original image 101 and / or information associated with the original image and / or information associated with the expanded image 102 to be generated.
[0058] For example, image information 402 may include a noise image. The noise image has a size corresponding to the expanded image 102 to be generated. For another example, image information 402 may also include mask information. The mask information indicates the area in the noise image that corresponds to the original image. For example, the mask information may be binary, where the area of the expanded image to be generated is, for example, 255, and the area corresponding to the original image is, for example, 0. In practice, this mask information can be calculated based on the ratio of expansion to the top, bottom, left, and right, or based on the position of the original image placed on a blank canvas.
[0059] The image information 402 is encoded by the encoding module 420 and then input to the third model 430 . The third model 430 processes the encoded image information 402 and the second skeleton information and inputs the processed data to the decoding module 440 for decoding to generate the expanded image 102 .
[0060] In some other embodiments, the third model 430 may also receive prompt words 403. The prompt words may, for example, describe details such as clothing and accessories of the target object 103. Based on the prompt words, the expanded image 102 may be rendered more accurately and vividly.
[0061] The disclosed solution can avoid the problem of inappropriate limb number and size ratio in the expanded human body. By extracting the first skeletal information of the human body part in the original image to determine the second skeletal information of the larger human body part and generating the expanded image based on this, the accuracy of the portrait expansion is significantly improved, enabling it to truly reflect the proportions of the human body and the corresponding body parts, thereby significantly improving the user experience.
[0062] Example Process
[0063] 5 shows a flow chart of a process 600 of image processing according to some embodiments of the present disclosure. The process 500 may be implemented at the electronic device 110. It should be understood that the process 500 may also be implemented at the server 120.
[0064] In block 510 , the electronic device 110 acquires a first image associated with a target object, where the first image presents a first area of the target object.
[0065] In block 520 , the electronic device 120 determines first skeleton information corresponding to the first region based on the first image.
[0066] In some embodiments, the electronic device 110 may process the first image using a first model to determine a first set of key points corresponding to the first area; and determine the first skeleton information based on the first set of key points.
[0067] In some embodiments, the first model includes a classification module, a coding book, and a decoding module, and the electronic device 110 can use the classification module to determine the classification feature representation of the first image; based on the classification feature representation, determine a target feature representation that matches the classification feature representation from a set of candidate feature representations in the coding book; and use the decoding module to process the target feature representation to determine the first group of key points corresponding to the first area.
[0068] In block 530 , the electronic device 110 determines second skeleton information corresponding to a second region of the target object based on the first skeleton information, where the second region has a larger range than the first region.
[0069] In some embodiments, the electronic device 110 may process the first skeleton information using a second model to determine a second set of key points corresponding to the second area; and determine the second skeleton information based on the second set of key points.
[0070] In some embodiments, the second model is trained based on multiple samples, and the sample pairs include first training skeleton information and corresponding second training skeleton information, and the second training skeleton information corresponds to a larger range of body parts.
[0071] In some embodiments, the second training skeleton information is determined based on a first training image corresponding to all body parts, the first skeleton information is determined based on a second training image, and the second training image is generated by intercepting a portion of the first training image.
[0072] In block 540 , the electronic device 110 generates a second image associated with the target object based on the second skeleton information and the first image, where the second image presents the second region of the target object.
[0073] In some embodiments, the electronic device 110 may provide the second skeleton information and the first image to a third model; and determine the second image based on a third image generated by the third model.
[0074] In some embodiments, the electronic device 110 may provide a noise image and mask information to the third model, wherein the noise image has a size corresponding to the second image to be generated, and the mask information indicates an area in the noise image corresponding to the first image.
[0075] In some embodiments, the electronic device 110 may provide a prompt word to the third model, so that the third model generates the third image based on the prompt word.
[0076] In some embodiments, the electronic device 110 may fuse the first image and the third image to determine the second image.
[0077] In some embodiments, the target object includes a human subject, and the first image corresponds to a half-body image of the human subject, and the second image corresponds to a full-body image of the human subject.
[0078] Example devices and equipment
[0079] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 6 shows a schematic structural block diagram of a device 600 for data processing according to some embodiments of the present disclosure.
[0080] As shown in Figure 6, the device 600 may include an image acquisition module 610, which is configured to acquire a first image associated with a target object, wherein the first image presents a first area of the target object. The device 600 may also include a first skeleton information determination module 620, which is configured to determine first skeleton information corresponding to the first area based on the first image. The device 600 may also include a second skeleton information determination module 630, which is configured to determine second skeleton information corresponding to a second area of the target object based on the first skeleton information, wherein the second area has a larger range than the first area. The device 600 may also include an image generation module 640, which is configured to generate a second image associated with the target object based on the second skeleton information and the first image, wherein the second image presents the second area of the target object.
[0081] In some embodiments, the first skeleton information determination module 620 may also include a first image processing module, which is configured to: process the first image using a first model to determine a first set of key points corresponding to the first area; and determine the first skeleton information based on the first set of key points.
[0082] In some embodiments, the first model includes a classification module, a coding book and a decoding module, and the first image processing module can also be configured to: use the classification module to determine the classification feature representation of the first image; based on the classification feature representation, determine the target feature representation that matches the classification feature representation from a set of candidate feature representations in the coding book; and use the decoding module to process the target feature representation to determine the first group of key points corresponding to the first area.
[0083] In some embodiments, the second skeleton information determination module 630 can also be configured to process the first skeleton information using a second model to determine a second set of key points corresponding to the second area; and determine the second skeleton information based on the second set of key points.
[0084] In some embodiments, the second model is trained based on multiple samples, and the sample pairs include first training skeleton information and corresponding second training skeleton information, and the second training skeleton information corresponds to a larger range of body parts.
[0085] In some embodiments, the second training skeleton information is determined based on a first training image corresponding to all body parts, the first skeleton information is determined based on a second training image, and the second training image is generated by intercepting a portion of the first training image.
[0086] In some embodiments, the image generation module 640 may be further configured to provide the second bone information and the first image to a third model; and determine the second image based on a third image generated by the third model.
[0087] In some embodiments, the image generation module 640 can also be configured to provide a noise image and mask information to the third model, wherein the noise image has a size corresponding to the second image to be generated, and the mask information indicates an area in the noise image corresponding to the first image.
[0088] In some embodiments, the image generation module 640 may be further configured to provide a prompt word to the third model, so that the third model generates the third image based on the prompt word.
[0089] In some embodiments, the apparatus 600 may further include a fusion module, which may be configured to fuse the first image and the third image to determine the second image.
[0090] In some embodiments, the target object includes a human subject, and the first image corresponds to a half-body image of the human subject, and the second image corresponds to a full-body image of the human subject.
[0091] The units included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 600 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0092] Figure 7 shows a block diagram of a computing device / server 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the computing device / server 700 shown in Figure 7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0093] As shown in FIG7 , computing device / server 700 is in the form of a general-purpose computing device. Components of computing device / server 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 760, and one or more output devices 760. Processing unit 710 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device / server 700.
[0094] The computing device / server 700 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device / server 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk or any other medium, which can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device / server 700.
[0095] The computing device / server 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0096] The communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device / server 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device / server 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0097] Input device 750 may be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 may be one or more output devices, such as a display, speaker, printer, etc. Computing device / server 700 may also communicate with one or more external devices (not shown) via communication unit 740, as needed, such as storage devices, display devices, etc., with one or more devices that allow users to interact with computing device / server 700, or with any device that allows computing device / server 700 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0098] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, wherein the one or more computer instructions are executed by a processor to implement the method described above.
[0099] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0100] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0101] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0102] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0103] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
1. An image processing method, comprising: Acquire a first image associated with a target object, wherein the first image presents a first area of the target object; Based on the first image, determining first bone information corresponding to the first region; Based on the first skeleton information, determining second skeleton information corresponding to a second area of the target object, where the second area has a larger range than the first area; as well as A second image associated with the target object is generated based on the second skeleton information and the first image, wherein the second image presents the second region of the target object.
2. The method according to claim 1, wherein determining first bone information corresponding to the first region based on the first image comprises: Processing the first image using a first model to determine a first set of key points corresponding to the first region; as well as Based on the first group of key points, the first skeleton information is determined.
3. The method of claim 2, wherein the first model comprises a classification module, a codebook, and a decoding module, and processing the first image using the first model comprises: Determining, using the classification module, a classification feature representation of the first image; Based on the classification feature representation, determining a target feature representation matching the classification feature representation from a set of candidate feature representations in the codebook; as well as The target feature representation is processed using the decoding module to determine the first group of key points corresponding to the first area.
4. The method according to claim 1, wherein determining second skeleton information corresponding to the second area of the target object based on the first skeleton information comprises: Processing the first skeleton information using a second model to determine a second set of key points corresponding to the second region; as well as Based on the second group of key points, the second skeleton information is determined.
5. The method according to claim 4, wherein the second model is trained based on multiple samples, the sample pairs include first training skeleton information and corresponding second training skeleton information, and the second training skeleton information corresponds to a larger range of body parts.
6. The method according to claim 5, wherein the second training skeleton information is determined based on a first training image corresponding to the entire body part, the first skeleton information is determined based on a second training image, and the second training image is generated by intercepting a portion of the first training image.
7. The method according to claim 1, wherein generating a second image associated with the target object based on the second skeleton information and the first image comprises: Providing the second bone information and the first image to a third model; as well as The second image is determined based on a third image generated by the third model.
8. The method according to claim 7, further comprising: A noise image and mask information are provided to the third model, wherein the noise image has a size corresponding to the second image to be generated, and the mask information indicates a region in the noise image corresponding to the first image.
9. The method according to claim 7, further comprising: The third model is provided with a cue word so that the third model generates the third image based on the cue word.
10. The method of claim 7, wherein determining the second image based on the third image generated by the third model comprises: The first image and the third image are fused to determine the second image.
11. An image processing device, comprising: An image acquisition module is configured to acquire a first image associated with a target object, wherein the first image presents a first area of the target object; A first skeleton information determination module is configured to determine first skeleton information corresponding to the first region based on the first image; A second skeleton information determination module is configured to determine second skeleton information corresponding to a second area of the target object based on the first skeleton information, the second area having a larger range than the first area; as well as The image generation module is configured to generate a second image associated with the target object based on the second skeleton information and the first image, wherein the second image presents the second area of the target object.
12. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
14. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1-10.
Citation Information
Patent Citations
Image complement method, device, terminal, and computer-readable storage medium
CN109255768A
Method, device and apparatus for generating whole body image and computer readable storage medium
CN110288532A
Picture composition method and device, and electronic equipment
CN111953907A
Three-dimensional modeling method and device, computer readable storage medium and computer equipment
CN113724378A
Skeleton recognition method, skeleton recognition program, and gymnastics scoring support system
JP2023003929A