Human body image super-division method and device, electronic equipment and storage medium

By iteratively using the face and body super-resolution diffusion model to process human images, the problem of unnatural transition at the junction of face and body images is solved, and efficient and natural super-resolution effects are achieved.

CN120707380AActive Publication Date: 2025-09-26HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410317545.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-26
Estimated Expiration
2044-03-19

AI Technical Summary

Technical Problem

During the super-resolution process of human body images, the resolution difference between the initial super-resolution face image and the human body image is large, resulting in an unnatural transition at the junction of the face area image and the human body area image.

Method used

The face super-resolution diffusion model and the body super-resolution diffusion model are used to iteratively super-resolution the initial body image. By learning the face noise and body noise, a super-resolution body image with a natural transition is obtained by fusion, including compressing the body image to obtain the feature map, cropping the face feature map, and performing multiple iterative fusion.

Benefits of technology

It improves the naturalness of the transition between the face area and the body area image, meets the clarity requirements of the face and body, reduces the amount of data calculation, and improves the super-resolution efficiency and image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707380A_ABST
    Figure CN120707380A_ABST
Patent Text Reader

Abstract

The invention discloses a human body image super-division method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring an initial human body image; performing first super-division processing on the initial human body image through a super-division model to obtain a first super-division human body image, the super-division model comprising a face super-division diffusion model and a human body super-division diffusion model; and taking the first super-division human body image obtained after the first super-division processing as an initial human body image in the second super-division processing, and repeatedly executing the process of the first super-division processing until the Tth super-division processing is completed, thereby obtaining a first super-division human body image of the Tth super-division processing. According to the method, in the human body image super-division process, iteration updating super-division processing is carried out on the initial human body image based on the human face super-division diffusion model and the human body super-division diffusion model, and finally transition of the junction of the human face area image and the human body area image in the super-division human body image obtained through multiple times of super-division iteration fusion is more natural.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a human body image super-resolution method, super-resolution device, electronic device and storage medium. Background Art

[0002] Super-resolution, also known as super-resolution, refers to inputting a low-resolution image and processing it to obtain a high-resolution image with the same content.

[0003] Currently, users require higher resolution for the facial portion of a human body image than for the overall human body image. When super-resolutioning a human body image, the face image is typically super-resolutioned separately to obtain an initial super-resolution face image and an initial super-resolution human body image. The initial super-resolution face image is then pasted to replace the face region image in the initial super-resolution human body image to obtain the target super-resolution human body image.

[0004] However, due to the large difference in resolution between the initial super-resolved face image and the initial super-resolved body image, there is an unnatural transition problem at the junction of the face region image and the body region image in the target super-resolved body image. Summary of the Invention

[0005] In view of this, embodiments of the present application provide a human body image super-resolution method, super-resolution device, electronic device and storage medium to overcome the above problems of the prior art.

[0006] In a first aspect, an embodiment of the present application provides a human body image super-resolution method, which includes: obtaining an initial human body image; performing a first super-resolution processing on the initial human body image through a super-resolution model to obtain a first super-resolution human body image; wherein the super-resolution model includes a face super-resolution diffusion model and a human body super-resolution diffusion model, and the process of the first super-resolution processing includes: obtaining an initial face feature map and an initial human body feature map according to the initial human body image; inputting the initial face feature map to the face super-resolution diffusion model to obtain a first super-resolution face image; inputting the initial human body feature map to the human body super-resolution diffusion model to obtain an initial super-resolution human body image, and the resolution of the first super-resolution face image is higher than the resolution of the initial super-resolution human body image; fusing the first super-resolution face image and the initial super-resolution human body image to obtain a first super-resolution human body image; using the first super-resolution human body image obtained after the first super-resolution processing as the initial human body image in the second super-resolution processing, and repeating the process of the first super-resolution processing until the T-th super-resolution processing is completed to obtain the first super-resolution human body image of the T-th super-resolution processing.

[0007] The solution provided by the present application is to learn the face noise of the initial human image and the human body noise of the initial human image based on the face super-resolution diffusion model, and fuse the face noise and the human body noise to obtain a first super-resolution human image during the super-resolution process of the human body image. The first super-resolution human body image is iteratively super-resolution processed, the face super-resolution diffusion model iteratively learns the face noise in the previous face super-resolution image, and the human body super-resolution diffusion model iteratively learns the human body noise in the previous human body super-resolution image, so that the transition at the junction of the face area image and the human body area image in the super-resolution human body image obtained by multiple super-resolution iterative fusions is finally made more natural.

[0008] Among them, in some optional embodiments, obtaining an initial facial feature map and an initial human body feature map based on the initial human body image includes: compressing the initial human body image to obtain a compressed human body image; and obtaining an initial facial feature map and an initial human body feature map based on the compressed human body image.

[0009] The solution provided in this embodiment reduces the size of data involved in the super-resolution process by compressing the initial human body image and performing super-resolution based on the compressed human body image, thereby improving the super-resolution efficiency of the human body image.

[0010] Among them, in some optional embodiments, obtaining an initial facial feature map and an initial human body feature map based on a compressed human body image includes: determining the compressed human body image as the initial human body feature map; and cropping the initial human body feature map according to the facial frame of the initial human body feature map to obtain the initial facial feature map.

[0011] The solution provided in this embodiment crops the initial human body feature map based on the face frame, and the cropped initial face feature map always contains face features, thereby improving the accuracy of obtaining the initial face feature map.

[0012] Among them, in some optional embodiments, the initial human feature map is cropped according to the face frame of the initial human feature map. Before obtaining the initial human feature map, the human image super-resolution method further includes: performing face detection on the initial human image to obtain first face frame coordinates; determining second face frame coordinates of the initial human feature map according to the first face frame coordinates and a mapping relationship between the initial human image and the initial human feature map; and determining the face frame of the initial human feature map according to the second face frame coordinates.

[0013] The solution provided in this embodiment determines the face frame of the initial human feature map based on the first face frame coordinates of the initial human image and the mapping relationship between the initial human image and the initial human feature map, and the obtained face frame has high accuracy.

[0014] Among them, in some optional embodiments, fusing the first super-resolution face image and the initial super-resolution body image to obtain the first super-resolution body image includes: selecting a second super-resolution face image corresponding to the second face frame coordinates in the initial super-resolution body image; replacing the second super-resolution face image with the first super-resolution face image in the initial super-resolution body image to obtain the first super-resolution body image.

[0015] The solution provided in this embodiment realizes the update of the face area image in the initial super-resolution human body image obtained based on human body super-resolution by using the first super-resolution face image obtained based on face super-resolution. The obtained first super-resolution human body image meets both the face clarity requirements and the human body clarity requirements, thereby improving the super-resolution experience of human body images.

[0016] Among them, in some optional embodiments, the first super-resolution face image includes a first super-resolution head image and a first super-resolution non-head image, and the initial super-resolution human body image includes an initial super-resolution head image and an initial super-resolution non-head image; in the initial super-resolution human body image, the second super-resolution face image is replaced with the first super-resolution face image to obtain a first super-resolution human body image, including: in the initial super-resolution human body image, the initial super-resolution head image is replaced with the first super-resolution head image to obtain a first super-resolution human body image.

[0017] The solution provided in this embodiment realizes the update of the initial super-resolution head image in the initial super-resolution human body image obtained based on human body super-resolution, reduces the amount of data calculation in the super-resolution processing process, and improves the super-resolution efficiency of human body images.

[0018] Among them, in some optional embodiments, the human body image super-resolution method also includes: decoding the first super-resolution human body image of the T-th super-resolution processing to obtain a target super-resolution human body image, and the resolution of the target super-resolution human body image is higher than the resolution of the initial human body image.

[0019] The solution provided in this embodiment decodes the first super-resolution human body image processed for the Tth super-resolution, and the obtained target super-resolution human body image has the same dimension as the initial human body image, thereby improving the image quality of the target super-resolution human body image.

[0020] Among them, in some optional embodiments, the initial face feature map includes multiple sub-face feature maps, and before the initial face feature map is input into the face super-resolution diffusion model to obtain the first super-resolution face image, the human body image super-resolution method also includes: performing sliding window blocking processing on the initial face feature map to obtain multiple sub-face feature maps; inputting the initial face feature map into the face super-resolution diffusion model to obtain the first super-resolution face image, including: inputting multiple sub-face feature maps into the face super-resolution diffusion model one by one, so that the face super-resolution diffusion model outputs the first super-resolution face image according to the multiple sub-face feature maps.

[0021] The solution provided in this embodiment, by sliding window blocking the initial facial feature map into multiple sub-face feature maps, and inputting the multiple sub-face feature maps one by one into the face super-resolution diffusion model for face super-resolution, avoids the increase in data calculation amount for face super-resolution caused by the large image data of the facial feature map input into the face super-resolution diffusion model at a single time, and improves the super-resolution efficiency of face super-resolution of the facial feature map.

[0022] Among them, in some optional embodiments, the initial human body feature map includes multiple sub-human body feature maps, and before the initial human body feature map is input into the human body super-resolution diffusion model to obtain the initial super-resolution human body image, the human body image super-resolution method also includes: performing sliding window blocking processing on the initial human body feature map to obtain multiple sub-human body feature maps; inputting the initial human body feature map into the human body super-resolution diffusion model to obtain the initial super-resolution human body image, including: inputting multiple sub-human body feature maps into the human body super-resolution diffusion model one by one, so that the human body super-resolution diffusion model outputs the initial super-resolution human body image according to the multiple sub-human body feature maps.

[0023] The solution provided in this embodiment divides the initial human feature map into multiple sub-human feature maps through sliding window block processing, and inputs the multiple sub-human feature maps one by one into the human super-resolution diffusion model for human super-resolution. This avoids the increase in the amount of data calculation for human super-resolution caused by the large image data of the human feature map input into the human super-resolution diffusion model at a single time, thereby improving the super-resolution efficiency of human feature maps.

[0024] Among them, in some optional embodiments, before the initial human body image is super-resolved for the first time by the super-resolution model to obtain the first super-resolved human body image, the human body image super-resolution method also includes: constructing a super-resolution diffusion model based on the diffusion model and the control network model; based on the first historical face feature map and the second historical face feature map, training the super-resolution diffusion model to obtain a face super-resolution diffusion model, the first historical face feature map and the second historical face feature map contain the same facial features, and the resolution of the first historical face feature map is higher than the resolution of the second historical face feature map; based on the first historical human body feature map and the second historical human body feature map, training the super-resolution diffusion model to obtain a human body super-resolution diffusion model, the first historical human body feature map and the second historical human body feature map contain the same human body features, and the resolution of the first historical human body feature map is higher than the resolution of the second historical human body feature map.

[0025] The solution provided in this embodiment constructs a super-resolution diffusion model based on a diffusion model and a control network model. The control network model serves as a conditional control generation model of the super-resolution diffusion model, so that the super-resolution process of the super-resolution of the human body feature map by the super-resolution diffusion model is controllable, thereby improving the user's super-resolution experience of the human body feature map.

[0026] Among them, in some optional embodiments, based on the first historical face feature map and the second historical face feature map, the super-resolution diffusion model is trained to obtain a face super-resolution diffusion model, including: inputting the first historical face feature map into the diffusion model in the super-resolution diffusion model, so that the diffusion model outputs the first historical face feature to the control network model according to the first historical face feature map; inputting the second historical face feature map into the control network model in the super-resolution diffusion model, so that the control network model is trained according to the first historical face feature and the second historical face feature map to obtain a face super-resolution diffusion model.

[0027] The solution provided in this embodiment realizes the training of the control network model in the super-resolution diffusion model based on the historical facial feature map, without the need to train the diffusion model in the super-resolution diffusion model, reducing the amount of data calculation during the training process of the super-resolution diffusion model and improving the training speed of the super-resolution diffusion model.

[0028] Among them, in some optional embodiments, based on the first historical human body feature map and the second historical human body feature map, the super-resolution diffusion model is trained to obtain the human body super-resolution diffusion model, including: inputting the first historical human body feature map into the diffusion model in the super-resolution diffusion model, so that the diffusion model outputs the first historical human body features to the control network model according to the first historical human body feature map; inputting the second historical human body feature map into the control network model in the super-resolution diffusion model, so that the control network model is trained according to the first historical human body features and the second historical human body feature map to obtain the human body super-resolution diffusion model.

[0029] The solution provided in this embodiment realizes the training of the control network model in the super-resolution diffusion model based on the historical human body feature map, eliminating the need to train the diffusion model in the super-resolution diffusion model, reducing the amount of data calculation during the training process of the super-resolution diffusion model, and improving the training speed of the super-resolution diffusion model.

[0030] In some optional embodiments, obtaining the initial human body image includes: obtaining an initial scene image containing the initial human body image; performing human body detection on the initial scene image to obtain a human body bounding box; and cropping the initial scene image according to the human body bounding box to obtain the initial human body image.

[0031] The solution provided in this embodiment crops the initial human image from the initial scene image based on the human body bounding box obtained by human body detection, thereby improving the accuracy of obtaining the initial human body image.

[0032] Among them, in some optional embodiments, the human body image super-resolution method also includes: segmenting the first super-resolution human body image of the T-th super-resolution processing to obtain a human body mask image; pasting the human body mask image back to the position corresponding to the initial human body image in the initial scene image to obtain the target scene image.

[0033] The solution provided in this embodiment realizes super-resolution processing of only the initial human body image in the initial scene image, avoiding super-resolution processing of the entire initial scene image resulting in low super-resolution efficiency, and improving the super-resolution efficiency of super-resolution of the human body image.

[0034] In some optional embodiments, the human body image super-resolution method further includes: performing corrosion processing on the boundary of the human body mask image in the target scene image.

[0035] The solution provided in this embodiment can reduce the unnatural transition between the boundary of the super-resolution human body image and the initial scene image by corroding the boundary of the super-resolution human body image in the target scene image, thereby improving the fusion degree between the super-resolution human body image and the initial scene image.

[0036] In some optional embodiments, the human body image super-resolution method further includes: performing Gaussian blur processing on the boundary of the human body mask image in the target scene image.

[0037] The solution provided in this embodiment can reduce the unnatural transition between the boundary of the super-resolution human image and the initial scene image by performing Gaussian blur processing on the boundary of the super-resolution human image in the target scene image, thereby improving the fusion degree between the super-resolution human image and the initial scene image.

[0038] In a second aspect, an embodiment of the present application provides a human body image super-resolution device, which includes: an initial image acquisition module for acquiring an initial human body image; a super-resolution processing module for performing a first super-resolution processing on the initial human body image through a super-resolution model to obtain a first super-resolution human body image; wherein the super-resolution model includes a face super-resolution diffusion model and a human body super-resolution diffusion model, and the process of the first super-resolution processing includes: obtaining an initial face feature map and an initial human body feature map according to the initial human body image; inputting the initial face feature map to the face super-resolution diffusion model to obtain a first super-resolution face image; inputting the initial human body feature map to the human body super-resolution diffusion model to obtain an initial super-resolution human body image, the resolution of the first super-resolution face image being higher than the resolution of the initial super-resolution human body image; fusing the first super-resolution face image and the initial super-resolution human body image to obtain the first super-resolution human body image; a repeated execution module for using the first super-resolution human body image obtained after the first super-resolution processing as the initial human body image in the second super-resolution processing, and repeatedly executing the process of the first super-resolution processing until the T-th super-resolution processing is completed to obtain the first super-resolution human body image of the T-th super-resolution processing.

[0039] In a third aspect, an embodiment of the present application provides an electronic device, which includes: one or more processors, and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and one or more processors call the computer instructions to enable the electronic device to execute the human body image super-resolution method provided in the first aspect above.

[0040] In a fourth aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to execute the human body image super-resolution method provided in the first aspect above.

[0041] In some optional embodiments, the chip system also includes a memory, and the memory is connected to one or more processors through circuits or wires.

[0042] In some optional embodiments, the chip system also includes a communication interface.

[0043] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes instructions. When the instructions are executed on an electronic device, the electronic device executes the human body image super-resolution method provided in the first aspect above.

[0044] In a sixth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the human body image super-resolution method provided in the first aspect above.

[0045] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 A structural diagram of a software system of an electronic device provided in an embodiment of the present application is shown.

[0048] Figure 2 A flow chart of a human body image super-resolution method provided in an embodiment of the present application is shown.

[0049] Figure 3A schematic flow chart of the super-resolution process in the human body image super-resolution method provided in an embodiment of the present application is shown.

[0050] Figure 4 A schematic diagram of a scenario of sliding window block processing in the human body image super-resolution method provided in an embodiment of the present application is shown.

[0051] Figure 5 A schematic diagram of a scenario flow of a human body image super-resolution method provided in an embodiment of the present application is shown.

[0052] Figure 6 A schematic diagram of a scenario flow of super-resolution processing in the human body image super-resolution method provided in an embodiment of the present application is shown.

[0053] Figure 7 A flow chart of the super-resolution model construction method provided in an embodiment of the present application is shown.

[0054] Figure 8 A schematic diagram of a scene of a diffusion model in the human image super-resolution method provided in an embodiment of the present application is shown.

[0055] Figure 9 A structural schematic diagram of a super-resolution diffusion model in the human body image super-resolution method provided in an embodiment of the present application is shown.

[0056] Figure 10 A structural schematic diagram of a super-resolution diffusion model used for reasoning in the human image super-resolution method provided in an embodiment of the present application is shown.

[0057] Figure 11 A structural block diagram of a human body image super-resolution device provided in an embodiment of the present application is shown.

[0058] Figure 12 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.

[0059] Figure 13 A functional block diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0060] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application.

[0061] The disclosure below provides many different embodiments or examples for implementing different structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, these are merely examples and are not intended to limit the present application. In addition, the present application may repeat reference numbers and / or reference letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0062] Super-resolution, also known as super-resolution, refers to inputting a low-resolution image and processing it to obtain a high-resolution image with the same content.

[0063] Currently, users require higher resolution for the facial portion of a human body image than for the overall human body image. When super-resolutioning a human body image, the face image is typically super-resolutioned separately to obtain an initial super-resolution face image and an initial super-resolution human body image. The initial super-resolution face image is then pasted to replace the face region image in the initial super-resolution human body image to obtain the target super-resolution human body image.

[0064] However, due to the large difference in resolution between the initial super-resolved face image and the initial super-resolved body image, there is an unnatural transition problem at the junction of the face region image and the body region image in the target super-resolved body image.

[0065] In response to the above problems, the embodiments of the present application provide a human body image super-resolution method, a super-resolution device, an electronic device and a storage medium, which obtain an initial human body image and perform a first super-resolution processing on the initial human body image through a super-resolution model to obtain a first super-resolution human body image, wherein the super-resolution model includes a face super-resolution diffusion model and a human body super-resolution diffusion model, and the first super-resolution processing process includes obtaining an initial face feature map and an initial human body feature map according to the initial human body image, and inputting the initial face feature map into the face super-resolution diffusion model to obtain a first super-resolution face image, and inputting the initial human body feature map into the human body super-resolution diffusion model to obtain an initial super-resolution human body image, the resolution of the first super-resolution face image is higher than the resolution of the initial super-resolution human body image, and the first super-resolution face image and the initial super-resolution human body image are fused to obtain a first super-resolution human body image, and the first super-resolution face image and the initial super-resolution human body image are fused to obtain a first super-resolution human body image, and the first super-resolution The first super-resolution human body image obtained after processing is used as the initial human body image in the second super-resolution processing. The process of the first super-resolution processing is repeated until the T-th super-resolution processing is completed, and the first super-resolution human body image of the T-th super-resolution processing is obtained. In the process of human body image super-resolution, the face noise of the initial human body image is learned based on the face super-resolution diffusion model and the human body noise of the initial human body image is learned based on the human body super-resolution diffusion model, and the face noise and the human body noise are fused to obtain the first super-resolution human body image. The first super-resolution human body image is iteratively super-resolution processed, the face super-resolution diffusion model iteratively learns the face noise in the previous face super-resolution image, and the human body super-resolution diffusion model iteratively learns the human body noise in the previous human body super-resolution image, and finally makes the transition at the junction of the face area image and the human body area image in the super-resolution human body image obtained by multiple super-resolution iterative fusion more natural.

[0066] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0067] The human image super-resolution method provided in the embodiments of the present application can be applied to electronic devices. The electronic devices may include various terminal devices, which may also be referred to as terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc.

[0068] The terminal device may be a mobile phone, a robot vacuum, a drone, a smart TV, a wearable device, a personal digital assistant (PDA), a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal used in industrial control, a wireless terminal used in self-driving, a wireless terminal used in remote medical surgery, a wireless terminal used in smart grids, a wireless terminal used in transportation safety, a wireless terminal used in smart cities, a wireless terminal used in smart homes, etc. The terminal device type is not limited here and can be set according to actual needs.

[0069] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to illustrate the software structure of the electronic device.

[0070] See also Figure 1 , which shows a schematic diagram of the structure of the software system of an electronic device provided by one embodiment of the present application. The software system includes several layers, each with a clear role and division of labor, and communication between layers via software interfaces. In some embodiments, the Android system is divided into four layers: application layer, application framework layer, system library layer, and kernel layer, from top to bottom.

[0071] The application layer may include a series of applications. For example, the application layer may include super-resolution applications, camera applications, gallery applications, call applications, Wireless Local Area Networks (WLAN) applications, video applications, Media Provider applications, Filesystem in Userspace (FUSE) applications, etc.

[0072] Among them, the super-resolution application program can be used to obtain a human body image, and perform super-resolution processing on the human body image to obtain a super-resolution human body image.

[0073] Media Provider is used to create multimedia files in FUSE or access multimedia files in FUSE. Each application in the application layer can create multimedia files in FUSE or access multimedia files in FUSE through MediaProvider.

[0074] FUSE is used to store the multimedia files that the media provider creates. Of course, in other embodiments, FUSE also can be used to store other data.

[0075] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include a window manager, content provider, resource manager, view system, package management service (PMS), and activity management service (AMS).

[0076] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0077] Content providers are used to store and retrieve data and make it accessible to applications. Data can include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0078] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.

[0079] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0080] The Package Manager service, as a package manager, is responsible for installing, managing, and uninstalling applications on Android devices. It scans designated directories on the system for files ending in APK, parses these files, and extracts all application information, storing it in packages.xml.

[0081] When a new application is installed, the package management service will identify all components of the application (such as activities, services, and broadcast receivers) and assign corresponding permissions to these components. At the same time, the package management service also monitors the status of installed applications to ensure the integrity and security of the applications.

[0082] The package management service also manages an app's Device Encrypted (DE) and Credential Encrypted (CE) data. The DE data key is only accessible after the system has performed a verified boot. The CE directory encrypts data using a key associated with a user's authentication method (e.g., pattern, password, etc.), which is only accessible after the user has authenticated.

[0083] The CE directory of the application may include the original uid of the application. The package management service is used to execute the DE data and CE data of the super-resolution application in this embodiment.

[0084] The Activity Management Service, as the activity manager service, is primarily responsible for managing and tracking the activity tasks and lifecycles of all applications. When an application is opened, the Activity Management Service starts the application's process and allocates processor resources and memory to it. When the application is no longer in the foreground or background, or when the system runs low on memory, the Activity Management Service terminates or kills the application's process.

[0085] For example, the Activity Management Service can be responsible for managing and tracking the active tasks and lifecycle of a super-resolution application. When a super-resolution application is opened, the Activity Management Service starts the application's process and allocates processor resources and memory to it for super-resolution processing of acquired human images. When the super-resolution application is no longer in the foreground or background, or when the system runs out of memory, the Activity Management Service terminates or kills the application's process.

[0086] System libraries may include Surface Manager, Media Libraries, Android Rruntime, etc.

[0087] The Android runtime consists of core libraries and a virtual machine (VM). The Android runtime is responsible for scheduling and management of the Android system. The core libraries consist of two parts: one containing the Java language's callable functions and the other the Android core library. The application layer and the application framework layer run in the VM. The VM executes the Java files in the application and framework layers as binary files. The VM is responsible for managing object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0088] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0089] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0090] The kernel layer can include modules such as camera driver, display driver, Wi-Fi driver, Bluetooth driver, and audio driver.

[0091] It is understandable that Figure 1 The layers in the illustrated software structure and the components contained in each layer do not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer layers than shown, and each layer may include more or fewer components, and this application does not limit this.

[0092] It should be noted that although the embodiment of the present application is described using the Android system as an example, its basic principles are also applicable to Electronic devices using operating systems such as Harmony.

[0093] See also Figure 2 , which shows a flow chart of a human image super-resolution method provided by an embodiment of the present application. In a specific embodiment, the human image super-resolution method can be applied to electronic devices. Figure 2 The process shown in FIG. 1 is described in detail. The human body image super-resolution method may include the following steps S110 to S130.

[0094] Step S110: Acquire an initial human body image.

[0095] In an embodiment of the present application, when a user needs to super-resolution a human body image, he or she may send a super-resolution instruction to an electronic device, and the electronic device receives and responds to the super-resolution instruction to obtain an initial human body image.

[0096] The human body image is an image of the entire body, and may include a face image and a torso image. The face image may include a head image and a non-head image. The torso image is an image of the body other than the face image in the human body image.

[0097] In some embodiments, when a user needs to super-resolution a human body image, a super-resolution instruction can be sent to an electronic device. The electronic device receives and responds to the super-resolution instruction, obtains an initial scene image containing an initial human body image, performs human body detection on the initial scene image, obtains a human body bounding box, and crops the initial scene image according to the human body bounding box to obtain an initial human body image. This realizes cropping the initial human body image from the initial scene image based on the human body bounding box obtained by human body detection, thereby improving the accuracy of obtaining the initial human body image.

[0098] As an embodiment, the electronic device may include a camera. When a user needs to super-resolution a human body image, the user may send a super-resolution instruction to the electronic device. The electronic device receives and responds to the super-resolution instruction, controls the camera to capture an environmental image of the human body's environment, obtains an initial scene image, performs human body detection on the initial scene image, obtains a human body bounding box, and crops the initial scene image according to the human body bounding box to obtain an initial human body image. The initial scene image includes the initial human body image.

[0099] The camera may be any one of a wide-angle camera, a macro camera, an ultra-wide-angle camera, a panoramic camera, etc. The type of camera is not limited here and can be set according to actual needs.

[0100] As an embodiment, the electronic device can be connected to the camera via a network and exchange data with the camera via the network. When a user needs to super-resolution a human body image, the electronic device can send a super-resolution instruction to the electronic device. The electronic device receives and responds to the super-resolution instruction, sends an acquisition instruction to the camera via the network, and the camera receives and responds to the acquisition instruction. It captures an environmental image of the human body's environment to obtain an initial scene image, sends the initial scene image to the electronic device via the network, and the electronic device receives the initial scene image returned by the camera. It performs human body detection on the initial scene image to obtain a human body bounding box, and crops the initial scene image according to the human body bounding box to obtain an initial human body image.

[0101] Among them, the network can be any one of a ZigBee network, a Bluetooth (BT) network, a Wireless Fidelity (Wi-Fi) network, a Thread network, a Long Range Radio (LoRa) network, a Low-Power Wide-Area Network (LPWAN), an infrared network, a Narrow Band Internet of Things (NB-IoT), a Controller Area Network (CAN), a Digital Living Network Alliance (DLNA) network, a Wide Area Network (WAN), a Local Area Network (LAN), a Metropolitan Area Network (MAN) or a Wireless Personal Area Network (WPAN), etc. The type of network is not limited here and can be set according to actual needs.

[0102] In some embodiments, the electronic device pre-stores an initial human body image. When the user needs to super-resolution the human body image, the user can send a super-resolution instruction to the electronic device, which receives and responds to the super-resolution instruction and reads the pre-stored initial human body image.

[0103] In some embodiments, the server pre-stores an initial human body image, and the server is connected to the electronic device via a network, and performs data exchange with the electronic device via the network.

[0104] When a user needs to super-resolution a human body image, he or she can send a super-resolution instruction to an electronic device. The electronic device receives and responds to the super-resolution instruction, sends an acquisition instruction to a server via the network, the server receives and responds to the acquisition instruction, sends a pre-stored initial human body image to the electronic device via the network, and the electronic device receives the initial human body image returned by the server.

[0105] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be any cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), big data or artificial intelligence platforms, etc. The type of server is not limited here and can be set according to actual needs.

[0106] In some embodiments, when a user needs to super-resolution a human body image, a super-resolution instruction can be sent to an electronic device. The electronic device receives and responds to the super-resolution instruction, generates an upload prompt message, and receives the initial human body image uploaded by the user according to the upload prompt message.

[0107] The upload prompt information can be used to prompt the user to upload the initial human body image to the electronic device according to the upload prompt information. The upload prompt information can be at least one of a sound prompt information, a text prompt information, or a light prompt information. The type of upload prompt information is not limited here and can be set according to actual needs.

[0108] In some embodiments, when a user needs to super-resolution a human body image, he or she may send a super-resolution instruction carrying an initial human body image to an electronic device. The electronic device receives and responds to the super-resolution instruction and obtains the initial human body image according to the super-resolution instruction.

[0109] In some embodiments, the electronic device may be provided with an input panel. When a user needs to perform super-resolution of a human body image, the user may input a super-resolution instruction based on the input panel of the electronic device, for example, by handwriting the super-resolution instruction on the input panel of the electronic device, or by pressing a button on the input panel of the electronic device to input the super-resolution instruction. The electronic device receives the super-resolution instruction through the input panel.

[0110] In some embodiments, the electronic device may be provided with a voice recognition module. When a user needs to perform super-resolution of a human body image, the user may send voice information within the voice collection range of the voice recognition module. The voice recognition module collects the voice information sent by the user and performs voice recognition on the collected voice information to obtain a voice recognition result. When it is determined that the voice recognition result contains keywords for instructing super-resolution of a human body image, for example, the keyword is "human body image super-resolution", or for example, the keywords are "human body image" and "super-resolution", etc., it is determined that the super-resolution instruction has been received.

[0111] As an example, the voice information sent by the user is: super-resolution of human body images, and the voice recognition result of the voice recognition contains the keywords "human body image" and "super-resolution", then it is determined that the super-resolution instruction is received.

[0112] In some implementations, the client is connected to the electronic device via a network, and performs data exchange with the electronic device via the network.

[0113] When a user needs to super-resolution a human body image, he or she can send a super-resolution instruction to the client. The client receives and responds to the super-resolution instruction, forwards the super-resolution instruction to the electronic device through the network, and the electronic device receives the super-resolution instruction forwarded by the client.

[0114] Among them, the client can be any one of a mobile client (for example, a mobile phone client, a personal digital assistant (PDA) client, a tablet computer (Tablet Personal Computer, Tablet PC) client, a laptop client, a smart watch client, a smart bracelet client or a wearable client, etc.) or a fixed client (for example, a desktop computer client, a smart panel client, etc.). The type of client is not limited here and can be set according to actual needs.

[0115] Step S120: performing a first super-resolution process on the initial human body image using the super-resolution model to obtain a first super-resolution human body image.

[0116] In an embodiment of the present application, after the electronic device obtains the initial human body image, it can perform a first super-resolution process on the initial human body image through a super-resolution model to obtain a first super-resolution human body image.

[0117] Among them, the super-resolution model can include a face super-resolution diffusion model and a human body super-resolution diffusion model. The face super-resolution diffusion model can be used to perform super-resolution processing on face images, and the human body super-resolution diffusion model can be used to perform super-resolution processing on human body images. Figure 3 As shown, the process of the first super-resolution processing may include steps S121 to S124.

[0118] Step S121: obtaining an initial face feature map and an initial body feature map according to the initial body image.

[0119] In an embodiment of the present application, the electronic device can determine the initial human body image as the initial human body feature map, and crop the initial human body feature map according to the face frame of the initial human body feature map to obtain the initial face feature map, thereby realizing the cropping of the initial human body feature map based on the face frame. The cropped initial face feature map always contains facial features, thereby improving the accuracy of obtaining the initial face feature map.

[0120] In some embodiments, the electronic device may determine the initial human body image as the initial human body feature map, and perform face detection on the initial human body image to obtain first face frame coordinates, and determine the second face frame coordinates of the initial human body feature map based on the first face frame coordinates and the mapping relationship between the initial human body image and the initial human body feature map, and determine the face frame of the initial human body feature map based on the second face frame coordinates, and crop the initial human body feature map based on the face frame of the initial human body feature map to obtain the initial face feature map, thereby realizing the determination of the face frame of the initial human body feature map based on the first face frame coordinates of the initial human body image and the mapping relationship between the initial human body image and the initial human body feature map, and the obtained face frame has high accuracy.

[0121] In some embodiments, the electronic device can compress the initial human body image to obtain a compressed human body image, and obtain an initial facial feature map and an initial human body feature map based on the compressed human body image. By compressing the initial human body image and performing super-resolution based on the compressed human body image, the size of data involved in the super-resolution process is reduced, and the super-resolution efficiency of the human body image is improved.

[0122] The compression processing may be a latent space coding processing, and the compressed human body image may be a latent space human body image.

[0123] The electronic device can determine the compressed human body image as the initial human body feature map, and crop the initial human body feature map according to the face frame of the initial human body feature map to obtain the initial face feature map. The initial human body feature map is cropped based on the face frame. The cropped initial face feature map always contains facial features, thereby improving the accuracy of obtaining the initial face feature map.

[0124] The electronic device can perform face detection on the initial human body image to obtain the first face frame coordinates, and determine the second face frame coordinates of the initial human body feature map based on the first face frame coordinates and the mapping relationship between the initial human body image and the compressed human body image, and determine the face frame of the initial human body feature map based on the second face frame coordinates, thereby realizing the determination of the face frame of the initial human body feature map based on the first face frame coordinates of the initial human body image and the mapping relationship between the initial human body image and the compressed human body image, and the obtained face frame has high accuracy.

[0125] Step S122: inputting the initial facial feature map into the face super-resolution diffusion model to obtain a first super-resolution face image.

[0126] In an embodiment of the present application, the electronic device may input an initial facial feature map into a facial super-resolution diffusion model, which receives and responds to the initial facial feature map and outputs a first super-resolution facial image. The resolution of the first super-resolution facial image is higher than the resolution of the initial human body image.

[0127] In some embodiments, the initial facial feature map may include multiple sub-facial feature maps. The electronic device may perform sliding window block processing on the initial facial feature map to obtain multiple sub-facial feature maps, and input the multiple sub-facial feature maps one by one into the face super-resolution diffusion model. The face super-resolution diffusion model receives and responds to the multiple sub-facial feature maps and outputs a first super-resolution face image. By performing sliding window block processing on the initial facial feature map into multiple sub-facial feature maps and inputting the multiple sub-facial feature maps one by one into the face super-resolution diffusion model for face super-resolution, the image data of the facial feature map input into the face super-resolution diffusion model at a single time is avoided, which leads to an increase in the amount of data calculation for face super-resolution, and improves the super-resolution efficiency of face feature map super-resolution.

[0128] The sliding window block processing process is the process of sliding window processing on the image according to the preset pixel block size and pixel overlap length. For example, the preset pixel block size can be 64*64, and the pixel overlap length can be 16; the preset pixel block size can also be 25*25, and the pixel overlap length can be 5, etc. The size of the preset pixel block size and the pixel overlap length are not limited here and can be set according to actual needs. Figure 4 As shown, it shows a schematic diagram of a scene of sliding window block processing of images.

[0129] Step S123: inputting the initial human body feature map into the human body super-resolution diffusion model to obtain an initial super-resolution human body image.

[0130] In an embodiment of the present application, the electronic device can input an initial human body feature map to the human body super-resolution diffusion model, and the human body super-resolution diffusion model receives and responds to the initial human body feature map and outputs an initial super-resolution human body image.

[0131] The resolution of the first super-resolution face image is higher than the resolution of the initial super-resolution body image, and the resolution of the initial super-resolution body image is higher than the resolution of the initial body image.

[0132] In some embodiments, the initial human body feature map may include multiple sub-human body feature maps. The electronic device may perform sliding window block processing on the initial human body feature map to obtain multiple sub-human body feature maps, and input the multiple sub-human body feature maps one by one into the human body super-resolution diffusion model. The human body super-resolution diffusion model receives and responds to the multiple sub-human body feature maps and outputs an initial super-resolution human body image. By sliding window block processing of the initial human body feature map into multiple sub-human body feature maps, and inputting the multiple sub-human body feature maps one by one into the human body super-resolution diffusion model for human body super-resolution, the image data of the human body feature map input into the human body super-resolution diffusion model at a single time is avoided, which leads to an increase in the amount of data calculation for human body super-resolution, and improves the super-resolution efficiency of the human body feature map.

[0133] Step S124: Fusing the first super-resolved face image and the initial super-resolved body image to obtain a first super-resolved body image.

[0134] In an embodiment of the present application, the electronic device can fuse the first super-resolution face image and the initial super-resolution body image to obtain a first super-resolution body image, and the resolution of the first super-resolution body image is higher than the resolution of the initial body image.

[0135] Specifically, the electronic device can select a second super-resolution face image corresponding to the second face frame coordinates in the initial super-resolution human body image, and replace the second super-resolution face image with the first super-resolution face image in the initial super-resolution human body image to obtain a first super-resolution human body image, thereby realizing the first super-resolution face image obtained based on face super-resolution, and updating the face area image in the initial super-resolution human body image obtained based on human body super-resolution. The obtained first super-resolution human body image meets both the face clarity requirements and the human body clarity requirements, thereby improving the super-resolution experience of human body images.

[0136] In some embodiments, the first super-resolved face image may include a first super-resolved head image and a first super-resolved non-head image, and the initial super-resolved body image may include an initial super-resolved head image and an initial super-resolved non-head image.

[0137] The electronic device can select the initial super-resolution human head image corresponding to the second face frame coordinates in the initial super-resolution human body image, and replace the initial super-resolution human head image with the first super-resolution human head image in the initial super-resolution human body image to obtain the first super-resolution human body image, thereby realizing the first super-resolution human head image obtained based on face super-resolution and updating the initial super-resolution human head image in the initial super-resolution human body image obtained based on human body super-resolution, reducing the amount of data calculation in the super-resolution processing process and improving the super-resolution efficiency of human body image super-resolution.

[0138] Step S130: Use the first super-resolution human body image obtained after the first super-resolution processing as the initial human body image in the second super-resolution processing, and repeat the process of the first super-resolution processing until the T-th super-resolution processing is completed, thereby obtaining the first super-resolution human body image of the T-th super-resolution processing.

[0139] It can be understood that in the embodiment of the present application, the electronic device can perform T super-resolution processing, and the process of each super-resolution processing is similar to that of the first super-resolution processing. However, for other super-resolution processing except the first super-resolution processing, the result obtained from the previous super-resolution processing is used as the input for the next super-resolution processing, and this iterative process is repeated until the T-th super-resolution processing is completed, and the first super-resolution human body image of the T-th super-resolution processing is obtained.

[0140] In an embodiment of the present application, the electronic device can use the first super-resolution human body image obtained after the first super-resolution processing as the initial human body image in the second super-resolution processing, repeatedly perform the first super-resolution processing process T times, until the T-th super-resolution processing is completed, and obtain the first super-resolution human body image of the T-th super-resolution processing, so that in the human body image super-resolution process, the face noise of the initial human body image is learned based on the face super-resolution diffusion model and the human body noise of the initial human body image is learned based on the human body super-resolution diffusion model, and the face noise and the human body noise are fused to obtain the first super-resolution human body image, the first super-resolution human body image is iteratively super-resolution processed, the face super-resolution diffusion model iteratively learns the face noise in the previous face super-resolution image, and the human body super-resolution diffusion model iteratively learns the human body noise in the previous human body super-resolution image, and finally makes the transition at the junction of the face area image and the human body area image in the super-resolution human body image obtained by the fusion of multiple super-resolution iterations more natural.

[0141] Wherein, T is a positive integer greater than or equal to 2.

[0142] In some embodiments, the electronic device can decode the first super-resolution human body image of the T-th super-resolution processing to obtain a target super-resolution human body image, and the resolution of the target super-resolution human body image is higher than the resolution of the initial human body image. By decoding the first super-resolution human body image of the T-th super-resolution processing, the obtained target super-resolution human body image has the same dimension as the initial human body image, thereby improving the image quality of the target super-resolution human body image.

[0143] In some embodiments, the electronic device can segment the first super-resolution human body image of the Tth super-resolution processing to obtain a human body mask image, and paste the human body mask image back to the position corresponding to the initial human body image in the initial scene image to obtain a target scene image, thereby achieving super-resolution processing of only the initial human body image in the initial scene image, avoiding super-resolution processing of the entire initial scene image resulting in low super-resolution efficiency, and improving the super-resolution efficiency of super-resolution of human body images.

[0144] In some embodiments, the electronic device can perform corrosion processing on the boundary of the human body mask image in the target scene image. By performing corrosion processing on the boundary of the super-resolution human body image in the target scene image, the unnatural transition between the boundary of the super-resolution human body image and the initial scene image can be reduced, and the fusion degree of the super-resolution human body image and the initial scene image can be improved.

[0145] In some embodiments, the electronic device can perform Gaussian blur processing on the boundary of the human body mask image in the target scene image. By performing Gaussian blur processing on the boundary of the super-resolution human body image in the target scene image, the unnatural transition between the boundary of the super-resolution human body image and the initial scene image can be reduced, and the fusion degree of the super-resolution human body image and the initial scene image can be improved.

[0146] In one application scenario, such as Figure 5 As shown, the human body image super-resolution method may include steps S201 to S210.

[0147] Step S201: Acquire an initial scene image.

[0148] Step S202: performing human body detection and face detection on the initial scene image to obtain a human body bounding box and a face bounding box.

[0149] Step S203: Obtain a latent space feature map based on the human body bounding box and the initial scene image.

[0150] Specifically, the initial scene image is cropped based on the human body bounding box to obtain an initial human body image, and the initial human body image is input into the encoder ε for encoding to obtain a latent space feature map.

[0151] Step S204: performing sliding window block processing on the latent space feature map to obtain multiple sub-latent space feature maps.

[0152] Step S205: Obtain human super-resolution image noise ∈ based on multiple sub-latent space feature maps g And face super-resolution image noise ∈ f .

[0153] Specifically, multiple sub-latent space feature maps are input into the human super-resolution diffusion model to obtain the human super-resolution image noise ∈ g , input the sub-latent space feature maps involving the face area in multiple sub-latent space feature maps into the face super-resolution diffusion model, and obtain the face super-resolution image noise ∈ f .

[0154] Step S206: Fusion of human super-resolution image noise ∈ g And face super-resolution image noise ∈ f , and get the super-resolved human body noise ∈.

[0155] Specifically, based on the mapping coordinates of the face bounding box in the latent space, the human super-resolution image noise ∈ g The noise of the face part is replaced by the face super-resolution image noise ∈ f , and get the super-resolved human body noise ∈.

[0156] Step S207: Repeat steps S204 to S206 T times to obtain the target super-resolved human body noise z0.

[0157] Specifically, from the second time to the Tth time, the super-resolved human body noise ∈ outputted in the previous time is used as the input of the next sliding window block processing to obtain the target super-resolved human body noise z0.

[0158] Step S208: Decode the target super-resolved human body noise z0 to obtain a target super-resolved human body image.

[0159] Specifically, the target super-resolved human body noise z0 is input to the decoder D, and the decoder D receives and responds to the target super-resolved human body noise z0, decodes the target super-resolved human body noise z0, and obtains the target super-resolved human body image.

[0160] Step S209: Fusing the target super-resolved human body image and the initial scene image to obtain the current scene image.

[0161] Specifically, the target super-resolved human body image is segmented to obtain a human body mask image, and the human body mask image is pasted back to the position corresponding to the initial human body image in the initial scene image to obtain the current scene image.

[0162] Step S210: performing corrosion processing and Gaussian blur processing on the boundary of the human body mask image in the target scene image to obtain the target scene image.

[0163] It should be noted that if Figure 6 As shown, when the initial scene image includes multiple human body images, steps S201 to S210 are performed on each human body image to obtain a target scene image.

[0164] Both ends of the diffusion model and the control network model are connected to an autoencoder ε, which adopts a variational autoencoder. The variational autoencoder downsamples the image through a coding factor f, where f is a coding factor value preset by the user, preferably f=8.

[0165] In addition, both the diffusion model and the control network model take text input and encode the user-entered text information into a text embedding through a text encoder. The text encoder can be the text encoder used in contrastive language-image pretraining (CLIP). The diffusion model uses a U-Net structure, which outputs a noise estimate.

[0166] This embodiment provides a solution, which obtains a first super-resolution human body image by obtaining an initial human body image and performing a first super-resolution processing on the initial human body image through a super-resolution model, wherein the super-resolution model includes a face super-resolution diffusion model and a human body super-resolution diffusion model, and the process of the first super-resolution processing includes obtaining an initial face feature map and an initial human body feature map according to the initial human body image, and inputting the initial face feature map into the face super-resolution diffusion model to obtain a first super-resolution face image, and inputting the initial human body feature map into the human body super-resolution diffusion model to obtain an initial super-resolution human body image, the resolution of the first super-resolution face image is higher than the resolution of the initial super-resolution human body image, and the first super-resolution face image and the initial super-resolution human body image are fused to obtain a first super-resolution human body image, and the first super-resolution human body image obtained after the first super-resolution processing is used as The initial human body image in the second super-resolution processing repeats the process of the first super-resolution processing until the T-th super-resolution processing is completed, and the first super-resolution human body image of the T-th super-resolution processing is obtained. In the process of human body image super-resolution, the face noise of the initial human body image is learned based on the face super-resolution diffusion model and the human body noise of the initial human body image is learned based on the human body super-resolution diffusion model, and the face noise and the human body noise are fused to obtain the first super-resolution human body image. The first super-resolution human body image is iteratively super-resolution processed, the face super-resolution diffusion model iteratively learns the face noise in the previous face super-resolution image, and the human body super-resolution diffusion model iteratively learns the human body noise in the previous human body super-resolution image, and finally makes the transition at the junction of the face area image and the human body area image in the super-resolution human body image obtained by multiple super-resolution iterative fusion more natural.

[0167] See also Figure 7 , which shows a flow chart of a super-resolution model construction method provided by an embodiment of the present application. In a specific embodiment, the super-resolution model construction method can be applied to electronic devices. Figure 7 The process shown in FIG3 is described in detail, and the super-resolution model construction method may include the following steps S310 to S330.

[0168] Step S310: constructing a super-resolution diffusion model based on the diffusion model and the control network model.

[0169] In this embodiment, the electronic device may fuse the diffusion model and the control network model to obtain a super-resolution diffusion model.

[0170] Among them, the diffusion model is a generative model that can learn to approximate unknown data distribution given independent and identically distributed sample data from the location data distribution.

[0171] The diffusion model mainly includes two processes: noise addition process and noise removal process, such as Figure 8The denoising process is the process of gradually adding Gaussian noise to the real image in the dataset; the denoising process is the process of gradually removing the noise from the noisy image to restore the real image.

[0172] As an example, given a data point x0 with probability distribution q(x0), the noise addition process gradually destroys the data structure of the data point x0 by repeatedly applying the following Markov diffusion kernel.

[0173] The Markov diffusion kernel is: Where t∈{1, 2,…, T}, is a predefined or learned noise variance change. Reasonable design can theoretically guarantee q(x t ) converges to a Gaussian distribution on the unit sphere.

[0174] The marginal distribution for any time t has the following analytical form:

[0175] The goal of the denoising process is to t to x t-1 Learn a transfer kernel defined as the following Gaussian distribution:

[0176] p θ (x t-1 x t )=N(x t-1 ;μ θ (x t , t), ∑ θ (x t ,t)), where θ is a learnable parameter.

[0177] Through such a learned transfer kernel, the data distribution q(x0) can be approximated by the following marginal distribution, marginal distribution where p(x T )=N(x T ; 0, I).

[0178] In one application scenario, such as Figure 9As shown, it shows a structural schematic diagram of the super-resolution diffusion model, which includes a diffusion model and a control network model. The diffusion model is a stable diffusion model Stable Diffusion, which includes a first autoencoder Auto Encoder_1, a Gaussian noise adder Add Gaussian Noise, a convolution layer Convolution Layer, a text encoder Text Encoder, a time encoder Time Encoder, multiple first encoder blocks SD Encoder Block, a first intermediate block SD Middle Block, multiple decoder blocks SD DecoderBlock and an autodecoder Auto Decoder. The control network model includes a second autoencoder Auto Encoder_2, multiple zero convolution layers Zero Convolution, multiple second encoder blocks and a second intermediate block.

[0179] Multiple first encoder blocks SD Encoder Block, first intermediate blocks SD Middle Block, and multiple decoder blocks SD Decoder Block form a Unet model, and each first encoder block can include multiple residual networks (Resnet).

[0180] During the training process of the super-resolution diffusion model, Auto Encoder_1 is used to encode the input high-definition image HRImage, Text Encoder is used to encode the input prompt word Prompt, Time Encoder is used to encode time Time, and the second auto encoder Auto Encoder_2 is used to encode the input low-definition image LR Image.

[0181] During the training process of the super-resolution diffusion model, the parameters of the stable diffusion model are frozen, and the stable diffusion model is not trained. The main training is to control the parameters of the network model to obtain the target super-resolution diffusion model for inference, such as Figure 10 shown.

[0182] Step S320: Based on the first historical face feature map and the second historical face feature map, the super-resolution diffusion model is trained to obtain a face super-resolution diffusion model.

[0183] In this embodiment, the electronic device can obtain a first historical facial feature map and a second historical facial feature map, and based on the first historical facial feature map and the second historical facial feature map, train the super-resolution diffusion model to obtain a facial super-resolution diffusion model, and construct the super-resolution diffusion model based on the diffusion model and the control network model. The control network model serves as a conditional control generation model of the super-resolution diffusion model, so that the super-resolution process of the super-resolution of the facial feature map by the super-resolution diffusion model is controllable, which can improve the user's super-resolution experience of the facial feature map.

[0184] The first historical facial feature map and the second historical facial feature map contain the same facial features, and the resolution of the first historical facial feature map is higher than the resolution of the second historical facial feature map.

[0185] Specifically, the electronic device can obtain the first historical face feature map and the second historical face feature map, and input the first historical face feature map to the diffusion model in the super-resolution diffusion model. The diffusion model receives and responds to the first historical face feature map, outputs the first historical face feature to the control network model, and the control network model receives the first historical face feature output by the diffusion model. The electronic device inputs the second historical face feature map to the control network model in the super-resolution diffusion model. The control network model receives the second historical face feature map input by the electronic device, and is trained according to the first historical face feature and the second historical face feature map to obtain a super-resolution diffusion model of the face. This realizes the training of the control network model in the super-resolution diffusion model based on the historical face feature map, and there is no need to train the diffusion model in the super-resolution diffusion model. This reduces the amount of data calculation in the training process of the super-resolution diffusion model and improves the training speed of the super-resolution diffusion model.

[0186] Step S330: Based on the first historical human body feature map and the second historical human body feature map, the super-resolution diffusion model is trained to obtain a human body super-resolution diffusion model.

[0187] In this embodiment, the electronic device can obtain a first historical human body feature map and a second historical human body feature map, and based on the first historical human body feature map and the second historical human body feature map, train the super-resolution diffusion model to obtain a human body super-resolution diffusion model, and construct the super-resolution diffusion model based on the diffusion model and the control network model. The control network model serves as a conditional control generation model of the super-resolution diffusion model, so that the super-resolution process of the human body feature map by the super-resolution diffusion model is controllable, which can improve the user's super-resolution experience of the human body feature map.

[0188] The first historical human body feature map and the second historical human body feature map contain the same human body features, and the resolution of the first historical human body feature map is higher than the resolution of the second historical human body feature map.

[0189] Specifically, the electronic device can obtain the first historical human body feature map and the second historical human body feature map, and input the first historical human body feature map to the diffusion model in the super-resolution diffusion model. The diffusion model receives and responds to the first historical human body feature map, outputs the first historical human body feature to the control network model, and the control network model receives the first historical human body feature output by the diffusion model. The electronic device inputs the second historical human body feature map to the control network model in the super-resolution diffusion model. The control network model receives the second historical human body feature map input by the electronic device, and is trained according to the first historical human body feature and the second historical human body feature map to obtain a super-resolution diffusion model of the human body. This realizes the training of the control network model in the super-resolution diffusion model based on the historical human body feature map, and there is no need to train the diffusion model in the super-resolution diffusion model. This reduces the amount of data calculation in the training process of the super-resolution diffusion model and improves the training speed of the super-resolution diffusion model.

[0190] This embodiment provides a solution, which constructs a super-resolution diffusion model based on a diffusion model and a control network model, and trains the super-resolution diffusion model based on a first historical face feature map and a second historical face feature map to obtain a face super-resolution diffusion model, and trains the super-resolution diffusion model based on a first historical body feature map and a second historical body feature map to obtain a body super-resolution diffusion model. This realizes the construction of a super-resolution diffusion model based on the diffusion model and the control network model, and the control network model serves as a conditional control generation model of the super-resolution diffusion model, so that the super-resolution process of the super-resolution diffusion model for face feature maps and for body feature maps is controllable, which can improve the user's super-resolution experience.

[0191] See also Figure 11 , which shows a human image super-resolution device 600 provided by an embodiment of the present application. The human image super-resolution device 600 can be applied to electronic devices. The following takes electronic devices as an example. Figure 11 The human body image super-resolution device 600 shown in FIG. 1 is described in detail. The human body image super-resolution device 600 may include an initial image acquisition module 610 , a super-resolution processing module 620 and a repeated execution module 630 .

[0192] The initial image acquisition module 610 can be used to acquire an initial human body image; the super-resolution processing module 620 can be used to perform a first super-resolution processing on the initial human body image through a super-resolution model to obtain a first super-resolution human body image; the repeated execution module 630 can be used to use the first super-resolution human body image obtained after the first super-resolution processing as the initial human body image in the second super-resolution processing, and repeatedly execute the first super-resolution processing process until the T-th super-resolution processing is completed, thereby obtaining the first super-resolution human body image of the T-th super-resolution processing.

[0193] Among them, the super-resolution model can include a face super-resolution diffusion model and a human body super-resolution diffusion model, and the super-resolution processing module 620 can include a first acquisition unit, a first input unit, a second input unit and a fusion unit.

[0194] The first acquisition unit can be used to obtain an initial facial feature map and an initial human body feature map based on the initial human body image; the first input unit can be used to input the initial facial feature map to the face super-resolution diffusion model to obtain a first super-resolution facial image; the second input unit can be used to input the initial human body feature map to the human body super-resolution diffusion model to obtain an initial super-resolution human body image, and the resolution of the first super-resolution facial image is higher than the resolution of the initial super-resolution human body image; the fusion unit can be used to fuse the first super-resolution facial image and the initial super-resolution human body image to obtain a first super-resolution human body image.

[0195] In some embodiments, the first acquisition unit may include a compression subunit and an acquisition subunit.

[0196] The compression subunit can be used to compress the initial human body image to obtain a compressed human body image; the acquisition subunit can be used to obtain an initial face feature map and an initial human body feature map based on the compressed human body image.

[0197] In some implementations, obtaining the subunit may include determining a secondary subunit and trimming the secondary subunit.

[0198] The determination sub-sub-unit can be used to determine the compressed human body image as the initial human body feature map; the cropping sub-sub-unit can be used to crop the initial human body feature map according to the face frame of the initial human body feature map to obtain the initial face feature map.

[0199] In some embodiments, the human image super-resolution apparatus 600 may further include a detection module, a first determination module, and a second determination module.

[0200] The detection module can be used to crop the secondary sub-unit according to the face frame of the initial human feature map, crop the initial human feature map, and perform face detection on the initial human image before obtaining the initial face feature map to obtain the first face frame coordinates; the first determination module can be used to determine the second face frame coordinates of the initial human feature map based on the first face frame coordinates and the mapping relationship between the initial human image and the initial human feature map; the second determination module can be used to determine the face frame of the initial human feature map based on the second face frame coordinates.

[0201] In some embodiments, fusing a unit may include selecting a subunit and replacing a subunit.

[0202] The selection subunit can be used to select a second super-resolution face image corresponding to the second face frame coordinates in the initial super-resolution face image; the replacement subunit can be used to replace the second super-resolution face image with the first super-resolution face image in the initial super-resolution face image to obtain the first super-resolution face image.

[0203] In some embodiments, the first super-resolved face image may include a first super-resolved head image and a first super-resolved non-head image, the initial super-resolved body image may include an initial super-resolved head image and an initial super-resolved non-head image; the replacement sub-unit may include a replacement sub-sub-unit.

[0204] The replacement sub-subunit may be configured to replace the initial super-resolved human head image with the first super-resolved human head image in the initial super-resolved human body image to obtain the first super-resolved human body image.

[0205] In some implementations, the human image super-resolution apparatus 600 may further include a decoding module.

[0206] The decoding module can be used to decode the first super-resolution human body image obtained by the T-th super-resolution processing to obtain a target super-resolution human body image, and the resolution of the target super-resolution human body image is higher than the resolution of the initial human body image.

[0207] In some embodiments, the initial facial feature map may include multiple sub-facial feature maps, and the human image super-resolution device 600 may further include a first sliding window module.

[0208] The first sliding window module can be used for the first input unit to input the initial face feature map to the face super-resolution diffusion model, and before obtaining the first super-resolution face image, the initial face feature map is subjected to sliding window block processing to obtain multiple sub-face feature maps.

[0209] In some embodiments, the first input unit may include a first input sub-unit.

[0210] The first input sub-unit can be used to input multiple sub-face feature maps to the face super-resolution diffusion model one by one, so that the face super-resolution diffusion model outputs a first super-resolution face image according to the multiple sub-face feature maps.

[0211] In some embodiments, the initial human feature map may include multiple sub-human feature maps, and the human image super-resolution device 600 may further include a second sliding window module.

[0212] The second sliding window module can be used for the second input unit to input the initial human feature map to the human super-resolution diffusion model, and before obtaining the initial super-resolution human image, the initial human feature map is subjected to sliding window block processing to obtain multiple sub-human feature maps.

[0213] In some embodiments, the second input unit may include a second input sub-unit.

[0214] The second input sub-unit can be used to input multiple sub-human feature maps to the human body super-resolution diffusion model one by one, so that the human body super-resolution diffusion model outputs an initial super-resolution human body image according to the multiple sub-human feature maps.

[0215] In some embodiments, the human image super-resolution apparatus 600 may further include a construction module, a first training module, and a second training module.

[0216] The construction module can be used for the super-resolution processing module 620 to perform the first super-resolution processing on the initial human body image through the super-resolution model, and before obtaining the first super-resolution human body image, a super-resolution diffusion model is constructed based on the diffusion model and the control network model; the first training module can be used to train the super-resolution diffusion model based on the first historical face feature map and the second historical face feature map to obtain a face super-resolution diffusion model, the first historical face feature map and the second historical face feature map contain the same facial features, and the resolution of the first historical face feature map is higher than the resolution of the second historical face feature map; the second training module can be used to train the super-resolution diffusion model based on the first historical human body feature map and the second historical human body feature map to obtain a human body super-resolution diffusion model, the first historical human body feature map and the second historical human body feature map contain the same human body features, and the resolution of the first historical human body feature map is higher than the resolution of the second historical human body feature map.

[0217] In some embodiments, the first training module may include a third input unit and a fourth input unit.

[0218] The third input unit can be used to input the first historical face feature map into the diffusion model in the super-resolution diffusion model, so that the diffusion model outputs the first historical face feature to the control network model according to the first historical face feature map; the fourth input unit can be used to input the second historical face feature map into the control network model in the super-resolution diffusion model, so that the control network model is trained according to the first historical face feature and the second historical face feature map to obtain a face super-resolution diffusion model.

[0219] In some implementations, the second training module may include a fifth input unit and a sixth input unit.

[0220] The fifth input unit can be used to input the first historical human body feature map to the diffusion model in the super-resolution diffusion model, so that the diffusion model outputs the first historical human body feature to the control network model according to the first historical human body feature map; the sixth input unit can be used to input the second historical human body feature map to the control network model in the super-resolution diffusion model, so that the control network model is trained according to the first historical human body feature and the second historical human body feature map to obtain a super-resolution diffusion model of the human body.

[0221] In some embodiments, the initial image acquisition module 610 may include a second acquisition unit, a detection unit, and a cropping unit.

[0222] The second acquisition unit can be used to acquire an initial scene image containing an initial human body image; the detection unit can be used to perform human body detection on the initial scene image to obtain a human body bounding box; and the cropping unit can be used to crop the initial scene image according to the human body bounding box to obtain an initial human body image.

[0223] In some embodiments, the human body image super-resolution device 600 may further include a segmentation module and a pasting module.

[0224] The segmentation module can be used to segment the first super-resolution human body image of the Tth super-resolution processing to obtain a human body mask image; the pasting module can be used to paste the human body mask image back to the position corresponding to the initial human body image in the initial scene image to obtain a target scene image.

[0225] In some embodiments, the human body image super-resolution apparatus 600 may further include an erosion module.

[0226] The erosion module can be used to erode the boundary of the human body mask image in the target scene image.

[0227] In some implementations, the human image super-resolution apparatus 600 may further include a blur module.

[0228] The blur module can be used to perform Gaussian blur processing on the boundary of the human mask image in the target scene image.

[0229] The solution provided by this embodiment obtains an initial human body image and performs a first super-resolution processing on the initial human body image through a super-resolution model to obtain a first super-resolution human body image, wherein the super-resolution model includes a face super-resolution diffusion model and a human body super-resolution diffusion model. The process of the first super-resolution processing includes obtaining an initial face feature map and an initial human body feature map according to the initial human body image, inputting the initial face feature map to the face super-resolution diffusion model to obtain a first super-resolution face image, and inputting the initial human body feature map to the human body super-resolution diffusion model to obtain an initial super-resolution human body image. The resolution of the first super-resolution face image is higher than the resolution of the initial super-resolution human body image, and the first super-resolution face image and the initial super-resolution human body image are fused to obtain a first super-resolution human body image, and the first super-resolution human body image obtained after the first super-resolution processing is used as For the initial human body image in the second super-resolution processing, the process of the first super-resolution processing is repeated until the T-th super-resolution processing is completed, and the first super-resolution human body image of the T-th super-resolution processing is obtained. In the process of human body image super-resolution, the face noise of the initial human body image is learned based on the face super-resolution diffusion model and the human body noise of the initial human body image is learned based on the human body super-resolution diffusion model, and the face noise and the human body noise are fused to obtain the first super-resolution human body image. The first super-resolution human body image is iteratively super-resolution processed, the face super-resolution diffusion model iteratively learns the face noise in the previous face super-resolution image, and the human body super-resolution diffusion model iteratively learns the human body noise in the previous human body super-resolution image, and finally makes the transition at the junction of the face area image and the human body area image in the super-resolution human body image obtained by multiple super-resolution iterative fusion more natural.

[0230] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from the other embodiments. The same or similar parts between the various embodiments can be referred to in detail. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. Any processing method described in the method embodiment can be implemented by the corresponding processing module in the device embodiment, and will not be repeated in detail in the device embodiment.

[0231] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0232] See also Figure 12 , which shows a hardware structure diagram of an electronic device 700 provided by an embodiment of the present application. Figure 12As shown, the electronic device 700 may include a processor 710, an external memory interface 720, an internal memory 721, a Universal Serial Bus (USB) interface 730, a charging management module 740, a power management module 741, a battery 742, an antenna 1, an antenna 2, a mobile communication module 750, a wireless communication module 760, an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, an earphone interface 770D, a sensor module 780, a button 790, a motor 791, an indicator 792, a camera 793, a display screen 794, and a Subscriber Identification Module (SIM) card interface 795, etc. Among them, the sensor module 780 may include a pressure sensor 780A, a gyroscope sensor 780B, an air pressure sensor 780C, a magnetic sensor 780D, an acceleration sensor 780E, a distance sensor 780F, a proximity light sensor 780G, a fingerprint sensor 780H, a temperature sensor 780J, a touch sensor 780K, an ambient light sensor 780L, a bone conduction sensor 780M, etc.

[0233] It should be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 700. In other embodiments of the present application, the electronic device 700 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0234] For example, Figure 12 The processor 710 shown may include one or more processing units, for example, the processor 710 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processor (NPU), etc. The different processing units may be independent devices or integrated into one or more processors.

[0235] The AP may be used to control and manage the super-resolution application. For example, the AP may control the super-resolution application to perform super-resolution processing on the acquired human body image.

[0236] The controller may be the nerve center and command center of the electronic device 700. The controller may generate an operation control signal based on the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0237] Processor 710 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 710 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 710. If processor 710 needs to use the same instruction or data again, it can directly retrieve it from the memory. This avoids duplicate accesses, reduces processor 710 latency, and thus improves system efficiency.

[0238] In some embodiments, the processor 710 may include one or more interfaces. The interfaces may include an Inter-Integrated Circuit (I2C) interface, an Inter-Integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI), a General Purpose Input / Output (GPIO) interface, a Subscriber Identity Module (SIM) interface, and / or a Universal Serial Bus (USB) interface.

[0239] Electronic device 700 implements display functionality through a GPU, display screen 794, and an application processor. A GPU is a microprocessor for image processing that connects display screen 794 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 710 may include one or more GPUs that execute program instructions to generate or modify display information.

[0240] Display screen 794 is used to display images, videos, etc. Display screen 794 includes a display panel. The display panel can be a liquid crystal display (LCD), organic light emitting diode (OLED), active-matrix organic light emitting diode (AMOLED), flexible light emitting diode (FLED), mini-LED, micro-LED, micro-OLED, quantum dot light emitting diode (QLED), etc. In some embodiments, electronic device 700 may include one or N display screens 794, where N is a positive integer greater than 1.

[0241] The electronic device 700 can realize the shooting function through the ISP, camera 793, video codec, GPU, display screen 794 and application processor.

[0242] The ISP processes data fed back by camera 793. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and transformed into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 793.

[0243] The camera 793 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 700 may include 1 or N cameras 793, where N is a positive integer greater than 1.

[0244] The external memory interface 720 can be used to connect an external memory card, such as a Secure Digital (SD) card, to expand the storage capacity of the electronic device 700. The external memory card communicates with the processor 710 via the external memory interface 720 to implement data storage functions. For example, files such as captured images and videos can be saved on the external memory card.

[0245] The internal memory 721 can be used to store computer executable program codes, and the executable program codes include instructions. The processor 710 executes various functional applications and data processing of the electronic device 700 by running the instructions stored in the internal memory 721. The internal memory 721 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as an audio acquisition function, an image shooting function, etc.), etc. The data storage area can store data (such as audio data, image data) created during the use of the electronic device 700, etc. In addition, the internal memory 721 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash memory (UFS), etc.

[0246] See also Figure 13 , which shows a functional block diagram of an electronic device 800 according to an embodiment of the present application. Figure 13 As shown, the electronic device 800 includes: one or more processors 810 ( Figure 13 Only one processor is shown in the figure) and a memory 820, the memory 820 is coupled to one or more processors 810, the memory 820 is used to store computer program code 830, the computer program code 830 includes computer instructions, and one or more processors 810 call the computer instructions to enable the electronic device 800 to implement the steps in any of the above methods.

[0247] Those skilled in the art will understand that Figure 13 This is merely an example of the electronic device 800 and does not limit the electronic device 800. In practice, the electronic device 800 may include more or fewer components than shown, or may combine certain components or different components. For example, it may also include input and output devices, network access devices, etc. The electronic device 800 may also be the same device as the electronic device 700 described in the above embodiment.

[0248] The processor 810 may be a central processing unit (CPU), other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0249] In some embodiments, the memory 820 may be an internal storage unit of the electronic device 800, such as a hard disk or memory of the electronic device 800. In other embodiments, the memory 820 may also be an external storage device of the electronic device 800, such as a plug-in hard disk equipped on the electronic device 800, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Alternatively, the memory 820 may include both an internal storage unit of the electronic device 800 and an external storage device. The memory 820 is used to store an operating system, application programs, a boot loader, data, and other programs, such as program code of a computer program. The memory 820 may also be used to temporarily store data that has been output or is about to be output.

[0250] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0251] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0252] An embodiment of the present application also provides a chip system, which is applied to a foldable electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the foldable electronic device to implement the steps of any of the above methods.

[0253] In some embodiments, the chip system further includes a memory, which is connected to one or more processors via circuits or wires.

[0254] In some embodiments, the chip system further includes a communication interface.

[0255] An embodiment of the present application also provides a computer-readable storage medium, which includes instructions. When the instructions are executed on a foldable electronic device, the foldable electronic device implements the methods described in the above-mentioned various method embodiments.

[0256] An embodiment of the present application also provides a computer program product. When the computer program product is run on a foldable electronic device, the foldable electronic device executes the above-mentioned related steps to implement the methods described in the above-mentioned various method embodiments.

[0257] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera / foldable electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0258] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0259] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0260] In the embodiments provided in this application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0261] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0262] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0263] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0264] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0265] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some implementations," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0266] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A human body image super-resolution method, characterized in that: The method comprises: Acquire an initial human body image; The initial human body image is subjected to a first super-resolution process using a super-resolution model to obtain a first super-resolution human body image; wherein the super-resolution model includes a face super-resolution diffusion model and a human body super-resolution diffusion model, and the first super-resolution process includes: Acquire an initial face feature map and an initial body feature map according to the initial body image; Inputting the initial facial feature map into the face super-resolution diffusion model to obtain a first super-resolution face image; Inputting the initial human feature map into the human super-resolution diffusion model to obtain an initial super-resolution human image, wherein the resolution of the first super-resolution face image is higher than the resolution of the initial super-resolution human image; fusing the first super-resolved face image and the initial super-resolved body image to obtain the first super-resolved body image; The first super-resolution human body image obtained after the first super-resolution processing is used as the initial human body image in the second super-resolution processing, and the process of the first super-resolution processing is repeated until the T-th super-resolution processing is completed, thereby obtaining the first super-resolution human body image of the T-th super-resolution processing.

2. The method according to claim 1, characterized in that The step of obtaining an initial face feature map and an initial body feature map according to the initial body image includes: Compressing the initial human body image to obtain a compressed human body image; An initial face feature map and an initial body feature map are obtained according to the compressed body image.

3. The method according to claim 2, characterized in that The step of obtaining an initial face feature map and an initial body feature map according to the compressed body image includes: Determining the compressed human body image as an initial human body feature map; The initial human body feature map is cropped according to the face frame of the initial human body feature map to obtain an initial human face feature map.

4. The method according to claim 3, characterized in that Before clipping the initial human body feature map based on the face frame of the initial human body feature map to obtain the initial human face feature map, the method further includes: Performing face detection on the initial human body image to obtain first face frame coordinates; determining second face frame coordinates of the initial body feature map according to the first face frame coordinates and a mapping relationship between the initial body image and the initial body feature map; A face frame of the initial human feature map is determined according to the second face frame coordinates.

5. The method according to claim 4, characterized in that The fusing of the first super-resolved face image and the initial super-resolved body image to obtain the first super-resolved body image includes: Selecting a second super-resolved face image corresponding to the second face frame coordinates in the initial super-resolved human body image; In the initial super-resolved human body image, the second super-resolved face image is replaced with the first super-resolved face image to obtain the first super-resolved human body image.

6. The method according to claim 5, characterized in that The first super-resolved face image includes a first super-resolved head image and a first super-resolved non-head image, and the initial super-resolved body image includes an initial super-resolved head image and an initial super-resolved non-head image; The method of replacing the second super-resolved face image with the first super-resolved face image in the initial super-resolved body image to obtain the first super-resolved body image includes: In the initial super-resolution human body image, the initial super-resolution human head image is replaced with the first super-resolution human head image to obtain the first super-resolution human body image.

7. The method according to any one of claims 2 to 6, characterized in that The method further comprises: The first super-resolution human body image obtained by the T-th super-resolution processing is decoded to obtain a target super-resolution human body image, wherein the resolution of the target super-resolution human body image is higher than the resolution of the initial human body image.

8. The method according to any one of claims 1 to 7, characterized in that The initial face feature map includes a plurality of sub-face feature maps. Before inputting the initial face feature map into the face super-resolution diffusion model to obtain the first super-resolution face image, the method further includes: Performing sliding window block processing on the initial facial feature map to obtain multiple sub-facial feature maps; The inputting the initial facial feature map to the face super-resolution diffusion model to obtain a first super-resolution face image includes: Input a plurality of sub-face feature maps to the face super-resolution diffusion model one by one, so that the face super-resolution diffusion model outputs a first super-resolution face image according to the plurality of sub-face feature maps.

9. The method according to any one of claims 1 to 8, characterized in that The initial human body feature map includes a plurality of sub-human body feature maps. Before inputting the initial human body feature map into the human body super-resolution diffusion model to obtain an initial super-resolution human body image, the method further includes: Performing sliding window block processing on the initial human body feature map to obtain multiple sub-human body feature maps; The inputting the initial human body feature map into the human body super-resolution diffusion model to obtain an initial super-resolution human body image includes: Inputting a plurality of sub-human feature maps into the human body super-resolution diffusion model one by one, so that the human body super-resolution diffusion model outputs an initial super-resolution human body image according to the plurality of sub-human feature maps.

10. The method according to any one of claims 1 to 9, characterized in that Before performing a first super-resolution process on the initial human body image using the super-resolution model to obtain a first super-resolution human body image, the method further includes: Construct a super-resolution diffusion model based on the diffusion model and control network model; Training the super-resolution diffusion model based on a first historical face feature map and a second historical face feature map to obtain the face super-resolution diffusion model, wherein the first historical face feature map and the second historical face feature map contain the same facial features, and a resolution of the first historical face feature map is higher than a resolution of the second historical face feature map; Based on a first historical human body feature map and a second historical human body feature map, the super-resolution diffusion model is trained to obtain the human body super-resolution diffusion model, wherein the first historical human body feature map and the second historical human body feature map include the same human body features, and the resolution of the first historical human body feature map is higher than the resolution of the second historical human body feature map.

11. The method according to claim 10, characterized in that The method of training the super-resolution diffusion model based on the first historical face feature map and the second historical face feature map to obtain the face super-resolution diffusion model includes: Inputting a first historical face feature map into the diffusion model in the super-resolution diffusion model, so that the diffusion model outputs the first historical face feature to the control network model according to the first historical face feature map; Input the second historical face feature map to the control network model in the super-resolution diffusion model, so that the control network model is trained according to the first historical face features and the second historical face feature map to obtain the face super-resolution diffusion model.

12. The method according to claim 10 or 11, characterized in that The method of training the super-resolution diffusion model based on the first historical human body feature map and the second historical human body feature map to obtain the human body super-resolution diffusion model includes: Inputting a first historical human body feature map to the diffusion model in the super-resolution diffusion model, so that the diffusion model outputs the first historical human body feature to the control network model according to the first historical human body feature map; The second historical human body feature map is input into the control network model in the super-resolution diffusion model, so that the control network model is trained according to the first historical human body feature and the second historical human body feature map to obtain the human body super-resolution diffusion model.

13. The method according to any one of claims 1 to 12, characterized in that The obtaining of the initial human body image comprises: Acquire an initial scene image including an initial human body image; Performing human body detection on the initial scene image to obtain a human body bounding box; The initial scene image is cropped according to the human body bounding box to obtain an initial human body image.

14. The method according to claim 13, characterized in that The method further comprises: Segmenting the first super-resolved human body image obtained by the T-th super-resolved processing to obtain a human body mask image; Paste the human body mask image back to the position corresponding to the initial human body image in the initial scene image to obtain a target scene image.

15. The method according to claim 14, characterized in that The method further comprises: Performing corrosion processing on the boundary of the human body mask image in the target scene image.

16. The method according to claim 14 or 15, characterized in that The method further comprises: Gaussian blur processing is performed on the boundary of the human body mask image in the target scene image.

17. A human body image super-resolution device, characterized in that: The device comprises: An initial image acquisition module, used for acquiring an initial human body image; The super-resolution processing module is configured to perform a first super-resolution processing on the initial human body image using a super-resolution model to obtain a first super-resolution human body image; wherein the super-resolution model includes a face super-resolution diffusion model and a human body super-resolution diffusion model, and the process of the first super-resolution processing includes: Acquire an initial face feature map and an initial body feature map according to the initial body image; Inputting the initial facial feature map into the face super-resolution diffusion model to obtain a first super-resolution face image; Inputting the initial human feature map into the human super-resolution diffusion model to obtain an initial super-resolution human image, wherein the resolution of the first super-resolution face image is higher than the resolution of the initial super-resolution human image; fusing the first super-resolved face image and the initial super-resolved body image to obtain the first super-resolved body image; A repeated execution module is used to use the first super-resolution human body image obtained after the first super-resolution processing as the initial human body image in the second super-resolution processing, and repeatedly execute the process of the first super-resolution processing until the T-th super-resolution processing is completed, thereby obtaining the first super-resolution human body image of the T-th super-resolution processing.

18. An electronic device, characterized in that: The electronic device includes: one or more processors and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 16.

19. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises instructions, which, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Pedestrian re-identification method and device fusing human face and human body features, and medium

    CN114898402A

  • Living body detection method and device, electronic equipment and storage medium

    CN116129537A

  • Image restoration method and device, electronic equipment and storage medium

    CN116167945A

  • Training image-processing neural networks by synthetic photorealistic indicia-bearing images

    US20200089998A1