Image processing methods and electronic devices
By identifying target individuals and their social circles in images and cropping accordingly, the user experience problem caused by the limitation of image display area is solved, and an image display effect that better meets user needs is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2026-04-03
AI Technical Summary
Due to the size limitations of the area used to display images, unreasonable image cropping methods in existing technologies result in cropped images that cannot effectively reflect the original image information, leading to a poor user experience.
The system identifies the target person's face in the image to be processed and crops the image based on social relationships and intimacy information to ensure that the cropped image meets the user's needs, including the target person's face, and uses depth-of-field effects to display the image of the preset area.
It improves the accuracy of image cropping and user experience, ensuring that the cropped image better matches the user's social context and display needs, thereby enhancing user satisfaction.
Smart Images

Figure CN120260088B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and more specifically, to an image processing method and an electronic device. Background Technology
[0002] As terminal devices become increasingly feature-rich, electronic devices have become indispensable communication and entertainment tools in people's work and life. People can use the photo and video recording functions of smartphones and other smart terminals to collect media data and view the collected image data. People can also use the editing functions of smart terminals to modify and edit media data, and display, play, or share media data with others through smart terminals.
[0003] In some cases, due to size limitations of the area used to display the image, it is necessary to crop the image and display the cropped image in that area. Improper image cropping may result in the displayed image not meeting the user's needs, leading to a poor user experience. Summary of the Invention
[0004] This application provides an image processing method and electronic device that enables cropped images to better meet user needs and improve user experience.
[0005] In a first aspect, an image processing method is provided, comprising: determining a target face image corresponding to a target person among at least one candidate face image in an image to be processed, wherein the target person is the same person among at least one candidate person corresponding to at least one candidate person and at least one preset person; cropping the image to be processed to obtain a target image, wherein the target image includes the target face image.
[0006] If the candidate objects in the image to be processed include at least one target person from the preset objects, the image is cropped. The cropped target image includes the face image of the target person, so that the content recorded in the target image is more in line with the user's needs and the user experience is improved.
[0007] In some possible implementations, the image to be processed includes multiple non-contact subject regions containing multiple subjects. The method further includes: performing face detection on the image to be processed to obtain a first detected face image; if the area of a candidate subject region is greater than or equal to a subject area threshold, taking the first detected face image in the candidate subject region as the candidate face image, wherein the candidate subject region is a subject region among the multiple subject regions that includes at least one of the first detected face images, and the subject area threshold is the product of the area of the largest subject region among the multiple subject regions and a preset subject ratio, wherein the preset subject ratio is greater than 0 and less than 1.
[0008] In images containing large buildings or other objects with significant volume, passersby may be captured in the image. For candidate subject regions that are relatively small compared to other candidate subject regions, the facial images recorded within those regions can be excluded from the candidate face images, thus allowing the target image determined based on the candidate face images to better meet the user's needs.
[0009] In some possible implementations, the method further includes: performing face detection on the image to be processed to obtain a plurality of second detected face images; determining a candidate face image from the plurality of second detected face images, wherein the area of the candidate face image is greater than or equal to a face area threshold, the face area threshold being the product of the area of the second detected face image with the largest area in the candidate subject region where the candidate face image is located and a preset face area ratio, wherein the candidate subject region is a subject region in the image to be processed where at least one subject is located and a second detected face image exists, different subject regions do not touch, and the preset face area ratio is greater than 0 and less than 1.
[0010] A candidate subject area might include accessories such as photo albums and pendants. These accessories may contain facial images. However, the area of these facial images differs significantly from other facial images in the candidate image. By setting a preset facial area ratio, the influence of items such as accessories that contain facial images is taken into account, making the determination of candidate objects more accurate. This results in the target image obtained by cropping based on the candidate objects better meeting the user's needs.
[0011] In one possible implementation, the number of the at least one preset character is multiple, and the at least one preset character has social relationships.
[0012] By using people with whom there is a social relationship as preset people, and recording target people in the image to be processed that are all or part of the same preset people, the target image obtained by cropping the image to be processed will record the target people. This makes the people recorded in the target image more consistent with the user's social situation, thereby increasing the user's satisfaction with the cropped target image and improving the user experience.
[0013] In one possible implementation, the method is applied to an electronic device, wherein the at least one preset person is determined based on multiple data images from a dataset of users of the electronic device.
[0014] The electronic device determines preset characters based on image data from a dataset of users, making the preset characters more suitable for user needs. Furthermore, by determining the preset characters from the dataset, users are not required to manually set them, thus improving the user experience.
[0015] In one possible implementation, the at least one preset person includes multiple data persons belonging to a target social circle within at least one social circle. The at least one social circle is determined based on intimacy information, which represents the intimacy between any two data persons recorded using the multiple data images. The intimacy information is determined based on image recording information, which represents the data persons recorded in each of the multiple data images. In each social circle, a data person can be connected to any other data person in the social circle via at least one edge. An edge between two data persons indicates that the intimacy between the two data persons is greater than or equal to a preset intimacy.
[0016] The closeness between two data figures can be positively correlated with the number of data images of those two figures.
[0017] The closeness between two data individuals can also be positively correlated with the uniformity of the distribution of data images of the two individuals in terms of time and / or location.
[0018] By analyzing multiple user data images, the closeness between each pair of data individuals recorded in those images is determined. Based on this closeness, a social circle is identified. The image to be processed is then cropped according to this social circle, ensuring that faces of individuals matching those individuals within the social circle are preserved in the cropped target image. In this way, the information retained in the target image is relevant to the user's social circle, better reflecting their actual social situation and improving the user experience.
[0019] Social circles can be identified based on multiple data images in an image set without relying on user annotations, demonstrating high intelligence and improving user experience.
[0020] In one possible implementation, the at least one preset person includes a central person among the plurality of data people. The temporal and locational distribution divergences of at least one data image of the central person satisfy preset conditions. The temporal distribution divergence is the distribution divergence of the image's shooting time, and the locational distribution divergence is the distribution divergence of the image's shooting location. The plurality of social circles are determined based on a decentralized person connection graph. The decentralized person connection graph includes multiple person nodes representing multiple non-central people, and edges connecting any two of the plurality of non-central people whose intimacy is greater than or equal to a preset intimacy. The plurality of non-central people are multiple people among the plurality of data people other than the central person.
[0021] Generally, in a person connection graph, the central person node has many edges connecting it to other nodes, and the central person has high relationships with other people, resulting in high weights for the edges connecting to the central person node. If the connection graph includes the central person node, all data related to the central person might be grouped into a single social circle. This is especially problematic when there are multiple central people, leading to low accuracy in determining the final social circles. Determining social circles by decentralizing the person connection graph improves the accuracy of the obtained social circles and enhances the user experience.
[0022] In one possible implementation, the target social circle is at least one social circle whose circle intimacy meets a preset intimacy condition, and the circle intimacy of each social circle is the average intimacy between the data figures in the social circle and the central figure.
[0023] By defining social circles with high intimacy as intimate social circles, users can easily and quickly select the social circles with whom they are most closely connected. When image cropping is required, the cropped target image includes the facial images of people from the intimate social circles, thus improving the user experience.
[0024] In one possible implementation, the at least one preset person includes the central person of a plurality of data image records of the image set, and the temporal distribution divergence and location distribution divergence of at least one of the data images recording the central person satisfy preset conditions, wherein the temporal distribution divergence is the distribution divergence of the image's shooting time, and the location distribution divergence is the distribution divergence of the image's shooting location.
[0025] Temporal and locational dispersion reflect the uniformity of the distribution of images corresponding to a person. The more uniform the distribution of images corresponding to a person, the stronger the correlation between the images and the user using the electronic device storing the image set.
[0026] The preset characters include a central character, so that during the cropping process, the face image of the central character who is closely related to the user is retained in the cropped target image, thus improving the user experience.
[0027] In one possible implementation, the method is applied to an electronic device that stores a collection of images.
[0028] The electronic device used to execute the method of this application stores a set of images. The determination of the preset person can be completed offline by the electronic device without relying on a network. Compared with identifying social circles through social behavior on social networks, the method provided by the embodiments of this application has wider applicability, and the identification results are more targeted to the users of the electronic device, thereby improving the user experience.
[0029] In one possible implementation, the size of the target image is preset, and the method further includes: displaying a preset image of the preset region and the target image with a depth-of-field effect when the ratio of the area of the target person's location in the target image to the area of the preset region of the target image is greater than 0 and less than or equal to a preset area ratio, wherein the preset area ratio is greater than 0 and less than 1.
[0030] When the area of the target person in the target image is smaller than the product of the area of the preset area and the preset area ratio, the target person occupies a smaller proportion of the preset area. The target image and the preset image of the preset area are displayed with a depth-of-field effect. The target person, as the foreground, occludes the preset image less, which better meets the user's display needs for the preset image and improves the user experience.
[0031] In one possible implementation, the size of the target image is preset, and the method further includes: when the area of the target subject region where the target person is located in the target image is greater than 0 in the edge region of the preset region of the target image, and the target subject region does not overlap with the center region of the preset region, displaying the preset image of the preset region and the target image with a depth-of-field effect, wherein the edge region does not overlap with the center region.
[0032] When the target subject area where the target person is located in the target image overlaps with the edge area of the preset area, but does not overlap with the center area of the preset area, the target image and the preset image of the preset area are displayed with a depth-of-field effect. The target person, as the foreground, only occludes the edge area of the preset image and does not occlude the center area, which better meets the user's display needs for the preset image and improves the user experience.
[0033] In a second aspect, an image processing apparatus is provided, including a unit for performing the method of the first aspect. This apparatus may be a terminal device or a chip within a terminal device. The image processing apparatus includes a unit for performing the method of the first aspect described above.
[0034] Thirdly, an electronic device is provided, including a memory and a processor, the memory for storing a computer program, and the processor for calling and running the computer program from the memory, causing the electronic device to perform the method of the first aspect.
[0035] The fourth aspect provides a chip including a processor, which, when executing instructions, performs the method of the first aspect.
[0036] Fifthly, a computer-readable storage medium is provided, the computer-readable storage medium storing computer program code for implementing the method of the first aspect.
[0037] In a sixth aspect, a computer program product is provided, the computer program product comprising: computer program code, the computer program code being used to implement the method of the first aspect. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of a hardware system for an electronic device applicable to this application;
[0039] Figure 2 This is a schematic diagram of a software system applicable to an electronic device of this application;
[0040] Figure 3 This is a schematic diagram of a graphical user interface;
[0041] Figure 4 This is a diagram illustrating image cropping;
[0042] Figure 5 This is a schematic diagram of another graphical user interface;
[0043] Figure 6 This is a schematic flowchart of an image processing method provided in an embodiment of this application;
[0044] Figure 7 This is a schematic flowchart of another image processing method provided in the embodiments of this application;
[0045] Figure 8 This is a schematic flowchart of another image processing method provided in the embodiments of this application;
[0046] Figure 9 This is a schematic flowchart of another image processing method provided in the embodiments of this application;
[0047] Figure 10 This is a schematic diagram of the preset composition rules provided in the embodiments of this application;
[0048] Figure 11 This is a schematic diagram of an image cropping according to a preset composition rule provided in an embodiment of this application;
[0049] Figure 12 This is a schematic diagram of a lock screen interface provided in an embodiment of this application;
[0050] Figure 13 This is a schematic flowchart of another image processing method provided in the embodiments of this application;
[0051] Figure 14 This is a schematic flowchart of a target subject determination method provided in an embodiment of this application;
[0052] Figure 15 This is a schematic flowchart of another image processing method provided in the embodiments of this application;
[0053] Figure 16 This is a schematic diagram of a person-image heterogeneous diagram provided in an embodiment of this application;
[0054] Figure 17 This is a schematic diagram of a character connection diagram provided in an embodiment of this application;
[0055] Figure 18 This is a schematic diagram of a decentralized person connection graph provided in an embodiment of this application;
[0056] Figure 19 This is a schematic structural diagram of an image processing apparatus provided in this application. Detailed Implementation
[0057] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0058] As terminal devices become increasingly feature-rich, electronic devices have become indispensable communication and entertainment tools in people's work and life. People can use the photo and video recording functions of smartphones and other smart terminals to collect media data and view the collected image data. People can also use the editing functions of smart terminals to modify and edit media data, and display, play, or share media data with others through smart terminals.
[0059] In some cases, due to size limitations of the area used to display the image, it is necessary to crop the image and display the cropped version in that area. Improper image cropping may result in the displayed image failing to effectively reflect the information of the original image, leading to a poor user experience.
[0060] Due to size limitations of the area used to display the image, the original image needs to be cropped. The cropped image can then be displayed within that area. By cropping the image, the proportion of the subject within the image can be increased, making the subject stand out more.
[0061] The subject can be understood as the main object of representation in an image. It is the content center, the structural center, and the visual center that attracts attention. The subject often appears in an image as a single object or a group of objects, forming a striking image and becoming the most eye-catching focal point of the entire picture.
[0062] However, improper image cropping methods may result in cropped images that fail to effectively reflect the information in the original images, leading to a poor user experience.
[0063] To address the aforementioned problems, this application provides an image processing method and an electronic device.
[0064] Figure 1 A hardware system for an electronic device applicable to this application is shown.
[0065] The method provided in this application can be applied to various network-connected electronic devices such as mobile phones, tablets, wearable devices, laptops, netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of electronic device.
[0066] Figure 1A schematic diagram of the structure of electronic device 100 is shown. Electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0067] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0068] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0069] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0070] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0071] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0072] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a minimized display, a microLED, a micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0073] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0074] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0075] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0076] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.
[0077] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.
[0078] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.
[0079] Figure 2 This is a software structure block diagram of an electronic device 100 according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the system libraries of the Android runtime, and the kernel layer. The application layer may include a series of application packages.
[0080] like Figure 2 As shown, the application package can include applications such as camera, calendar, call, map, navigation, WLAN, Bluetooth, music, video, SMS, wallpaper, gallery, and multimedia editor.
[0081] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0082] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, wallpaper frame, etc.
[0083] The window manager manages window programs. It can obtain the screen size, determine the presence of a status bar, lock the screen, and capture screenshots. The content provider stores and retrieves data, making this data accessible to applications. The view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon can include views for displaying text and views for displaying images. The phone manager provides communication functionality for the electronic device 100. The resource manager provides various resources for applications. The notification manager allows applications to display notification information in the status bar, which can be used to convey informational messages and can disappear automatically after a short pause without user interaction.
[0084] Wallpaper frames are used to display, schedule, and store images based on user actions within the user interface of applications such as wallpaper, gallery, and multimedia editors.
[0085] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.
[0086] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0087] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0088] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0089] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0090] The media library supports playback and recording of various common audio and video formats, as well as still image files. It also supports multiple audio and video encoding formats.
[0091] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0092] A 2D graphics engine is a graphics engine for 2D drawing.
[0093] The system library can include modules for subject detection, face detection, key figures detection, aesthetic scoring, image quality enhancement, and image cropping.
[0094] Alternatively, an algorithm engine layer for the Android runtime can be set up between the application framework layer and the system libraries of the Android runtime. This algorithm engine layer can include modules for subject detection, face detection, key person detection, aesthetic scoring, image quality enhancement, and image cropping, among others.
[0095] The key person detection module is used to determine whether there is a person in the image who is the same as a preset person.
[0096] The image cropping module is used to crop images.
[0097] The subject detection module is used to detect whether a subject exists in an image. If a subject is present in the image, the subject detection module also detects the region where the subject is located.
[0098] The key person detection module determines whether there is a person in the image to be processed who is the same as a preset person. This determination can be based on candidate face images in the image to be processed, which are detected by the face detection module. In other words, the key person detection module can be used to identify the candidates recorded in the image to be processed based on the candidate face images in the image to be processed, and to determine whether the candidates are the same person as the preset person.
[0099] When a candidate and a preset person are the same person, the key person detection module can identify the same person as the target person. The image cropping module can crop the image to be processed so that the cropped image includes the target person's face image from the candidate face image.
[0100] The key figures detection module can also be used to process images stored in the gallery to identify the central figure and at least one social circle. Preset figures can include the central figure, or figures from one or more social circles.
[0101] The aesthetic scoring module is used to score the aesthetics of images.
[0102] The image enhancement module is used to enhance the image quality. For example, it can enhance the image quality of portraits, animals, plants, buildings, landscapes, etc.
[0103] The image cropping module may include a composition rule cropping module, which is used to determine whether to crop the image according to a preset composition rule.
[0104] A hardware abstraction layer (HAL) can also be set up between the system libraries and the kernel layer. The HAL is an interface layer located between the operating system kernel layer and the hardware circuitry, and its purpose is to abstract the hardware. The hardware abstraction layer can include camera modules, sensor modules, audio modules, etc.
[0105] The kernel layer is the layer between hardware and software. The kernel layer can include driver modules such as display drivers, camera drivers, audio drivers, and sensor drivers.
[0106] The following is combined with Figures 3 to 5 The application scenarios of the image processing method provided in the embodiments of this application will be described.
[0107] Figure 3Image (a) shows a graphical user interface (GUI) of an electronic device, which is the desktop 310 of the electronic device. When the electronic device detects that the user clicks the album icon 311 of the photo album application (APP) on the desktop 310, it can launch the photo album application and display the following: Figure 3 Another GUI is shown in (b) above. Figure 3 The GUI shown in (b) can be referred to as album interface 320. Album interface 320 may include multiple image group covers 321 and 322. The multiple image group covers correspond to multiple image groups. Each image group may include one or more images. Each image group cover may display one image from the image group corresponding to that image group cover. It should be understood that the electronic device may store the multiple image groups.
[0108] An image group cover is an image from the image group corresponding to the image group cover. Alternatively, an image group cover can be obtained by cropping an image from the image group corresponding to the image group cover.
[0109] Since the area size of the image group cover is generally smaller than the size of the images in the image group corresponding to the image group cover, and the aspect ratio of the area of the image group cover may not be the same as the aspect ratio of the images in the image group corresponding to the image group cover, the image group cover can also display an image obtained by cropping, scaling or other processing of a certain image in the image group corresponding to the image group cover, so that the image before cropping can be displayed in the image group cover.
[0110] Image dimensions refer to both the width and height of an image, and can be expressed as "width × height". The width of an image can be understood as the number of pixels along its width direction. The height of an image can be understood as the number of pixels along its height direction. For example, an image with dimensions of 1980 × 1080 means that the image's width is 1980 pixels and its height is 1080 pixels. The height of an image can also be understood as its length.
[0111] like Figure 4 As shown in (a), the image displayed on the image group cover 321 can be an image obtained by cropping the image 410 stored in the electronic device according to the cropping frame 411. Figure 4 As shown in (b), the image displayed in conjunction with the cover 322 may be obtained by cropping the image 420 stored in the electronic device according to the cropping frame 421 and scaling the cropped image.
[0112] Figure 5Image (a) shows a GUI for an electronic device, which may be referred to as a lock screen interface 510. The lock screen interface 510 includes a clock control 511 and a lock screen wallpaper 512. The lock screen wallpaper 512 may be a... Figure 5 The image obtained after cropping and scaling the original image shown in (b) is shown in the figure.
[0113] The following is combined with Figures 6 to 18 The image processing method provided in the embodiments of this application will be described in detail. The execution subject of the method provided in this application can be an electronic device or a software / hardware module in an electronic device capable of image processing. For ease of explanation, the following embodiments use an electronic device as an example.
[0114] Figure 6 This is a schematic flowchart of an image processing method provided in an embodiment of this application. The method may include steps S610 to S620, which are described in detail below.
[0115] Step S610: In at least one candidate face image in the image to be processed, determine the target face image corresponding to the target person. The target person is the same person as at least one preset person among at least one candidate person corresponding to at least one candidate face image.
[0116] Before proceeding to step S610, it can be determined whether the image to be processed contains a face image. If the image to be processed contains a face image, all or part of the face image can be used as candidate face images.
[0117] Face detection algorithms are used to detect faces in images to determine whether they contain human faces. These algorithms are used to detect the presence and location of faces in an image. Face detection algorithms include neural network algorithms and cascaded classifier algorithms.
[0118] All face images in the image to be processed can be used as candidate face images. Alternatively, candidate face images can be determined from the face images in the image to be processed based on a preset face area ratio.
[0119] The distance between an object and the camera is negatively correlated with the area of the region where the object is located in the image captured by the camera.
[0120] In some embodiments, when the image to be processed includes multiple face images, the ratio between the area of each face image and the maximum area among the multiple face images can be calculated. Face images whose ratio is greater than or equal to a preset face area ratio can be used as candidate face images. The preset face area ratio is greater than 0 and less than 1. That is, the ratio between the area of each candidate face image and the maximum area among the at least one candidate face images is greater than or equal to the preset face area ratio.
[0121] Generally, when capturing images, users focus more on people closer to the camera. If the image to be processed contains people close to the camera, people farther away from the camera can be excluded from the cropped image, ensuring that the cropped target image contains people who are more relevant to the user's focus when capturing the image.
[0122] In other embodiments, the image to be processed may include at least one subject region, and different subject regions do not touch. Candidate face images can be determined based on the size relationship of the subject regions.
[0123] Before step S610, face detection is performed on the image to be processed to obtain a first detected face image. The first detected face image can be any detected face image.
[0124] Among the at least one subject region, the region including the first detected face image can be used as a candidate subject region. If the area of the candidate subject region is greater than or equal to a subject area threshold, the first detected face image within the candidate subject region can be used as a candidate face image. The subject area threshold is the product of the area of the largest subject region among multiple subject regions and a preset subject ratio. The preset subject ratio is greater than 0 and less than 1.
[0125] In other words, if the ratio of the area of the candidate subject region to the area of the largest subject region in the image to be processed is greater than or equal to the preset subject ratio, the first detected face image in the candidate subject region can be used as the candidate face image.
[0126] In images containing large buildings or other objects with significant volume, passersby may be captured in the image. For candidate subject regions that are relatively small compared to other candidate subject regions, the facial images recorded within those regions can be excluded from the candidate face images, thus allowing the target image determined based on the candidate face images to better meet the user's needs.
[0127] In some other embodiments, candidate face images can be determined based on the size relationship between the face areas in a candidate subject region.
[0128] Before step S610, face detection is performed on the image to be processed, which can obtain multiple second detected face images. The second detected face images may include all or part of the detected face images.
[0129] Among multiple second-detection face images, candidate face images can be determined. The area of a candidate face image is greater than or equal to a face area threshold. The face area threshold is the product of the area of the second-detection face image with the largest area in the candidate subject region where the candidate face image is located and a preset face area ratio. The preset face area ratio is greater than 0 and less than 1.
[0130] In other words, for each second detected face image, the face area threshold can be expressed as the product of the area of the largest second detected face image in the candidate subject region where the second detected face image is located and a preset face area ratio. If the ratio of the area of a second detected face image to the area of the largest second detected face image in the candidate subject region is greater than or equal to the preset face area ratio, then the second detected face image can be used as a candidate face image.
[0131] A candidate subject area might include accessories such as photo albums and pendants. These accessories may contain facial images. However, the area of these facial images differs significantly from other facial images in the candidate image. By setting a preset facial area ratio, the influence of items such as accessories that contain facial images is taken into account, making the determination of candidate objects more accurate. This results in the target image obtained by cropping based on the candidate objects better meeting the user's needs.
[0132] In some other embodiments, candidate face images can be determined based on the size relationship of the main body regions and the size relationship between the face areas in the candidate main body regions.
[0133] Multiple second detected face images obtained by face detection on the image to be processed can include all detected face images.
[0134] A second detected face image whose area is greater than or equal to a face area threshold can be used as a first detected face image. That is, if the ratio of the area of a second detected face image to the area of the second detected face image with the largest area in the candidate subject region is greater than or equal to a preset face area ratio, then the second detected face image can be used as a first detected face image.
[0135] For different second detected face images, the face area threshold can be expressed as the product of the area of the second detected face image with the largest area in the candidate subject region where the second detected face image is located and the preset face area ratio. In other words, the first detected face image can be the face image with the larger area in a candidate face image.
[0136] If the area of a candidate subject region is greater than or equal to a subject area threshold, the first detected face image in that subject region can be used as the candidate face image. The subject area threshold can be expressed as the product of the area of the largest subject region among multiple subject regions and a preset subject proportion.
[0137] Before proceeding to step S610, the candidate object corresponding to each candidate face image can also be determined. Using face recognition technology, the candidate object corresponding to each candidate face image can be determined.
[0138] Facial recognition is a biometric technology that identifies individuals based on their facial features. Facial recognition algorithms identify and verify faces by matching and comparing extracted facial features. These algorithms include Support Vector Machine (SVM) algorithms, neural network algorithms, clustering algorithms, template-based recognition algorithms, facial feature point-based recognition algorithms, and whole-image-based recognition algorithms.
[0139] Preset characters, also known as key characters, can be understood as characters that users are interested in.
[0140] In some embodiments, the preset character may include a user-specified character. Exemplarily, for performing... Figure 6 The electronic device described in this method can store facial images of multiple individuals. The user can designate one or more facial images as preset individuals.
[0141] In other embodiments, Figure 6 The method shown can be applied to electronic devices, where the preset character can be determined based on multiple data images from the user's dataset of the electronic device.
[0142] The preset characters are determined based on the user's data set, so that the determined preset characters better meet the user's needs.
[0143] For example, based on the data figures recorded in multiple data images, one or more data figures that are recorded most frequently in the multiple data images can be used as preset figures.
[0144] Alternatively, based on the amount of user interaction with each data image in the dataset, determine the degree of user preference for the data person recorded in each data image. The degree of user preference for each data person is positively correlated with the amount of user interaction with at least one data image of that data person. One or more data persons with the highest degree of preference can be used as preset persons.
[0145] For example, a user can interact with a data image in one or more ways. The interactions between the user and the data image can be statistically analyzed separately for each different interaction method. For instance, the various interaction methods between a user and a data image could include clicking, browsing, editing, sharing, etc.
[0146] After statistically analyzing the interaction volume corresponding to various interaction methods, the user's interaction volume for each data image can be obtained by weighted summing of the interaction volume of various interaction paradigms for that data image. The interaction methods can correspond to multiple weights.
[0147] The degree of preference for a data person can be the sum of the number of interactions a user makes with the data images of that data person across multiple data images, or the sum of the ranking of the number of interactions a user makes with the data images of that data person across multiple data images.
[0148] In some other embodiments, the number of preset characters can be multiple, and these multiple preset characters have social relationships.
[0149] By using people with whom there is a social relationship as preset people, and recording target people in the image to be processed that are all or part of the same preset people, the target image obtained by cropping the image to be processed will record the target people. This makes the people recorded in the target image more consistent with the user's social situation, thereby increasing the user's satisfaction with the cropped target image and improving the user experience.
[0150] The social relationships between characters can be preset by the user.
[0151] For example, execute Figure 6 The electronic device described in this method stores a set of images. This device can identify the people recorded in the image set. Users can set up multiple social circles, where people within each circle have social relationships. Users can also specify whether social relationships exist between different people.
[0152] The social relationships between pre-defined characters can also be obtained by processing a set of images.
[0153] For example, each person in a social circle has a social relationship with at least one person in that social circle. Preset people may include those belonging to the target social circle. The target social circle may contain one or more social circles.
[0154] Electronic devices that process image sets, and those that perform... Figure 6 The electronic devices used in the illustrated image processing methods can be the same or different. For example, before performing S610, a method is used to perform... Figure 6 The electronic device using the image processing method shown can receive at least one social circle sent by other electronic devices, each social circle including multiple people.
[0155] Multiple data images in an image set can be used for... Figure 6 The method shown involves using an electronic device to store user-saved images. Multiple data images may include images taken by the user. The user can store images taken using multiple electronic devices on a server. The server sends an image set to one of the user's electronic devices, causing the electronic device to process the image set to obtain at least one social circle, and proceed to step S610 based on the at least one social circle. Alternatively, the server can process the image set to obtain at least one social circle and then send it to the user's electronic device. Figure 6 The method shown involves an electronic device sending data to at least one social circle.
[0156] If multiple electronic devices are logged into with the same account, it can be determined that the users of the multiple electronic devices are the same. Alternatively, electronic devices with the same identity verification information can have the same user.
[0157] Alternatively, multiple data images in an image set can be used for... Figure 6 The data was acquired by an electronic device using the method shown. Used for... Figure 6 The images captured by the electronic device using the method shown can be understood as images taken by the user of the electronic device and stored on the electronic device.
[0158] When the preset person is determined based on a set of images, the preset person may include a central person and / or multiple data persons in the target social circle.
[0159] The target social circle may include one or more social circles, which may be determined based on intimacy information. Intimacy information represents the intimacy between every two data individuals recorded using the multiple data images. This intimacy information is determined based on image recording information. Image recording information represents the data individuals recorded in each of the multiple data images in the image set.
[0160] Multiple data figures within each social circle can have social relationships. This can be understood as any one of these data figures being able to connect to any other data figure through at least one edge. An edge between two data figures indicates that the intimacy between them is greater than or equal to a preset intimacy level.
[0161] By analyzing multiple user data images, the closeness between each pair of data individuals recorded in those images is determined. Based on this closeness, a social circle is identified. The image to be processed is then cropped according to this social circle, ensuring that faces of individuals matching those individuals within the social circle are preserved in the cropped target image. In this way, the information retained in the target image is relevant to the user's social circle, better reflecting their actual social situation and improving the user experience.
[0162] Social circles can be identified based on multiple data images in an image set without relying on user annotations, demonstrating high intelligence and improving user experience.
[0163] In progress Figure 6 In the case of social circle identification using the method shown, the identification is based on images stored in the electronic device, and therefore can be completed offline without relying on a network. This is in contrast to identifying social circles through social behaviors on social networks. Figure 6 The method shown is more applicable and the recognition results are more targeted to users of electronic devices, thereby improving the user experience.
[0164] When a social circle is derived from processing a set of images, it can be determined based on intimacy information. Intimacy information represents the degree of closeness between any two individuals recorded in multiple images within the dataset.
[0165] The affinity between two data characters can be greater than or equal to 0. If the affinity between two data characters is greater than a preset affinity, a relationship can be established between them. The preset affinity can be greater than 0. That is, if the affinity between two data characters is not 0 even when the preset affinity is infinitely close to 0, a relationship can be established between them.
[0166] The closeness between two data figures can be positively correlated with the number of data images of those two figures.
[0167] The closeness between two data individuals can also be positively correlated with the uniformity of the temporal and / or geographical distribution of data images containing those two individuals. The temporal uniformity of the distribution of data images containing two individuals can be represented by the temporal Kullback-Leibler divergence. The geographical uniformity of the distribution of data images containing two individuals can be represented by the geographical Kullback-Leibler divergence.
[0168] KL divergence, also known as relative entropy or information divergence, is an asymmetric measure of the difference between two probability distributions. Relative entropy can be used to measure the distance between two random distributions. When two random distributions are identical, their relative entropy is zero; as the difference between the two random distributions increases, their relative entropy also increases.
[0169] Image recording information can be represented by a person-image heterogeneity graph. The person-image heterogeneity graph represents the correspondence between multiple data images and multiple data persons. Each data image records the person corresponding to that data image. For example, if a data image includes a face image of a data person, it can be determined that the data image records that data person.
[0170] A person-image heterogeneous graph can include multiple image nodes and multiple person nodes. Multiple image nodes correspond one-to-one with multiple data images. Multiple person nodes correspond one-to-one with multiple data people. By removing the multiple image nodes from the person-image heterogeneous graph and connecting the person nodes corresponding to two people with a proximity greater than or equal to a preset proximity value through edges, a person connection graph can be obtained. In other words, the person connection graph includes multiple person nodes representing multiple data people, and edges between every two data people that are related. The proximity between related data people is greater than or equal to a preset proximity value.
[0171] Social circles can be identified based on the person connection diagram.
[0172] In some embodiments, a central person can be determined based on image recording information, as well as the shooting time and location of each data image. The temporal and locational distribution divergences of at least one data image recording the central person satisfy preset conditions. The temporal distribution divergence is the distribution divergence of the image's shooting time. The locational distribution divergence is the distribution divergence of the image's shooting location.
[0173] Based on the image recording information, as well as the shooting time and location of each data image, the central person can be identified without obtaining the user's facial recognition information, thus not affecting the security of the user's private information and improving the user experience.
[0174] To determine the central figure, the temporal and locational dispersion of the data image containing the data figure are calculated for each data figure, and the data figures that meet the preset conditions are taken as the central figures.
[0175] The number of central figures can be one or more.
[0176] By analyzing the capture times of the data images corresponding to a specific data person across multiple preset time intervals, a first time distribution can be obtained. Based on the divergence between this first time distribution and the uniform time distribution, the time distribution divergence of the data images corresponding to that data person can be determined.
[0177] For example, the number of matching images within each preset time interval is counted, and a first time distribution can be determined based on the number of matching images within each preset time interval. Matching images within the first preset time interval refer to data images corresponding to the data person whose shooting time matches the first preset time interval. The first preset time interval can be any one of multiple preset time intervals. The first time distribution can be determined based on the ratio of the number of matching images within each preset time interval to the total number of images corresponding to the first person.
[0178] For example, if the timing information in the shooting time of the first target image belongs to the target preset time interval, and the second target image has not been counted, then the number of matching images in the target preset time interval is incremented by one; the first target image is any one of the images corresponding to the first person, and the second target image is an image corresponding to the first person whose timing information belongs to the target preset time interval and whose date information is the same as the date information in the shooting time of the first target image.
[0179] When statistically analyzing matching images, images with the same date and belonging to the same preset time interval are counted only once. In other words, if multiple matching images exist within a given preset time interval on the same day, only one is counted. This is because the first time distribution is used to determine the time distribution dispersion (i.e., uniformity) of the images corresponding to the first person, and the dispersion of the time distribution of the images corresponding to the first person determines whether the first person is the central figure. The more uniform the time distribution, the greater the likelihood that the first person is the central figure (e.g., the owner of the device or a family member of the owner). In actual use, even non-central figures may take photos within a specific preset time interval on a given day, generating multiple images. Including photos taken within a specific preset time interval in the time distribution statistics would lead to an uneven time distribution of the images corresponding to that person, potentially causing the person to be mistakenly identified as the central figure. Therefore, the method provided by this implementation counts only one image from a concentrated period of photos, discarding images generated from concentrated periods to prevent affecting the accuracy of the time distribution statistics of the images corresponding to the first person, thus making the identification of the central figure more accurate.
[0180] Based on the shooting locations of the data images corresponding to a specific data person, the distribution of these shooting locations across multiple known locations is statistically analyzed to obtain a first location distribution. These known locations include the shooting locations of all data images in the image set. By determining the divergence between the first location distribution and the uniform location distribution, the location distribution divergence of the images corresponding to that data person can be determined.
[0181] For example, a marker value is determined for each known location, and the distribution of the first location can be determined based on the marker values for each known location. The marker value corresponding to the first known location is used to mark whether there is an image in the image corresponding to the data person whose shooting location is the first known location, where the first known location is any one of a plurality of known locations.
[0182] Considering that the first location distribution is used to determine the location distribution divergence (i.e., uniformity) of the target co-occurrence images, and that the location distribution divergence is used to determine whether the first person is the central person. The more uniform the location distribution, the greater the probability that the first person is the central person. However, in actual use, even non-central people may have multiple images taken at a certain location when taking pictures. Including images taken at a certain location in the location distribution statistics will result in an uneven location distribution of the images corresponding to the first person, which may lead to the first person being mistakenly identified as the central person. Therefore, in this implementation, the location distribution of the images corresponding to the first person is analyzed by statistically analyzing whether there are images corresponding to the first person at each known location, without considering the number of images taken at each known location. This prevents images taken at a certain location from affecting the accuracy of the first location distribution statistics, thereby making the identification of the central person more accurate.
[0183] The preset conditions may include: the divergence sum being in the top Q1 positions of the first ranking; the divergence sum is the sum of the time distribution divergence and the location distribution divergence, and the first ranking is obtained by ranking the multiple data individuals in ascending order of their divergence sums, where Q1 is a positive integer; or, the divergence ranking sum being in the top Q2 positions of the second ranking; the divergence ranking sum is the sum of the time divergence ranking and the location divergence ranking, and the second ranking result is obtained by ranking the multiple data individuals in ascending order of their divergence ranking sums, where Q2 is a positive integer.
[0184] Temporal and locational dispersion reflect the uniformity of the distribution of images corresponding to a person. The more uniform the distribution of images corresponding to a person, the stronger the correlation between the person and the user using an electronic device storing the image set. By using the sum of dispersion or the sum of dispersion rankings, it is possible to easily and quickly select the central person.
[0185] For details on how the central figure is determined, please refer to the description in the following embodiments.
[0186] By removing the central node corresponding to the central figure in the character connection graph, as well as the edges connected to that central figure node, we can obtain a character-free connection graph.
[0187] In other words, the decentralized person connection graph includes multiple person nodes representing multiple non-central persons, and edges connecting any two of these non-central persons that are related. These multiple non-central persons are those other than the central person among the multiple data persons.
[0188] Social circles can also be identified by removing the connections between people in the graph.
[0189] Community detection can be performed on person-connected graphs or decentralized person-connected graphs to identify social circles. The methods for community detection can be found in the following embodiments.
[0190] When determining the social circle based on the removed person connection graph, the preset people can include data people in the target social circle, as well as the central person.
[0191] Generally, in a person connection graph, the central person node has many edges connecting it to other nodes, and the central person has high relationships with other people, resulting in high weights for the edges connecting to the central person node. If the connection graph includes the central person node, all data related to the central person might be grouped into a single social circle. This is especially problematic when there are multiple central people, leading to low accuracy in determining the final social circles. Determining social circles by decentralizing the person connection graph improves the accuracy of the obtained social circles and enhances the user experience.
[0192] In other words, in each social circle, any data person can be connected to any other data person in that social circle through at least one edge.
[0193] Step S620: Crop the image to be processed to obtain the target image, which includes the target face image.
[0194] When there are multiple social circles, each social circle can be used as a target social circle. That is, in step S610, at least one candidate person can be compared with at least one preset person corresponding to different social circles to obtain the social circle person corresponding to each social circle. When the social circle person corresponding to each social circle is determined, the social circle person corresponding to each social circle can be used as the target person, and step S620 can be performed.
[0195] In other words, at least one set of preset characters is determined for each social circle, and the preset characters in the set of preset characters for each social circle include multiple data characters in that social circle.
[0196] Before step S610, in the at least one candidate face image, social circle figures corresponding to at least one social circle can be identified. Each social circle face figure is the same person from a preset set of figures corresponding to that social circle among the at least one candidate figure. In step S610, the social circle figures corresponding to the at least one social circle can be designated as target figures, thereby determining target face images corresponding to the at least one social circle. Therefore, in step S620, the image to be processed is cropped to obtain at least one target image corresponding to the at least one social circle. The target image corresponding to each social circle includes the target face image corresponding to that social circle.
[0197] The number of people in a given social circle can be one or more. When identifying target individuals across different social circles, the person in the social circle with the largest number of such individuals can also be chosen as the target individual. In other words, the target individual can be the person in the social circle with the largest number of individuals corresponding to at least one social circle.
[0198] Alternatively, the target social circle can be at least one social circle with the same relationship type.
[0199] Multiple social circles can correspond to multiple relationship types. Social circles corresponding to different relationship types can each serve as target social circles. Therefore, different target individuals can be identified based on different target social circles.
[0200] The relationship type corresponding to a social circle can be set by the user. Alternatively, the relationship type can be determined based on the person information recorded in the data images of the image set and the image information of the data images. For details on how to determine the relationship type of a social circle based on the image set, please refer to the description in the following embodiments.
[0201] Alternatively, the target social circle can be at least one social circle whose circle intimacy meets the preset intimacy condition. The circle intimacy of each social circle is the average intimacy between the data figures in that social circle and the central figure.
[0202] A social circle whose intimacy level meets at least one preset intimacy condition can also be called an intimate social circle. The preset intimacy condition can be that the circle intimacy level is greater than a preset circle intimacy threshold. Alternatively, the preset intimacy condition can be that the circle intimacy level ranks in the top X positions. The ranking of circle intimacy can be obtained by ranking multiple social circles in descending order of circle intimacy level, where X is a positive integer.
[0203] By defining social circles with high intimacy as intimate social circles, users can easily and quickly select the social circles with whom they are most closely connected. When image cropping is required, the cropped target image includes the facial images of people from the intimate social circles, thus improving the user experience.
[0204] For the target image, it can be determined whether it can be displayed with a depth-of-field effect.
[0205] The target image can have a preset size. A preset area can be set within the target image. This preset area can be used to display a preset image. The preset image can include pictures or symbols such as text or numbers. If the target image is a wallpaper, the preset area can be used to display clock controls or other controls. If the target image is the cover of an image group, the preset area can be used to display the name of the image group, such as "Outdoors" or "Family." In other words, the name of the image group can be displayed in a preset area on the cover of the image group.
[0206] In some embodiments, when the ratio of the area of the target person's location within a preset area to the area of the preset area in the target image is greater than 0 and less than or equal to a preset area ratio, a depth-of-field effect can be achieved by displaying the preset image and the target image within the preset area. The preset area ratio is greater than 0 and less than 1. Specifically, in the target image, the area containing the target person can be used as the image in the main layer, and other areas besides the target person can be used as the image in the background layer. The preset image can be used as the image in the intermediate layer between the main layer and the background layer. In other words, the main layer is relatively more prominent than the intermediate layer, and the image in the main layer can be understood as the foreground of the image in the intermediate layer.
[0207] The area where the target person is located in the target image can also be called the target subject area.
[0208] When the ratio of the area of the target person in the target image to the area of the preset area is less than the product of the area of the preset area and the preset area ratio, the target person occupies a smaller proportion of the preset area. The target image and the preset image of the preset area are displayed with a depth-of-field effect. The target person, as the foreground, occludes the preset image less, which better meets the user's display needs for the preset image and improves the user experience.
[0209] In other embodiments, when the area of the target subject region at the edge of the preset region is greater than 0 and the target subject region does not overlap with the center region of the preset region, the preset image and the target image of the preset region can be displayed with a depth-of-field effect, and the edge region does not overlap with the center region.
[0210] The edge region of the preset area can be located at the edge of the preset area. For example, the distance between each point in the edge region and the center of the preset area can be greater than or equal to a preset distance. The edge region can be located at the top or bottom of the preset area, or it can be an area near the left or right edge. The edge region can also surround or semi-surround the preset area.
[0211] In some embodiments, the subject layer can be understood as a foreground image. In other embodiments, other layers may be included before the subject layer as foreground for the subject layer.
[0212] Before step S620, the edges of the target person in the image to be processed can be set in the edge region. The area of the edge region is less than or equal to a preset area ratio, and part of the edge of the edge region can coincide with the edge of the preset region. Therefore, in step S620, the image to be processed is cropped, and the resulting target image can be displayed with a depth-of-field effect.
[0213] When the target subject area where the target person is located in the target image overlaps with the edge area of the preset area, but does not overlap with the center area of the preset area, the target image and the preset image of the preset area are displayed with a depth-of-field effect. The target person, as the foreground, only occludes the edge area of the preset image and does not occlude the center area, which better meets the user's display needs for the preset image and improves the user experience.
[0214] The method provided in this application embodiment performs image cropping when the candidate objects in the image to be processed include at least one target person from a preset group. The cropped image includes the face image of the target person, making the cropping of the image to be processed more in line with user needs and improving user experience.
[0215] The following is combined with Figures 7 to 15 Taking wallpaper setting scenes as an example, for Figure 6 The image processing method shown is explained below. In the wallpaper setting scene, the preset area can be the clock area, and the edge area can be located at the bottom of the clock area in the clock control.
[0216] Figure 7 This is a schematic flowchart of an image processing method provided in an embodiment of this application. Figure 7 The image processing method 700 shown may include steps S701 to S708, which are described in detail below.
[0217] Before proceeding to step S701, the image to be processed can be acquired.
[0218] The methods for acquiring the image to be processed may include reading the image from memory, receiving an image sent by another electronic device, or preprocessing an initial image to obtain the image to be processed. The initial image may be read from memory or received from another electronic device.
[0219] Preprocessing of the initial image can include watermark removal, noise reduction, and other similar processes.
[0220] Preprocessing of the initial image may also include determining whether the size of the initial image, whether it has undergone watermarking, noise reduction, or other processing, meets the specifications.
[0221] Specifications may include an initial image size greater than a preset minimum size. The preset minimum size can be expressed as a preset minimum width and a preset minimum height. In other words, specifications may state that the width of the initial image is greater than or equal to the preset minimum width, and the height of the initial image is greater than or equal to the preset minimum height.
[0222] The ratio of the preset minimum length to the preset minimum width can be equal to the ratio of the length to the width of the cropping frame. The size of the cropping frame can be equal to the size of the screen of the electronic device used to display the wallpaper. The ratio of the preset minimum width to the width of the cropping frame can be greater than, less than, or equal to 1.
[0223] If the initial image is too small and does not meet the specifications, the wallpaper generated from it may have lower resolution, resulting in a poor user experience. Therefore, initial images that do not meet the specifications can be discarded. Figure 7 The image processing shown.
[0224] If the initial image size does not meet the specifications, a prompt message can be output to remind the user that the initial image does not meet the specifications.
[0225] If the size of the initial image meets the specifications, the initial image can be used as the image to be processed, and step S701 can be performed.
[0226] It should be understood that watermark removal, noise reduction, and other processing can also be performed if the size of the initial image meets the specifications. In other words, if the size of the initial image meets the specifications, watermark removal, noise reduction, and other processing can be performed, and the processed image can be used as the image to be processed in step S701.
[0227] Step S701: Determine whether there is a subject in the image to be processed.
[0228] If a subject exists in the image to be processed, step S702 can be performed. If no subject exists in the image to be processed, the following steps can be performed: Figure 8 The image processing method shown is 800.
[0229] Step S702: Determine whether the clock style of the lock screen interface of the electronic device is the default style.
[0230] With the clock style set to the default, the clock control is located at the top center of the lock screen and includes hour and minute information. The clock control may or may not include date information. When the clock control includes date information, the date information can be positioned above the hour and minute information. The hour and minute information can be represented using numbers and symbols, with numbers displayed in Arabic numerals.
[0231] If the clock style is the default style, S703 can be performed. If the clock style is not the default style, the following can be performed: Figure 9 The image processing method shown is 900.
[0232] Step S703: Determine whether the number of subjects in the image to be processed is one.
[0233] If the number of subjects in the image to be processed is one, proceed to step S704. If the number of subjects in the image to be processed is multiple instead of one, proceed to step S705.
[0234] Different objects can correspond to different subjects. Alternatively, a subject can include multiple objects whose regions are connected. In other words, there are no connections between the regions where different subjects reside.
[0235] The object can be a person or a thing. By treating multiple connected objects in the same area as a single subject, it is no longer necessary to identify each object separately, making the determination of the subject much simpler.
[0236] Step S704: Designate this subject as the target subject.
[0237] Step S705: Identify the target subject among multiple subjects.
[0238] For methods of identifying the target subject among multiple subjects, please refer to [link / reference]. Figure 14 Explanation.
[0239] After the target subject is determined through step S704 or S705, step S706 can be performed.
[0240] Step S706: Set the highest point of the target subject at the edge of the preset area in the cropping frame.
[0241] The size of the cropping frame can be the same as the size of the electronic device's display screen. The preset area can be the clock area on the electronic device's display screen used to display the clock control. Both the preset area and the edge area can be rectangles or approximately rectangles; for example, the preset area and the edge area can be rectangles with one or more rounded corners. The width of the edge area is the same as the preset area, the bottom of the edge area coincides with the bottom of the preset area, and the ratio between the height of the edge area and the height of the preset area is a preset overlap ratio.
[0242] In other words, the edge region can be the area near the bottom of the clock region, and the height of the edge region is the product of the height of the clock region and a preset overlap ratio. The preset overlap ratio is greater than 0 and less than 1, for example, it can be 30%, 35%, or 40%, etc.
[0243] Step S707: Scale the image to be processed according to various scaling ratios to obtain multiple depth-of-field cropped images within the cropping frame.
[0244] During the scaling process, the position of the highest point of the target subject within the cropping frame remains unchanged. In other words, in step S707, scaling is performed with the highest point of the target subject as the fixed point.
[0245] By scaling the image to be processed, the cropping box can be completely filled by the scaled image, ensuring that there are no blank areas within the cropping box. Furthermore, by scaling the image to be processed at various scaling ratios, multiple depth-of-field cropped images can be obtained.
[0246] Step S708: Determine whether there is at least one suitable image among the multiple depth-of-field cropped images that satisfies the depth-of-field wallpaper condition.
[0247] If a suitable image with the depth-of-field wallpaper condition exists among multiple depth-of-field cropped images, proceed to step S709.
[0248] In the case where no suitable image meets the depth-of-field wallpaper condition among multiple depth-of-field cropped images, the following steps are performed: Figure 9The image processing method shown is 900.
[0249] The depth-of-field wallpaper conditions include one or more of the following conditions: (1) The ratio of the area of the target subject in the depth-of-field cropped image to the area of the target subject in the image to be processed is greater than or equal to the preset subject retention ratio; (2) The ratio of the area of the target subject in the depth-of-field cropped image to the area of the cropping frame is greater than or equal to the preset subject area ratio; (3) When the image to be processed is magnified to obtain the depth-of-field cropped image, the magnification factor is less than or equal to the preset magnification factor; (4) When the target subject includes a person and the image to be processed includes the face image of the target subject, the depth-of-field cropped image includes the face image of the target subject.
[0250] The area of the subject in the depth-of-field cropped image within the image to be processed can be obtained by dividing the area of the subject in the depth-of-field cropped image by the scaling factor of the depth-of-field cropped image relative to the image to be processed.
[0251] The preset subject retention ratio is greater than 0 and less than 1, for example, it can be 60%, 70%, or 80%. The preset subject area ratio is greater than 0 and less than 1, for example, it can be 10%, 15%, 20%, or 25%. The preset magnification is greater than 1, for example, it can be 1.5, 2, 2.5, or 3.
[0252] The magnification factor of the image to be processed can be understood as the ratio of the height (or width) of the magnified image to the height (or width) of the image to be processed.
[0253] The depth-of-field wallpaper conditions may also include: if the target subject is a person or an animal, the target image records the facial region of the animal in the image to be processed.
[0254] For a person, the facial region can be the area that records the eyes, nose and mouth, and the facial region can also record one or more of the ears, forehead and chin.
[0255] For animals, the batter area can be the area where eyes, nose, mouth, and ears are recorded, or it can be the area where forehead, chin, etc. are recorded.
[0256] Step S709: Determine the target image from at least one suitable image.
[0257] If there is only one qualified image, the qualified image can be used as the target image.
[0258] When there are multiple images that meet the criteria, the target image can be selected from among them.
[0259] The target image can be the qualifying image with the highest aesthetic score. Alternatively, it can be the qualifying image with the largest area of the retained target subject. Another option is the qualifying image with the largest area of the retained target subject among qualifying images with aesthetic scores greater than a preset score threshold.
[0260] In cases where a subject can include multiple connected objects within its region, the region containing a subject in the image to be processed can include multiple face images. In this case, priority should be given to ensuring that the cropped image includes multiple face images, and if the cropped image includes multiple face images, the center of the subject should be located as centrally as possible in the cropped image. If the target image cannot include multiple face images, the cropping of the image to be processed should ensure that the cropped image includes face images corresponding to the same person as the preset character from among the multiple face images in the image to be processed.
[0261] Figure 8 This is a schematic flowchart of an image processing method provided in an embodiment of this application. Figure 8 The image processing method shown may include steps S801 to S806, which are used to determine the target image based on the image to be processed when there is no subject in the image to be processed. These steps are described in detail below.
[0262] Step S801: Crop the image to be processed according to the default cropping method to obtain multiple default cropped images.
[0263] The default cropping method includes one or more of the following: sliding window cropping, default center cropping, etc.
[0264] During the sliding window cropping process, the cropping frame can be used as a sliding window to crop the image to be processed, so as to obtain multiple default cropped images.
[0265] During the default center cropping process, the center of the image to be processed is used as the center of the cropping frame to crop the image.
[0266] Before cropping the image to be processed, it can be scaled up to various ratios. Therefore, in step S801, the images obtained by scaling up to multiple ratios can be cropped using the default cropping method.
[0267] When an image is enlarged and then cropped using the default method, the enlargement ratio can be less than or equal to a preset magnification factor. Enlarging the image can also involve upsampling. Upsampling can be achieved through interpolation or super-resolution processing. Keeping the enlargement ratio less than or equal to the preset magnification factor helps avoid excessive enlargement that could result in a blurry image.
[0268] In some embodiments, when the image to be processed is scaled and then cropped using the default method, it can be determined whether the ratio of the area of the cropped image (i.e., the image within the cropping frame) to the area of the scaled image to be processed is greater than or equal to a preset ratio threshold. If the ratio of the area of the cropped image to the area of the scaled image to be processed is greater than or equal to the preset ratio threshold, the cropped image is used as the default cropped image.
[0269] If the ratio of the cropping box area to the scaled area of the image to be processed is too small, the amount of information that the graph correlation within the cropping box cannot reflect is far less than the amount of information in the image to be processed. By setting a preset ratio threshold, it is possible to avoid the cropped image failing to reflect the effective information of the image to be processed.
[0270] Step S802: Aesthetic scoring is performed on multiple default cropped images to obtain an aesthetic score for each default cropped image.
[0271] For example, an image can be aesthetically scored based on multiple dimensions such as its aesthetic composition and expert prior information.
[0272] Step S803: Determine whether there is a target default cropped image among the multiple default cropped images whose aesthetic score is greater than or equal to a preset aesthetic score threshold.
[0273] If there is a target default cropped image among the multiple default cropped images with an aesthetic score greater than or equal to a preset aesthetic score threshold, step S804 can be performed. Conversely, if there is no target default cropped image among the multiple default cropped images with an aesthetic score greater than or equal to the preset aesthetic score threshold, step S805 can be performed.
[0274] The preset aesthetic score threshold can be set based on experience, or it can be the aesthetic score obtained by aesthetically evaluating the image to be processed.
[0275] Step S804: Among at least one target default cropped image, the default cropped image with the highest aesthetic score is selected as the target image.
[0276] Step S805: Center-crop the image to be processed to obtain the target image.
[0277] In step S805, the center of the image to be processed can be set at the center of a cropping frame with the same size as the display screen, and the image to be processed can be resized so that the resized image fills the entire area of the cropping frame, and the opposite side edges of the resized image coincide with the edges of the cropping frame. Thus, by cropping the image to be processed according to the cropping frame, the target image can be obtained.
[0278] Figure 9 This is a schematic flowchart of an image processing method provided in an embodiment of this application. Figure 9 The image processing method shown may include steps S901 to S906, used when the clock style is not the default style, or Figure 7 In step S707, if no conditional image satisfying the depth-of-field wallpaper condition exists among the multiple cropped images, the target image is determined based on the image to be processed. These steps are described in detail below.
[0279] Step S901: Determine whether the number of subjects in the image to be processed is one.
[0280] When the number of subjects in the image to be processed is multiple, rather than one, steps S902 to S903 can be performed. When the number of subjects in the image to be processed is one, step S904 can be performed.
[0281] Step S902: Identify the target subject among multiple subjects.
[0282] For methods of identifying the target subject among multiple subjects, please refer to [link / reference]. Figure 14 Explanation.
[0283] Step S903: Set the target subject at the center of the cropping frame and crop the image to be processed to obtain the target image.
[0284] The target image is the image obtained by placing the target subject in the image to be processed at the center of the cropping frame, and the target image includes all or part of the target subject in the image to be processed.
[0285] During the S903 process, after centering the target subject within the cropping frame, the image to be processed can be scaled. Then, the image can be cropped according to the cropping frame to obtain the target headshot.
[0286] When the image to be processed is magnified, the magnification factor can be less than or equal to the preset magnification factor.
[0287] Dividing the area of the target subject in the multi-subject cropped image by the scaling factor corresponding to that multi-subject cropped image yields the initial cropped area of the target subject's portion in the image to be processed. The ratio of the initial cropped area to the area of the target subject in the image to be processed can be greater than or equal to a preset subject retention ratio.
[0288] By setting a preset subject retention ratio, the main part of the target subject can be preserved in the cropped multi-subject image.
[0289] When the target subject includes a person and the image to be processed includes the face image of the target subject, the multi-subject cropped image includes the face image of the target subject.
[0290] Step S904: Designate this subject as the target subject.
[0291] Step S905: Determine whether the area where the target subject is located includes a face image.
[0292] If the area containing the target subject includes a face image, step S903 can be performed. If the area containing the target subject does not include a face image, step S906 can be performed.
[0293] Since step S905 is performed when the image to be processed contains only one subject, the determination of whether the area where the target subject is located contains a face image can also be the determination of whether the image to be processed contains a face image.
[0294] Step S906: Determine whether cropping the image to be processed according to the preset composition rules can result in a cropped image.
[0295] Preset composition rules include the center rule, the rule of thirds, etc. Cropping the image to be processed according to the preset composition rules will result in the image being positioned as indicated by the preset composition rules.
[0296] The center rule indicates the location as the center of the clipping box. For example... Figure 10 (a) shows the location indicated by the centering rule. Line 1012 bisects the width of the cropping frame 1010, and line 1014 bisects the length of the cropping frame 1010. The intersection of lines 1012 and 1014 is the center of the cropping frame 1010. According to the centering rule, the center of the target body 1011 can be set at the center of the cropping frame.
[0297] The rule of thirds indicates the location of the cropping box by its thirds. For example... Figure 10As shown in (b), the rule of thirds indicates the location of the cut-off lines. Lines 1022 and 1023 divide the width of the cut-off frame 1020 into thirds, and lines 1024 and 1025 divide the length of the cut-off frame 1020 into two equal parts. According to the centering rule, the center of the target body 1011 can be set on lines 1022, 1023, 1024, or 1025. For example, the center of the target body 1011 can be set at the intersection of any two lines among lines 1022, 1023, 1024, and 1025.
[0298] The image can be cropped according to the rules, or it can be an image that meets or does not meet the subject protection conditions.
[0299] Subject protection conditions may include a ratio between the area of the target subject in the cropped image and the area of the target subject in the image to be processed that is greater than or equal to a preset subject retention ratio, or a cropped image that includes the target subject's face image when the target subject is a person and the image to be processed includes the target subject's face image.
[0300] After cropping the image to be processed according to preset composition rules, it can be determined whether there is an image in the cropped image that meets the subject protection conditions. If there is an image in the cropped image that meets the subject protection conditions, then the image with the subject protection conditions is used as the rule-cropped image. If there is no image in the cropped image that meets the subject protection conditions, then the cropped image is used as the rule-cropped image.
[0301] Before cropping the image to be processed according to preset composition rules, the image can be scaled. In other words, the image to be processed, before or after scaling, can be cropped according to preset composition rules.
[0302] When the image to be processed is magnified, the magnification factor can be less than or equal to the preset magnification factor.
[0303] Dividing the area of the target subject in the regularly cropped image by the scaling factor corresponding to the multi-subject cropped image yields the initial cropped area of the target subject in the image to be processed. The ratio of the initial cropped area to the area of the target subject in the image to be processed can be greater than or equal to a preset subject retention ratio.
[0304] In some cases, because the target subject is too close to the edge of the image to be processed, or due to the shape of the target subject, the regular cropped image may not meet the requirements of the preset magnification or the preset subject retention ratio. That is, cropping the image to be processed according to the preset composition rules cannot result in a regular cropped image.
[0305] The preset magnification factor requirement means that, assuming the cropped image is obtained by cropping the image to be processed after magnification, the magnification factor of the image to be processed must be less than or equal to the preset magnification factor. The preset subject retention ratio requirement means that the ratio of the area of the target subject in the cropped image to the total area of the target subject in the image to be processed must be greater than or equal to the preset subject retention ratio.
[0306] In other words, after cropping the image to be processed, it can be determined whether the cropped image meets the requirements of the preset magnification and / or the preset subject retention ratio. Images that meet the preset magnification and / or the preset subject retention ratio requirements can be used as regular cropped images.
[0307] like Figure 11 The image 1811 to be processed shown in (a) is cropped according to the centering rule, and the center of the target subject 1812 in the image to be processed is located at the center of the cropping region 1813. The image in the cropping region 1813 is used to represent the position of the cropped image in the image 1811 to be processed. That is, the cropped image can be the image in the cropping region 1813, or it can be obtained by scaling the image in the cropping region 1813.
[0308] For the image to be processed 1811, the area of the target subject in the cropping region 1813 is the initial cropping area of the subject in the image to be processed. For the image in the cropping region 1813, if the ratio of the initial cropping area of the subject to the area of the target subject in the image to be processed is less than the preset subject retention ratio, that is, the cropped image corresponding to the cropping region 1813 does not meet the requirement of the preset subject retention ratio, and the cropped image corresponding to the cropping region 1813 can no longer be used as a rule-processed image.
[0309] like Figure 11 Image 1821 to be processed, shown in (b), is cropped according to the rule of thirds. The center of the target subject 1822 in the image to be processed is located on the third line 1824 to the right of the cropping area 1823. The image in the cropping area 1823 is used to represent the position of the cropped image in the image to be processed 1821. Since the size of the cropping area 1823 is too small, that is, the magnification of the cropped image corresponding to the cropping area 1823 relative to the image to be processed 1821 is greater than the preset magnification, and the requirement of the preset magnification is not met, the cropped image corresponding to the cropping area 1823 can no longer be used as the image for regular processing.
[0310] If cropping the image to be processed according to the preset composition rules can yield a regularly cropped image, steps S907 to S908 can be performed. Conversely, if cropping the image to be processed according to the preset composition rules cannot yield a regularly cropped image, steps S910 to S912 can be performed.
[0311] Step S907: Perform aesthetic scoring on the regular cropped image to obtain the aesthetic score of the subject condition image.
[0312] Step S908: Determine whether the aesthetic score of the rule-cropped image is greater than or equal to the preset aesthetic score threshold.
[0313] If the aesthetic score of the cropped image is greater than or equal to the preset aesthetic score threshold, step S909 can be performed.
[0314] Step S909: Use the rule-cropped image as the target image.
[0315] If the aesthetic score of the subject condition image is less than the preset aesthetic score threshold, steps S910 to S912 can be performed.
[0316] Step S910: Crop the image to be processed according to the default cropping method to obtain multiple irregularly cropped images.
[0317] Step S911: Aesthetic scoring is performed on multiple irregularly cropped images to obtain an aesthetic score for each irregularly cropped image.
[0318] Before performing the aesthetic evaluation in step S907 or step S911, the regularly cropped images or irregularly cropped images can be screened to determine the screened images that meet the face protection conditions and then perform the aesthetic evaluation.
[0319] Regularly cropped or irregularly cropped images can be used as images to be screened to determine whether they meet the face protection criteria. The face protection criteria may include: in at least one facial region recorded in the image to be processed, each face is completely located within or outside the image to be screened.
[0320] By determining whether an image meets the facial protection criteria, incomplete facial areas can be avoided in the filtered images, thereby improving the user experience.
[0321] Facial protection conditions may also include: for people or animals whose facial regions are completely outside the image to be screened, and whose body regions are outside the image to be screened.
[0322] This avoids the situation where the filtered images only record the torso from the image to be processed, without recording the facial area that best reflects the characteristics of the person or animal, thus improving the user experience.
[0323] It should be understood that facial regions in the image to be processed can be obtained through face detection. Body regions in the image to be processed can be obtained through body recognition. There can be a correspondence between facial and body regions; each facial region and its corresponding body region belong to a person or animal.
[0324] Step S912: Select the irregularly cropped image with the highest aesthetic score as the target image.
[0325] pass Figure 7 The resulting target image can be used as wallpaper and displayed according to depth-of-field effects. (Through...) Figure 8 The obtained target image cannot be displayed according to the depth-of-field effect when used as wallpaper. Figure 9 Whether the obtained target image can be displayed according to the depth-of-field effect when used as wallpaper can be further determined.
[0326] Depth of field effect is a 3D visual layering effect that can highlight the main subject of the wallpaper. The main subject can be called the foreground of the wallpaper, such as people, mountains, flowers, animals, etc. In the embodiments of this application, when the lock screen wallpaper is displayed with depth of field effect, part of the controls on the lock screen interface (such as clock controls) is obscured by the foreground.
[0327] For example, see Figure 12 This is an illustration of a lock screen interface using a depth-of-field wallpaper. Figure 12 As shown in (a), the lock screen interface 1101 includes a lock screen wallpaper 1102 and a clock control 1103. The lock screen wallpaper 1102 includes a background 1102a and a foreground 1102b.
[0328] like Figure 12 As shown in (b), the lock screen interface 1101 includes at least three layers: the layer containing the background 1102a, the layer containing the clock control 1103, and the layer containing the foreground 1102b. The layer containing the clock control 1103 is located above the layer containing the background 1102a, and the layer containing the foreground 1102b is located above the layer containing the clock control 1103. When these three layers overlap, as shown... Figure 12 As shown in (a), a portion of the clock control 1103 is obscured by the foreground 1102b. Additionally, as... Figure 12 As shown in (b), the background 1102a includes all the content on the lock screen wallpaper 1102, and the foreground 1102b includes only the main object on the lock screen wallpaper 1102, namely the person on the lock screen wallpaper 1102.
[0329] It should be noted that, Figure 12 The dashed lines in the diagram are only used to identify areas and do not actually exist. For example, clock control 1103 is a view control and does not include the dashed lines around it. Additionally, clock controls can also be called time indicators or other names.
[0330] Figure 7 or Figure 9 In the method shown, the target subject in the image to be processed can be taken as the main object in the target image.
[0331] like Figure 12 As shown in (c) above, by Figure 7 The method shown allows setting a cropping frame 1112 in the image to be processed 1111. The highest point of the target subject in the image to be processed is located at the edge region 1113 of the preset region 1114 in the cropping frame 1112. The edge region 1113 is located at the bottom of the preset region 1114, and the ratio between the height of the edge region 1113 and the height of the preset region 1114 is a preset overlap ratio. The preset region can be understood as a clock region.
[0332] The highest point of the target subject remains unchanged within the cropping frame 1112. The image to be processed is scaled, and the scaled image is then cropped according to the cropping frame to obtain a depth-of-field cropped image. Since the highest point of the target subject remains unchanged within the cropping frame 1112, its position within the clock region 1114 also remains unchanged. Even after scaling at various ratios, the highest point of the target subject remains located within the edge region 1113.
[0333] The target image can be obtained by cropping the image based on the depth of field. The cropped image can be used as the target image, or, the cropped image that meets the depth-of-field wallpaper criteria can be used as the target image.
[0334] Therefore, in the target image, the highest point of the subject is located in the edge region, specifically near the bottom of the clock area. When the target image is used as wallpaper, the clock area is used to display the clock controls. Part of the clock controls are obscured by the subject, thus... Figure 7 The resulting target image can be used as wallpaper and displayed with a depth-of-field effect. In other words, the lock screen will include the target image displayed with a depth-of-field effect.
[0335] pass Figure 9 The obtained target image can be processed according to... Figure 13 The image processing method shown is used to process the target image to determine whether it can be displayed according to the depth-of-field effect.
[0336] Figure 13 This is a schematic flowchart of an image processing method provided in an embodiment of this application. Figure 13The image processing method 1200 shown may include steps S1201 to S1206, which are described in detail below.
[0337] Step S1201: Determine whether the subject in the target image exists in the clock area where the clock control is located and whether the proportion of the subject in the clock area is less than or equal to the first preset proportion.
[0338] The first preset ratio is greater than 0 and less than 1.
[0339] If the result of S1201 is negative, it is determined that the target image cannot be displayed according to the depth-of-field effect. If the result of S1201 is positive, proceed to step S1202.
[0340] Step S1202: Determine whether the proportion of the part of the subject in the target image where each digit of the time-division information is located in the time-division digital region is less than or equal to the second preset proportion.
[0341] The second preset ratio is greater than 0 and less than 1. The first preset ratio and the second preset ratio can be equal or unequal. For example, the first preset ratio and the second preset ratio can be equal to the preset overlap ratio.
[0342] If the result of S1202 is negative, it is determined that the target image cannot be displayed according to the depth-of-field effect. If the result of S1202 is positive, proceed to step S1203.
[0343] Step S1203: Determine whether the subject in the target image has a portion located in the date region where the date is located.
[0344] If the subject in the target image contains a portion within the date region, then the target image cannot be displayed according to the depth-of-field effect. If the subject in the target image does not contain a portion within the date region, then the target image can be displayed according to the depth-of-field effect.
[0345] In other words, regarding the passage Figure 9 The obtained target image can be used to determine whether the target image meets the following depth display conditions: (1) The part of the subject in the clock area exists in the target image and the proportion of the clock area is less than or equal to the first preset proportion; (2) The part of the subject in the time and minute number area where each number of the time and minute information is located in the target image occupies the time and minute number area less than or equal to the second preset proportion; (3) The subject in the target image does not exist in the date area where the date is located.
[0346] The first preset ratio and the second preset ratio can be equal or unequal. For example, the first preset ratio and the second preset ratio can both be 40%.
[0347] If the target image meets the depth-of-field display conditions, the target image can be displayed according to the depth-of-field effect.
[0348] Figure 14 This is a schematic flowchart of a target subject determination method provided in an embodiment of this application. Figure 14 The target subject determination method 1300 shown may include steps S1301 to S1306 for determining the target subject among multiple subjects. These steps are described in detail below.
[0349] Step S1301: Determine whether the image to be processed contains a human face image.
[0350] If the image to be processed does not contain a face image, proceed to step S1302. If the image to be processed includes a face image, proceed to step S1303.
[0351] Step S1302: The subject with the largest area in the image to be processed is taken as the target subject.
[0352] Step S1303: Determine whether the area ratio of the main body of the face image to the largest main body is greater than or equal to a preset main body ratio. The preset main body ratio is greater than 0 and less than 1, for example, it can be 20%, 30%, or 40%.
[0353] The region in the image to be processed where the subject of the face image is located can be called the candidate subject region.
[0354] If the area ratio of the subject containing the face to the largest subject is less than a preset subject ratio, proceed to step S1302. If the area ratio of the subject containing the face to the largest subject is greater than or equal to the preset subject ratio, proceed to step S1304.
[0355] Based on the judgment in S1303, in subsequent processing, if the area of the subject including the face image is not very small compared to other subjects in the image to be processed, that is, the area ratio of the subject containing the face to the subject with the largest area is greater than or equal to the preset subject ratio, then the target subject is determined among the subjects including the face image.
[0356] Step S1304: Determine whether the candidate person corresponding to the face image and at least one preset person include the same target person.
[0357] The preset character can be user-defined or determined based on a set of images. For example, the preset character can be a central figure determined based on a set of images, or a character with social relationships determined based on a set of images.
[0358] The method for identifying the central figure and figures with social relationships based on a set of images can be found in [reference needed]. Figure 15 Explanation.
[0359] If the candidate corresponding to the face image does not share the same target person as at least one preset person, proceed to step S1305. If the candidate corresponding to the face image does share the same target person as at least one preset person, proceed to step S1306.
[0360] Step S1305: Take the subject containing the face image corresponding to the target person as the target subject.
[0361] Step S1306: Select the subject containing the largest face image as the target subject.
[0362] Through steps S1301 to S1306, the target subject is determined among multiple subjects. The target subject can be understood as the subject that the user is most interested in among the multiple subjects recorded in the image to be processed.
[0363] Figure 15 This is a schematic flowchart of an image processing method provided in an embodiment of this application. Figure 15 The method includes steps S1401 to S1406, for determining people with social relationships based on a set of images.
[0364] Image collections can be stored in the execution Figure 15 In an electronic device that performs the method shown in 14. Alternatively, the electronic device performing the method shown in 14 can receive a set of images and then process the set of images.
[0365] Step S1401: Determine the person recorded in each of the multiple images in the image set.
[0366] The images in an image set can also be called data person images. The people recorded in the images of an image set can also be called data people.
[0367] By using face clustering algorithms to perform face recognition and clustering on images in an image set, face clustering results can be obtained. The face clustering results include at least one class, and each class corresponds to one person.
[0368] Face clustering algorithms perform face recognition on each image to obtain face images within the image. Optionally, the algorithm can label each recognized face image in an image to distinguish multiple face images within the same image. Then, the algorithm clusters the face images from multiple images in the image set, grouping faces with similar features into the same class, resulting in multiple classes. The algorithm can label each class, i.e., each person. For example, it can assign a number to each person, setting a corresponding ID. In other words, each person can be represented by a unique ID.
[0369] For example, a face clustering algorithm can cluster face images identified from an image together with historical faces to obtain multiple classes. Historical face images refer to face images identified by the face clustering algorithm within a historical time period.
[0370] For example, a particular face image may not be clustered with other face images. For a given face image, there may be no other face images with similar features. In this case, the face clustering algorithm can set the person ID corresponding to that face image as a preset identifier. During subsequent algorithm execution, people with the preset identifier can be excluded from the calculation, and only people with IDs not matching the preset identifier are considered. This improves the accuracy of social circle and social circle type recognition, thus enhancing the user experience.
[0371] For example, the face clustering algorithm or other related modules can also analyze the gender, age, etc. of a person based on the recognized face information. Information such as the person's gender and age can be saved to the image library data storage module, or it can be transferred along with the face clustering results in subsequent data transfers; this application embodiment does not impose any limitations on this.
[0372] Based on the face clustering results, the person recorded in each of the multiple images in the image set can be identified.
[0373] For example, based on the face clustering results, a person-image heterogeneity map can be generated. The person-image heterogeneity map is used to represent the correspondence between people and images containing people.
[0374] A heterogeneous graph is a graph in which the sum of the node types and edge types is greater than 2. That is, a graph where the sum of node types and edge types is greater than 2. It can be understood that heterogeneous graphs belong to graph data, and their representation can be an image, a list, a vector, a class, etc. This application does not impose any limitations on these representations.
[0375] In the face clustering results, each person can be considered a person node. Each image containing a face can be considered an image node. An image node containing a face is connected to its corresponding person node via an edge, resulting in a person-image heterogeneous graph. For example, if image A contains face 1, and according to the face clustering results, face 1 corresponds to person 'a', then the nodes of image A are connected to the nodes of person 'a' via an edge. The same applies to other images and other people in the face clustering results. This yields a person-image heterogeneous graph. In other words, the person-image heterogeneous graph contains both person nodes and image nodes. An edge between a person node and an image node indicates that the image contains that person, or in other words, an edge between a person node and an image node indicates that the person belongs to that image.
[0376] Optionally, the person-image heterogeneous graph may also include node information for person nodes and image nodes. Specifically, the information for a person node is the information of the person corresponding to that node, including but not limited to person ID, gender, and age. The information for an image node is the information of the image corresponding to that node, including but not limited to image hash ID, image capture time, image capture location, and the camera used to capture the image (e.g., whether it was captured by a front-facing camera).
[0377] Figure 16 This is a schematic diagram of a person-image heterogeneous diagram provided in an embodiment of this application. Figure 16 In the diagram, icons with circular borders represent character nodes, while icons with square borders represent image nodes. This is understandable. Figure 16 This is merely an example to illustrate the composition and structure of a person-image heterogeneous graph, and is not intended to be limiting or represent actual data.
[0378] Step S1402: Determine the intimacy between each pair of people based on the people recorded in each image in the image set.
[0379] The degree of connection between two individuals can also be referred to as the intimacy or closeness between them. This degree of connection is positively correlated with the number of facial images of those two individuals in an image set.
[0380] For example, a set of co-occurring images can be determined based on the person-image heterogeneity graph. Based on the image information of each image in the co-occurring image set, the affinity between pairs of people in the face clustering results can be determined.
[0381] Co-occurrence images are images in which two or more people co-occur in the same image. In other words, a co-occurrence image contains at least two people. For example, if image A includes people a, b, and c, then image A is a co-occurrence image. If image B only includes person a, then image B is not a co-occurrence image.
[0382] Based on the person-image heterogeneity graph, the number of person nodes connected to image nodes can be determined. If an image node is connected to two or more person nodes, the image corresponding to that image node is a co-occurring image. A co-occurring image set is generated from the identified co-occurring images. The co-occurring image set includes multiple co-occurring images from the dataset.
[0383] Furthermore, for each pair of characters in a co-occurring image, that co-occurring image can be called the co-occurrence image of those two characters. For example, if image A includes characters a, b, and c, then image A is the co-occurrence image of characters a and b, the co-occurrence image of characters a and c, and the co-occurrence image of characters b and c.
[0384] For ease of description, in the following embodiments, the co-occurrence image of character a and character b will be referred to as the character ab co-occurrence image, and the co-occurrence image of character a and character c will be referred to as the character ac co-occurrence image.
[0385] Based on the person-image heterogeneity graph, the co-occurrence images of any two people can be determined. The person nodes connected to each image node in the person-image heterogeneity graph are the person nodes corresponding to the people who co-occur in the same image. In other words, if two person nodes are connected to a certain image node, then the image corresponding to that image node is the co-occurrence image of the two people corresponding to those two person nodes.
[0386] Based on the image information of each co-occurring image in the co-occurring image set, the closeness between each pair of people in the face clustering results can be determined. The image information includes information such as the shooting time and location.
[0387] The method for calculating the intimacy level between two characters will be further explained in subsequent embodiments.
[0388] Step S1403: Determine the character connection diagram based on the intimacy between each pair of characters.
[0389] In the character connection graph, two character nodes with a non-zero affinity are connected by an edge. Different character nodes represent different characters.
[0390] The graph connecting the characters is an isomorphic graph. An isomorphic graph is a graph in which both the node type and the edge type have only one type. Similarly, it can be understood that an isomorphic graph is also a type of graph data, and its representation can be an image, a list, a vector, a class, etc. This application does not limit this in any way.
[0391] Based on the person-image heterogeneous graph, image nodes are removed. Then, two person nodes with a non-zero affinity are connected by an edge, with affinity used as the weight of the edge. This yields the person connection graph. Image information is no longer required in the person connection graph. Information such as person gender and age may or may not be recorded in the person connection graph.
[0392] In one embodiment, for Figure 16 The simplified person-image heterogeneous graph shown can be used to obtain the person connection graph as follows: Figure 17 As shown.
[0393] Step S1404: Based on the people recorded in each image in the image set, determine the central person among the people.
[0394] A central figure refers to the person at the center of social interaction within the face clustering results. Generally, a central figure has numerous and close relationships with other individuals. There can be one or more central figures. A central figure can be the owner of an electronic device or a non-owner, such as a family member.
[0395] Images containing a central figure may be generated at multiple times and locations, resulting in a relatively even distribution of these images in terms of time and location. In other words, the temporal and locational dispersion of images containing a central figure is small. Compared to images containing non-central figures, images containing a central figure are generally not concentrated in one or a few times, and / or, not concentrated in one or a few locations. For example, if the central figure is the phone owner, the owner could take a photo with their phone at any time, generating an image containing themselves. However, for someone other than the owner, such as a friend, an image might only be generated during a short period of time spent with the owner, using the owner's phone. Therefore, the distribution of images containing the owner is more even than that of images containing their friends.
[0396] Based on this, in this application, the temporal and locational distribution divergences of the images of each person can be determined according to the person-image heterogeneity graph, thereby identifying the central person.
[0397] Among the people included in the face clustering results, those whose temporal and location distribution dispersions of corresponding images meet the first preset condition can be used as central people. Here, an image of a particular person is also called the image corresponding to that person, or an image containing that person; it refers to an image that includes that person. For example, an image of person 'a' is also called the image corresponding to person 'a', meaning it contains person 'a'; an image of person 'b' is also called the image corresponding to person 'b', meaning it contains person 'b', and so on.
[0398] Optionally, the temporal distribution divergence of each person's image can be characterized by temporal KL divergence, and the locational distribution divergence of each person's image can be characterized by location KL divergence.
[0399] The method for determining the central figure will be further explained in subsequent embodiments.
[0400] Step S1405: Remove the node representing the central figure from the figure connection graph to obtain the decentralized figure connection graph.
[0401] Based on the person connection graph, the node corresponding to the identified central person is removed, and the edges connecting the central person node to other person nodes are also removed, resulting in a decentralized person connection graph.
[0402] by Figure 17 Taking the figure connection diagram shown as an example, if the central figure is Figure 17 For the character nodes filled with solid black, the resulting decentered character connection graph, after removing the central character node from the character connection graph, can be obtained as follows: Figure 18 .
[0403] Step S1406: Based on the decentralized person connection graph, determine at least one social circle, and there are social relationships between the people in each social circle.
[0404] Each social circle includes multiple (at least two) individuals.
[0405] Based on the Louvain algorithm, community detection can be performed on decentralized person-connected graphs to identify at least one social circle. Specifically, this can include the following steps:
[0406] 1) Use the decentralized person-connected graph as the target connected graph.
[0407] 2) Treat each community in the target connection graph as a node, and perform modularity optimization operation on each node in the target connection graph to obtain the updated connection graph.
[0408] Specifically, the modularity optimization operation includes: for any node x, try to add node x to each adjacent community and calculate the corresponding modularity benefit; if the maximum modularity benefit among all modularity benefits is negative, then do not change node x, that is, do not add node x to any adjacent community; if the maximum modularity benefit among the modularity benefits of adjacent communities is not negative, add node x to the adjacent community corresponding to the maximum modularity benefit.
[0409] Here, the neighboring communities of node x refer to the communities of the nodes adjacent to node x. It should be noted that for a single node (i.e., a node that has not joined a community), it is a community itself. In other words, the community of a single node includes only one node, which is the node itself.
[0410] After traversing all nodes in the target connectivity graph through modularity optimization operations, an updated connectivity graph is obtained.
[0411] Specifically, when node x is added to any neighboring community (taking neighboring community C as an example), the corresponding modularity gain calculation process can be as follows:
[0412] a. Calculate the modularity of the target connectivity graph when node x joins the neighboring community C, and obtain the first modularity.
[0413] b. Calculate the modularity of the target connectivity graph when node x has not joined the neighboring community C, and obtain the second modularity.
[0414] c. Calculate the difference between the first modularity and the second modularity to obtain the modularity benefit corresponding to the adjacent community C.
[0415] The first and second modularity in steps a and b can be calculated using the following formula:
[0416]
[0417] In formula (1), Q represents the modularity of the target connectivity graph. W represents the sum of the weights of all edges in the target connectivity graph. i and j are both node numbers in the target connectivity graph, where i represents node i and j represents node j. i s represents the sum of the weights of all edges connected to node i in the target connectivity graph. j This represents the sum of the weights of all edges connected to node j in the target connected graph. A i,j δ(C) represents the weight of the edge between node i and node j. i C j ) is an indicator function, C i C represents the community where node i is located. j This indicates the community where node j belongs. When node i and node j belong to the same community (i.e., C...), the community is considered the same. iWith C j When the values are the same, δ(C) i C j The value of ) is 1; when node i and node j do not belong to the same community (i.e., C). i With C j When they are not the same, δ(C) i C j The value of ) is 0.
[0418] 3) Use the updated connection graph as the target connection graph, and treat each community in the updated connection graph as a node. Return to step 2) until all the nodes included in all communities no longer change, and obtain at least one community.
[0419] In other words, during the first execution of step 2), the target connectivity graph is a decentralized person connectivity graph, where each node is a single node, i.e., a node not yet part of a community. Therefore, each node is treated as a community. In the second and subsequent executions of step 2), the target connectivity graph becomes an updated connectivity graph, which includes both single nodes and discovered communities. Therefore, the communities in the updated connectivity graph are treated as nodes, and modularity optimization operations are performed on each. During the modularity optimization operation, for a single node in the updated connectivity graph, its community only includes itself. This process is iterated repeatedly, changing the communities to which nodes belong and optimizing the modularity of each community, until the modularity gain of each community no longer increases and the number of nodes contained in each community no longer changes.
[0420] In this embodiment, the Louvain algorithm is used for community detection, which has low computational time complexity and stable community partitioning results, ensuring stable and accurate social circle generation. Of course, in some embodiments, other algorithms can also be used for community detection to obtain social circles. For example, an infomap algorithm can be used for community detection. This application does not limit the specific method used for community detection.
[0421] 4) Identify at least one social circle based on at least one community.
[0422] In one embodiment, community discovery is performed based on the Louvain algorithm. After obtaining at least one community, the set of people corresponding to all nodes in each community can be directly determined as a social circle, thus obtaining the social circles corresponding to each community.
[0423] For example, after community discovery based on the Louvain algorithm, the obtained communities can be filtered to obtain the final social circle. Filtering the obtained communities can be understood as removing the N person nodes with the smallest sum of weights from the community.
[0424] For any community, calculate the sum of the weights of the edges connecting each character node in the community to other character nodes in the community; remove the N character nodes with the smallest sum of weights from the community to obtain the de-edge community; and determine the set of characters corresponding to all nodes in the de-edge community as a social circle.
[0425] Among them, the N person nodes with the smallest sum of weights can be called marginal person nodes. That is, by removing marginal person nodes from the community, we get the de-marginalized community.
[0426] Optionally, the social circles obtained from community discovery can be represented as a connection graph, i.e., a social circle graph is generated. The social circle graph is used to represent the people included in each social circle, as well as the relationships and intimacy between people.
[0427] It should be noted that the community discovery process described above can disregard node information and rely solely on the weights of nodes and edges. Therefore, in the character connection graph generated in step S1403, node information can be omitted, and only nodes, edges, and edge weights (i.e., affinity) can be included. This further simplifies the structure of the character connection graph and improves the algorithm's efficiency.
[0428] Furthermore, in this embodiment, community discovery is performed based on a decentralized person connection graph to obtain the social circle. This is because: generally, in a person connection graph, the central person node has many edges connecting it to other nodes, and the central person has a high degree of intimacy with other people; therefore, the weight of the edges connecting to the central person node will be very high. If the central person node is included in community discovery, it is easy to assign the central person node to a certain community in a certain iteration. Thus, in the next iteration, when the community where the central person node is located is used as a node, it has many edges with other nodes, and / or the sum of the weights of the edges pointing to the community where the central person node is located will be very high. Therefore, it will assign other points that do not belong to that community to that community, resulting in an inaccurate social circle. Therefore, in this embodiment, community discovery is performed based on a decentralized person connection graph to obtain the social circle, which can improve the accuracy of community discovery, thereby improving the accuracy of the obtained social circle and improving the user experience.
[0429] After defining your social circles, you can also define the type of each social circle. Multiple social circles can correspond to multiple types.
[0430] In some embodiments, the type of each social circle can be determined based on the person-image heterogeneity graph. The type of social circle represents the type of relationship between the people in the social circle and the central person.
[0431] A person-image heterogeneous graph can include multiple person nodes, multiple image nodes, and node information for person nodes and / or image nodes. Node information for person nodes can include information such as the person's gender and / or age. Node information for image nodes can include one or more of the following: image capture time, image capture location, and whether the image was captured by a front-facing camera.
[0432] By processing the person-image heterogeneous graph using a relational model, the relationship type between each person node and the central person node can be obtained. The relational model can be a neural network model. It can include feature extraction models and multilayer perceptron models. The feature extraction model is used to process the person-image heterogeneous graph to obtain the node features of each node in both the person and image nodes. The multilayer perceptron model is used to process the node features of each node to obtain the relationship type between each person and social figures.
[0433] The relationship model can be trained using training data. The training data includes a training person-image heterogeneity graph and label information. The label information is used to represent the type of labeled relationship between each person node in the training person-image heterogeneity graph and the central person in the training person-image heterogeneity graph.
[0434] The relationship type between the person node determined by the relation model and the central person can include, but is not limited to, family members, colleagues, and classmates.
[0435] Count the number of people corresponding to each relationship type in each social circle, and determine the relationship type with the most people as the type of social circle.
[0436] After determining the relationship type for each social circle, you can also change the relationship type of people in each social circle whose relationship type is different from that of that social circle (i.e., the nodes corresponding to the people to be corrected) to the relationship type of that social circle.
[0437] The process of determining the intimacy level between two characters is explained below.
[0438] In a specific embodiment, taking any two people—person a and person b—in the co-occurring image set as an example, the process of calculating the affinity between person a and person b is explained. In the implementation of the above step "S203, determining the affinity between pairs of people in the face clustering result based on the image information of each image in the co-occurring image set," determining the affinity between pairs of people in the co-occurring image set based on the image information of each image in the co-occurring image set includes:
[0439] 1) Calculate the pointwise mutual information (PMI) between person a and person b.
[0440] Mutual information is used to measure the correlation between two things. The greater the mutual information, the greater the correlation between the two things; conversely, the smaller the mutual information, the smaller the correlation between the two things.
[0441] The mutual information PMI(a,b) between person a and person b can be expressed as:
[0442]
[0443] Among them, pic a This represents the number of co-occurring images containing person 'a' in the set of co-occurring images. b This represents the number of co-occurring images containing person b in the set of co-occurring images. a,b This indicates the number of images containing characters a and b that co-occur in the set of co-occurring images. `total_pics` represents the total number of all co-occurring images in the set.
[0444] In one embodiment, the intimacy between person a and person b can be directly represented by PMI(a,b), that is, the value of PMI(a,b) is the intimacy between person a and person b.
[0445] In other embodiments, the affinity between person a and person b can be determined based on PMI(a,b) in combination with other parameters. As one possible implementation, the affinity between person a and person b can be determined based on PMI(a,b) in combination with the temporal and locational dispersion of the image, following the process described below. Specifically, in addition to step 1) above, steps 2) to 6) may also be included:
[0446] 2) Based on the shooting time of the co-occurrence images of people ab, the temporal distribution of the co-occurrence images of people ab is statistically analyzed (hereinafter referred to as the co-occurrence time distribution of people ab).
[0447] The co-occurrence time distribution of figures a and b is used to characterize the distribution of the shooting times of co-occurring images of figures a and b across multiple preset time intervals. The length and number of preset time intervals can be determined based on actual conditions.
[0448] As one possible implementation, a day (00:00 to 24:00) can be divided into multiple preset time intervals based on timing information, where the duration of each preset time interval is less than 24 hours. In a specific embodiment, a day can be divided into 24 preset time intervals, and the duration of each preset time interval is 1 hour. For example, the 24 preset time intervals may include: [00:00, 01:00), [01:00, 02:00), [02:00, 03:00)...[23:00, 24:00].
[0449] It should be noted that preset time intervals can also be divided in other ways. For example, January (from 00:00 on the 1st to 24:00 on the 31st) can be divided into multiple preset time intervals based on daily and time information. This application does not impose any restrictions on the way preset time intervals are divided. The following explanation uses the division of a day into multiple preset times based on time information as an example. The execution process of other division methods is similar and will not be repeated.
[0450] In a specific embodiment, the co-occurrence time distribution of characters a and b can be statistically analyzed according to the following process:
[0451] a) Count the number of matching co-occurring images of people ab within each preset time interval.
[0452] In this context, a matching co-occurrence image within a preset time interval refers to an image among the co-occurrence images of person ab whose timing information belongs to that preset time interval (i.e., the shooting time matches that preset time interval). For example, a matching image within the preset time interval [08:00, 09:00) refers to a co-occurrence image among the co-occurrence images of person ab whose timing information belongs to [08:00, 09:00). For instance, if a certain person ab co-occurrence image x (referred to as image x) was shot at 08:18:12 on April 12, 2023, and the timing information of image x is 08:18:12, which belongs to [08:00, 09:00), then image x is a matching co-occurrence image within the preset time interval [08:00, 09:00). Following this method, matching co-occurrence images within each preset time interval are determined, and the number of matching co-occurrence images within each preset time interval is counted.
[0453] Optionally, during the statistical process, for multiple co-occurring images with the same date information within the same preset time interval, the count can be performed only once. That is, images of people a and b co-occurring with the same date information and timing information belonging to the same preset time interval are counted as a single image. For example, image x was taken at 08:18:12 on April 12, 2023, and image y (referred to as image y) was taken at 08:25:10 on April 12, 2023. Image x and image y have the same date information, both being April 12, 2023. The timing information of both images x and image y belongs to the preset time interval [08:00, 09:00), making them co-occurring images within this preset time interval. Therefore, during the statistical analysis, images x and y are treated as a single image, and the count is performed only once. This is because the co-occurrence time distribution of people a and b will later be used to determine the divergence (i.e., uniformity) of the time distribution of people a and b, and the time distribution divergence reflects the closeness between people a and b. The more evenly the time distribution, the closer the relationship between person A and person B. However, in actual use, even two very close people might take photos together within a certain preset time interval on a particular day, resulting in multiple images. Including photos taken within a preset time interval in the time distribution statistics can lead to uneven co-occurrence time distribution between the two people, potentially resulting in lower intimacy between two closer people and higher intimacy between two less close people who did not have photos taken together. Therefore, the method provided in this implementation counts the images taken together only once, discarding those generated during the concentrated photo period, thus preventing them from affecting the accuracy of the co-occurrence time distribution statistics for people A and B, and thereby improving the accuracy of intimacy calculation.
[0454] In a specific embodiment, taking multiple preset time intervals including: [00:00, 01:00), [01:00, 02:00), [02:00, 03:00)...[23:00, 24:00) as an example, the process of counting the number of matching co-occurring images in each preset time interval can be expressed by the following formula:
[0455] Through vector t ab ∈R 24×1 This represents the number of matching co-occurring images in each preset time interval. t ab [i] represents vector t abThe element t represents the preset time interval number, i = 1, 2, 3…24. Specifically, when i = 1, the preset time interval 1 can be [00:00, 01:00); when i = 2, the preset time interval 2 can be [01:00, 02:00); when i = 3, the preset time interval 3 can be [02:00, 03:00); and when i = 24, the preset time interval 24 can be [23:00, 24:00]. ab [i] represents the number of matched images within the preset time interval i.
[0456] First, let vector t ab All elements within the array are initialized to 0. Then, the number of matching co-occurring images within each preset time interval is counted based on the following formula:
[0457]
[0458] Here, m represents the co-occurrence image m of people a and b (or simply image m). The function f(m) represents the preset time interval to which the timing information of image m belongs. t ab [f(m)] represents the number of matching images within the preset time interval f(m). g(m) = 1 indicates that images with the same date information as image m and belonging to the same preset time interval as person ab have already been counted once. In this case, they will not be counted again. Therefore, t ab The value of [f(m)] remains unchanged, i.e., t ab [f(m)]=t ab [f(m)]. Correspondingly, g(m) = 0 indicates that images with the same date information as image m and belonging to the same preset time interval have not been counted. In this case, the image needs to be included in the statistics. Therefore, t ab The value of [f(m)] is increased by 1, i.e., t ab [f(m)]=t ab [f(m)]+1.
[0459] M represents the set of images of person ab co-occurring, M = {M1, M2, ..., Mx}. M1, M2, ..., Mx all represent images of person ab co-occurring, and x represents the total number of images of person ab co-occurring.
[0460] It is understandable that when multiple preset time intervals include [00:00, 01:00), [01:00, 02:00), [02:00, 03:00)...[23:00, 24:00), determining which preset time interval an image's timing information belongs to can be done solely based on the image's hour information, without needing to obtain minute and second information. For example, if the hour information of image m is 10, then image m belongs to the preset time interval [10:00, 11:00). Therefore, in this embodiment, dividing the multiple preset time intervals into 24 preset time intervals in the above manner, with the lower limit of each preset time interval being an integer, can reduce the computational load during time distribution statistics and improve the algorithm's operating efficiency.
[0461] b. Determine the co-occurrence time distribution of characters ab based on the number of matching co-occurring images within each preset time interval.
[0462] Based on the number of matching co-occurring images in each preset time interval determined above, the proportion of matching co-occurring images in each preset time interval among all co-occurring images of person ab is determined, and the co-occurrence time distribution of person ab is obtained.
[0463] For ease of description, the proportion of the number of matching co-occurring images within a certain preset time interval to the total number of co-occurring images of all characters (ab) will be referred to as the proportion of co-occurring images corresponding to that preset time interval.
[0464] Optionally, the co-occurrence time distribution of characters a and b can be represented as:
[0465] d ab,t ∈R 24×1 ,
[0466]
[0467] Where, vector d ab,t ∈R 24×1 d represents the co-occurrence time distribution of characters a and b. ab,t [i] represents any element in the vector, d ab,t [i] represents the proportion of co-occurring images corresponding to the preset time interval i.
[0468] 3) Determine the KL divergence between the co-occurrence time distribution and the uniform time distribution of characters a and b, and obtain the time KL divergence of characters a and b.
[0469] Uniform temporal distribution refers to the even distribution of co-occurring images of all figures (a, b, and c) across multiple preset time intervals. In other words, the proportion of co-occurring images in each preset time interval is the same, which is 1 / 24 of the number of preset time intervals. For example, when the number of preset time intervals is 24, uniform temporal distribution means that the proportion of co-occurring images in each preset time interval is 1 / 24.
[0470] The temporal KL divergence of figures a and b is used to characterize the difference between the co-occurrence temporal distribution and the uniform temporal distribution of figures a and b. In other words, the temporal KL divergence of figures a and b is used to characterize the uniformity of the distribution of co-occurring images of figures a and b across multiple preset time intervals, thus reflecting the uniformity of the temporal co-occurrence of figures a and b.
[0471] The time distribution can be represented as a vector d. u,t ∈R 24×1 The elements in this set are all 1 / 24. Therefore, the time KL divergence of character ab can be expressed as:
[0472]
[0473] Among them, D ab,t The time KL divergence of character ab is represented by d. u,t [i] represents vector d u,t any element in, d u,t [i] = 1 / 24.
[0474] 4) Statistical analysis of the location distribution of co-occurring images of people ab (hereinafter referred to as the co-occurrence location distribution of people ab).
[0475] The co-occurrence location distribution of people ab is used to characterize the distribution of the shooting locations of images of people ab co-occurring across multiple known locations. These known locations can be obtained from the shooting locations of the images in the image library data storage module.
[0476] In a specific embodiment, the co-occurrence location distribution of person ab can be statistically analyzed according to the following process:
[0477] a. Determine the set of shooting locations, which includes the shooting locations of all images in the image library data storage module.
[0478] The set of shooting locations can be represented as N = {N1, N2, ..., Ny}, where N1, N2, ..., Ny are the elements of set N, representing shooting locations. y represents the total number of shooting locations.
[0479] Optionally, at any time before performing step a, a set of shooting locations can be generated by searching for the shooting locations of all images stored in the image library data storage module, and then stored in the image library data storage module. Furthermore, when images in the image library data storage module are updated, this set of shooting locations can be updated so that the method provided in this application embodiment can be retrieved as needed.
[0480] For ease of description, the elements in the set of shooting locations will be referred to as location elements below. That is to say, the set of shooting locations includes multiple location elements, and each location element is a shooting location of an image in the image library data storage module.
[0481] b. Count the number of images of people (ab) that co-occur for each location element.
[0482] For any location element Ni in set N, the number of co-occurring ab images of people corresponding to that location element refers to the number of co-occurring ab images of people whose shooting location is Ni. For example, the number of co-occurring ab images of people corresponding to the location element "Beijing" refers to the number of co-occurring ab images of people whose shooting location is "Beijing".
[0483] c. Determine the distribution of locations where characters a and b co-occur based on the number of co-occurring images corresponding to each location element.
[0484] Based on the number of co-occurring images of person ab corresponding to each location element determined above, the proportion of the number of co-occurring images of person ab corresponding to each location element in all co-occurring images of person ab is determined, thus obtaining the co-occurrence location distribution of person ab.
[0485] For ease of description, the proportion of the number of co-occurring images of people corresponding to a certain location element in all co-occurring images of people is referred to as the proportion of co-occurring images corresponding to that location element.
[0486] Optionally, when determining the proportion of co-occurring images corresponding to each location element, if the number of co-occurring images of person ab corresponding to a location element is 0, then the proportion is determined to be 0; if the number of co-occurring images of person ab corresponding to a location element is not 0 (i.e., greater than 0), then the proportion of co-occurring images corresponding to that location element is determined to be 1 / the number of non-zero location elements. Here, non-zero location elements refer to location elements whose corresponding number of co-occurring images of person ab is not 0. For example, the set of shooting locations includes 10 location elements, N = {N1, N2, ..., N10}. Statistically, the number of co-occurring images of person ab corresponding to 5 location elements N1, N3, N5, N7, and N10 in this set is not 0, so the proportion of co-occurring images corresponding to these 5 location elements is 1 / 5 = 0.2, and the proportion of co-occurring images corresponding to other location elements is 0. Therefore, the distribution of co-occurring locations of person ab can be represented as: d ab,t ∈R 10 ×1 = [0.2,0,0.2,0,0.2,0,0.2,0,0.2,0,0,0.2].
[0487] As another optional implementation, in the process of statistically analyzing the location distribution of co-occurring images of people a and b, steps b and c above can be replaced with: counting whether co-occurring images of people a and b exist for each location element; and determining the location distribution of co-occurrence of people a and b based on the statistical results.
[0488] Optionally, the label set includes label values corresponding to each location element. These label values are used to indicate whether an image with that location element exists within the set of images showing co-occurrence of people (a, b, and c). Specifically, for any location element Ni, it is determined whether an image with location Ni exists within the set of images showing co-occurrence of people (a, b, and c). If yes, the label value corresponding to location element Ni is set to 1 (i.e., the first value); otherwise, the label value is set to 0 (i.e., the second value). Then, the total number of location elements with a label value of 1 is counted to obtain the number of non-zero location elements (i.e., the total number of the first value). Finally, the proportion of co-occurring images corresponding to location elements with a label value of 1 is set to 1 / the number of non-zero location elements.
[0489] Specifically, the above process can be expressed by the following formula:
[0490] Through vector l ab ∈R y×1 This represents the set of tag values corresponding to each location element. ab [j] represents vector l ab The element in the array, where j is the index of the element at location j, j = 1, 2, 3…y. ab [j] represents the tag value corresponding to the location element j.
[0491] First, vector l ab All elements in the vector are initialized to 0. Then, the vector is shifted according to the following formula. ab Assigning values to each element:
[0492]
[0493] Where m represents the co-occurrence image m of people a and b (or simply image m). The function h(m) represents the location where image m was taken. ab [h(m)] represents the tag value corresponding to the shooting location h(m). q(m) = 0 indicates that the shooting location h(m) has not been counted, and the tag value of the location element with the same shooting location h(m) is currently 0. In this case, since image m was captured by h(m), the tag value corresponding to the location element with the same shooting location h(m) is assigned to 1, i.e., l ab[h(m)] = 1. Correspondingly, q(m) = 1 indicates that the shooting location h(m) has been counted once, meaning that a co-occurrence image of person ab at shooting location h(m) has been confirmed, and the tag value of location elements with the same shooting location h(m) has been assigned a value of 1. In this case, l ab The value of [h(m)] remains unchanged, i.e., l ab [h(m)]=l ab [h(m)]. M represents the set of images of characters a and b co-occurring, which will not be elaborated further.
[0494] The co-occurrence locations of characters a and b can then be represented as follows:
[0495] d ab,l ∈R y×1
[0496]
[0497] Where, vector d ab,l ∈R y×1 d represents the distribution of co-occurrence locations of characters a and b. ab,l [j] represents any element in the vector, d ab,l [j] represents the proportion of co-occurring images corresponding to location element j.
[0498] In this implementation, the co-occurrence location distribution of characters a and b is used to determine the dispersion (i.e., uniformity) of their location distribution, reflecting the closeness between them. A more uniform location distribution indicates a closer relationship between characters a and b. However, in practice, even very close characters might take multiple photos at a particular location. Including these concentrated photos in the location distribution statistics would lead to uneven co-occurrence location distribution, potentially lowering the closeness of two close characters while increasing the closeness of two less close characters who didn't have concentrated photos taken at the same location. Therefore, this implementation analyzes the location distribution of characters a and b by statistically analyzing whether co-occurring images exist for each location element, without considering the number of images taken at each location element. This prevents concentrated photos from affecting the accuracy of the co-occurrence location distribution statistics, thus improving the accuracy of the closeness calculation.
[0499] 5) Determine the KL divergence between the co-occurrence location distribution and the uniform location distribution of characters a and b, and obtain the location KL divergence of characters a and b.
[0500] Uniform location distribution means that all co-occurring images of people (ab) are evenly distributed among all location elements in the set of shooting locations. In other words, the proportion of co-occurring images corresponding to each location element is the same, which is 1 / 10 = 0.1. For example, when the total number of location elements is 10, uniform location distribution means that the proportion of co-occurring images corresponding to each location element is 1 / 10 = 0.1.
[0501] The location KL divergence of people a and b is used to characterize the difference between the co-occurrence location distribution of people a and b and the uniform location distribution. In other words, the location KL divergence of people a and b is used to characterize the uniformity of the distribution of co-occurring images of people a and b across multiple preset time intervals, thus reflecting the uniformity of the location co-occurrence of people a and b.
[0502] A uniform location distribution can be represented as a vector d. u,l ∈R y×1 The elements in this set are all 1 / y. Therefore, the location KL divergence of characters a and b can be expressed as:
[0503]
[0504] Among them, D ab,l The location KL divergence of person ab is represented by d. u,l [j] represents vector d u,l any element in, d u,l [j] = 1 / y.
[0505] 6) Determine the intimacy between character a and character b based on the time KL divergence of character a and character b, the location KL divergence of character a and character b, and PMI(a,b).
[0506] In a specific embodiment, the intimacy between character a and character b can be expressed as:
[0507] closure(a,b)=PMI(a,b)×(1-D ab,t )×(1-D ab,l )
[0508] Where closure(a,b) represents the intimacy level between character a and character b, D ab,t D represents the time KL divergence of characters a, b, and c. ab,l The location KL divergence of person ab.
[0509] As can be seen from formula (11), D ab,t The smaller the value, the larger the closure(a,b); D ab,l The smaller the value, the larger the closure(a,b). D ab,t The smaller the value, the more evenly the co-occurrence time distribution of characters a and b is; D ab,lThe smaller the value, the more evenly the co-occurrence locations of characters a and b are distributed. In other words, according to formula (11), the more evenly the co-occurrence time distribution of characters a and b is, the greater the intimacy between characters a and b, and the closer characters a and b are; moreover, the more evenly the co-occurrence time and location of characters a and b are, the greater the intimacy between characters a and b, and the closer characters a and b are.
[0510] In this embodiment, when determining the intimacy between person a and person b, the co-occurrence time KL divergence and co-occurrence location KL divergence of persons a and b are further added to the PMI(a,b). The co-occurrence time KL divergence of persons a and b can characterize the uniformity of the temporal distribution of their co-occurrence images. The co-occurrence location KL divergence of persons a and b can characterize the uniformity of the geographical distribution of their co-occurrence images. The higher the intimacy between persons a and b, the more uniform the temporal and geographical distribution of their co-occurrence images. Therefore, in this embodiment, the co-occurrence time KL divergence and co-occurrence location KL divergence of persons a and b are integrated into the intimacy calculation, improving the accuracy of intimacy calculation and thus improving the accuracy of social circle identification.
[0511] It is understood that steps 2) to 5) above are also called divergence calculation operations, and steps 1) to 6) above are also called intimacy calculation operations.
[0512] The method for identifying the central figure will be explained below.
[0513] In a specific embodiment, step S205 above, identifying the central person based on the person-heterogeneous graph, may specifically include:
[0514] 1) Determine the image set for each person based on the person-image heterogeneity diagram.
[0515] As described in the above embodiments, an image containing a person can be referred to as the image of that person.
[0516] Specifically, based on the person-image heterogeneity graph, all image nodes connected to a given person node can be identified, and a set of images corresponding to these image nodes can be generated, thus obtaining the image set of the person for that person node. Taking any person 'a' as an example, the connection matrix can be determined according to the connection relationship matrix shown in Table 4. The image nodes in the connection matrix that include person node 'a' are identified as image node A, image node C, and image node N, etc. A set of images corresponding to these image nodes is generated, resulting in the image set of person 'a'.
[0517] 2) Based on the image information of each image in the image set of person a, determine the time KL divergence of person a's image (hereinafter referred to as the time KL divergence of person a) and the location KL divergence of person a's image (hereinafter referred to as the location KL divergence of person a).
[0518] Here, person 'a' is any person in the person-image heterogeneous graph. That is, for each person in the person-image heterogeneous graph, the time KL divergence and location KL divergence are determined according to the method in this step, thus obtaining the time KL divergence and location KL divergence of each person.
[0519] In a specific embodiment, the time KL divergence and location KL divergence of person a can be determined according to the following procedure:
[0520] a. Based on the shooting time of each image in the image set of person a, calculate the time distribution of the images of person a (hereinafter referred to as the time distribution of person a).
[0521] The time distribution of person a is used to characterize the distribution of the shooting time of person a's image across multiple preset time intervals. The preset time intervals are similar to those in the above embodiment and will not be described again.
[0522] The statistical analysis of the time distribution of person 'a' is similar to that of the statistical analysis of the co-occurrence time distribution of person 'ab' in the above embodiment. The difference lies in that: in this embodiment, the scope of the statistics is the set of images of person 'a', while in the above embodiment, the scope of the statistics is all co-occurring images of person 'ab'. In this embodiment, the object of the statistics is the images of person 'a', while in the above embodiment, the object of the statistics is the co-occurring images of person 'ab'. In this embodiment, images in the set of images of person 'a' whose shooting time matches a certain preset time interval can be called matching images of that preset time interval. In the above embodiment, images whose shooting time matches a certain preset time interval can be called matching co-occurring images of that preset time interval. The specific implementation process and beneficial effects of this embodiment are the same as those in the above embodiments and will not be repeated here.
[0523] b. Determine the KL divergence between the time distribution of person a and the uniform time distribution, and obtain the time KL divergence of person a.
[0524] The uniform time distribution is the same as in the above embodiments. The temporal KL divergence of person a is used to characterize the difference between the time distribution of person a and the uniform time distribution. In other words, the temporal KL divergence of person a is used to characterize the uniformity of the distribution of the image of person a across multiple preset time intervals, that is, to reflect the uniformity of the appearance of person a over time.
[0525] The calculation method for the time KL divergence of character a is similar to that for characters a and b, and will not be repeated here.
[0526] c. Based on the shooting locations of each image in the image set of person a, calculate the location distribution of the images of person a (hereinafter referred to as the location distribution of person a).
[0527] The location distribution of person a is used to characterize the distribution of the locations where images of person a were taken.
[0528] In a specific embodiment, the co-occurrence location distribution of person ab can be statistically analyzed according to the following process:
[0529] The location distribution of person 'a' is statistically analyzed similarly to the co-occurrence location distribution of person 'ab' in the above embodiment. The difference is that in this embodiment, the scope of the statistics is the set of images of person 'a', not the set of co-occurring images of person 'ab', and the object of the statistics is the images of person 'a', not the co-occurring images of person 'ab'. The specific implementation process and beneficial effects are detailed in the above embodiment and will not be repeated here.
[0530] d. Determine the KL divergence between the location distribution of person a and the uniform location distribution, and obtain the location KL divergence of person a.
[0531] The uniform location distribution is the same as in the above embodiments. The location KL divergence of person a is used to characterize the difference between the location distribution of person a and the uniform location distribution. In other words, the location KL divergence of person a is used to characterize the uniformity of the distribution of the image of person a across multiple preset time intervals, thereby reflecting the uniformity of the location of person a.
[0532] The calculation method for the location KL divergence of character a is similar to that for characters a and b, and will not be repeated here.
[0533] 3) Determine the central figure based on the time KL divergence and location KL divergence of all figures.
[0534] In one embodiment, the time KL divergence and location KL divergence of each person can be summed to obtain the divergence sum of each person. Then, the divergence sums of all people are ranked (i.e., the first ranking), and the Q1 person with the smallest divergence sum is selected as the center person, where Q1 is an integer greater than or equal to 1. If the ranking is based on the divergence sum in ascending order, the top Q1 person in the divergence sum ranking is selected as the center person. If the ranking is based on the divergence sum in descending order, the last Q1 person in the divergence sum ranking is selected as the center person. It is understood that the KL divergence sum can be calculated directly or by weighted summation; this embodiment does not impose any limitations on this.
[0535] In another embodiment, the time KL divergence of all characters can be ranked in ascending order, and the location KL divergence can also be ranked in ascending order. Then, the time KL divergence rankings and location KL divergence rankings of each character are summed to obtain the divergence ranking sum for each character. The character with the smallest Q2 divergence ranking sum among all characters is selected as the center character, where Q2 is an integer greater than or equal to 1. For example, characters A, B, and C have time KL divergence rankings of {1,2,3} and location KL divergence rankings of {2,3,1}. Summing these two rankings yields a divergence ranking sum of {3,5,4} for characters A, B, and C. If Q2 is 1, then character A, with the smallest divergence ranking sum, is selected as the center character. It should be noted that when determining the center character, if multiple characters have the same divergence ranking sum and Q2 center characters cannot be directly obtained, then the character with the smaller time KL divergence and / or location KL divergence is selected as the center character. In addition, when summing the rankings, it can be either direct summation or weighted summation, and this application embodiment does not limit this in any way.
[0536] In other words, the person whose time KL divergence and location KL divergence satisfy the first preset condition is determined as the central person. The first preset condition is any one of the following conditions: ① The divergence and ranking are in the first Q1 positions of the first ranking, which is obtained by ranking the people according to the divergence and ranking in ascending order; ② The divergence, ranking, and sum are in the first Q2 positions of the second ranking, which is obtained by ranking the people according to the divergence, ranking, and sum in ascending order; ③ The divergence and ranking are in the last Q1 positions of the fourth ranking, which is obtained by ranking the people according to the divergence and ranking in descending order; ④ The divergence, ranking, and sum are in the last Q2 positions of the fifth ranking, which is obtained by ranking the people according to the divergence, ranking, and sum in descending order.
[0537] It should be understood that the above examples are provided to help those skilled in the art understand the embodiments of this application, and are not intended to limit the embodiments of this application to the specific values or scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or changes based on the above examples, and such modifications or changes also fall within the scope of the embodiments of this application.
[0538] The above text combined Figures 1 to 18 The image processing method of the embodiments of this application is described in detail below, and will be combined with the following. Figure 19 This document describes in detail the apparatus embodiments of this application. It should be understood that the image processing apparatus in the embodiments of this application can execute the various image processing methods described in the foregoing embodiments of this application. That is, the specific working processes of the various products described below can be referred to the corresponding processes in the foregoing method embodiments.
[0539] Figure 19 This is a schematic diagram of the image processing apparatus provided in the embodiments of this application.
[0540] It should be understood that the image processing device 1900 can perform... Figure 6 or Figure 17 The image processing method shown; the image processing apparatus 1900 includes: a processing unit 1910 and a cropping unit 1920.
[0541] The processing unit 1910 is configured to determine, in at least one candidate face image in the image to be processed, a target face image corresponding to a target person, wherein the target person is the same person among at least one candidate person corresponding to at least one candidate person and at least one preset person.
[0542] The cropping unit 1920 is used to crop the image to be processed to obtain a target image, the target image including the target face image.
[0543] Optionally, the method is applied to an electronic device, wherein the at least one preset person is determined based on multiple data images from a dataset of users of the electronic device.
[0544] Optionally, the at least one preset person includes multiple data persons belonging to a target social circle in at least one social circle. The at least one social circle is determined based on intimacy information, which represents the intimacy between any two data persons among the multiple data persons recorded using the multiple data images. The intimacy information is determined based on image recording information, which represents the data persons recorded in each of the multiple data images. In each social circle, a data person can be connected to any other data person in the social circle through at least one edge. An edge between two data persons indicates that the intimacy between the two data persons is greater than or equal to a preset intimacy.
[0545] Optionally, the at least one preset person includes a central person among the plurality of data people, and the temporal distribution divergence and location distribution divergence of at least one data image of the central person satisfy preset conditions, wherein the temporal distribution divergence is the distribution divergence of the image's shooting time, and the location distribution divergence is the distribution divergence of the image's shooting location.
[0546] The multiple social circles are determined based on a decentralized person connection graph, which includes multiple person nodes representing multiple non-central persons, and edges connecting any two of the multiple non-central persons whose intimacy is greater than or equal to a preset intimacy. The multiple non-central persons are multiple persons other than the central person among the multiple data persons.
[0547] Optionally, the target social circle is at least one of the social circles whose circle intimacy meets a preset intimacy condition, and the circle intimacy of each social circle is the average intimacy between the data people in the social circle and the central person.
[0548] Optionally, the at least one preset person includes the central person of multiple data image records of the image set, and the temporal distribution divergence and location distribution divergence of at least one data image of the central person satisfy preset conditions, wherein the temporal distribution divergence is the distribution divergence of the image's shooting time, and the location distribution divergence is the distribution divergence of the image's shooting location.
[0549] Optionally, the electronic device stores the image set.
[0550] Optionally, the size of the target image is preset.
[0551] The device 1900 further includes a display unit, which is used to display a preset image of the preset area and the target image with a depth-of-field effect when the ratio of the area of the target person in the target image to the area of the preset area of the target image is greater than 0 and less than or equal to a preset area ratio.
[0552] Optionally, the size of the target image is preset.
[0553] The device 1900 further includes a display unit, which is used to display a preset image of the preset area and the target image with a depth-of-field effect when the area of the target subject region where the target person is located in the target image is greater than 0 in the edge region of the preset area of the target image and the target subject region does not overlap with the center region of the preset area.
[0554] It should be noted that the aforementioned image processing device 1900 is embodied in the form of a functional unit. The term "unit" here can be implemented in software and / or hardware, without specific limitations.
[0555] For example, a "unit" can be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components that support the described functions.
[0556] Therefore, the units of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0557] This application also provides a chip, which includes a data interface and one or more processors. When the one or more processors execute instructions, they read instructions stored in a memory through the data interface to implement the image processing method described in the above method embodiments.
[0558] The one or more processors can be general-purpose processors or special-purpose processors. For example, the one or more processors can be central processing units (CPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, such as discrete gate, transistor logic devices, or discrete hardware components.
[0559] The chip can be used as a component of a terminal device or other electronic device. For example, the chip can be located in electronic device 100.
[0560] Processors and memory can be configured separately or integrated together. For example, processors and memory can be integrated onto a system-on-a-chip (SoC) in a terminal device. That is, the chip can also include memory.
[0561] The memory may store a program, which can be run by a processor to generate instructions, causing the processor to execute the image processing method described in the above method embodiments according to the instructions.
[0562] Optionally, the memory may also store data. Optionally, the processor may also read data stored in the memory, which may be stored at the same memory address as the program, or the data may be stored at a different memory address than the program.
[0563] For example, the memory can be used to store related programs of the image processing method provided in the embodiments of this application, and the processor can be used to call the related programs of the image processing method stored in the memory to implement the image processing method of the embodiments of this application. For example, in at least one candidate face image in the image to be processed, a target face image corresponding to a target person is determined, and the target person is the same person as at least one preset person among at least one candidate person corresponding to at least one candidate face image; the image to be processed is cropped to obtain a target image, and the target image includes the target face image.
[0564] This chip can be installed in electronic devices.
[0565] This application also provides a computer program product that, when executed by a processor, implements the touch recognition method described in any of the method embodiments of this application.
[0566] The computer program product can be stored in memory, for example, it is a program. The program is eventually converted into an executable object file that can be executed by the processor after processes such as preprocessing, compilation, assembly and linking.
[0567] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a computer, implements the touch recognition method described in any of the method embodiments of this application. The computer program may be a high-level language program or an executable object program.
[0568] The computer-readable storage medium is, for example, memory. Memory can be volatile or non-volatile, or it can include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0569] The embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).
[0570] In the description of this application, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance, or a specific order or sequence. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0571] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0572] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0573] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0574] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0575] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and other division methods may exist in actual implementation; for example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0576] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0577] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0578] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An image processing method, characterized in that, include: Perform face detection on the image to be processed to obtain at least one first detected face image; Among the multiple subjects recorded in the image to be processed, at least one candidate subject is identified, and each candidate subject includes at least one first detected face image; If, among the at least one candidate subject, there is a candidate subject whose area is greater than or equal to a subject area threshold, at least one candidate face image is determined in the at least one first detected face image. The subject area threshold is the product of the area of the largest subject in the image to be processed and a preset subject ratio, wherein the preset subject ratio is greater than 0 and less than 1. Among the at least one candidate face image, a target face image corresponding to the target person is determined, wherein the target person is the same person among the at least one candidate person corresponding to the at least one candidate person and the at least one preset person. The image to be processed is cropped to obtain a target image, the target image including the target face image; If none of the candidate subjects has an area greater than or equal to a subject area threshold, the image to be processed is cropped based on the subject with the largest area to obtain the target image.
2. The method according to claim 1, characterized in that, The number of the at least one first detected face images is multiple. For any first detection image, if the first ratio is greater than or equal to the preset face area ratio, then the first detection image is a candidate face image. The preset face area ratio is greater than 0 and less than 1. The first ratio is the ratio of the area of the first detection image to the area of the first detected face image with the largest area among the candidate subjects in which the first detection image is located.
3. The method according to claim 1 or 2, characterized in that, The size of the target image is preset, and the method further includes: when the ratio of the area of the part of the target subject where the target person is located in the target image located in the preset area of the target image to the area of the preset area is greater than 0 and less than or equal to a preset area ratio, displaying the target image and the preset image located in the preset area with a depth-of-field effect, wherein the preset area ratio is greater than 0 and less than 1, and the at least one candidate subject includes the target subject.
4. The method according to claim 1 or 2, characterized in that, The size of the target image is preset, and the method further includes: when the area of the part of the target subject in the target image where the target person is located is greater than 0 in the edge region of the preset region of the target image, and the target subject does not overlap with the center region of the preset region, the target image and the preset image located in the preset region are displayed with a depth effect, the edge region does not overlap with the center region, and the at least one candidate subject includes the target subject.
5. The method according to any one of claims 1-4, characterized in that, The method is applied to an electronic device, wherein the at least one preset character is determined based on multiple data images from an image set of the user of the electronic device.
6. The method according to claim 5, characterized in that, The at least one preset person includes multiple data persons belonging to a target social circle within at least one social circle. The at least one social circle is determined based on intimacy information, which represents the intimacy between any two data persons recorded using the multiple data images. The intimacy information is determined based on image recording information, which represents the data persons recorded in each of the multiple data images. Any data person in each social circle can be connected to any other data person in the social circle through at least one edge. An edge between two data persons indicates that the intimacy between the two data persons is greater than or equal to a preset intimacy.
7. The method according to claim 6, characterized in that, The at least one preset person includes the central person among the plurality of data people, and the time distribution divergence and location distribution divergence of at least one data image of the central person satisfy preset conditions, wherein the time distribution divergence is the distribution divergence of the image's shooting time, and the location distribution divergence is the distribution divergence of the image's shooting location. The multiple social circles are determined based on a decentralized person connection graph, which includes multiple person nodes representing multiple non-central persons, and edges connecting any two of the multiple non-central persons whose intimacy is greater than or equal to a preset intimacy. The multiple non-central persons are multiple persons other than the central person among the multiple data persons.
8. The method according to claim 7, characterized in that, The target social circle is at least one social circle whose circle intimacy meets the preset intimacy condition, and the circle intimacy of each social circle is the average intimacy between the data person in the social circle and the central person.
9. The method according to claim 5, characterized in that, The at least one preset person includes the central person of multiple data image records of the image set. The temporal distribution divergence and location distribution divergence of at least one of the data images recording the central person satisfy preset conditions. The temporal distribution divergence is the distribution divergence of the image's shooting time, and the location distribution divergence is the distribution divergence of the image's shooting location.
10. The method according to any one of claims 5-9, characterized in that, The electronic device stores the image set.
11. An electronic device, characterized in that, The device includes a processor and a memory, the memory being used to store a computer program, and the processor being used to retrieve and run the computer program from the memory, causing the electronic device to perform the method of any one of claims 1 to 10.
12. A chip, characterized in that, It includes a processor and a data interface, wherein the processor reads instructions stored in memory through the data interface to implement the method as described in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for implementing the method of any one of claims 1 to 10.
Citation Information
Patent Citations
Photo processing method and device, storage medium and electronic equipment
CN111339330A
Image cutting method and device, terminal equipment and computer readable storage medium
CN116543004A
Image processing device, operation method of image processing device and operation program of image processing device
JP2021157621A