Information presentation method and device, storage medium and electronic device

By recognizing and adjusting the posture of the subject being photographed, the problem of users being unable to pose appropriately when taking photos with high-resolution electronic devices is solved, thus improving the image quality.

CN115223238BActive Publication Date: 2026-03-24PING AN BANK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Users are unable to pose appropriately when taking photos with high-resolution electronic devices, resulting in low-quality photos.

Method used

By recognizing the posture information of the subject to be photographed, determining the matching status with the preset posture, and calculating the prompt time based on the target distance and the current time, the system outputs a prompt message to adjust the posture so that the subject can adjust its posture.

Benefits of technology

It improves the quality of captured images by using a flexible prompting mechanism to guide the subject into a suitable pose, thereby enhancing the image quality of the photograph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223238B_ABST
    Figure CN115223238B_ABST
Patent Text Reader

Abstract

The application discloses an information prompting method and device, a storage medium and an electronic device. The method comprises the following steps: identifying the posture of a to-be-photographed object in a shooting scene to obtain object posture information of the to-be-photographed object; when the object posture information does not match preset posture information, determining a target distance between the to-be-photographed object and the electronic device; obtaining a current time and, according to the target distance and the current time, obtaining a target prompting time; and when the target prompting time is reached, outputting prompting information for prompting the to-be-photographed object to adjust the posture. The application can improve the quality of an image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of electronics, and particularly relates to an information prompting method and device, a storage medium and an electronic device. BACKGROUND

[0002] With the continuous development of electronic devices, the camera pixels on electronic devices such as smart phones are getting higher and higher, so more and more users tend to use electronic devices such as smart phones for taking pictures. In order to meet the needs of users for taking pictures, major electronic device manufacturers continuously update and upgrade the hardware of electronic devices to improve the picture-taking pixels of electronic devices. However, in order to take high-quality photos, not only does the camera of the electronic device need to have a high resolution, but the photographed object also needs to pose in a suitable pose. However, most photographed objects cannot pose in a suitable pose, so high-quality photos cannot be taken. SUMMARY

[0003] The embodiments of the present application provide an information prompting method and device, a storage medium and an electronic device, which can improve the quality of images.

[0004] In a first aspect, the embodiments of the present application provide an information prompting method applied to an electronic device, comprising:

[0005] identifying the pose of a to-be-photographed object in a shooting scene to obtain object pose information of the to-be-photographed object;

[0006] determining a target distance between the to-be-photographed object and the electronic device when the object pose information does not match preset pose information;

[0007] obtaining a current time, and obtaining a target prompting time according to the target distance and the current time;

[0008] outputting prompting information for prompting the to-be-photographed object to adjust the pose when the target prompting time is reached.

[0009] In a second aspect, the embodiments of the present application provide an information prompting device applied to an electronic device, comprising:

[0010] an information obtaining module configured to identify the pose of a to-be-photographed object in a shooting scene to obtain object pose information of the to-be-photographed object;

[0011] a distance determining module configured to determine a target distance between the to-be-photographed object and the electronic device when the object pose information does not match preset pose information;

[0012] a time obtaining module configured to obtain a current time, and obtain a target prompting time according to the target distance and the current time;

[0013] The information output module is configured to output prompt information for prompting the object to be photographed to adjust the posture when the target prompt time is reached.

[0014] In a third aspect, an embodiment of the present application provides a storage medium having a computer program stored thereon, which, when executed by a computer, causes the computer to perform the information prompting method provided by the embodiments of the present application.

[0015] In a fourth aspect, an embodiment of the present application further provides an electronic device, which comprises a memory and a processor. The processor is configured to execute the information prompting method provided by the embodiments of the present application by invoking a computer program stored in the memory.

[0016] In the embodiments of the present application, when the object posture information of the object to be photographed does not match the preset posture information, the target prompt time is obtained according to the target distance between the object to be photographed and the electronic device and the object attribute information of the object to be photographed. When the target prompt time is reached, the prompt information for prompting the object to be photographed to adjust the posture is output. Therefore, the posture of the object to be photographed can be adjusted based on the output prompt information to make the object to be photographed assume a suitable posture, and thus the quality of the image obtained by photographing the object to be photographed is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] The technical solutions of the present application and the beneficial effects thereof will become apparent through the following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings.

[0018] Figure 1 FIG. 1 is a flowchart of an information prompting method provided by an embodiment of the present application.

[0019] Figure 2 FIG. 2 is a first scenario diagram of the information prompting method provided by an embodiment of the present application.

[0020] Figure 3 FIG. 3 is a second scenario diagram of the information prompting method provided by an embodiment of the present application.

[0021] Figure 4 FIG. 4 is a structural diagram of an information prompting apparatus provided by an embodiment of the present application.

[0022] Figure 5 FIG. 5 is a first structural diagram of an electronic device provided by an embodiment of the present application.

[0023] Figure 6 FIG. 6 is a second structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0024] Please refer to the illustrations, where the same component symbols represent the same components. The principles of this application are illustrated by example in a suitable computing environment. The following description is based on the specific embodiments of this application illustrated, and should not be construed as limiting other specific embodiments not detailed herein.

[0025] It should be noted that the terms "first," "second," and "third," etc., used in this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules, but some embodiments also include steps or modules not listed, or some embodiments also include other steps or modules inherent to these processes, methods, products, or devices.

[0026] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0027] This application provides an information prompting method, an information prompting device, a storage medium, and an electronic device. The entity executing the information prompting method can be the information prompting device provided in this application, or an electronic device integrating the information prompting device. The information prompting device can be implemented in hardware or software. The electronic device can be a smartphone, tablet computer, PDA, laptop computer, or other device equipped with a processor and possessing information prompting capabilities.

[0028] Please see Figure 1 , Figure 1 This is a flowchart illustrating the information prompting method provided in an embodiment of this application. The process may include:

[0029] In 101, the posture of the object to be photographed in the shooting scene is identified to obtain the object posture information.

[0030] With the continuous development of electronic devices, the camera pixels on devices such as smartphones are getting higher and higher, leading more and more users to prefer using smartphones and other electronic devices for taking pictures. To meet users' photography needs, major electronic device manufacturers are constantly updating and upgrading the hardware of their devices to improve camera pixel counts. However, taking high-quality photos requires not only a high-resolution camera but also a suitable pose for the subject. Unfortunately, most subjects cannot pose appropriately, making it impossible to take high-quality photos.

[0031] In this embodiment, the electronic device identifies the posture of the object to be photographed in the shooting scene and obtains the object posture information of the object to be photographed.

[0032] In this context, when an electronic device launches a camera application (such as the system application "Camera") based on user input, the scene that its camera is pointing at is the shooting scene. For example, after a user launches the "Camera" application by tapping its icon on the electronic device, if the user points the camera at a particular scene, that scene is the shooting scene. Based on the above description, those skilled in the art should understand that the shooting scene does not refer to a specific scene, but rather to the scene that the camera is pointing at in real time.

[0033] The subject to be photographed is the object that needs to be photographed or recorded, such as a person, cat, or dog.

[0034] Object pose information can include the pose of the object to be photographed. See also... Figure 2 When the subject to be photographed is posed as follows Figure 2 When the pose shown is shown, the object's pose information can be standing with both hands holding a hat.

[0035] Since key points often reflect the posture of the subject, subject posture information can also include the positions of key points. Key points can correspond to joints with a certain degree of freedom in the human body, such as the neck, shoulders, elbows, wrists, waist, knees, and ankles. Key points can also include objects such as the mouth, nose, left eye, right eye, left ear, right ear, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle.

[0036] In step 102, when the object's posture information does not match the preset posture information, the target distance between the object to be photographed and the electronic device is determined.

[0037] The preset pose information can be collected in advance and stored in electronic devices. For example, professional photographers can design different poses, and the different poses, or the key point positions corresponding to the different poses, can be used as candidate pose information.

[0038] Before acquiring the pose information of the subject to be photographed, the user, or the subject itself, can select one of several potential poses as the preset pose information. Alternatively, the electronic device can determine one of several potential poses as the preset pose information. After acquiring the pose information of the subject, the electronic device matches the subject pose information with the preset pose information.

[0039] If the object's pose information does not match the preset pose information, the electronic device determines the target distance between the object to be photographed and the electronic device. For example, the electronic device can obtain the target distance between the object to be photographed and the electronic device through infrared ranging or time-of-flight (TOF) ranging.

[0040] If the object's pose information matches the preset pose information, the electronic device can directly capture the image of the object and obtain the corresponding image. Alternatively, the electronic device can capture the image of the object after receiving a photo-taking command.

[0041] In step 103, obtain the current time and, based on the target distance and the current time, obtain the target prompt time.

[0042] Different target distances correspond to different target prompt times. For example, target distance and target prompt time can be inversely correlated; that is, the farther the target distance, the earlier the target prompt time. For instance, if the target distance between the subject and the electronic device is 10cm and the current time is 10:20, then the target prompt time could be 10:22; if the target distance is 20cm and the current time is 10:20, then the prompt time could be 10:21.

[0043] In an optional embodiment, a preset mapping relationship between distance and prompt time can also be set in advance. For example, distance D1 corresponds to t1 minutes after the current time (if the current time is T, then the prompt time is T+t1), distance D2 corresponds to t2 minutes after the current time, and distance D3 corresponds to t3 minutes after the current time. Then, assuming the target distance between the object to be photographed and the electronic device is D1, the current time is 10:10, and t1 is 5, then the electronic device can determine the target prompt time as 10:15.

[0044] Understandably, the target prompt time cannot be earlier than the current time.

[0045] In step 104, when the target prompt time is reached, a prompt message is output to prompt the subject to adjust its posture.

[0046] In this embodiment, when the target prompt time is reached, the electronic device outputs a prompt message to prompt the subject to adjust its posture, so that the subject can adjust its posture and thus match the subject's posture information with the preset posture information.

[0047] For example, electronic devices can display preset posture information on the screen, so the subject can adjust its posture according to the preset posture information displayed on the screen, thus making it convenient for the user to adjust their posture.

[0048] For example, electronic devices can also output corresponding prompts on how to make adjustments through display or voice broadcast, such as "Please raise your left hand 1 centimeter", etc.

[0049] In this embodiment, when the object posture information of the object to be photographed does not match the preset posture information, the target prompt time is obtained based on the target distance between the object to be photographed and the electronic device and the object attribute information of the object to be photographed. When the target prompt time is reached, prompt information is output to prompt the object to be photographed to adjust its posture. Thus, the posture of the object to be photographed can be adjusted based on the output prompt information so that the object to be photographed can pose in a suitable posture, thereby improving the quality of the image obtained by photographing the object to be photographed.

[0050] Furthermore, since this application can determine the corresponding target prompt time based on the target distance between the object to be photographed and the electronic device, and output prompt information to adjust the posture when the target prompt time is reached, the flexibility of the output prompt information can be improved.

[0051] In an optional embodiment, obtaining the target prompt time based on the target distance and the current time includes:

[0052] Determine the target age value of the subject to be photographed;

[0053] The target prompt time is obtained based on the target distance, current time, and target age value.

[0054] In this embodiment, the electronic device can also determine the target age value of the object to be photographed, and obtain the target prompt time based on the target distance between the object to be photographed and the electronic device, the current time, and the target age value of the object to be photographed.

[0055] Different target age values ​​can correspond to different target prompt times.

[0056] Considering that younger subjects may lack patience while older subjects may have more patience, at the same target distance, a prompt can be issued earlier for younger subjects and later for older subjects. For example, the electronic device can obtain multiple candidate prompt times based on the target distance. Then, based on the target's age, the electronic device can determine the target prompt time from these candidate times.

[0057] For example, suppose that based on the target distance and the current time (10:18), there are multiple candidate prompt times including 10:20, 10:21 and 10:22. If the target age is 10 years old, the target prompt time can be 10:20; if the target age is 20 years old, the target prompt time can be 10:22.

[0058] In an optional embodiment, a mapping relationship between distance and candidate prompt times can be preset. For example, distance D1 corresponds to candidate prompt times T11, T12, and T13; distance D2 corresponds to candidate prompt times T21, T22, and T23; and distance D3 corresponds to candidate prompt times T31, T32, and T33. A mapping relationship between age values ​​and prompt times can also be preset. For example, age value E1 corresponds to prompt time T11, age value E2 corresponds to prompt time T12, age value E3 corresponds to prompt time T14, and so on. Therefore, assuming the target distance is D1 and the target age value is E1, the electronic device can determine the target prompt time as T11.

[0059] In an optional embodiment, determining the target age value of the subject to be photographed includes:

[0060] Acquire the facial image of the subject to be photographed;

[0061] An age recognition model is used to identify the age of a person in a face image, thus obtaining the target age value of the subject to be photographed.

[0062] For example, an electronic device can take a picture of the subject to be photographed, obtaining a preview image. Then, the electronic device can crop the face region from the preview image to obtain the face image of the subject. Subsequently, the electronic device can input the face image into a pre-trained age recognition model to perform age recognition on the face image and obtain the target age value of the subject to be photographed.

[0063] In an optional embodiment, before acquiring the facial image of the subject to be photographed, the following may be included:

[0064] Establish an age recognition model.

[0065] In this embodiment, multiple different face images are collected as training data, and in order to improve the accuracy of the established age recognition model, the collected face images can correspond to as many different age values ​​as possible.

[0066] Furthermore, based on the acquired training data, deep learning can be performed using relevant technologies to obtain a trained age recognition model.

[0067] In an optional embodiment, considering the influence of the external environment on age recognition, an age correction parameter can be used to represent the parameters that affect the external environment on age recognition. In subsequent age recognition, the target age value is corrected based on the influence of the age correction parameter on age recognition.

[0068] Therefore, after establishing the age recognition model, it can also include:

[0069] Statistical tests were conducted on the training data based on the age correction parameter to determine the correspondence between the correction parameter value and the correction age value.

[0070] The age correction parameters include at least one of the following: illumination parameters, facial expression parameters, and facial pose parameters.

[0071] Taking lighting parameters as an example, for the same face, if the external lighting is strong, that is, the light intensity value in the face image is high, the age value obtained by age recognition will be lower than the actual age. Conversely, if the external lighting is weak, that is, the light intensity value in the face image is low, the age value obtained by age recognition will be higher than the actual age.

[0072] Compared to facial expression parameters, facial expressions can include anger, sadness, happiness, etc. For example, if a face's facial expression is happy, the age value obtained through age recognition will be younger than the actual age. If a face's facial expression is sad, the age value obtained through age recognition will be older than the actual age.

[0073] Facial pose can also affect age recognition results. For example, if different pitch angle values ​​are obtained when estimating the pose of the same face, the age value obtained from age recognition may be larger or smaller than the actual age.

[0074] Of course, in addition to the parameters mentioned above, age correction parameters can also include other parameters that affect the age recognition results. The purpose of this application is to determine the correction values ​​of different age correction parameter values ​​for facial age, thereby improving the accuracy of age estimation.

[0075] In this embodiment of the application, after establishing the age recognition model, statistical tests can be performed on all face images used as training data based on age correction parameters to obtain the degree of influence of age correction parameters on age recognition. Optionally, the degree of influence can be described by the corrected age value.

[0076] Assuming the age correction parameters only include illumination parameters, statistical tests can be performed on all face images used as training data based on these illumination parameters. Illumination detection is performed on the face images using relevant techniques to test different illumination parameter values. The corrected age value for age recognition is then determined based on these tested illumination parameter values. Since the target age value for each face image can already be determined using the pre-trained age recognition model, and the actual age value for each face image is already determined before collecting training data, subtracting the estimated age value from the actual age value yields the corrected age value for the face image based on the illumination parameters.

[0077] For example, when the illumination parameter is a1, the target age value determined by the age recognition model for face image A is b1, and the actual age value of face image A, which was determined during the collection of training data, is b2. Therefore, when the illumination parameter is a1, the corrected age value caused by age recognition will be (b2-b1).

[0078] Furthermore, in the embodiments of this application, the corresponding corrected age value can be determined based on different facial expression parameter values ​​or facial pose parameter values.

[0079] In this embodiment of the application, statistical testing is performed according to the above process, and finally, an estimated function between multiple different age correction parameter values ​​and correction age values ​​can be obtained, as follows:

[0080] h(x) = θ0 + θ1x1 + θ2x2 + ... + θ n x n .

[0081] Where h(x) is the corrected age value, x1, x2, ... x n These are the correction parameter values ​​for different age correction parameters, θ1, θ2…θ n θ0 is the coefficient corresponding to the age correction parameter, and θ0 is the offset of all age correction parameters from the corrected age value.

[0082] For example, x1 can represent the illumination parameter value in the age correction parameters, x2 can represent the facial expression parameter value in the age correction parameters, and x3 can represent the facial pose parameter value in the age correction parameters. n This can represent the parameter value corresponding to the nth parameter in the age correction parameters. θ1, θ2…θ nThese correspond to the lighting parameters, facial expression parameters, facial pose parameters, and the nth parameter, respectively, and θ0, θ1, θ2…θ n Once determined, the numerical value of the identification model relative to the same age will not change.

[0083] Optionally, the range of the corrected age value can be between [-100, 100].

[0084] After training the face age recognition model through the above process and determining the estimation function, when performing age recognition based on the same age recognition model, it is no longer necessary to repeat the above steps of establishing the face age model and statistically testing the training data based on the age correction parameters to determine the correspondence between the correction parameter values ​​and the correction age values.

[0085] When age recognition is required, a facial image of the subject is acquired. This image is then input into a pre-trained age recognition model to obtain the target age value of the subject. The target correction parameter value for the age correction parameter in the facial image is determined. Based on the estimated age value and the target correction parameter value, the corrected age value of the subject is determined. The target prompt time is obtained based on the target distance, current time, and corrected age value. The specific implementation of obtaining the target prompt time based on the target distance, current time, and corrected age value can be found in the previous section on obtaining the target prompt time based on the target distance, current time, and target age value, and will not be repeated here. The target age value can be in the range [0, 100].

[0086] In this embodiment, the age correction parameters include at least one of the following: illumination parameters, facial expression parameters, and facial pose parameters.

[0087] The following sections determine the target correction parameter values ​​for different age correction parameters in the face images of the subjects to be photographed.

[0088] The process of determining the target illumination parameter values ​​is as follows:

[0089] Light detection can be performed on the facial image of the subject to be photographed using relevant technologies to obtain the target lighting parameter values, such as the light intensity value in the facial image of the subject.

[0090] The process for determining the target facial expression parameter values ​​is as follows:

[0091] It can perform facial expression recognition on the facial image of the subject to be photographed using relevant technologies to obtain the target facial expression parameter value.

[0092] The process for determining the target face pose parameter values ​​is as follows:

[0093] The face image of the subject to be photographed can be used to perform face pose detection according to relevant technologies to obtain the target face pose parameter values, such as the rotation angle value and pitch angle value of the target face.

[0094] Once the target correction parameter value is determined, the target corrected age value is calculated based on the pre-determined correspondence between the age correction parameter value and the corrected age value, i.e., the estimation function mentioned above. For example, the target correction parameter value can be substituted into the estimation function to calculate the target corrected age value. The target corrected age value also falls within the range of [-100, 100].

[0095] Once the target corrected age value is determined, the sum of the target age value and the target corrected age value is used as the corrected age value of the subject to be photographed. For example, the sum of the target age value and the target corrected age value can be directly calculated, and the final result is used as the corrected age value of the subject to be photographed.

[0096] In an optional embodiment, obtaining the target prompt time based on the target distance and the current time includes:

[0097] Determine the gender of the subject to be photographed;

[0098] Get the target prompt time based on the target distance, current time, and gender.

[0099] Different genders can correspond to different target prompt times.

[0100] For example, an electronic device can obtain multiple candidate prompt times based on the target distance and the current time. Then, the electronic device can determine the target prompt time from the multiple candidate prompt times based on gender.

[0101] For example, suppose that the multiple candidate prompt times determined based on the target distance and the current time (10:15) include 10:20 and 10:21. If the gender is female, the target prompt time can be 10:21; if the gender is male, the target prompt time can be 10:20.

[0102] In an optional embodiment, a mapping relationship between distance and candidate prompt times can be preset. For example, distance D1 corresponds to candidate prompt times T11 and T12, distance D2 corresponds to candidate prompt times T21 and T22, and distance D3 corresponds to candidate prompt times T31 and T32. A mapping relationship between gender and prompt times can also be preset. For example, female gender corresponds to prompt times T11, T21, and T31, and male gender corresponds to prompt times T12, T22, and T32. Therefore, assuming the target distance is D2 and the gender is female, the electronic device can determine the prompt time as T21.

[0103] In an optional embodiment, determining the gender of the subject to be photographed includes:

[0104] Acquire the facial image of the subject to be photographed;

[0105] A gender recognition model is used to identify the gender of a person in a face image, thus determining the gender of the subject to be photographed.

[0106] For example, an electronic device can take a picture of the subject to be photographed, obtaining a preview image. Then, the electronic device can crop the face region from the preview image to obtain the face image of the subject. Subsequently, the electronic device can input the face image into a pre-trained gender recognition model to perform gender recognition on the face image and obtain the gender of the subject to be photographed.

[0107] In an optional embodiment, before acquiring the facial image of the subject to be photographed, the process may further include:

[0108] Acquire multiple training face images;

[0109] The first neural network of the preset model is used to obtain the high-level feature set of multiple training face images, and the second neural network of the preset model is used to obtain the low-level feature set of multiple training face images.

[0110] The high-level feature set and the low-level feature set are fused to obtain the fused feature set;

[0111] The fused feature set is used as training data and input into the prediction module of the preset model for training to obtain the gender recognition model.

[0112] The face database includes multiple face images, including frontal, side, and multi-angle images. The multi-angle images include multiple top-down, bottom-up, and side-view images, etc. Therefore, multiple training face images can be obtained from the face database.

[0113] Multiple facial images can be used, where one facial image corresponds to one user, or multiple facial images correspond to the same user. Furthermore, multiple facial images corresponding to the same user can be facial images from multiple angles.

[0114] A face database can contain multiple face images, ranging from high-resolution to low-resolution, and from images with varying degrees of noise to images in multiple poses. These multiple pose images include smiling faces, serious faces, and faces displaying various expressions.

[0115] Facial databases can be built by users themselves, such as by collecting a large number of facial images online, collecting facial images of themselves and their relatives and friends, or taking a large number of facial images on the street.

[0116] Face databases can also utilize existing face databases, such as the CelebA database.

[0117] It should be noted that the facial images in the facial database correspond to their gender characteristics.

[0118] Multiple training face images can be randomly selected from a face database, or multiple training face images can be selected based on user information. For example, images can be selected based on user-captured photos. Specifically, if the user's photos show a high proportion of East Asian faces, then the selected training face images will have a higher proportion of East Asian faces. If the user's photos show a high proportion of children's faces, then the selected training face images will have a higher proportion of children's faces. If the user's photos show a high proportion of female faces, then the selected training face images will have a higher proportion of female faces.

[0119] Images have three main underlying features: color, texture, and shape. These underlying features also include color, brightness, orientation, texture, and edge features. The set of underlying features consists of a collection of various underlying features from multiple training face images.

[0120] High-level features are extracted from lower-level features to reveal more advanced semantic information about an image. They can also be understood as features constructed using specific algorithms (such as convolutional neural networks) based on lower-level features, generally referring to more complex features such as the contours of objects in an image. Compared to simply extracting the raw information of an image from lower-level features, high-level features are more expressive, fully considering the contextual information of the scene. A high-level feature set includes a collection of various high-level features from multiple training face images.

[0121] Specifically, the first neural network of the preset model, such as the deep convolutional neural network of the preset model, can be used to obtain the high-level feature set of multiple training face images, and the second neural network of the preset model, such as the shallow convolutional neural network of the preset model, can be used to obtain the low-level feature set of multiple training face images.

[0122] Feature fusion can be understood as integrating features from different sources and removing redundancy; the resulting integrated information will be beneficial for our subsequent analysis and processing. Specifically, feature fusion can be implemented through algorithms, such as algorithms based on Bayesian decision theory, algorithms based on sparse representation theory, and algorithms based on deep learning theory.

[0123] The process involves obtaining a high-level feature set and a low-level feature set, then fusing them to obtain a fused feature set. Specifically, high-level and low-level features corresponding to the same input information are fused to obtain fused features. The same input information can be the same face image or a specific feature within the same face image, such as skin color. After inputting multiple inputs, multiple high-level and low-level features are obtained. Based on these multiple high-level and low-level features, multiple fused features are derived, forming a fused feature set.

[0124] Convolutional Neural Networks (CNNs) are a type of feedforward neural network. Their artificial neurons can respond to surrounding units within a certain coverage area, making them excellent for large-scale image processing. A CNN consists of convolutional layers and pooling layers. Specifically, deep CNNs contain more convolutional and pooling layers than shallow CNNs. Deep CNNs can be used to obtain high-level feature sets of an image, while shallow CNNs can be used to obtain low-level feature sets.

[0125] After obtaining the fused feature set, it is used as training data to input into the prediction module of the preset model for training. The prediction module of the preset model learns and optimizes the various calculation parameters within the preset model based on the training data, thus obtaining the trained preset model.

[0126] Specifically, a training face image can be obtained first. Then, the low-level feature set and high-level feature set of this training face image are obtained. These two sets are then fused and input into the prediction module of a preset model, such as a Logistic Regression Classifier, for training and learning to obtain a prediction result. If the prediction result is correct, the computational parameters of the preset model after training are retained; if the prediction result is incorrect, the computational parameters of the preset model are modified and training continues until the prediction result is correct. Then, the above steps are repeated with other training face images until all training face images have been predicted once. Then, all training face images are re-predicted one or more times until the prediction result no longer changes, resulting in the final optimized computational parameters. The preset model with the final optimized computational parameters is the trained gender recognition model.

[0127] Multiple training face images can also be input into a preset model. The preset model predicts the gender of each training face image, obtaining a prediction result. For example, if the probability of predicting male is 70% and the probability of predicting female is 30%, then the prediction result is considered male, with a probability of 70%. The model is scored based on pre-set correct results: if the correct result is male, the prediction is correct; if the correct result is female, the prediction is incorrect. The calculation parameters of the prediction model are adjusted based on the correct probability of the prediction results. When the accuracy of the adjusted prediction model cannot be further improved, and the prediction probability for each training face image cannot be improved overall, then the calculation parameters of the prediction model at this point are considered optimal. The preset model with these optimal calculation parameters is the trained gender recognition model.

[0128] It should be noted that during the training process, the computational parameters of the prediction module can be changed, as can the computational parameters of the first and second neural networks, to optimize all computational parameters in the entire preset model.

[0129] It should be noted that the above process refers to the training of a preset model. The training process can be performed on a server. After training, the trained gender recognition model is ported to an electronic device such as a smartphone. The electronic device then uses the trained gender recognition model to determine the facial image of the subject to be photographed. Alternatively, the training process can also be performed on the electronic device. After training, the electronic device directly uses the trained gender recognition model to determine the facial image of the subject to be photographed. Furthermore, the training process can also be performed on a server. When the electronic device needs to determine the facial image of the subject to be photographed, it sends the subject's facial image to the server, the server performs the determination, and then sends the determination result back to the electronic device.

[0130] In some embodiments, after determining the gender of the subject's facial image, the preview image (the image displayed in real-time on the screen of an electronic device) can be optimized based on that gender. For example, if the facial image is determined to be male, a low level of beautification is applied to the preview image. If the facial image is determined to be female, a high level of beautification is applied to the preview image. Different optimization strategies can also be set based on gender; for example, for females, whitening, skin smoothing, dark circle removal, and adding decorations can be performed.

[0131] In an optional embodiment, after outputting a prompt message to prompt the subject to adjust its posture when the target prompt time is reached, the method further includes:

[0132] When the pose information of the object to be photographed matches the preset pose information, the object to be photographed is photographed to obtain the target image.

[0133] In this embodiment, when the object posture information of the object to be photographed matches the preset posture information, the electronic device can automatically photograph the object and obtain the target image.

[0134] In an optional embodiment, when the object posture information of the object to be photographed matches the preset posture information, the user can perform the corresponding photo-taking operation, such as issuing the corresponding voice, such as "Please take a photo", so that the electronic device receives the photo-taking instruction. In response to the photo-taking instruction, the electronic device can take a picture of the object to be photographed and obtain the target image.

[0135] In an optional embodiment, obtaining the object pose information of the object to be photographed includes:

[0136] Acquire human images of the subject to be photographed;

[0137] A key point detection model is used to detect key points in human images to obtain multiple key points corresponding to the subject to be photographed.

[0138] The pose information of the subject is determined based on multiple key points corresponding to the subject.

[0139] For example, an electronic device can acquire a preview image of the object to be photographed and use a pre-trained human detection model to perform human detection on the preview image, obtaining a human bounding box. The electronic device can then crop the human image from the preview image based on this bounding box, using it as the human image of the object to be photographed. The electronic device can then use a pre-trained keypoint detection model to detect keypoints on the human image, obtaining the positions of multiple keypoints corresponding to the object to be photographed. Based on these multiple keypoint positions, the electronic device can determine the pose information of the object to be photographed. For example, the electronic device can use the positions of multiple keypoints corresponding to the object to be photographed as the object pose information. Alternatively, the electronic device can use a pre-trained pose recognition model to perform pose recognition based on these multiple keypoint positions, obtaining the pose of the object to be photographed. The electronic device can then use the pose of the object to be photographed as its object pose information.

[0140] In an optional embodiment, after obtaining a human image of the object to be photographed, the electronic device inputs the human image into a preset key point detection model to obtain multiple corresponding heatmaps. Based on the multiple heatmaps, the electronic device obtains the coordinates of multiple key points of the object to be photographed, wherein one heatmap corresponds to one key point coordinate.

[0141] For example, electronic devices can pre-train a pre-defined Cascaded Pyramid Network (CPN) model and use the trained CPN model as a pre-defined keypoint detection model. After obtaining a human image, the electronic device can input the human image into the pre-defined keypoint detection model to obtain multiple corresponding heatmaps.

[0142] After obtaining multiple heatmaps, the electronic device can locate the position of the pixel with the highest probability on each heatmap. The position of the pixel with the highest probability on each heatmap is the key point coordinate of that heatmap, thus obtaining the coordinates of multiple key points of the object to be photographed. The electronic device can then use these multiple key point coordinates of the object to be photographed as the object's pose information.

[0143] The number of keypoints can be 14, 17, or 21, etc., without specific restrictions. The keypoint coordinates include x-coordinates and y-coordinates, meaning that each keypoint coordinate can be represented by a set of (x, y) coordinates.

[0144] It is understood that in the embodiments of this application, the heatmaps and key point coordinates are in one-to-one correspondence. For example, if there are 17 heatmaps, 17 key point coordinates can be obtained; if there are 21 heatmaps, 21 key point coordinates can be obtained.

[0145] In some embodiments, the electronic device obtains multiple corresponding heatmaps based on the human image input into a preset keypoint detection model, which may include:

[0146] The electronic device inputs the human body image into a preset key point detection model to obtain multiple sets of corresponding feature maps, where each set of feature maps includes multiple feature maps of different sizes;

[0147] The electronic device fuses the feature maps in each group of feature maps to obtain multiple corresponding heat maps, where one group of feature maps corresponds to one heat map.

[0148] For example, an electronic device can arrange multiple feature maps of different scales in each set of feature maps in descending order. Then, the electronic device identifies the middle feature map in each set as the first feature map. Next, using this first feature map as a standard, the electronic device can upsample or downsample the other feature maps in each set so that the size of the upsampled or downsampled feature maps is the same as the size of the first feature map. Subsequently, the electronic device can fuse the first feature map and the upsampled or downsampled feature maps to obtain multiple corresponding heatmaps.

[0149] It is understood that upsampling enlarges the size of the feature map, while downsampling reduces the size of the feature map. In the embodiments of this application, feature maps smaller than the first feature map in each group of feature maps can be upsampled, and feature maps larger than the first feature map in each group of feature maps can be downsampled.

[0150] In an optional embodiment, the electronic device can input a human image into a preset keypoint detection model, and obtain multiple sets of corresponding second feature maps through the residual blocks of multiple convolutional layers (such as convolutional layers c2, c3, c4, and c5) of the preset keypoint detection model. Each set of second feature maps includes multiple second feature maps, and each convolutional layer corresponds to one of the second feature maps in each set. The depth of convolutional layer c2 is less than the depth of convolutional layer c3, the depth of convolutional layer c3 is less than the depth of convolutional layer c4, and the depth of convolutional layer c4 is less than the depth of convolutional layer c5. Then, the electronic device can connect the multiple second feature maps in each set of second feature maps with different numbers of bottleneck blocks to obtain multiple sets of corresponding feature maps. Each set of feature maps includes multiple feature maps of different sizes. The deeper the convolutional layer, the more bottleneck blocks are connected to the corresponding feature map. Next, the electronic device can upsample the feature maps in each set of feature maps to a unified dimension and then perform a fusion process, such as adding the upsampled and unified dimension feature maps pixel by pixel to obtain multiple corresponding heatmaps.

[0151] In an optional embodiment, after the electronic device obtains multiple corresponding heatmaps from a preset keypoint detection model based on the human image, it may further include:

[0152] The electronic device performs Gaussian filtering on each heatmap to obtain multiple target heatmaps.

[0153] The electronic device obtains the coordinates of multiple key points of the object to be photographed based on multiple heat maps, which may include:

[0154] The electronic device obtains the coordinates of multiple key points of the object to be photographed based on multiple target heat maps, where one target heat map corresponds to one key point coordinate.

[0155] For example, since each of the multiple heatmaps obtained by the electronic device contains some noise, the device can perform Gaussian filtering on each heatmap to remove noise and obtain multiple target heatmaps. Then, based on these target heatmaps, the electronic device can obtain the coordinates of multiple key points of the object to be photographed. Each target heatmap corresponds to one key point coordinate.

[0156] It should be noted that noise refers to points that interfere with the determination of key points; that is, the presence of noise may lead to inaccurate determination of key points.

[0157] Understandably, determining key point coordinates based on the target heatmap is more accurate than determining them based on the heatmap. However, obtaining the target heatmap also requires certain processor resources. Therefore, when processor resources are sufficient, key point coordinates can be determined based on the target heatmap; when processor resources are insufficient, key point coordinates can be determined based on the heatmap.

[0158] In some embodiments, prior to acquiring the human image of the subject to be photographed, the method may further include:

[0159] Electronic devices acquire multiple sample human body images;

[0160] Electronic devices acquire the coordinates of multiple key points corresponding to the human body in each sample human body image;

[0161] Electronic devices use multiple sample human images and the coordinates of multiple key points corresponding to the human body in each sample human image to train a pre-set neural network model.

[0162] Electronic devices use the trained neural network model as a preset key point detection model.

[0163] For example, an electronic device can retrieve multiple sample human body images stored in a database or other device. Each sample human body image is marked with multiple keypoint coordinates. These keypoint coordinates correspond to the human body in each sample human body image. In this embodiment, the electronic device can obtain the multiple keypoint coordinates marked in each sample human body image, that is, the multiple keypoint coordinates corresponding to the human body in each sample human body image.

[0164] After obtaining multiple sample human body images and the coordinates of multiple key points corresponding to the human body in each sample human body image, the electronic device can use these multiple sample human body images and the coordinates of multiple key points corresponding to the human body in each sample human body image to train a preset neural network model. The trained neural network model is the preset key point detection model.

[0165] In some embodiments, the electronic device may also train a preset neural network model using the plurality of sample human body images, the coordinates of multiple key points corresponding to the human body in each sample human body image, and a preset loss function. The trained neural network model is the preset key point detection model.

[0166] It's important to note that the loss function is typically used to estimate the degree of discrepancy between the model's predictions (such as the predicted coordinates of keypoints) and the true values ​​(such as the actual labeled coordinates of keypoints). It is a non-negative real-valued function. Generally, the smaller the loss function, the better the model's robustness. The loss function can be set according to specific needs.

[0167] The preset neural network model can be a cascaded pyramid network model. This cascaded pyramid network model may include a GlobalNet network and a RefineNet network. The GlobalNet network can be used for coarse training on all key points of the human body. The RefineNet network can refine the key points that are difficult to train as reflected by the GlobalNet network.

[0168] In some embodiments, the pre-defined neural network model may include an Inception-v4 network or an Attention ResNet network and a RefineNet network. The Inception-v4 network or Attention ResNet network can be used for coarse training on all key points of the human body. The RefineNet network can refine the key points that are difficult to train, as reflected by the GlobalNet network.

[0169] In some embodiments, prior to acquiring the human image of the subject to be photographed, the method may further include:

[0170] The electronic device acquires multiple sets of key point coordinates, where each set of key point coordinates includes multiple key point coordinates;

[0171] The electronic device acquires the human posture corresponding to the coordinates of each set of key points.

[0172] Electronic devices use multiple sets of key point coordinates and the human posture corresponding to each set of key point coordinates to train a preset shallow neural network model.

[0173] Electronic devices use the trained shallow neural network model as the preset pose recognition model.

[0174] For example, electronic devices can acquire multiple sets of keypoint coordinates and the corresponding human posture for each set of keypoint coordinates. Each set of keypoint coordinates includes multiple keypoint coordinates.

[0175] Having obtained multiple sets of keypoint coordinates and the corresponding human pose for each set of keypoint coordinates, the electronic device can train a pre-defined shallow neural network model using these coordinates. The trained shallow neural network model can then be used as a pre-defined pose recognition model.

[0176] In some embodiments, the electronic device may also train a preset shallow neural network model using multiple sets of keypoint coordinates, the human pose (real human pose) corresponding to each set of keypoint coordinates, and a preset loss function. The trained shallow neural network model can then be used as a preset pose recognition model.

[0177] It's important to note that the loss function is typically used to estimate the degree of discrepancy between the model's predictions (such as the model's predicted human pose) and the true values ​​(such as the actual human pose). It is a non-negative real-valued function. Generally, the smaller the loss function, the better the model's robustness. The loss function can be set according to specific needs.

[0178] The preset shallow neural network model can be the ResNet 18 network model.

[0179] In some embodiments, since the coordinate representations of two people in the same pose at different positions in an image are very different, in order to control this variable, the electronic device can normalize the keypoint coordinates among multiple sets of keypoint coordinates after acquiring them. For example, the following formula can be used to normalize the keypoint coordinates:

[0180]

[0181] In this formula, N2 represents the normalized x-coordinate or y-coordinate. N1 represents the original x-coordinate or y-coordinate. min This represents the x-coordinate or y-coordinate with the smallest value among multiple sets of keypoint coordinates. N max This represents the x-coordinate or y-coordinate with the largest value among multiple sets of keypoint coordinates. A is a constant, and the value of A can be 240, 264, 293, 320, 335, 370, etc.

[0182] In other embodiments, to demonstrate the correlation between the x and y coordinates of the same keypoint, the x and y coordinates of the same keypoint can be placed at the same position in different channels for training. For example, suppose a set of keypoints includes 5 keypoints with coordinates (x1, y1), (x2, y2), (x3, y3), (x4, y4), and (x5, y5), and the human posture corresponding to this set of keypoints is "standing". The training data to be input into the preset shallow neural network model is (a, b). Then, [x1, x2, x3, x4, x5] and [y1, y2, y3, y4, y5] can be used as a, and the human posture "standing" can be used as b.

[0183] In an optional embodiment, after determining the coordinates of multiple key points of the object to be photographed, the electronic device can input these coordinates into a preset pose recognition model to identify the pose of the object. The electronic device can then use the pose of the object as its pose information. The preset pose recognition model is a trained model.

[0184] In an optional embodiment, before determining the target distance between the object to be photographed and the electronic device when the object pose information does not match the preset pose information, the method further includes:

[0185] Obtain scene information about the location of the subject being photographed;

[0186] Based on the scene information, the preset posture information is determined from multiple candidate posture information.

[0187] The scene information describes the environment in which the subject is being photographed. For example, if the subject is at the beach, the scene information would be "beach". If the subject is at the foot of a mountain, the scene information would be "foot of the mountain". If the subject is next to a building, the scene information would be "next to the building".

[0188] The electronic device can pre-collect multiple candidate pose information and multiple scene information, and store the corresponding candidate pose information and scene information in a one-to-one correspondence. The correspondence between candidate pose information and scene information can be determined by professional photographers. Alternatively, professional photographers can determine a portion of the correspondence between candidate pose information and scene information, and then the electronic device can perform machine learning based on the determined correspondence to determine the correspondence between the remaining candidate pose information and scene information. One candidate pose information can correspond to one or more scene information, and one scene information can also correspond to one or more candidate pose information.

[0189] In this embodiment, the electronic device acquires scene information of the scene where the object to be photographed is located, and determines preset posture information from multiple candidate posture information based on the scene information. This step can be performed before, after, or simultaneously with the step of acquiring the object posture information of the object to be photographed.

[0190] For example, suppose the candidate posture information P1 corresponds to the scene information S1, the candidate posture information P2 corresponds to the scene information S2, and the candidate posture information P3 corresponds to the scene information S3. If the scene information of the scene where the subject is to be photographed is S1, the electronic device can determine the preset posture information as P1. If the scene information of the scene where the subject is to be photographed is S2, the electronic device can determine the preset posture information as P2. If the scene information of the scene where the subject is to be photographed is S3, the electronic device can determine the preset posture information as P3.

[0191] Understandably, if there are multiple candidate pose information corresponding to the scene information of the scene where the subject is located, the electronic device can determine any one of the candidate pose information corresponding to the scene information of the scene where the subject is located as the preset pose information. Alternatively, the electronic device can determine the candidate pose information with the highest usage frequency corresponding to the scene information of the scene where the subject is located as the preset pose information.

[0192] In an optional embodiment, before determining the target distance between the object to be photographed and the electronic device when the object pose information does not match the preset pose information, the method further includes:

[0193] The preset posture information is determined from multiple candidate posture information based on the object's posture information.

[0194] In order to ensure that the subject poses appropriately and in accordance with their preferences, the subject can first pose in their preferred manner. The electronic device can then acquire the subject's pose information and determine the preset pose information from multiple candidate pose information based on this information.

[0195] For example, an electronic device can select the candidate posture information that is closest to the object's posture information from multiple candidate posture information as the pre-selected posture information.

[0196] Taking the candidate posture information as an example, which includes multiple key point positions, the electronic device acquires the multiple key point positions included in the object posture information; the electronic device compares the multiple key point positions included in each candidate posture information with the multiple key point positions included in the object posture information one by one, and determines the candidate posture information that is closest to the object posture information from the multiple candidate posture information as the preset posture information. For example, the candidate posture information with the smallest Euclidean target distance between the included key point positions is used as the preset posture information, or the candidate posture information with the largest number of key point positions with the smallest target distance between each included key point position is used as the preset posture information.

[0197] For example, suppose there are multiple candidate posture information, including candidate posture information P1 and candidate posture information P2. Candidate posture information P1 includes multiple keypoint positions: left elbow position L11, right elbow position L12, and neck position L13. Candidate posture information P2 includes multiple keypoint positions: left elbow position L21, right elbow position L22, and neck position L23. The object posture information T includes multiple keypoint positions: left elbow position LT1, right elbow position LT2, and neck position LT3. The electronic device calculates the Euclidean target distance D1 between the left elbow position L11, right elbow position L12, and neck position L13 and the left elbow position LT1, right elbow position LT2, and neck position LT3, and calculates the Euclidean target distance D2 between the left elbow position L21, right elbow position L22, and neck position L23 and the left elbow position LT1, right elbow position LT2, and neck position LT3. Assuming D1 is greater than D2, the electronic device can determine that the preset posture information is P2.

[0198] For example, suppose there are multiple candidate pose information including candidate pose information P1 and candidate pose information P2. The candidate pose information P1 includes multiple key point positions, namely the left elbow position L11, the right elbow position L12 and the neck position L13. The candidate pose information P2 includes multiple key point positions, namely the left elbow position L21, the right elbow position L22 and the neck position L23. The object pose information T includes multiple key point positions, namely the left elbow position LT1, the right elbow position LT2 and the neck position LT3. Among them, the target distance between L11 and LT1 is greater than the target distance between L21 and LT1, the target distance between L12 and LT2 is greater than the target distance between L22 and LT2, and the target distance between L13 and LT3 is greater than the target distance between L23 and LT3. Since the target distances between key point positions L21, L22, and L23 and key point positions LT1, LT2, and LT3 are all greater than the target distances between key point positions L11, L12, and L13 and key point positions LT1, LT2, and LT3, it can be determined that the candidate posture information P2 is closest to the object posture information T. Therefore, the candidate posture information P2 can be determined as the preset posture information.

[0199] In an optional embodiment, before acquiring the object pose information of the object to be photographed, the method further includes:

[0200] Displays multiple candidate attitude information;

[0201] In response to a selection operation for multiple candidate posture information, the candidate posture information selected by the selection operation is determined as the preset posture information.

[0202] In this embodiment, the user can select a preset posture from multiple candidate posture information. For example, the electronic device can display multiple candidate posture information on the screen, such as constructing a corresponding contour based on each candidate posture information and displaying the constructed contour on the screen. The user can issue a corresponding voice or perform a corresponding touch on the screen, thereby receiving a selection operation for multiple candidate posture information from the electronic device. Then, the electronic device can determine the candidate posture information selected by the selection operation as the preset posture information.

[0203] For example, such as Figure 3 As shown, assuming the electronic device displays candidate posture information P1, P2, P3, P4, and P5 on the screen, if the user clicks on candidate posture information P1, then the preset posture information is P1. If the user clicks on candidate posture information P3, then the preset posture information is P3. If the user says "Select the first candidate posture information," then the preset posture information is P1. If the user says "Select the fourth candidate posture information," then the preset posture information is P4.

[0204] In an alternative embodiment, the electronic device may determine preset posture information from a plurality of candidate posture information based on scene information and object posture information.

[0205] For example, an electronic device can first determine multiple candidate posture information corresponding to scene information from multiple candidate posture information, and then determine the preset posture information from multiple candidate posture information corresponding to scene information based on object posture information.

[0206] For example, an electronic device can first determine multiple candidate posture information corresponding to the object posture information from multiple candidate posture information (such as multiple candidate posture information with the smallest Euclidean target distance), and then determine the preset posture information from multiple candidate posture information corresponding to the object posture information based on scene information.

[0207] In practical applications, it is sometimes necessary to obtain the customer's image when they conduct banking business. Obtaining this image typically requires the customer to assume a specific pose to verify their identity. In such cases, the information prompting method provided in this application can be used to obtain the customer's image. For example, the required pose is used as preset pose information; the customer's object pose information is obtained and matched with the preset pose information. When the object pose information does not match the preset pose information, the target distance between the object to be photographed and the electronic device is determined; based on the target distance, a target prompt time is obtained; when the target prompt time is reached, a prompt message to adjust the pose is output, allowing the customer to adjust their pose and thus assume the appropriate pose to prove their identity.

[0208] Please seeFigure 4 , Figure 4 This is a schematic diagram of the structure of the information prompting device provided in an embodiment of this application. The information prompting device 200 includes: an information acquisition module 201, a distance determination module 202, a time acquisition module 203, and an information output module 204.

[0209] The information acquisition module 201 is used to identify the posture of the object to be photographed in the shooting scene and obtain the object posture information of the object to be photographed.

[0210] The distance determination module 202 is used to determine the target distance between the object to be photographed and the electronic device when the object posture information does not match the preset posture information;

[0211] The time acquisition module 203 is used to acquire the current time and, based on the target distance and the current time, acquire the target prompt time;

[0212] The information output module 204 is used to output a prompt message to prompt the subject to be photographed to adjust its posture when the target prompt time is reached.

[0213] In an optional embodiment, the time acquisition module 203 can be used to: determine the target age value of the object to be photographed; and acquire the target prompt time based on the target distance, the current time, and the target age value.

[0214] In an optional embodiment, the time acquisition module 203 can be used to: acquire a facial image of the subject to be photographed; and use an age recognition model to perform age recognition on the facial image to obtain a target age value of the subject to be photographed.

[0215] In an optional embodiment, the time acquisition module 203 can be used to: determine the gender of the object to be photographed; and acquire the target prompt time based on the target distance, the current time, and the gender.

[0216] In an optional embodiment, the time acquisition module 203 can be used to: acquire a facial image of the subject to be photographed; and use a gender recognition model to perform gender recognition on the facial image to obtain the gender of the subject to be photographed.

[0217] In an optional embodiment, the information prompting device 200 may further include a shooting module, which is used to: when the object posture information of the object to be photographed matches the preset posture information, to take a picture of the object to be photographed and obtain a target image.

[0218] In an optional embodiment, the information acquisition module 201 can be used to: acquire a human image of the object to be photographed; perform key point detection on the human image using a key point detection model to obtain multiple key points corresponding to the object to be photographed; and determine the object posture information of the object to be photographed based on the multiple key points corresponding to the object to be photographed.

[0219] It should be noted that the information prompting device provided in this application embodiment and the information prompting method in the above embodiment belong to the same concept, and the specific implementation process can be found in the above embodiments, which will not be repeated here.

[0220] This application provides a storage medium storing a computer program. When the stored computer program is executed on the processor of the electronic device provided in this application, the processor of the electronic device performs any of the steps in the information prompting method suitable for the electronic device. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0221] This application also provides an electronic device, please refer to Figure 5 The electronic device 300 includes components such as a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, electronic device 300 may also include a screen or a camera, etc.

[0222] The processor 301 is the control center of the electronic device. It connects various parts of the electronic device through various interfaces and lines. By running or executing the application program stored in the memory 302 and calling the data stored in the memory 302, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.

[0223] Memory 302 can be used to store applications and data. The applications stored in memory 302 contain executable code. Applications can be composed of various functional modules. Processor 301 executes various functional applications and data processing by running the applications stored in memory 302.

[0224] In this embodiment, the processor 301 in the electronic device loads the executable code corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 302 runs the applications stored in the memory 302, thereby realizing the process:

[0225] The pose of the object to be photographed in the shooting scene is identified to obtain the object pose information of the object to be photographed;

[0226] When the object's posture information does not match the preset posture information, the target distance between the object to be photographed and the electronic device is determined.

[0227] Obtain the current time, and based on the target distance and the current time, obtain the target prompt time;

[0228] When the target prompt time is reached, a prompt message is output to prompt the subject to adjust its posture.

[0229] Please see Figure 6 , Figure 6 This is a schematic diagram of a second structure of an electronic device provided in an embodiment of this application.

[0230] The electronic device 300 may include components such as a memory 302, a processor 301, an input unit 303, and an output unit 304.

[0231] The processor 301 is the control center of the electronic device. It connects various parts of the electronic device through various interfaces and lines. By running or executing the application program stored in the memory 302 and calling the data stored in the memory 302, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.

[0232] Memory 302 can be used to store applications and data. The applications stored in memory 302 contain executable code. Applications can be composed of various functional modules. Processor 301 executes various functional applications and data processing by running the applications stored in memory 302.

[0233] The input unit 303 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.

[0234] The output unit 304 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic devices. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. The output unit may include a display panel.

[0235] In this embodiment, the processor 301 in the electronic device loads the executable code corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302, thereby realizing the process:

[0236] The pose of the object to be photographed in the shooting scene is identified to obtain the object pose information of the object to be photographed;

[0237] When the object's posture information does not match the preset posture information, the target distance between the object to be photographed and the electronic device is determined.

[0238] Obtain the current time, and based on the target distance and the current time, obtain the target prompt time;

[0239] When the target prompt time is reached, a prompt message is output to prompt the subject to adjust its posture.

[0240] In some implementations, when the processor 301 executes the process of obtaining the target prompt time based on the target distance and the current time, it may perform the following: determining the target age value of the object to be photographed; and obtaining the target prompt time based on the target distance, the current time, and the target age value.

[0241] In some implementations, when the processor 301 executes the process of determining the target age value of the subject to be photographed, it may perform the following: acquire a facial image of the subject to be photographed; and use an age recognition model to perform age recognition on the facial image to obtain the target age value of the subject to be photographed.

[0242] In some implementations, when the processor 301 executes the process of obtaining the target prompt time based on the target distance and the current time, it may perform the following: determining the gender of the object to be photographed; and obtaining the target prompt time based on the target distance, the current time, and the gender.

[0243] In some implementations, when the processor 301 determines the gender of the object to be photographed, it may perform the following: acquire a facial image of the object to be photographed; and use a gender recognition model to perform gender recognition on the facial image to obtain the gender of the object to be photographed.

[0244] In some implementations, after the processor 301 outputs a prompt message to prompt the object to be photographed to adjust its posture when the target prompt time is reached, it may also perform the following: when the object posture information of the object to be photographed matches the preset posture information, the processor takes a picture of the object to be photographed to obtain a target image.

[0245] In some implementations, when the processor 301 executes the process of acquiring the object pose information of the object to be photographed, it may perform the following: acquiring a human image of the object to be photographed; using a key point detection model to perform key point detection on the human image to obtain multiple key points corresponding to the object to be photographed; and determining the object pose information of the object to be photographed based on the multiple key points corresponding to the object to be photographed.

[0246] The above provides a detailed description of an information prompting method, apparatus, storage medium, and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. An information prompting method, applied to electronic devices, characterized in that, include: The pose of the object to be photographed in the shooting scene is identified to obtain the object pose information of the object to be photographed; When the object's posture information does not match the preset posture information, the target distance between the object to be photographed and the electronic device is determined. Obtain the current time, and based on the target distance and the current time, obtain the target prompt time; The target notification time is later than the current time; When the target prompt time is reached, a prompt message is output to prompt the subject to adjust its posture; The step of obtaining the target prompt time based on the target distance and the current time includes: Determine the target age value of the object to be photographed, and obtain the target prompt time based on the target distance, the current time, and the target age value; Alternatively, determine the gender of the object to be photographed, and obtain the target prompt time based on the target distance, the current time, and the gender; Determining the target age value of the subject to be photographed includes: Acquire the facial image of the subject to be photographed; The age of the face image is determined by using an age recognition model to obtain the target age value of the subject to be photographed. The target age value is corrected according to age correction parameters; the age correction parameters include at least one of illumination parameters, facial expression parameters, and facial pose parameters. The estimation function between the age correction parameter value and the corrected age value is as follows: h(x)=θ0+θ1x1+θ2x2+… +θ n x n ; Where h(x) is the corrected age value, x1, x2, ... x n These are the correction parameter values ​​for different age correction parameters, θ1, θ2…θ n θ0 is the coefficient corresponding to the age correction parameter, and θ0 is the offset of all age correction parameters from the corrected age value.

2. The information prompting method according to claim 1, characterized in that, Determining the gender of the subject to be photographed includes: Acquire the facial image of the subject to be photographed; The gender of the subject to be photographed is determined by using a gender recognition model to identify the gender of the face image.

3. The information prompting method according to claim 1, characterized in that, After outputting a prompt message to remind the subject to adjust its posture when the target prompt time is reached, the method further includes: When the pose information of the object to be photographed matches the preset pose information, the object to be photographed is photographed to obtain the target image.

4. The information prompting method according to any one of claims 1 to 3, characterized in that, The step of recognizing the posture of the object to be photographed in the shooting scene to obtain the object posture information of the object to be photographed includes: Acquire human images of the subject to be photographed; The key point detection model is used to detect key points in the human image to obtain multiple key points corresponding to the object to be photographed. Based on multiple key points corresponding to the object to be photographed, the object posture information of the object to be photographed is determined.

5. An information prompting device, applied to electronic devices, characterized in that, include: The information acquisition module is used to identify the posture of the object to be photographed in the shooting scene and obtain the object posture information of the object to be photographed. The distance determination module is used to determine the target distance between the object to be photographed and the electronic device when the object's posture information does not match the preset posture information. The time acquisition module is used to acquire the current time and, based on the target distance and the current time, acquire the target prompt time. The target notification time is later than the current time; The information output module is used to output a prompt message to prompt the subject to adjust its posture when the target prompt time is reached; Specifically, the time acquisition module is used for: Determine the target age value of the object to be photographed, and obtain the target prompt time based on the target distance, the current time, and the target age value; Alternatively, determine the gender of the object to be photographed, and obtain the target prompt time based on the target distance, the current time, and the gender; Determining the target age value of the subject to be photographed includes: Acquire the facial image of the subject to be photographed; The age of the face image is determined by using an age recognition model to obtain the target age value of the subject to be photographed. The target age value is corrected according to age correction parameters; the age correction parameters include at least one of illumination parameters, facial expression parameters, and facial pose parameters. The estimation function between the age correction parameter value and the corrected age value is as follows: h(x)=θ0+θ1x1+θ2x2+… +θ n x n ; Where h(x) is the corrected age value, x1, x2, ... x n These are the correction parameter values ​​for different age correction parameters, θ1, θ2…θ n θ0 is the coefficient corresponding to the age correction parameter, and θ0 is the offset of all age correction parameters from the corrected age value.

6. A storage medium, characterized in that, The storage medium stores a computer program that, when run on a computer, causes the computer to execute the information prompting method according to any one of claims 1 to 4.

7. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program, and the processor executing the information prompting method according to any one of claims 1 to 4 by calling the computer program stored in the memory.

Citation Information

Patent Citations

  • Control method and device, computer equipment and storage medium

    CN109558008A

  • Intelligent adjustment method, device and equipment for application program and storage medium

    CN111045774A

  • Image acquisition method and device, storage medium and electronic equipment

    CN111131702A