Image acquisition methods, electronic devices and readable storage media

By evaluating the screen structure and the pose of the target object in the preview screen on the terminal, the image quality problem when shooting easily moving objects is solved, and high-quality images are acquired.

CN119255092BActive Publication Date: 2026-01-06HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410110721.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2026-01-06
Estimated Expiration
2044-01-25

AI Technical Summary

Technical Problem

It is difficult to obtain high-quality images when photographing subjects that are easy to move, especially for subjects that are difficult to control, such as animals and children.

Method used

By displaying a preview screen on the terminal, the system detects target objects and evaluates the screen structure and the pose of the target objects. It uses region size parameters, visual balance parameters, and rule of thirds composition parameters for evaluation, and combines action semantic feature scoring to determine whether the preset conditions are met before acquiring the image.

Benefits of technology

It improves shooting quality, especially for difficult-to-control subjects such as animals and children, ensuring high-quality images are obtained under optimal shooting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119255092B_ABST
    Figure CN119255092B_ABST
Patent Text Reader

Abstract

The application discloses an image acquisition method, an electronic device and a readable storage medium. The method comprises the following steps: displaying a preview picture of image acquisition on a terminal; in response to a target object detected in the preview picture, acquiring a first evaluation result of a picture structure in the preview picture; in the case that the first evaluation result meets a first preset condition, acquiring a second evaluation result of a posture of the target object in the preview picture; and acquiring an image of the preview picture according to the second evaluation result. The embodiment of the application can improve the quality of the image acquired by photographing or the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to an image acquisition method, an electronic device, and a readable storage medium. Background Technology

[0002] With the development of computer and terminal technologies, mobile devices have become ubiquitous. Almost all mobile devices have camera functions, so mobile device users are generally accustomed to taking photos and videos with their devices. However, in some situations, the subject is easily movable, making it difficult to obtain high-quality images. Summary of the Invention

[0003] This application provides an image acquisition method, an electronic device, and a readable storage medium. The technical solution is as follows:

[0004] In a first aspect, an image acquisition method is provided, comprising: displaying a preview screen of image acquisition on a terminal; in response to a target object detected in the preview screen, acquiring a first evaluation result of the screen structure in the preview screen; if the first evaluation result satisfies a first preset condition, acquiring a second evaluation result of the pose of the target object in the preview screen; and acquiring an image of the preview screen based on the second evaluation result.

[0005] In this embodiment, for a target object in the preview screen of the terminal, the screen structure of the preview screen and the pose of the target object are evaluated. Based on the evaluation results, an image of the preview screen is obtained, thereby improving the quality of the photograph from the aspects of screen structure and the pose of the target object.

[0006] In one implementation, after displaying the preview screen on the terminal, the method further includes: acquiring an object in the preview screen; determining whether the object's type is a target object and obtaining a determination result; displaying an inquiry control on the terminal based on the determination result; and, in response to the user's click operation on the inquiry control, proceeding to the step of acquiring a first evaluation result of the screen structure in the preview screen.

[0007] After detecting an object in the preview screen, the user is asked whether to proceed to the first evaluation operation. That is, the user decides whether to trigger the process of taking a picture of the target object provided in this application embodiment, thereby improving the controllability of the picture taking process.

[0008] In one possible implementation, obtaining a first evaluation result of the image structure in the preview image includes: obtaining image structure evaluation parameters of the preview image based on the coordinates of the target object in the preview image; the image structure evaluation parameters include at least one of region size parameters, visual balance parameters, and rule of thirds composition parameters; and obtaining a first evaluation result based on the image structure evaluation parameters.

[0009] In this embodiment, the image structure of the preview image is evaluated by using region size parameters, visual balance parameters, and rule of thirds composition parameters. The preview image structure is measured from multiple dimensions to determine whether it meets the requirements, and a first evaluation result is obtained. The photographed image obtained based on the first evaluation result has a better visual effect in terms of image structure.

[0010] In one implementation, the image structure evaluation parameters include region size parameters, visual balance parameters, and rule-of-thirds composition parameters. Based on the coordinates of the target object in the preview image, the image structure evaluation parameters of the preview image are obtained, including: determining the proportion of the target object's region size in the preview image based on the coordinates; using the proportion as the region size parameter; determining the vertical and horizontal positions of the target object in the preview image based on the coordinates; using the horizontal and vertical positions as visual balance parameters; determining the minimum distance between the target object and the rule-of-thirds nodes in the preview image based on the coordinates; the rule-of-thirds nodes are nodes in the preview image that conform to the rule-of-thirds composition rules; and using the minimum distance as the rule-of-thirds composition parameter.

[0011] In this embodiment, based on the coordinates of the target object in the preview screen, relatively objective and accurate evaluation data can be obtained in terms of screen structure, visual balance and general composition rules, so as to achieve a comprehensive evaluation of the screen structure of the preview screen.

[0012] In one implementation, obtaining a second evaluation result of the pose of a target object in a preview screen includes: obtaining the action semantic features of the target object; determining the score corresponding to the action semantic features; and obtaining a second evaluation result based on the score if the score is higher than a preset threshold.

[0013] In this embodiment, a second evaluation result of the target object's posture is obtained based on the semantic features of the target object's actions. This allows for the determination of whether the preview image is suitable from the perspective of the target object's actions and postures, which is beneficial for obtaining images with better effects in terms of actions and postures.

[0014] In one implementation, if the score is higher than a preset threshold, a second evaluation result is obtained based on the score, including: assigning a category label of the corresponding action semantic feature to the image frame of the preview screen if the score is higher than the preset threshold; sorting the scores corresponding to the image frames including the category labels in the consecutive image frames of the preview screen; determining the target category label in the category labels based on the sorting result; and using the target category label as the second evaluation result.

[0015] In this embodiment, consecutive image frames are sorted by pose type or action type using scores, and then the target category label of consecutive image frames is determined based on the sorting, which helps to select the image with the best pose or action dimension from consecutive image frames.

[0016] In one implementation, acquiring an image of the preview screen based on a second evaluation result includes: displaying a photo prompt control on the terminal when the second evaluation result indicates that an image frame with a target category label exists in the preview screen; and acquiring an image frame with an action target category label as the image of the preview screen in response to a user's click operation on the photo prompt control.

[0017] In this embodiment, by receiving information from the user's manually issued click operation, the preview image is obtained, thereby improving the user's control over the photo-taking process and avoiding the acquisition of images that the user does not want to keep.

[0018] In one implementation, acquiring an image of the preview screen based on a second evaluation result includes: if the second evaluation result indicates that an image frame with a target category label exists in the preview screen, acquiring the image frame with the target category label as the image of the preview screen.

[0019] Images with target category labels have better effects in the pose or action dimension. In this embodiment, when the target object presents a better pose or action effect in the preview screen, the image of the preview screen is acquired, so as to capture a better quality image when the action of the target object is not easily controlled by the photographer or is irregular.

[0020] In one implementation, the target is at least one of an animal or a human being within a preset age range.

[0021] In this embodiment, the movements of animals are generally not easily controlled by the photographer, and the movements of humans in special age groups or in some special situations are also not easily controlled by the photographer. Therefore, for special types of target objects such as animals and humans in preset age groups, the image acquisition method provided in this application embodiment can be used to take pictures, which helps to improve the picture quality.

[0022] Secondly, a computer-readable storage medium is provided, which stores instructions that, when executed on a computer, enable the computer to perform the method described in the first aspect.

[0023] Thirdly, a computer program product containing instructions is provided that, when run on a computer, causes the computer to perform the method described in the first aspect.

[0024] The technical effects achieved by the second and third aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0025] Figure 1This is a schematic diagram illustrating an application scenario according to an embodiment of this application;

[0026] Figure 2 This is a schematic diagram illustrating another application scenario of an embodiment of this application;

[0027] Figure 3 This is a schematic diagram illustrating another application scenario of this application embodiment;

[0028] Figure 4 This is a schematic flowchart of an image acquisition method provided in an embodiment of this application;

[0029] Figure 5A This is a schematic diagram of the camera interface startup in an embodiment of this application;

[0030] Figure 5B This is a schematic diagram of the camera interface according to another embodiment of this application;

[0031] Figure 6 This is a schematic diagram of the camera interface in another example of this application;

[0032] Figure 7 This is a schematic diagram illustrating the implementation process of an image acquisition method in one example of this application;

[0033] Figure 8 This is a schematic diagram of the interactive interface before the photo-taking operation is performed, as shown in another example of this application.

[0034] Figure 9A This is a schematic diagram of the interactive interface before the photo-taking operation is performed, as shown in another example of this application.

[0035] Figure 9B This is a schematic diagram of the interactive interface before the photo-taking operation is performed in yet another example of this application;

[0036] Figure 9C This is a schematic diagram of the interactive interface before the photo-taking operation is performed in yet another example of this application;

[0037] Figure 9D This is a schematic diagram of the interactive interface before the photo-taking operation is performed in yet another example of this application;

[0038] Figure 10A This is a schematic diagram of an interactive interface presented before taking a picture, as shown in one example of this application.

[0039] Figure 10B This is a schematic diagram of the interactive interface presented before taking a picture, in another example of this application;

[0040] Figure 10C This is a schematic diagram of the interactive interface presented before taking a picture, in yet another example of this application;

[0041] Figure 11AThis is a schematic diagram illustrating the process of obtaining scores from image frames of a preview screen in one example of this application;

[0042] Figure 11B This is a schematic diagram illustrating the process of obtaining a screen structure score in one example of this application;

[0043] Figure 11C This is a schematic diagram illustrating the process of obtaining attitude scores in one example of this application;

[0044] Figure 12 This is a schematic diagram of an image acquisition device according to an embodiment of this application;

[0045] Figure 13 This is a schematic diagram of the software structure of an electronic device according to an embodiment of this application;

[0046] Figure 14 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0048] It should be understood that "multiple" as mentioned in this application refers to two or more. In the description of this application, unless otherwise stated, " / " indicates "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, to facilitate a clear description of the technical solutions of this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and that "first," "second," etc., do not necessarily imply differences.

[0049] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0050] The camera function has become one of the basic functions of mobile devices, and most non-mobile devices also have shooting capabilities. Meanwhile, the development of specialized shooting devices is accelerating, with a wide variety emerging. Users photograph a diverse range of subjects, some of which are prone to movement and difficult to maintain in a good shooting state, making it challenging to obtain satisfactory images using either the device or a specialized shooting device. For example, users often need to use their device's camera or a specialized camera to photograph animals and children. Because animals and children are difficult to control during the shooting process, they often move or change expressions at the moment the camera captures the image, resulting in distortion and blurring.

[0051] As users' lives become more diverse, they often desire to capture high-quality images in situations involving children and animals. Therefore, this application proposes an image acquisition method for obtaining high-quality images of specific target objects.

[0052] Figure 1 This is a schematic diagram of an application scenario of the image acquisition method provided in the embodiments of this application. Figure 1 In the scene shown, the subjects being photographed include a child (101) and a pet (102). Figure 1 In step a, while using mobile terminal 103, the user turns on the camera of mobile terminal 103, activates the camera function of mobile terminal 103, and points the camera at the child object 11 to take a picture. Then, as... Figure 1 In case b, the user wants to use the mobile terminal 103 to take photos and record the beautiful moments of the pet 102's life, so they open the camera function of the mobile terminal 103 and point the camera at the pet 102 to take pictures. Then, as... Figure 1 In the case of c, the child 101 and the pet 102 are within the same shooting range, and the user opens the camera of the mobile terminal 103 to simultaneously shoot the child 101 and the pet 102.

[0053] Figure 2 This is a schematic diagram illustrating another application scenario of the image acquisition method provided in the embodiments of this application. Figure 2 In the scenario shown, a user is video chatting with other users through a terminal 21 equipped with a camera. During the chat, other users point their terminal cameras at a child 22, and the user wants to take a screenshot on the terminal 21 to capture and record a precious moment of the child 22 during the chat. Therefore, the user activates the terminal's screenshot tool and obtains a screenshot image 23 of the current chat interface.

[0054] Figure 3This is a schematic diagram illustrating an application scenario of an image acquisition method provided in another embodiment of this application. A user pre-captures a video file 31 of the subject using a camera device 32. The user then edits the video file 31 using another terminal 33. During the editing process, the user aims to obtain images of the subject from the video file 31. Therefore, the user filters the video frames in the video file 31 and captures the desired image when it appears. The subject can be any person, animal, vehicle, starry sky, or other object that can be photographed, regardless of age.

[0055] The embodiments of this application are also applicable to many other scenarios, which will not be described one by one here.

[0056] exist Figures 1 to 3 Based on the scenario shown, this application embodiment also provides an image acquisition method. Figure 4 This is a schematic flowchart of an image acquisition method provided in an embodiment of this application. One embodiment of the image acquisition method provided in this application includes... Figure 4 Steps S41-S44 are shown.

[0057] Step S41: Display the preview screen on the terminal.

[0058] In this embodiment, the terminal can be an electronic device with a display screen that can display images or screens, such as a desktop computer, laptop computer, tablet computer, mobile phone, handheld computer, wearable smart device, smart furniture, or smart cabinet.

[0059] Displaying a preview screen on the terminal could be a preview of the camera interface, such as... Figure 5A As shown in a, the user clicks the camera application icon 51 on the terminal to launch the camera application and open... Figure 5A The camera interface 52, shown as b in Figure 52, includes a preview of the subject being photographed. (Through...) Figure 5A The preview screen shown in b allows users to observe the position of the subject and the desired image quality before taking the picture. If the subject's appearance in the preview screen matches the user's expectations, the user can click the camera control 53 to obtain the preview image of the subject, which will then be used as the final image.

[0060] In another possible implementation, the preview screen displayed on the terminal may be the chat interface during the chat process. In the chat interface, at least two different users of the two clients that have established a chat link through the network can see the screen of the other user on their own client. This screen may also be the preview screen in step S41.

[0061] Alternatively, in another possible implementation, the preview displayed on the terminal could be a video playback application or webpage playing a video file, where video frames serve as previews. During video playback, clicking the pause button could display a preview of a single video frame. Alternatively, during video playback, clicking a preview control or issuing a preview command could display a preview of the video file. Or, during video file editing, based on positioning commands, a video segment of the video file could be determined, and video frames representing that segment's content could be displayed as previews.

[0062] Step S42: In response to the target object detected in the preview screen, obtain the first evaluation result of the screen structure in the preview screen.

[0063] In this embodiment of the application, before step S42, the method further includes: performing target detection on the preview screen (i.e. detecting the target included in the image frame corresponding to the preview screen), obtaining the target in the preview screen, determining whether the obtained target is a preset type, and if so, taking the target detected in the preview screen as the target object.

[0064] The target object can be any of several preset object types, such as people, animals, and vehicles. After the preview screen is opened on the terminal, a neural network can be used to detect objects in the preview screen, determining the probability that the target belongs to any of the preset types. Based on the probability of the target belonging to any of the preset types, it can be determined whether the target is the target object. For example, if a neural network model is used to detect objects in the preview screen and a target is found, and a binary classification neural network is used on that target, the probabilities that the target belongs to the preset types of people, animals, and vehicles are 0.1, 0.2, and 0.7 respectively. Therefore, the detected target can be determined to be a vehicle.

[0065] In one possible implementation, the target object can be an object with fast-moving properties, such as young children, animals, vehicles, airplanes, high-speed trains, shooting stars, flowing water, waves, etc. Among these, young children may be children within a set age range, such as children aged 0-3 years. When users photograph young children or animals, it may be difficult to control the subject to maintain the desired state until the end of the shot. Therefore, it is challenging to capture suitable images of young children and animals, potentially resulting in unclear images and difficulty in keeping the target object within the lens's field of view throughout the shooting process. Furthermore, since objects such as vehicles, airplanes, high-speed trains, shooting stars, flowing water, and waves may generally have fast-moving properties, this embodiment of the application detects the image structure in the preview screen for these types of objects that are likely to move quickly under normal conditions, evaluates the image structure, and determines that the current lens has appropriately captured the object being photographed.

[0066] In one implementation, the screen structure includes, when the target object is located in the preview screen, the structure formed between the portion occupied by the target object in the preview screen and the remaining portion of the preview screen. For example, when the target object is a young child, the screen structure includes the portion occupied by the young child in the preview screen and the portion occupied by the remaining content in the preview screen.

[0067] In another implementation, the screen structure includes: the portion occupied by the target object in the preview screen, the portions occupied by other targets in the preview screen, and the structure formed by the remaining portions in the preview screen. For example, if the target object includes multiple people, one of whom is a young child, the screen structure may include: the portion occupied by the person in the preview screen, the portion occupied by the young child in the preview screen, and the portion occupied by the remaining content in the preview screen. As another example, if the target object includes vehicles, roads, and buildings, the screen structure may include the portions occupied by the vehicles in the preview screen, the portions occupied by the roads in the preview screen, the portions occupied by the buildings in the preview screen, and the portions occupied by the remaining content in the preview screen.

[0068] In another implementation, the evaluation criteria for the image structure can differ depending on the target object. For example, when photographing a moving car, a flying airplane, or a meteor, the average person may be too far away from the car, airplane, or meteor to make it difficult for them to occupy a major part of the image. Therefore, a more lenient evaluation criterion can be used to assess the image structure for cars, airplanes, and meteors.

[0069] In another possible implementation, the target object could be a rapidly moving target detected in the preview screen. Among the targets identified in the preview screen, if a target is moving at a speed within a set threshold range, then that target is considered the target object. Thus, when a human body is moving, dancing, or jumping, it may also be identified as a target object. Furthermore, when a user jumps in a shooting scene and wants to capture the jumping motion in an image, they can also evaluate the preview screen to capture an image that satisfies them.

[0070] Step S43: If the first evaluation result meets the first preset condition, obtain the second evaluation result of the pose of the target object in the preview screen.

[0071] In this embodiment, the first evaluation result may be a score, and the first preset condition may be a score threshold; or, the first evaluation result may be a type identifier, and the first preset condition may be at least one of multiple type identifiers. For example, if the first evaluation result is one of five categories: A, B, C, D, and E, and the first preset condition is B or C, then when the first evaluation result is B or C, the first evaluation result is considered to meet the first preset condition.

[0072] In one possible implementation, image quality is related to multiple factors, and the target object does not necessarily have to conform to a certain structural requirement to achieve a good user experience. Based on this, the first evaluation result might be a score corresponding to a type identifier, while the first pre-condition might be a threshold range for a score corresponding to one type identifier, or a combination of threshold ranges for scores corresponding to two or more type identifiers. In other words, the first evaluation result may include scores across multiple dimensions, or a score for a single dimension, while the first pre-condition may include threshold ranges for scores in at least one dimension. For example, the first evaluation result might include scores for five categories: A, B, C, D, and E, with scores of A: 10 points, B: 10 points, C: 50 points, D: 60 points, and E: 30 points, respectively. Simultaneously, the first pre-condition is a total score of 100 points or higher across all categories, or a score of 30 points or higher for category C and 50 points or higher for category D. In this case, the first evaluation result can be determined to meet the first pre-condition.

[0073] By using a combination of multiple types as the first preset condition, and with each type corresponding to a scoring threshold range, more accurate judgments can be made on target objects in special states. For example, a child wearing an animal-shaped coat may be identified as belonging to both the animal and child types in the preview image. Different judgment criteria based on image structure are set for the child and animal types of target objects, and a comprehensive judgment is made to determine that the first evaluation result meets the first preset condition.

[0074] When a special background is detected in the preview image, the user may want to capture an image with a different composition than a typical image of a person. For example, the user is taking a picture in front of a landscape, and the landscape occupies a large part of the preview image. Although the composition of the preview image may not be optimal depending on the proportions of the human body and the background, the image needs to be captured according to the composition of the current preview image in order to capture the complete landscape as the background. Therefore, in such scenarios, a first preset condition corresponding to the landscape background can be set.

[0075] The pose of a target object can include at least one of its posture, state, appearance, style, demeanor, and actions. In one possible implementation, the pose of the target object primarily includes its actions. The pose can include the target object's body posture; for example, if the target object is an animal or a person, the set animal or person posture is considered to meet the second preset condition. Similarly, if the target object is an object, the object's pose can be the state it presents in the image. If the pose of certain target objects does not change significantly whether in motion or at rest, such as vehicles or airplanes, then the second preset condition can include the general state of the target object in motion or at rest, or it can be assumed that the pose of this type of target object meets the second preset condition under all circumstances. For example, in the case of target objects such as vehicles or airplanes, if the target object's pose is assumed to meet the second preset condition, then when target recognition is performed in the preview image, and the target object in the preview image is determined to be a vehicle, airplane, or other target object of a set type, the result is obtained that the target object's pose meets the second preset condition.

[0076] In one implementation, when a target object exists in the preview screen, a first evaluation operation is performed on the preview screen to obtain a first evaluation result. If the first evaluation result meets a first preset condition, a second evaluation operation is performed on the preview screen to obtain a second evaluation result.

[0077] In another implementation, when a target object is present in the preview screen, a first evaluation result and a second evaluation result are obtained. If the first evaluation result meets the first preset condition and the second evaluation result meets the second preset condition, the subsequent step S44 is performed.

[0078] In this embodiment, a second evaluation result can be determined using a combination of scores from multiple dimensions or a single dimension. For example, the posture of the target object includes scores from multiple dimensions such as pose, state, style, appearance, demeanor, and movement, which can be used as the second evaluation result. In one possible implementation, the second evaluation result can be a default value for some target objects. For example, if the target object is a car or an airplane, the posture of this type of target object may not change, so the second evaluation result can be a default value.

[0079] When the target object is a person or animal, the second evaluation result of the target object's posture in the preview screen can include the result of evaluating the degree to which the action in the preview screen matches the set action. Alternatively, an exclusion method can be used, where a lower evaluation is given as the second evaluation result when the target object's action meets the exclusion criteria.

[0080] In one possible implementation, obtaining the second evaluation result of the target object's posture in the preview screen may include: obtaining the type of the target object's action, and obtaining the second evaluation result based on the type of the target object's action. Some reference action types can be pre-defined. When the target object's action type matches a reference action type, the degree of conformity between the target object and the reference action type can be scored. For example, if the reference action type is a human jumping motion, a relatively low score can be given if the target object's jumping motion extension is detected as half-open (just starting to jump) in the preview screen, and a relatively high score can be given if the target object's jumping motion extension is detected as fully open (at the highest point of the jump).

[0081] In another possible implementation, when a target object is present in the preview screen, a first evaluation result and a second evaluation result are obtained. If the first evaluation result meets the first preset condition and the second evaluation result meets the second preset condition, the subsequent step S44 is performed.

[0082] Step S44: Based on the second evaluation result, obtain the image of the preview screen.

[0083] In this embodiment of the application, acquiring the image of the preview screen based on the second evaluation result may include: acquiring the image in the preview screen when the second evaluation result meets the second preset conditions; or, when the second evaluation result meets the second preset conditions, displaying shooting prompt information in the preview screen, and capturing the image in the preview screen when detecting that the user clicks the shooting operation control according to the shooting prompt information.

[0084] In one possible implementation, acquiring an image from the preview screen may include: if the preview screen at the current time point meets the image acquisition conditions, acquiring and storing the image from the preview screen. Thus, in the case of automatic image acquisition, if the preview screen meets the second evaluation result, the image from the preview screen is automatically acquired, and the acquired image includes at least one image frame from the preview screen.

[0085] Therefore, in the embodiments of this application, as Figure 5B As shown, the preview screen is preview screen 54 in the camera interface. If preview screen 54 meets the first and second preset conditions, a prompt icon 55 can be generated in preview screen 54. Prompt icon 55 indicates that the current image quality is good and the camera can be taken. Alternatively, the preview screen 54 can display the prompt text "The current image quality is very good, you can take a picture now," or other similar content, so that the user can understand the meaning of prompt icon 55.

[0086] In one possible implementation, the first use on the user's terminal Figure 4 In the case shown, when the preview screen meets the first preset condition and the second preset condition, a prompt icon can be generated, and a layer can be generated on the preview screen. The prompt layer displays the function introduction information about the prompt icon, so that the user can know how to operate after the prompt icon is generated based on the function introduction information.

[0087] In this embodiment, a preview screen containing a target object can be evaluated to determine whether the preview screen meets the preset conditions (including the aforementioned first preset condition and second preset condition). If the preview screen meets the preset conditions, an image of the preview screen can be acquired, thereby improving the quality of the acquired image. In particular, for target objects with fast-moving attributes, it can help users acquire higher quality images.

[0088] When applied to photography scenarios, this application embodiment can perform multiple judgments on target objects such as young children and animals that are not easy to fix in their photography posture. When the target object is in a better or optimal photography state, the image of the target object is acquired, thereby helping users to take pictures of certain types of targets. It can also automatically detect high-quality images that are difficult for the user's naked eye to perceive. When the target object is in a better photography state, it can promptly capture images that are difficult for the user to complete in time by combining visual observation with manual operation, or prompt the user to capture images in time to generate high-quality images.

[0089] In the photography scenario, this embodiment of the application sets up a three-level trigger for subjects with distinct characteristics such as children, cats, and dogs. First, the preview screen is detected at the detection end (terminal). The first-level trigger uses artificial intelligence (AI) detection and recognition technology to detect target objects in the preview screen, identifying whether subjects such as children, cats, and dogs appear in the scene corresponding to the preview screen. If subjects such as children, cats, or dogs appear in the scene, the second-level trigger is activated. The second-level trigger, based on the image rating detection algorithm for subjects such as children, cats, and dogs, executes the evaluation operation related to step S42 mentioned above. The second-level trigger performs a first evaluation operation by real-time monitoring of the comprehensive state information of the subjects in the preview screen and the overall preview screen, obtaining a first evaluation result. If the first evaluation result meets the expected index of the rating algorithm, i.e., the threshold of the second-level trigger reaches the expected judgment condition, the third-level trigger is activated. The third-level trigger, based on the pet action detection algorithm using AI technology, executes the evaluation operation related to step S43 mentioned above. The system identifies the actions of the subject in the current scene, outputs the action evaluation index, sends it to the decision engine, calculates decision factors, and if the set threshold standard is reached, triggers the automatic photo-taking process or displays a prompt message to remind the user to take a photo.

[0090] The aforementioned level 1, level 2, and level 3 triggers can be triggered in a hierarchical manner, with the triggering of the level 1 trigger being a condition for the subsequent triggering of the level 2 and level 3 triggers. Specifically, the level 1 trigger executes the first part of the detection operation provided in this application embodiment, namely, detecting and identifying objects in the preview image. The level 2 trigger executes the second part of the detection operation provided in this application embodiment, namely, scoring the image structure. The level 3 trigger executes the third part of the detection operation provided in this application embodiment, namely, scoring the pose of the target object. By using hierarchical triggers, when it is determined that subsequent operations cannot be performed, the detection is re-executed on a new image of the preview image, thereby effectively solving the balance between algorithm power consumption, memory usage, and performance during preview, and improving the success rate of capturing the subject.

[0091] When applied to screenshot scenarios in this application embodiment, during user chat, the system can automatically detect moments that meet quality requirements and provide a prompt at those moments. Alternatively, after the user sends a screenshot command, it can capture the image frame with the best relative quality from multiple preview frames generated before and after the screenshot quality is determined, and use this frame as the image obtained from the screenshot operation. This helps users obtain higher-quality images when taking screenshots, avoiding situations where it is difficult to capture high-quality images during the screenshot process.

[0092] When applied to video processing scenarios in the embodiments of this application, the first preset condition, the second preset condition, and the evaluation condition for the third evaluation result can be determined based on the video being processed; the target object can also be determined based on the video being processed. The method provided by the embodiments of this application enables the detection of multiple image frames in a video during the video processing process, assisting users in obtaining the desired image frames from the video.

[0093] In one implementation, after the preview screen is displayed on the terminal, the image acquisition method further includes: acquiring an object in the preview screen; determining whether the object type is a target object and obtaining a determination result; displaying an inquiry control on the terminal based on the determination result; and, in response to the user's click operation on the inquiry control, proceeding to the step of acquiring a first evaluation result of the screen structure in the preview screen.

[0094] When the method provided in the embodiments of this application is applied to a photography scenario, such as Figure 6 As shown in 'a', when the user clicks the camera app icon on the terminal, and then enters... Figure 6The camera interface 61 shown in Figure b can display a mode selection control 62 (i.e., an inquiry control) if a target object is detected. The user can use the mode selection control 62 to change the shooting mode to one specifically targeting the target object, and then perform an evaluation operation on the preview screen. The evaluation operation performed on the preview screen can include at least one of the aforementioned first, second, and third evaluation operations. If the user does not operate the mode selection control 62, the default setting allows the user to choose whether or not to continue taking photos in the current shooting mode; this default setting can be changed according to the user's settings.

[0095] In another possible implementation, the objects in the preview screen can be displayed directly in the preview screen when the photo is taken, without checking the preview screen. Figure 6 The mode selection control 62 shown in b allows for the activation of a specific photo-taking function for a target object even when no target object is detected. If in Figure 6 In the example shown, the user operates the mode selection control 62, causing the terminal's camera application to enter target object shooting mode. In target object shooting mode, if the target object's state is determined to meet at least one of a first preset condition and a second preset condition, a prompt icon indicating that the image quality meets the preset conditions can be displayed at the location of the mode selection control 62. The prompt icon indicating that the image quality meets the preset conditions can be, for example,... Figure 5B The prompt icon shown is 55.

[0096] If the query control is triggered after a target object is detected, and at least one target object is detected among multiple targets in the preview screen, the query control can be displayed in the preview screen.

[0097] Similarly, when users establish video calls with other terminals, watch live streams, or process videos through their terminals, a query control can be displayed, allowing users to analyze the preview of the video call, the live stream, or the video processing to obtain high-quality images from the video call, live stream, or video.

[0098] This application embodiment can receive user click operations through an inquiry control and initiate a function to evaluate the quality of the preview screen according to the user's intention, thereby improving the user's control over the image acquisition function.

[0099] In one implementation, obtaining a first evaluation result of the image structure in the preview image includes: obtaining image structure evaluation parameters of the preview image based on the coordinates of the target object in the preview image; the image structure evaluation parameters include at least one of region size parameters, visual balance parameters, and rule of thirds composition parameters; and obtaining a first evaluation result based on the image structure evaluation parameters.

[0100] In this embodiment, the coordinates of the target object in the preview screen can include the coordinates of the edge of the target object in the coordinate system of the preview screen. The area of ​​the target object can be determined based on the coordinates of the edge of the target object, which serves as the area size parameter in the screen structure evaluation parameters of the preview screen.

[0101] When the image structure evaluation parameters include the area size parameter, the area size parameter can be determined based on the area surrounded by the edge of the target object and other areas outside the target object in the preview image, and the first evaluation result can be obtained based on the area size parameter.

[0102] Visual balance parameters may include: the relative position of the target object in the preview image, determined by its coordinates. Rule of thirds composition parameters may include: parameters indicating the degree to which the target object conforms to the rule of thirds. The rule of thirds composition may include: the photographic rule of thirds. The photographic rule of thirds is a compositional technique frequently used in photography, painting, design, and other arts, sometimes also called the grid composition. The rule of thirds means dividing the image horizontally into thirds, with the main subject placed in the center of each third. This composition is suitable for subjects with multiple parallel focal points.

[0103] This application embodiment evaluates the image structure of the preview image by using at least one of the region size parameter, visual balance parameter, and rule of thirds composition parameter, which helps to improve the image structure quality and visual harmony of the obtained image.

[0104] In one implementation, the image structure evaluation parameters include region size parameters, visual balance parameters, and rule of thirds composition parameters. The image structure evaluation parameters of the preview image are obtained based on the coordinates of the target object in the preview image, including: determining the proportion of the target object's region size in the preview image based on the coordinates; using the proportion as the region size parameter; determining the vertical and horizontal positions of the target object in the preview image based on the coordinates; and using the horizontal and vertical positions as visual balance parameters.

[0105] Based on the coordinates, determine the minimum distance between the target object and the rule of thirds nodes in the preview image; the rule of thirds nodes are the nodes in the preview image that conform to the rule of thirds composition rules; use the minimum distance as the rule of thirds composition parameter.

[0106] In this embodiment, the edge position of the target object in the preview screen can be determined based on its coordinates, thereby determining the size of the target object's region. The proportion of the target object's region size in the preview screen can be expressed as the ratio obtained by dividing the target object's region size by the total size of the preview screen.

[0107] In this embodiment, the horizontal and vertical positions in the preview screen can be determined based on the parallel direction of the rectangular side of the preview screen when the preview screen is rectangular, and then the horizontal and vertical positions of the target object can be determined. When the preview screen is rectangular or other shapes, the horizontal and vertical positions of the target object can be determined based on the vertical and horizontal orientation of the terminal display screen.

[0108] In this embodiment, the trisection nodes in the preview screen can be points included on the dividing lines formed by dividing the preview screen using the trisection method. When dividing the preview screen using the trisection method, the preview screen can be divided horizontally and vertically to form horizontal and vertical dividing lines, and the intersection of the horizontal and vertical dividing lines can be used as trisection nodes. Alternatively, points on the horizontal and vertical dividing lines can be used as trisection nodes.

[0109] In this embodiment, the preview image can be evaluated based on the size of the target object, the position of the target object in the preview image, and whether the distribution of the target object in the image conforms to certain segmentation rules, thereby making the obtained image more in line with the composition rules.

[0110] In another possible implementation, any one or any two of the region size parameter, visual balance parameter, and rule of thirds composition parameter can be used as the image structure evaluation parameter. Then, the image structure evaluation parameter is calculated using the method provided in the above embodiment to obtain the first evaluation result.

[0111] In one implementation, obtaining a second evaluation result of the pose of the target object in the preview screen includes: obtaining the action semantic features of the target object; determining the score corresponding to the action semantic features; and obtaining a third evaluation result based on the score if the score is higher than a preset threshold.

[0112] In this embodiment, a neural network model can be used to analyze the preview image to obtain the semantic features of the target object's actions. These semantic features can include the semantic characteristics of the actions themselves. Semantic features can include semantic components or semantic elements, specifically semantic components of different dimensions, which can be expressed using vectors. Different dimensions can include, for example, the attributes of the action, the classification of the action, the function of the action, and the relationship between actions. Therefore, through multiple different dimensions, the semantic features of the action can be categorized into: attribute features of the action, classification features of the action, functional features of the action, and relational features of the action.

[0113] Action semantic features can include semantic features of various action types. If the score corresponding to any action semantic feature is higher than a preset threshold, the current image frame in the preview screen can be considered a valid image frame, and a third evaluation operation is then performed to obtain the third evaluation result. If the scores of all action semantic features are not higher than the preset threshold, the current image frame in the preview screen can be considered an invalid image frame, and the third evaluation operation can be skipped. Instead, the first, second, and third evaluation operations are performed again on the next time sequence image frame.

[0114] In this embodiment of the application, the score corresponding to the action semantic feature may include scores corresponding to multiple different types of action semantic features, and the preset threshold may be the score corresponding to multiple different types of action semantic features in a set combination.

[0115] For example, a preset threshold refers to the score threshold for any one type of action semantic feature, and the preset threshold is 0. If the scores of multiple types of action semantic features are all 0, then the condition of a score higher than the preset threshold is not met. If, among multiple different types of action semantic features, there is one type of action semantic feature with a score higher than 0, then the score can be considered higher than the preset threshold. Alternatively, a preset threshold refers to the score threshold for any two types of action semantic features, and the preset threshold is N. Then, if, among multiple action semantic features, there are at least two preset dimensions with scores greater than N, then the score can be considered higher than the preset threshold.

[0116] If the score is higher than a preset threshold, a third evaluation result is obtained based on the score. This can be achieved by confirming a valid third evaluation result and outputting it if the score of any type of action semantic feature is higher than the preset threshold. Alternatively, a preset threshold can be set for each type of action semantic feature; if the score of any type of action semantic feature is higher than the corresponding preset threshold, a valid third evaluation result is confirmed and output. Or, a preset threshold can be set for all types of action semantic features; if the sum of the scores of all types of action semantic features is greater than the corresponding preset threshold, a valid third evaluation result is confirmed.

[0117] If the third evaluation result is valid, the current image frame corresponding to the preview screen can be considered a valid image frame; otherwise, the current image frame corresponding to the preview screen can be determined to be an invalid image frame and can be discarded.

[0118] In one possible implementation, different target objects correspond to different types of action semantic features, which are the semantic features of the actions the target object may perform. For example, if the target object is a young child, the types of actions the target object may perform might include: crawling, walking, running, jumping, raising hands. In some cases, actions may include facial expressions, so the types of actions a young child might perform could also include: smiling, crying, sobbing, laughing, making faces, frowning, looking blankly, etc. If the target object is an animal, the types of actions the target object might perform might be the same as those of a young child, or they might have other types of actions different from those of a young child. When the target object is flowing water, waves, or other objects, the action types might differ from those of people and animals. For each type of action that a target object might perform, a separate action semantic feature can be set.

[0119] If there are two or more target objects in the preview screen, and one of the target objects has a score for action semantic features that is higher than a preset threshold, then the current image frame of the preview screen can be considered to have a valid third evaluation result, or the current image frame of the preview screen can be considered to be a valid image frame.

[0120] In another possible implementation, if the current image frame of the preview screen does not have a valid third evaluation result, output data without a valid third evaluation result can be generated.

[0121] In one implementation, if the score is higher than a preset threshold, a second evaluation result is obtained based on the score, including: assigning a category label of the corresponding action semantic feature to the image frame of the preview screen if the score is higher than the preset threshold; sorting the scores corresponding to the image frames including the category labels in the consecutive image frames of the preview screen; determining the target category label in the category labels based on the sorting result; and using the target category label as the third evaluation result.

[0122] In this embodiment, each type of action semantic feature has its own category label. That is, each action type corresponds to one type of action semantic feature, and each type of action semantic feature corresponds to one category label, thus each action type also corresponds to one category label. For example, when the target is a young child, the types of actions that a young child may perform include: crawling, walking, running, jumping, raising hands, etc. Among them, the action semantic feature of crawling has a crawling category label, the action semantic feature of walking has a walking category label, and so on. Each action type's action semantic feature corresponds to one category label.

[0123] Therefore, if each type of action semantic feature is configured with a preset threshold for scoring, then if the scores of two or more action semantic features exceed the preset threshold among multiple types of action semantic features, the image frame of the preview screen will have more than two category labels.

[0124] For example, the types of actions that young children may perform include crawling, walking, running, jumping, raising their hands, laughing, and crying. If the scores of the semantic features of the actions corresponding to crawling, laughing, and crying are all higher than the preset threshold for a single image frame in the preview, then the image frame has three category labels: crawling, crying, and laughing.

[0125] In this embodiment, determining the target category label from the category labels based on the sorting results can include: taking the top N category labels as the target category labels, where N is a positive integer. Using the target category labels as the third evaluation result can include using all target category labels as the second evaluation result.

[0126] In a series of consecutive image frames in the preview screen, there may be multiple image frames with target category labels. Therefore, target category labels can be labeled for multiple image frames.

[0127] In this embodiment, sorting the scores corresponding to image frames including category labels in a series of image frames in the preview screen can be done by sorting the scores corresponding to image frames including category labels in a series of image frames in the preview screen within a set number (or a set time period).

[0128] In another possible implementation, image frames from the preview screen can be extracted at set time intervals, and multiple image frames can be extracted as consecutive image frames within a set total time period. Alternatively, image frames from the preview screen can be extracted at set intervals based on the number of image frames, and multiple image frames can be extracted as consecutive image frames within a set total time period.

[0129] In one implementation, acquiring an image of the preview screen based on a second evaluation result includes: displaying a photo-taking prompt control on the terminal when the second evaluation result indicates that an image frame with a target category label exists in the preview screen; and acquiring an image frame with a target category label as the image of the preview screen in response to a user's click operation on the photo-taking prompt control.

[0130] In this embodiment, the photo-taking notification control can be an icon. For example, when a young child is detected in the preview image, a smaller icon of the young child is displayed; or, when an animal is detected in the preview image, a smaller icon of the animal is displayed. Alternatively, when a target object of any defined type is detected in the preview image, an icon indicating that the target object has been detected is displayed.

[0131] In one possible implementation, the photo-taking prompt control can be a control used to receive photo-taking commands issued by the user. For example, when it is determined that the target object in the current preview screen is in a suitable state for taking a photo, the state of the photo-taking control on the photo-taking interface can be changed to indicate that the user can issue a photo-taking command.

[0132] In another possible implementation, due to the special nature of the object being photographed, it may not be possible to obtain a good evaluation result in all three evaluation operations before acquiring the preview image. For example, during the execution of the third evaluation operation, the target object's actions may not fully conform to the preset action type, making it difficult to obtain a high third evaluation result. In this case, if a better second evaluation result exists, the preview image can still be acquired after the third evaluation operation is performed.

[0133] When a user clicks on the camera prompt control, the best-performing image frame from a set time period or a range of image frames before and after the click can be captured and used as the preview image.

[0134] In one implementation, acquiring an image of the preview screen based on a third evaluation result includes: if the third evaluation result indicates that an image frame with a target category label exists in the preview screen, acquiring the image frame with the target category label as the image of the preview screen.

[0135] In one possible implementation, the image in the preview screen can be obtained automatically, or the user can be prompted to take a picture, and the image in the preview screen can be obtained according to the user's photo-taking command.

[0136] In one implementation, the target is at least one of an animal or a human being within a preset age range.

[0137] The preset age range for human subjects can be young children, such as those aged 0-13, or those aged 0-18, or other age groups.

[0138] The animals in this application embodiment can be pets or wild animals, and the human body within the preset age range can include adult human bodies.

[0139] In one specific implementation, algorithms such as subject detection, gender and age detection, face detection, and scene detection can be used to detect target objects such as animals and human bodies within a preset age range. Subject detection determines whether an animal or human body is present in the preview image. If a human body is detected, a gender and age detection algorithm can be applied to estimate its age. Furthermore, if the target is a young child, since the proportions of a child's facial features, face, and body differ from those of a mature adult, face detection and human body detection algorithms can be used to assist in age detection. This allows for the determination of at least one of the following data: facial structure information, facial feature information, face-to-body ratio, and facial feature proportions, thus aiding in age determination.

[0140] When the target object includes pets, scene detection combined with motion evaluation algorithms can be used to determine whether the target in the preview image includes pets. Non-domesticated animals photographed by users in the wild, zoos, and other similar locations can also be identified as pets.

[0141] According to user survey statistics, the number of photography scenarios featuring human bodies (especially young children) and pets such as cats and dogs (i.e., the targets in the preview screens of other embodiments of this application) is showing a year-on-year growth trend. At the same time, photographers face some challenges in these scenarios, which means that in order to achieve a better photography experience for users, the requirements for image acquisition devices in children's and pets' photography scenarios need to be improved.

[0142] A major problem encountered when photographing children, cats, dogs, and other pets is that these subjects are difficult for the photographer to control. They are prone to movement and instability, resulting in blurry subjects, chaotic compositions, and a low success rate in photographing such subjects. The image acquisition method provided in the application embodiment can address these issues in shooting scenarios with young children or pets, improving the photographing experience and increasing the success rate of photographing children and pets.

[0143] In a specific example of this application, a first-level trigger, a second-level trigger, and a third-level trigger can be set to assist in the execution of the image acquisition method steps.

[0144] like Figure 7As shown, the user can use the mode selection control 79 to determine whether to use the automatic shooting mode or the manual shooting mode for the target object. The first-level trigger 71 retrieves the object in the preview screen and determines whether the object in the preview screen is the target object. Different target objects can be detected using different detection algorithms. For example, if a human body 77 needs to be detected as the target object, a child detection algorithm is used. If a pet (e.g., a human body 77) needs to be detected as the target object, the algorithm will be different. Figure 7 If cats (76) and dogs (74) are selected as target objects, then a pet detection algorithm is used for detection. Here, "pet" can include animals in a broad sense, such as... Figure 7 Wild animals such as pandas (78 species) and captive animals in zoos.

[0145] The first-level trigger 71 uses a detection algorithm to examine the preview screen to determine if the subject in the preview screen is a child, a human body, a pet, or another target object. If a target object is found, the coordinates of the subject in the preview screen are calculated and output, and the target object among the subjects is used as the input to the second-level trigger 72. If there are multiple subjects in the preview screen, and one of them is a target object, then the subsequent second-level trigger 72 or third-level trigger 73 can be triggered automatically or in response to a detected user command.

[0146] As an example, a secondary trigger 72 could include a screen structure scorer. The screen structure scorer calculates screen structure evaluation parameters based on the coordinates provided by the primary trigger 71. The screen structure scorer can calculate the centroid and main body line of the target object in the preview screen based on the coordinates. The centroid of the target object can be the centroid of the geometric shape of the area occupied by the target object in the preview screen. The main body line of the target object can be the line connecting the centroids of the target object. It can analyze whether the centroid and main body line conform to a preset first rule.

[0147] Simultaneously, structural reference lines can be generated for the preview screen. These reference lines include the rule of thirds, the golden ratio, and spirals. These reference lines are used to determine whether the centroid or main body line of the target object in the preview screen conforms to the second rule. The centroid, main body line, and structural reference lines of the target object can be shown to the user in the preview screen or not. The content displayed in the preview screen can be independent of the terminal's internal calculations. For the target object in the preview screen, face frames and / or body frames can also be calculated and / or displayed. These face frames and / or body frames can be used to assist in determining whether the screen structure meets pre-set requirements.

[0148] After the second-level trigger 72 is triggered, the image structure scorer calculates the area size parameters, visual balance parameters, and rule of thirds composition parameters (ThirdRulescore in this example) based on the coordinate data provided by the first-level trigger 71.

[0149] This example uses RegionSizescore to measure whether the proportion of the target object's size (area or size) in the overall preview image is appropriate. Based on a prior composition knowledge base, two thresholds, α and β, are set for the region size. α is set as the small region threshold, and β as the large region threshold; both are relative values. The ratio between the target object's size and the current preview image size is calculated and compared with α and β. In this example, the ratio between the preview image size and the target object's size may fall into three categories: less than α, greater than β, or between α and β. When multiple target objects exist in the preview image, at least one ratio can be calculated for each target object. The ratios for each target object are normalized to obtain the RegionSizescore. This score is then weighted using prior weights to obtain the region size parameter, which serves as one dimension of the overall image structure score in subsequent calculations.

[0150] In this example, a VisualBalancescore is calculated to measure the appropriateness of the target object's position in the preview. The VisualBalancescore is determined by calculating the relative positions of the target object's centroid in the horizontal and vertical directions. Multiple target objects may exist in the preview, and a VisualBalancescore can be calculated for each. After normalizing the VisualBalancescores of all target objects in the preview, a VisualBalancescore is obtained. This VisualBalancescore is then weighted according to pre-set prior weights to obtain visual balance parameters, which serve as another dimension of the overall score in the image structure scorer calculation.

[0151] In this example, the Third Rule score is used to measure how well a target object in the preview conforms to the rule of thirds composition. By calculating the position of the target object relative to the corresponding rule of thirds nodes in the image, the position information of the nearest node is found, normalized, and then the Third Rule score is calculated. Prior weights are then used for weighted calculation to obtain the rule of thirds composition parameters, which serve as another dimension of data in the overall score calculation of the image structure scorer.

[0152] After obtaining the region size parameters, visual balance parameters, and rule of thirds composition parameters, the data from the three dimensions are combined to obtain normalized data. Then, combined with prior information, the normalized data is weighted and calculated to obtain the final evaluation score of the picture structure scorer.

[0153] The final evaluation score obtained by the image structure scorer included in the secondary trigger 72 has a wide range of applications, which can expand the applicable scenarios of image acquisition methods. It can be used as the image structure evaluation parameter in the table of embodiments of this application as the first evaluation result of the secondary trigger 72, and it can also be used as reference data when the subsequent tertiary trigger 73 performs the second evaluation operation. At the same time, the image structure scorer also has the scalability of scoring dimensions.

[0154] Figure 7 The three-level trigger 73 shown includes a posture scorer that scores the target object. Based on the output of the second-level trigger, it detects whether the posture of the target object, such as a pet or young child, matches a preset a priori "wonderful moment." In the user-selected automatic image acquisition mode, the posture scoring data is used as a second evaluation result and sent to the decision engine 75 as decision information. The decision engine 75 can then automatically generate a decision to acquire an image from the preview screen based on the second evaluation result. This decision is then sent to the enhancement engine 710, which can capture a higher-quality image from a series of image frames in the preview screen as the captured image. In the user-selected manual image acquisition mode, the score information or the detection results of the three triggers are displayed in real-time on the photo-taking interface to prompt the user to take a photo. The user triggers the photo-taking action by issuing a photo-taking command. In response to the user's command, the enhancement engine 710 can capture a higher-quality image from a series of image frames in the preview screen as the captured image. In one possible implementation, the enhancement engine 710 can also perform preset image processing on the captured image to improve image quality.

[0155] In one possible implementation, when the final evaluation score of the second-level trigger 72 for the preview screen is greater than the set threshold, it is necessary to evaluate and classify the action of the current target object. Since the action of the target object of the pet type is not easy to define, it is relatively difficult to directly detect and classify the action of the target object of the pet type. Therefore, the pose scorer of the target object can be designed based on the multimodal model.

[0156] For target objects categorized as young children, common action types can be defined as semantic categories for the secondary trigger 72, including semantic categories for limb extension movements, jumping movements, running movements, head turning movements, and laughing movements. Customized actions can also be set; for example, using pre-captured reference images, actions can be extracted from the target objects and used as customized actions, generating semantic information for these customized actions based on the reference images. For actions such as limb extension, jumping, running, head turning, and laughing, customized action amplitudes can be set. When acquiring photos, image frames exhibiting the same action are associated, and the image frame with the closest action amplitude to the customized action amplitude is saved. For target objects within the bounding box output by the primary trigger 71, the secondary trigger 72 can combine the defined semantic categories to determine the action type and corresponding score one by one. Considering the continuity of actions, multiple preceding and following image frames are needed as supplementary frames, and action detection and scoring are performed on each image frame.

[0157] In one possible implementation, the secondary trigger 72 performs a second evaluation operation on the target object in the preview screen by first scoring it and then determining its action category. The secondary trigger 72 first estimates the semantic category score corresponding to the action of the target object in the current preview screen. If the scores of all semantic categories in an image frame are 0, the current image frame is discarded. If an image frame has a non-zero semantic category score, it can further determine whether the non-zero semantic category score is greater than the determination threshold of the corresponding action semantic category, and assign an action label to the semantic category. Finally, it determines whether the number of action labels in the image frame is greater than 0. If the number of action labels in an image frame is greater than 0, then this image frame is considered a valid image frame. If an image frame has multiple action labels, the scores of the semantic categories corresponding to the action labels are sorted, and the action label corresponding to the highest score is taken as the final action label and output to the decision engine 75.

[0158] After the user selects manual mode, the system can display the pet's detection status information, image rating score, and other data in real time on the preview screen, and generate a prompt message to remind the user to press the manual shutter button.

[0159] If the user takes a photo in automatic mode, the camera app will display the first evaluation result of the target object, the second evaluation result of the target object, the scores obtained by the image structure scorer and the posture scorer in real time on the preview screen. If the user takes a photo in manual mode, the first evaluation result, the second evaluation result of the target object, the scores obtained by the image structure scorer and the posture scorer in the preview screen will still be displayed on the preview screen. After performing the second or third evaluation operation, the decision engine 75 will automatically decide to take the photo and issue the shutter information based on the current information of the second and third triggers, automatically completing the photo taking function. In automatic shooting mode, the preview screen can display a prompt message indicating that the photo has been taken and that the user has completed a quick shot of the target object and can view the result in the gallery.

[0160] In one possible implementation, after the user initiates a photo-taking action on the preview screen, the underlying system uses the information from the preview screen to complete a series of quick-shooting strategies that ensure basic capture, basic performance, and basic effects. This ensures that the subject has the most suitable exposure, color temperature, and sharpness at the moment of taking the photo, thereby providing the user with the best photography experience.

[0161] In one possible implementation, the target object may include a moving and uncontrollable subject. A layered triggering and three-level detection decision evaluation mechanism is adopted for one or more image frames of the target object in the preview screen. At the same time, an enhancement mechanism for the subject in the preview screen is added for the decision result. The best image frame is selected from multiple image frames corresponding to the preview screen as the final photo. This provides an intelligent photography solution that automatically or manually improves the success rate and effect of taking pictures of subjects with moving behavior.

[0162] In another example of this application, the preview process during the execution of the image acquisition method starts from... Figure 8 To begin, may include Figure 8 , Figures 9A to 9C The process. (Refer to...) Figure 8 As shown in 'a', after the user clicks the camera app icon 81 on the terminal desktop, the camera interface is launched. Figure 8As shown in b, the camera interface includes a preview screen 82, a shutter control 83, and other controls. The preview screen 82 is used to preview the image frames to be captured. When a camera captures a photo based on a user command or when the capture is performed automatically, the image frames in the preview screen 82 are captured as the photograph. Through the preview screen 82, the user can preview the photo to be taken. The content displayed on the preview screen is consistent with the camera lens's field of view and changes in real time as the content within the camera lens changes.

[0163] If a target object is detected in the preview screen, an icon corresponding to the type of the target object can be displayed on the camera interface. When the target object includes young children or pets, if a young child falls within the camera's field of view, such as... Figure 9A As shown in a, Figure 7 In the example shown, the first-level trigger 71 is activated, and a child icon 91 appears on the camera interface, prompting the user to enter a camera mode specifically for that target. Alternatively, as... Figure 9A As shown in b, when a pet is detected in the preview, Figure 7 In the example shown, the first-level trigger 71 is triggered, and a pet icon 92 appears in the photo-taking interface, prompting the user to enter a photo-taking mode for a specific target object.

[0164] like Figure 9B As shown in 'a', the user can click on the child icon 91 that appears in the camera interface, causing the child icon 91 to... Figure 9A Entering the unselected state indicated by 'a' Figure 9B The selected state shown as 'a' activates the photo-taking mode for young children. For example... Figure 9B As shown in b, users can also click on the pet icon 92 in the camera interface to activate the photo mode for pet-type targets, making the pet icon 92 appear in the background. Figure 9A The unselected state shown by 'b' indicates the entry point. Figure 9B As shown in b in the diagram.

[0165] Or, except Figure 9A , Figure 9B In addition to the examples shown, there are also examples such as Figure 10A As shown in 'a', when a child is detected in the preview screen, a thumbs-up icon 1001 is displayed. Icon 1001 indicates that a target object has been detected in the preview screen. Figure 10A As shown in b, when a pet is detected in the preview, icon 1001 is also displayed to indicate that a target object has been detected in the preview. Users can... Figure 10A Click on icon 1001 to enter the photo-taking mode for the target object.

[0166] In another possible implementation, if a target object is detected in the preview screen, the system can automatically enter a mode to take a picture of the target object. Users can configure this feature to disable the automatic entry into target object mode upon detection in the preview screen. Alternatively, users can configure the system to enter target object mode based on user commands when a target object is detected in the preview screen.

[0167] In another possible implementation, if a target object is detected in the preview screen, controls for automatically taking a picture of the target object and for manually taking a picture of the target object can be displayed. The user can click the corresponding control to enter either the automatic or manual picture mode for the target object. Figure 10B In the "a" field, regardless of the type of target object detected, icons corresponding to various target objects can be displayed in the preview screen, such as... Figure 10B The image shows icons such as child icon 1002 and pet icon 1003. When a child is detected in the preview screen, child icon 1002 is used as a trigger control to activate the automatic photo-taking mode for children. After the user clicks child icon 1002, the automatic photo-taking mode is activated for the child-type target object. For example... Figure 10B In the example 'b', an animal is detected in the preview screen, but the child icon 1002 can still be used as a trigger control for the manual photo-taking mode for children. In automatic photo-taking mode, the image of the target object is automatically acquired; in manual photo-taking mode, the image of the target object is acquired based on the user's manually issued command. When the target object is a pet, two different pet icons can be displayed, serving as trigger controls for both the automatic and manual photo-taking modes for pets.

[0168] In another possible implementation, the target object's icon is displayed regardless of whether it exists in the preview screen; only when the target object is detected is its icon changed to a different state. For example... Figure 10C As shown in 'a', when no young child is detected in the preview screen, the child icon 1002 on the camera interface is unselected (or in the first display state). Figure 10C As shown in b, when a young child is detected in the preview screen, the child icon 1002 on the camera interface is selected (or in the second display state). When the target is a pet, the same method can be used... Figure 10C Similar pet icons. When the target object is of other types, a similar pet icon can be used. Figure 10COther similar types of icons. Alternatively, if all different types of target images are displayed using the same icon, then that icon could be... Figure 10C The two different states in the preview indicate whether the target object has been detected.

[0169] In another possible implementation, when the user clicks... Figure 9B After selecting the child icon 91 or the pet icon 92, you can activate them individually or sequentially. Figure 7 The secondary trigger 72 and the tertiary trigger 73 shown obtain the score information of the image frames in the preview screen and can display the score information in the camera interface, such as... Figure 9C The score of 2.1 corresponding to the image structure evaluation parameter shown in 'a' and the score of 6.3 regarding the pose of the target object are as follows: Figure 9C The image structure evaluation parameter shown in b has a score of 2.2 and a score of 7.5 for the pose of the target object. If the target object is present in the preview image, the target object's detection box 93 can also be displayed. If the target object includes a face, the face's detection box 94 can also be displayed.

[0170] When the second evaluation result of the three-level trigger determines that the target object in the current preview screen meets the conditions for taking a picture, it can display as follows: Figure 9D The camera interface is shown below. In automatic shooting mode, as shown... Figure 9D As shown in 'a', the decision engine within the camera app can execute the decision to take a picture and obtain 95% of the preview image. In manual shooting mode, as... Figure 9D As shown in b, the score of the preview screen (including the score corresponding to the second evaluation result and / or the score corresponding to the third evaluation result) is displayed in real time in the photo-taking interface, and the preview image 96 is obtained in response to the user's click operation on the photo-taking control 97.

[0171] In another example, based on Figure 7 As shown in the example, the second-level flip-flop 72 and the third-level flip-flop 73 may include, for example: Figure 11A , Figure 11B , Figure 11C The modules shown can be used by both the second-level flip-flop 72 and the third-level flip-flop 73.

[0172] like Figure 11AAs shown, the image structure scorer and pose scorer may include a Best Moment Estimate Algorithm module 1101, a Human Body Detect module 1102, a Best Moment Detect module 1103, an Action Score module 1104, and a Best Moment Score module 1105. The Best Moment Estimate Algorithm module 1101 can determine whether the current moment is the optimal moment. If the target object is in the photo preparation stage, the Best Moment Estimate Algorithm module 1101 can determine that the current moment is not the optimal moment. After the target object is ready, the Best Moment Estimate Algorithm module 1101 can output a determination that the current moment is the optimal moment. In response to the determination of the optimal moment, the Human Body Detection module 1102 outputs a list of human body detection bounding boxes (human_bbox_list, human body box list). Based on the list of human body detection bounding boxes, the human body detection boxes (rectangles) and their scores can be obtained. Among them, the human body detection module 1102, the human body detection box information list, and the human body detection box score only use the word "human body" in their names. In the application process, they can be used to detect other types of target objects. For example, the human body detection module 1102 can obtain the pet detection box of the pet type target object and the corresponding score of the pet detection box.

[0173] Based on the human detection box obtained by the human detection box module 1102, the centroid and primary line of the human detection box can be obtained. Based on the centroid and primary line, region size parameters, visual balance parameters, and rule of thirds composition parameters can be obtained, and the composition score of the image frame of the preview screen can be obtained. The composition score can be weighted and calculated to obtain the first evaluation result of the aforementioned embodiment, or the composition score can be used as the first evaluation result of the aforementioned embodiment.

[0174] At the same time, still refer to Figure 11A Based on the data output by the human body detection module 1102, the optimal moment monitoring module 1103 can calculate and obtain an action score. The action score can be used to obtain the second evaluation result of the aforementioned embodiment, or the action score can be used as the second evaluation result of the aforementioned embodiment. Combining the image structure score and the action score yields the optimal moment score. Based on the optimal moment score, it can be determined whether to acquire the image frame of the current preview image.

[0175] In one possible implementation, refer to Figure 11B Based on the detection bounding box 1106 and its score output by the human detection module, an algorithm for obtaining the composition score can be executed. When the target object is a person, the centroid 1107 and main body line 1108 of the human detection bounding box can be obtained based on the bounding box and its corresponding score. Based on the centroid 1107 and main body line 108, the region size score, visual balance score, and rule-of-thirds composition score can be obtained. The region size score can be determined for a single detection bounding box 1106, based on the human body's detection bounding box 1106, or the relative region size of multiple target object detection bounding boxes 1106 in the preview image, and the preview image itself. The visual balance score of the preview image can be determined based on the target object's detection bounding box 1106, centroid 1107, and main body line 1108. The rule-of-thirds composition score can be obtained based on the horizontal rule-of-thirds lines 1109 and vertical rule-of-thirds lines 1110 of the preview image. The image structure score can be obtained by considering the area size score, visual balance score, and rule of thirds composition score.

[0176] If the image structure score reaches the preset image structure score threshold, Figure 11A The optimal moment detection module 1103 shown is triggered, calling the action scoring module 1104 to score the action of the target object in the preview screen. The process of scoring the action is as follows: Figure 11C As shown. Taking an animal as the target object as an example, when the target object is detected in the preview screen, firstly, multiple image frames 1111 corresponding to the preview screen are determined. For the predefined action semantics of the target object, including limb extension, jumping, running, head turning, and laughing, the score threshold for each category is 0.5. For multiple image frames 1111, the score corresponding to each action semantic is output in the order of limb extension, jumping, running, head turning, and laughing. For example... Figure 11CAs shown, the action scores for the semantic actions of limb extension, jumping, running, turning the head, and laughing in the first image frame 1111 are 0.3 / 0.4 / 0.0 / 0.0 / 0.0, respectively. This means that the action score for limb extension in the first image frame 1111 is 0.3, the action score for jumping is 0.4, the action score for running is 0.0, and the action score for laughing is 0.0. In other words, in the first image frame 1111, the probability that the target object performs the action of limb extension and jumping is 0.3 and 0.4, respectively, while the probability that the target object performs the action of running, turning the head, and laughing is 0. Similarly, the action scores for the second image frame 1111 corresponding to the action semantics of limb extension, jumping, running, turning the head, and laughing are 0.5 / 0.2 / 0.0 / 0.0 / 0.0 respectively; the action scores for the third image frame 1111 corresponding to the action semantics of limb extension, jumping, running, turning the head, and laughing are 0.7 / 0.2 / 0.0 / 0.0 / 0.0 respectively; the action scores for the third image frame 1111 corresponding to the action semantics of limb extension, jumping, running, turning the head, and laughing are 0.2 / 0.3 / 0.0 / 0.0 / 0.0 respectively. If the judgment threshold for each action semantic category is 0.5, then... Figure 11C The category label for the second image frame is "limbs extended". Figure 11C The category label for the third image frame is "limbs extended". Through score verification, the category label with the highest action score is "limbs extended", and the corresponding image frame is [image frame number missing]. Figure 11C The third image frame 1111 is selected as the image frame at the optimal moment. If the automatic shooting mode is in progress, the third image frame 1111 can be used as the photo obtained by taking the picture. If the manual shooting mode is in progress, the third image frame 1111 can be obtained as the photo obtained by taking the picture in response to the shooting command sent by the user.

[0177] This application embodiment further provides an image acquisition device, such as... Figure 12 As shown, it includes: a preview screen display module 1201, a first evaluation result module 1202, a second evaluation result module 1203, and an acquisition module 1204.

[0178] The preview screen display module 1201 is used to display the preview screen of the acquired image on the terminal.

[0179] The first evaluation result module 1202 is used to obtain the first evaluation result of the screen structure in the preview screen in response to the target object detected in the preview screen.

[0180] The second evaluation result module 1203 is used to obtain the second evaluation result of the pose of the target object in the preview screen when the first evaluation result meets the first preset condition.

[0181] The acquisition module 1204 is used to acquire the image of the preview screen based on the second evaluation result.

[0182] It should be noted that the image acquisition device provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0183] Meanwhile, the image acquisition device of this application embodiment may also include other functional modules to implement any steps and functions in the image acquisition method provided in any embodiment of this application.

[0184] In this application embodiment, a user can take a picture using an electronic device serving as a user terminal, and obtain an image of the target object by applying the image acquisition method provided in any embodiment of this application. Before providing a detailed description of the process of the image acquisition method provided in this application embodiment, a brief introduction to the software architecture of the electronic device involved in this application embodiment will be given first. Figure 13 This is a schematic diagram of the software architecture of an electronic device according to an exemplary embodiment. Figure 13 The electronic device shown runs the Android operating system. In other embodiments, the electronic device may be equipped with other types of operating systems.

[0185] Figure 13 This is a block diagram of a software system for an electronic device to which the method provided in the embodiments of this application is applied. See also... Figure 13 A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system comprises multiple layers, from top to bottom: the application layer, the application framework layer, the hardware abstraction layer, the driver layer, and the hardware layer. When the electronic device is a touchscreen, the user can send operation commands for schedule information layer by layer from the kernel layer to the application layer via the touchscreen.

[0186] The application layer can include a series of application packages. For example... Figure 13 As shown, the application package can include applications such as camera and gallery. These applications can all send information to each other. For example, the gallery application can store images captured when taking photos, or it can be used to cache the image frame sequence generated by the camera application when taking photos.

[0187] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example... Figure 13 As shown, the application framework layer can include camera management, camera devices, etc. Camera management and camera devices can access the camera application in the application layer through the camera access interface. When users operate the camera application through hardware such as touch screen, mouse, keyboard, touchpad, etc., the user's operation commands can be transmitted through the corresponding modules in the application framework layer.

[0188] The hardware abstraction layer is a layer between the camera hardware and software system, such as... Figure 13 As shown, the hardware abstraction layer includes a camera hardware abstraction layer, which includes the camera device corresponding to the terminal camera, such as... Figure 13 The camera device includes camera device 1 and camera device 2, etc. When the camera application runs, it can call algorithms from the camera algorithm library to perform a photo-taking operation. When the camera application calls algorithms from the camera algorithm library, it can implement the image acquisition method provided in any embodiment of this application, acquiring images by taking photos. Furthermore, the image acquisition method provided in the embodiments of this application, before automatically acquiring images or acquiring images in response to user instructions, can have its implementation process concentrated in the Hardware Abstraction Layer (HAL).

[0189] The driver layer includes drivers for hardware devices, such as camera devices (e.g., shown in Figure 13) that drive camera device 1, camera device 2, etc. The driver layer also includes digital signal processor (DSP) drivers and graphics processor (GPU) drivers.

[0190] Below the aforementioned four-layer architecture, the receiving device also includes a hardware layer. This hardware layer may include the aforementioned electronic device hardware components, such as sensors, image signal processors, digital signal processors, and graphics processors. The digital signal processors and graphics processors can call algorithms from the camera algorithm library to implement the steps in the methods described in this application's embodiments.

[0191] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.

[0192] The image acquisition device and image acquisition method embodiments provided in the above embodiments belong to the same concept. The specific working process and technical effects of the units and modules in the above embodiments can be found in the method embodiment section, and will not be repeated here.

[0193] As an example of this application, the aforementioned electronic device is capable of accessing a base station and also has the ability to access a wireless local area network. For example, the electronic device may be a mobile phone, tablet, smartwatch, or laptop. Please refer to [link / reference]. Figure 14 , Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0194] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0195] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0196] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0197] The processor 110 may also include a memory for storing data such as instructions and images. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly, such as repeatedly performing recording operations on target information.

[0198] In some embodiments, the processor 110 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0199] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0200] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, electronic device 100 may include one or N displays screens 194, where N is an integer greater than 1. During image acquisition, the display screen can be used to display a preview image or the acquired image.

[0201] Touch sensor 180K, also known as "touch panel". Touch sensor 180K can be set on display screen 194. Touch sensor 180K and display screen 194 together form touch screen, also known as "touch screen".

[0202] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0203] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. Available media can be magnetic media (such as floppy disks, hard disks, and magnetic tapes), optical media (such as Digital Versatile Discs (DVDs)), or semiconductor media (such as Solid State Disks (SSDs)).

[0204] The above-described embodiments are optional embodiments provided by this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the technical scope disclosed in this application should be included within the protection scope of this application.

Claims

1. An image acquisition method, characterized in that, The method comprises the following steps: displaying a preview picture of image acquisition on a terminal; obtaining a first evaluation result of a picture structure in the preview picture in response to a target object detected in the preview picture; obtaining a second evaluation result of a pose of the target object in the preview picture in a case where the first evaluation result meets a first preset condition; obtaining an image of the preview picture according to the second evaluation result.

2. The method of claim 1, wherein, After the preview picture is displayed on the terminal, the method further comprises the following steps: obtaining an object in the preview picture; judging whether the type of the object is the target object to obtain a judgment result; displaying an inquiry control on the terminal according to the judgment result; entering the step of obtaining the first evaluation result of the picture structure in the preview picture in response to a click operation of a user on the inquiry control.

3. The method according to claim 1 or 2, characterized in that, The step of obtaining the first evaluation result of the picture structure in the preview picture comprises the following steps: obtaining a picture structure evaluation parameter of the preview picture according to a coordinate of the target object in the preview picture; the picture structure evaluation parameter comprises at least one of a region size parameter, a visual balance parameter and a rule of thirds composition parameter; obtaining the first evaluation result according to the picture structure evaluation parameter.

4. The method of claim 3, wherein, The picture structure evaluation parameter comprises the region size parameter, the visual balance parameter and the rule of thirds composition parameter, and the step of obtaining the picture structure evaluation parameter of the preview picture according to the coordinate of the target object in the preview picture comprises the following steps: determining a proportion of a region size of the target object in the preview picture according to the coordinate; the proportion is taken as the region size parameter; determining a longitudinal position and a transverse position of the target object in the preview picture according to the coordinate; the transverse position and the longitudinal position are taken as the visual balance parameter; determining a minimum distance between the target object and a rule of thirds node in the preview picture according to the coordinate; the rule of thirds node is a node in the preview picture meeting a rule of thirds composition rule; the minimum distance is taken as the rule of thirds composition parameter.

5. The method of claim 1, wherein, The step of obtaining the second evaluation result of the pose of the target object in the preview picture comprises the following steps: obtaining a motion semantic feature of the target object; determining a score corresponding to the motion semantic feature; obtaining the second evaluation result according to the score in a case where the score is higher than a preset threshold.

6. The method of claim 5, wherein, The step of obtaining the second evaluation result according to the score in the case where the score is higher than the preset threshold comprises the following steps: in the case where the score is higher than the preset threshold, assigning a category label of the corresponding motion semantic feature to an image frame of the preview picture; sorting scores corresponding to image frames including the category label in continuous image frames of the preview picture; determining a target category label in the category label according to a sorting result; taking the target category label as the second evaluation result.

7. The method of claim 1, wherein, The step of obtaining the image of the preview picture according to the second evaluation result comprises the following steps: in a case where the second evaluation result indicates that there is an image frame with the target category label in the preview picture, displaying a photographing prompt control on the terminal. In response to a click operation of the user on the photographing prompt control, an image frame with the target category label of the action is acquired as the image of the preview picture.

8. The method of claim 1, wherein, The acquiring the image of the preview picture according to the second evaluation result comprises: In a case where the second evaluation result indicates that the image frame with the target category label exists in the preview picture, the image frame with the target category label is acquired as the image of the preview picture.

9. The method of claim 1, wherein, The target object is at least one of an animal and a human body in a preset age range.

10. An electronic device, comprising: The electronic device comprises a processor and a memory. The memory is configured to store a program for the electronic device to execute the method according to any one of claims 1-9, and store data involved in the method according to any one of claims 1-9. The processor is configured to execute the program stored in the memory.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores instructions, and when the instructions run on the computer, the computer executes the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Image shooting method and device

    CN105323491A

  • Shooting method and device

    CN115086556A