Shooting method, terminal equipment and computer readable storage medium
By collecting preview streams in the terminal device and automatically saving images that meet the conditions, the problem of missing wonderful moments due to delayed reaction time is solved, and the user experience is improved, especially in the recording of human pets and parent-child scenarios.
Patent Information
- Application Number
- CN202311574529.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-11-22
AI Technical Summary
When shooting sports scenes, terminal devices often miss exciting moments due to delayed reaction time, resulting in distortion or failure of recorded images, and poor user experience.
By setting up a camera module in the terminal device, the preview stream of the current scene is collected in real time, and the target image is automatically saved according to preset conditions, including multiple subjects and in the presence of children, people and animals.
It realizes that it automatically captures wonderful images without relying on user response time, improves user experience, and meets users' recording needs for human pets and parent-child scenarios.
Smart Images

Figure CN120075594A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of terminals, and in particular, to a shooting method, a terminal device, and a computer-readable storage medium. Background Art
[0002] The camera on a terminal device, such as a mobile phone, can often record beautiful scenes. When a user wants to freeze a beautiful scene, there is still a certain reaction time between pressing the shutter to record. And often during this reaction time, the wonderful moment ends. Especially for shooting some moving subjects, wonderful moments are often missed, and even the recorded images may be distorted or blurry failed images, resulting in a poor user experience. Summary of the Invention
[0003] Based on this, the present application provides a shooting method, a terminal device, and a computer-readable storage medium, which solve the problem of missing the wonderful moment of an image due to the reaction time, ensure the output of wonderful images, and improve the user experience.
[0004] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0005] In a first aspect, the present application provides a shooting method, which can be applied to a terminal device including a camera module. In this method, the terminal device receives an operation of the user to open the camera application, and according to this operation, collects a preview stream of the current scene through the camera module. When the current scene meets the first condition, the target image in the preview stream of the collected current scene is saved. The first condition includes that the current scene includes multiple subjects and there is a child among the multiple subjects; or the first condition includes that the current scene includes multiple subjects and there is a person and an animal among the multiple subjects.
[0006] In the above method, after identifying that the current scene is a qualified scene, such as including multiple subjects and there is a child, or including multiple subjects and there is a person and an animal, the target image in the preview stream is automatically saved, thus realizing automatic capture. There is no need to consider the reaction time and miss the record of wonderful moments. At the same time, it also covers the current user's recording needs for human-pet scenes and parent-child scenes, ensuring the user experience.
[0007] In addition, the preview stream collection is realized through the camera module set on the terminal device, so as to realize automatic capture. Among them, the camera module can include the front camera and / or the rear camera of the terminal device.
[0008] In this implementation manner, the user's need for automatic capture through the front camera and / or the rear camera can be met, freeing the hands and satisfying the recording of wonderful pictures of the user taking pictures with their own children or animals.
[0009] In an implementable manner of the first aspect, determining whether the current scene is a qualified scene mainly involves determining the number of entities included in the current scene and whether the current scene includes a person. When the number of entities is greater than the first threshold and the current scene includes a person, it is further determined whether the current scene includes a child or an animal. The first threshold can be an integer greater than 1.
[0010] By determining whether the number of entities is greater than the first threshold and whether the current scene includes a person, and when the above-mentioned number of entities meets the first threshold and whether the current scene includes a person, determining whether the current scene includes a child or an animal, it is relatively accurate to determine whether the current scene is a qualified scene.
[0011] Among them, determining whether the current scene includes a person can be obtained by performing entity detection on the images in the preview stream, and determining whether the current scene includes a child or an animal can be obtained by performing face detection on the images in the preview stream. As an example, the above-mentioned entity detection and face detection can be performed on each frame of the image in the preview stream, or can be performed according to the images in the preview stream at intervals of several frames, so as to save power consumption. In addition, the number of frames detected by entity detection and face detection can be the same or different. For example, entity detection can be performed once every four frames. Face detection can be performed once every five frames.
[0012] In an implementable manner of the first aspect, the information of the current scene can be determined based on a deep learning network. For example, the images included in the preview stream are input into the deep learning network for processing, and then image feature information is output. The image feature information characterizes the information of the current scene. In one implementation, the image feature information can be used to indicate the number of entities included in the current scene and the entity type. The entity type is used to indicate whether the entities included in the current scene are children or animals. In this way, the number of entities included in the current scene, whether the current scene includes a person, and whether the current scene includes a child or an animal can be determined based on the image feature information output by the deep learning network. In another implementation, the image feature information can be used to indicate the number of entities included in the current scene. When the current scene includes a person, the image feature information can also be used to indicate the age characteristics of the person included in the current scene, for determining whether the current scene includes a child. When the current scene includes an animal, the image feature information can also be used to indicate whether the entity included in the current scene is an animal, for determining whether the current scene includes an animal.
[0013] In this implementation manner, only by inputting the image into the deep learning network, the information of the current scene can be output, which is used to determine whether the current scene is a qualified scene, and the efficiency is higher.
[0014] In an implementable manner of the first aspect, when the current scene is a qualified scene, that is, when the current scene includes multiple subjects and there is a child among the multiple subjects; or when the current scene includes multiple subjects and there is a human and an animal among the multiple subjects, before saving the target image in the preview stream of the currently captured scene, the faces of the subjects in the current scene can also be matched, that is, it is determined whether the face angles of the subjects in the images included in the preview stream are within a preset range.
[0015] In this implementable manner, when the current scene is a qualified scene, by detecting the face angles of the subjects in the current scene and performing the operation of saving the target image in the preview stream when the face angles of the subjects meet the preset range, it can be ensured that the captured image can include the faces of the subjects and ensure the wonderful presentation of the image.
[0016] In an implementable manner of the first aspect, when the current scene is a qualified scene, that is, when the current scene includes multiple subjects and there is a child among the multiple subjects; or when the current scene includes multiple subjects and there is a human and an animal among the multiple subjects, the operation of saving the target image in the preview stream can be performed when it is determined that the exposure time and / or image sharpness of the images included in the preview stream, that is, the exposure time of the image is less than a second threshold and / or the image sharpness of the image is greater than a third threshold.
[0017] In this implementable manner, by triggering the photographing decision logic when it is determined that the exposure time and image sharpness meet the conditions, the captured image can be ensured to be clear and accurate.
[0018] In an implementable manner of the first aspect, when the photographing decision logic is triggered, multiple frames of images can be cached, and the target image can be selected from the cached multiple frames of images according to that the face angles of the images of the subjects in the images are within a preset range; and / or the images of the subjects in the images are located at preset positions in the images.
[0019] In this implementable manner, the image parameters in the multiple frames of images need to meet the preset angles and preset positions, so as to ensure that the ratio of the subject image in the captured image to the image is appropriate, ensure the wonderfulness of the target image, and thus meet the needs of users.
[0020] In a second aspect, the present application provides a photographing device, and the device has the function of implementing the behaviors of the terminal device in the method of the first aspect described above. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, for example, an input unit or module, a display unit or module, and a processing unit or module.
[0021] In a third aspect, a terminal device is provided. The electronic device includes: a processor; a memory; a camera module; and a computer program. The computer program is stored in the memory. When the computer program is executed by the processor, the terminal device is caused to execute the shooting method in the first aspect and any of its implementation manners as described above.
[0022] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes a computer program. When the computer program runs on an electronic device, the electronic device can be caused to execute the method described in any item of the first aspect.
[0023] In a fifth aspect, a computer program product containing instructions is provided. When it runs on an electronic device, the electronic device can be caused to execute the shooting method in the first aspect and any of its implementation manners.
[0024] In a sixth aspect, an embodiment of the present application provides a chip. The chip includes a processor. The processor is used to call a computer program in a memory to execute the shooting method in the first aspect and any of its implementation manners.
[0025] It can be understood that for the beneficial effects that can be achieved by the device described in the second aspect, the terminal device described in the third aspect, the computer-readable storage medium described in the fourth aspect, the computer program product described in the fifth aspect, and the chip described in the sixth aspect, reference can be made to the beneficial effects in the first aspect and any possible implementation manner thereof, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A schematic diagram of an automatic capture provided for the related art Figure 1 ;
[0027] Figure 2 A schematic diagram of an automatic capture provided for the related art Figure 2 ;
[0028] Figure 3 A schematic diagram of an automatic capture provided for the related art Figure 3 ;
[0029] Figure 4 A schematic structural diagram of a terminal device provided for an embodiment of the present application;
[0030] Figure 5 A schematic flowchart of a shooting method provided for an embodiment of the present application;
[0031] Figure 6 A schematic diagram of a shooting interface provided for an embodiment of the present application;
[0032] Figure 7Schematic diagram of the overall process of a shooting method provided by an embodiment of the present application;
[0033] Figure 8 Schematic diagram of the process of determining whether the current scene is a qualified scene provided by an embodiment of the present application;
[0034] Figure 9A Scene illustration of a shooting method provided by an embodiment of the present application Figure 1 ;
[0035] Figure 9B Scene illustration of a shooting method provided by an embodiment of the present application Figure 2 ;
[0036] Figure 10 Another schematic diagram of the process of determining whether the current scene is a qualified scene provided by an embodiment of the present application;
[0037] Figure 11 Schematic diagram of the process of determining a target image provided by an embodiment of the present application;
[0038] Figure 12 Schematic diagram of the structure of a chip system provided by an embodiment of the present application. Detailed implementation manners
[0039] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise stated, the meaning of "a plurality" is two or more.
[0040] Since its birth, shooting has created many wonderful moments and brought different visual enjoyments to people. With the development of technology, camera modules can be set on a variety of terminal devices, providing convenience for people. Initially, users could use the camera module on the terminal device to freeze wonderful scene moments, such as shooting some static scenes. As users' demands become higher and higher, they also need to shoot some moving scenes to record beautiful moments. When capturing a moving scene, users often need to operate the shooting button on the terminal device. When a user sees a wonderful moment they want to freeze, there is often a certain reaction time from seeing it to pressing the shutter to record, and the wonderful moment passes by in a flash, so it is often very difficult to freeze and record. As Figure 1 shown, the user wants to capture the state of the shooting subject (such as Figure 1 the person shown in Figure 1The third image in the sequence. During this process, if the capture button is pressed when the human eye sees the scene to be captured, due to the reaction time of the human, which is generally 50 ms, the actually captured image often does not match the scene seen by the human eye. For example, the captured image may be the one with the subject extended to the left.
[0041] Such as Figure 2 As shown, during the shooting process of a terminal device, the related solutions mainly involve identifying a moving subject within the shooting range through the rear camera on the terminal device. After the subject is identified, the dynamic image of the subject is automatically captured. However, currently, the scenarios for capturing moving scenes through the rear camera are limited, only applicable to some scenes with relatively large movement amplitudes, such as a ball-playing scene. It is difficult to meet the requirements for some scenes with relatively small movement amplitudes. Moreover, the current automatic capture can only be applied to the rear camera of the terminal device, and the front camera of the terminal device cannot perform automatic capture. When taking a self-portrait with the front camera, for some scenes, it is also impossible to record wonderful moments. For example Figure 3 In the human-pet scene shown, the animal is relatively active. While an adult needs to control the animal and press the capture button to take a self-portrait, it is very difficult to ensure the quality of the image.
[0042] Based on the above, the embodiments of the present application provide a shooting method. After the terminal device receives an operation to start shooting, if it determines that the current scene is a human-pet scene or a parent-child scene, it can use the image in the preview stream of the currently captured scene that meets the judgment conditions as the target image to be finally output. The above shooting method can automatically capture high-quality images during the shooting process, thereby meeting the user's capture requirements for various running scenes, such as those with relatively small movement amplitudes, and improving the user experience.
[0043] Exemplarily, the terminal device in the embodiments of the present application can be a mobile phone, a tablet computer, a smart watch, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc., which are devices including a camera module. The embodiments of the present application do not impose special restrictions on the specific form of the electronic device.
[0044] Exemplarily, taking a mobile phone as the above terminal device 400 as an example, Figure 4 shows a schematic structural diagram of the mobile phone. As Figure 4As shown in the figure, the mobile phone may include a processor 410, an external memory interface 420, an internal memory 421, a mobile communication module 430, a wireless communication module 440, a charging management module 450, a power management module 460, a battery 470, an antenna 1, an antenna 2, an audio module 480, a speaker 480A, a receiver 480B, a microphone 480C, a headphone jack 480D, a camera 490, a display screen 491, etc.
[0045] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the mobile phone. In some other embodiments of the present application, the mobile phone may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. For example, the mobile phone may further include: a subscriber identification module (SIM) card interface, a sensor module, etc. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0046] The processor 410 may include one or more processing units. For example, the processor 410 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0047] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0048] A memory may also be provided in the processor 410 for storing instructions and data. In some embodiments, the memory in the processor 410 is a cache memory. This memory may save the instructions or data that the processor 410 has just used or recycled. If the processor 410 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 410, and thus improves the efficiency of the system. In some embodiments, the processor 410 may include one or more interfaces.
[0049] The charging management module 450 is used to receive a charging input from a charger. Here, the charger can be a wireless charger or a wired charger. While charging the battery 470, the charging management module 450 can also supply power to the terminal device through the power management module 450.
[0050] The power management module 460 is used to connect the battery 470, the charging management module 460, and the processor 410. The power management module 460 receives inputs from the battery 470 and / or the charging management module 450 and supplies power to the processor 410, the internal memory 421, the display screen 400, the camera 490, the wireless communication module 440, etc. The power management module 460 can also be used to monitor parameters such as the battery capacity, the number of battery charge cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 460 can also be disposed in the processor 410. In some other embodiments, the power management module 460 and the charging management module 450 can also be disposed in the same device.
[0051] The wireless communication function of the mobile phone can be implemented through Antenna 1, Antenna 2, the mobile communication module 430, the wireless communication module 440, the modulation and demodulation processor, and the baseband processor, etc.
[0052] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the mobile phone can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, Antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0053] The mobile communication module 430 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the mobile phone. The mobile communication module 430 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 430 can receive electromagnetic waves through Antenna 1, perform filtering, amplification, etc. on the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation.
[0054] The wireless communication module 440 can provide wireless communication solutions applied to mobile phones, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0055] In some embodiments, antenna 1 of the mobile phone is coupled to the mobile communication module 430, and antenna 2 is coupled to the wireless communication module 440, enabling the terminal device 400 to communicate with the network and other devices through wireless communication technologies.
[0056] The terminal device 200 realizes the display function through the GPU, the display screen 491, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 491 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 410 may include one or more GPUs, which execute program instructions to generate or change display information.
[0057] The display screen 491 is used to display images, videos, etc. The display screen 491 includes a display panel. For example, the display screen 491 can be a touch screen. In some embodiments of this application, after the user opens the camera application, the display screen 491 can be used to display the preview stream collected by the camera 490.
[0058] The camera 490 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transfers the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard formats such as RGB and YUV. In some embodiments, the mobile phone may include one or N cameras 490, where N is a positive integer greater than 1. In some embodiments of the present application, after the mobile phone receives the operation of the user to open the camera application, the mobile phone can control the camera 490 to turn on. After the camera 490 is turned on, the camera 490 can be used to collect the preview stream of the current scene. Additionally, in some embodiments of the present application, when the mobile phone includes multiple cameras 490, these multiple cameras 490 may include a front camera and / or a rear camera.
[0059] The internal memory 421 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 421 may include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the mobile phone (such as audio data, phone book, etc.). In addition, the internal memory 421 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 410 enables the mobile phone to implement audio functions such as music playback and recording through the audio module 480 and the processor 410 by running the instructions stored in the internal memory 421 and / or the instructions stored in the memory provided in the processor. Among them, the audio module 480 includes a speaker 480A, a receiver 480B, a microphone 480C, and a headphone jack 480D.
[0060] The audio module 480 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into digital audio signals. The audio module 480 can also be used to encode and decode audio signals. In some embodiments, the audio module 480 may be provided in the processor 410, or some functional modules of the audio module 480 may be provided in the processor 410.
[0061] The speaker 480A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The terminal device 400 can listen to music or a hands-free call through the speaker 480A.
[0062] The receiver 480B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the terminal device 400 answers a call or a voice message, the voice can be listened to by placing the receiver 480B close to the human ear.
[0063] The microphone 480C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak with the mouth close to the microphone 480C to input the sound signal into the microphone 480C. The mobile phone can be provided with at least one microphone 480C. In some other embodiments, the mobile phone can be provided with two microphones 480C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the mobile phone can also be provided with three, four or more microphones 480C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0064] The headphone jack 480D is used to connect a wired headphone. The headphone jack 480D can be a USB interface 410, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0065] The following describes the specific process for the terminal device to implement shooting. As Figure 5 shown, Figure 5 FIG. is a schematic flowchart of a shooting method provided by an embodiment of the present application. The method is applied to a terminal device including a camera module, taking a mobile phone as an example. The method includes S501 - S503.
[0066] S501. The mobile phone receives an operation by the user to open the camera application.
[0067] Among them, a camera application can be installed in the mobile phone for the user to take pictures. In some embodiments, when the automatic capture mode is default enabled, the operation to open the camera application can be an operation on the icon of the camera application. Exemplarily, the desktop of the mobile phone can include an icon of the camera application, and the user can operate on the icon of the camera application displayed on the desktop, such as a click operation. Correspondingly, the mobile phone can receive the operation by the user on the icon of the camera application.
[0068] In some other embodiments, when the automatic capture mode is not enabled by default, the operation of opening the camera application may include an operation on the icon of the camera application and an operation of enabling the automatic capture mode. Exemplarily, the user may operate on the icon of the camera application displayed on the desktop. Correspondingly, the mobile phone may receive the operation of the user on the icon of the camera application. In response, the mobile phone may display the interface of the camera application. The interface of the camera application may include a control for enabling the automatic capture mode. The user may operate on this control to enable the automatic capture mode. It should be noted that this embodiment is described by taking the control for enabling the automatic capture mode as included in the interface of the camera application as an example. The control for enabling the automatic capture mode may also be set in other interfaces, such as the interface of the settings application. The embodiments of the present application do not make specific limitations in this regard. It can be understood that in this embodiment, the photographing decision logic may be triggered after the user enables the automatic capture mode, which can save the device power consumption.
[0069] S502. In response to the operation, the mobile phone collects a preview stream of the current scene through the camera module.
[0070] Exemplarily, after the mobile phone receives the operation of the user to open the camera application, in response to this operation, it may display the interface of the camera application, such as a preview interface. Additionally, in response, the mobile phone may also turn on the camera module of the mobile phone to collect a preview stream of the current scene. For example, after the mobile phone receives the operation of the user to open the camera, the camera module is turned on, and the camera module can receive the light of the current scene to form a preview stream of the current scene that can be seen by the user in real time. The mobile phone may also display the preview stream of the current scene collected by the camera module in the above preview interface.
[0071] Among them, the so-called preview stream is a video stream composed of multiple preview images, that is, the dynamic picture of the current scene is recorded in real time through the camera module. What is displayed on the preview interface of the mobile phone is the video stream of the dynamic picture of the current scene, waiting to be photographed. The so-called camera module generally has basic functions such as video shooting / transmission and static image capture. After the image is collected by the lens, the photosensitive component circuit and the control component in the camera module process the image and convert it into a digital signal that can be recognized by the terminal device, and then the image is restored after internal processing. The restored image can be displayed to the user through the display screen of the mobile phone. For example, combined Figure 6 , when the current scene is a little dog walking forward, at this time, the preview stream of the current scene collected by the mobile phone through the camera module is a video stream of the little dog walking forward composed of multiple images, that is, the preview stream. The display screen of the mobile phone may display a preview interface as shown in Figure 6 . The preview interface includes the video stream of the little dog walking forward collected by the camera module for the user to view. It can be understood that the preview interface is the viewfinder interface before the mobile phone starts automatic capture. The preview stream included in this preview interface is the real-time picture of the real scene.
[0072] In the above example, the preview stream captured by the camera module can be obtained through the rear camera of the terminal device or through the front camera of the terminal device.
[0073] S503. In response to the current scene satisfying the first condition, save the target image in the preview stream of the currently captured scene; wherein, the first condition includes: the current scene includes multiple subjects, and there is a child among the multiple subjects; or the first condition includes: the current scene includes multiple subjects, and there is a person and an animal among the multiple subjects.
[0074] Exemplarily, since users often record wonderful moments with their children, but most children are hyperactive and restless by nature. Even if the camera module of the terminal device is relatively good, it is possible that the user cannot take a satisfactory photo. Moreover, when taking a self-portrait with the front camera of the terminal device or taking a group photo with the rear camera, the user also needs to control the actions of the child. Therefore, it is necessary to use the automatic capture function of the terminal device to free the hands of adults and meet the user's usage requirements. Similarly, like children, animals are also hyperactive and restless by nature, and even animals are more flexible than children, and there are also situations that are difficult to control. When wanting to take a self-portrait with the front camera or a group photo with the rear camera with one's own animal, it is also necessary to use the automatic capture function of the terminal device. In order to meet the user's need to capture a group photo with their own animal or child during shooting, the terminal device needs to determine eligible scenes from numerous scenes.
[0075] Among them, to determine eligible scenes among numerous scenes, it is mainly determined through two conditions. In some embodiments, one of them is to determine whether the current scene contains multiple subjects, and the other is whether there is a child among the multiple subjects. When both of these conditions are met, the current scene can be determined as an eligible scene, such as a parent-child scene of an adult and a child, a group photo scene of children, a human-pet scene of a child and an animal, etc. In other embodiments, one of them is to determine whether the current scene contains multiple subjects, and the other is whether there is a person and an animal among the multiple subjects. When both of these conditions are met, the current scene can be determined as an eligible scene, such as a human-pet scene of a person and an animal. Among them, the person here can be a child or an adult.
[0076] In the embodiments of the present application, a subject refers to a person or an animal in the current scene captured by the camera module. For example, an adult, a child, a cat, a dog, etc. For example, assume that the current scene is "the user squats down outdoors and touches their pet dog". In the preview stream obtained by the camera module, in addition to the person and the pet dog, there are also the road and the scenery by the roadside. At this time, the subjects refer to the person and the pet dog in this scene.
[0077] Such as Figure 7As shown, after the mobile phone captures the preview stream, it can perform scene perception calculation on the images included in the preview stream to determine the number of subjects included in the current scene and whether the current scene includes a person; when the number of subjects is greater than the first threshold and the current scene includes a person, it further determines whether the current scene includes a child or an animal. Herein, the first threshold may be an integer greater than 1. As an example, the preview stream can be subject-detected to determine the number of subjects in the current scene and whether the current scene includes a person. The specific judgment process is as Figure 8 shown, the preview stream can be subject-detected to determine whether the current scene includes multiple subjects and whether the current scene includes a person. For example, it can be determined whether the number of subjects included in the current scene is greater than the first threshold (such as determining whether the number of subjects in the current scene is greater than or equal to 2) to determine whether the condition of multiple subjects is met. After the mobile phone determines that the condition of multiple subjects is met and determines that the current scene also includes a person, the mobile phone can further determine whether the current scene includes a child or an animal. Herein, the mobile phone can first determine whether a child is included, and if a child is not included, it further determines whether an animal is included. That is to say, when making a logical judgment, the priority of judging whether the subject is a child is higher than the priority of judging whether the subject is an animal. For example, continuing to combine Figure 8 it can be determined whether a child is included in the current scene by performing face detection on the preview stream (such as determining whether the number of children included in the current scene is greater than 0). At this time, if a child is included (such as the number of children included in the current scene is greater than 0), it can be determined as a qualified scene, such as being determined as a parent-child scene. If a child is not included, it can continue to determine whether an animal is included in the current scene (such as determining whether the number of animals included in the current scene is greater than 0). If an animal is included (such as the number of animals included in the current scene is greater than 0), it can be determined as a qualified scene, such as being determined as a human-pet scene. If it is determined that the number of subjects is less than 2, or it is determined that the number of subjects is greater than or equal to 2 but there is no person, or it is determined that the number of subjects is greater than or equal to 2 and there is a person but there is no child or animal among the subjects included in the current scene, the current process can be ended.
[0078] It should be noted that in the above example, the priority of children is higher than that of animals, and it can also be set that the priority of animals is higher than that of children. The order is not limited herein.
[0079] Taking the judgment of the current scene as a parent-child scene and the current scene as a human-pet scene as examples below.
[0080] In some examples, taking the current scene as a parent-child scene as an example. Suppose the current parent-child scene includes an adult and a child. Combining Figure 8As shown, after the mobile phone captures the preview stream through the camera module, it performs object detection on the images included in the preview stream to obtain the result that the number of objects is 2 and the current scene includes people. After that, the mobile phone continues to perform face detection based on the images in the preview stream to determine that there is a child among the objects included in the current scene. Among them, face detection refers to obtaining the facial feature information of the object, and it can be determined whether there is a child in the object according to the facial feature information. Then, it can be determined that the current scene is a parent-child scene, that is, it can be determined that the current scene is a qualified scene. The mobile phone can then automatically save the target image in the preview stream of the currently captured scene to achieve automatic capture. As Figure 9A shown Figure 9A shows that the current scene is a parent-child scene.
[0081] In some other examples, taking the current scene as a human-pet scene as an example, assume that the current human-pet scene includes an adult and a pet. Combining Figure 8 shown, after the mobile phone captures the preview stream through the camera module, it performs object detection on the images included in the preview stream to obtain the result that the number of objects is 2 and the current scene includes people. After that, the mobile phone continues to perform face detection based on the images in the preview stream and determines that there is no child among the objects included in the current scene. Then the mobile phone continues to determine that there is an animal among the objects included in the current scene. Then, it can be determined that the current scene is a human-pet scene, that is, it can be determined that the current scene is a qualified scene. The mobile phone can then automatically save the image in the preview stream of the currently captured scene to achieve automatic capture. As Figure 9B shown Figure 9B shows that the current scene is a human-pet scene.
[0082] Furthermore, the above determination of the number of objects and the current scene including people and the further determination of whether the current scene includes children or animals are determined based on the output results after processing by the deep learning network. Specifically, after the mobile phone captures the preview stream through the camera module, the preview stream can be processed by the deep learning network to output the image feature information of the images in the preview stream, and this image feature information indicates the number of objects and the object types included in the current scene. Among them, the object type is used to indicate whether the objects included in the current scene are children or animals; among them, whether the current scene includes people and whether the current scene includes children or animals are determined based on the object type. Based on the results output by the deep learning network, the mobile phone can determine whether the current scene is a qualified scene. For example, the camera module captures the current scene to obtain the preview stream of the current scene, and inputs the images included in the preview stream into the deep learning network for processing to obtain the processed output results. Among them, the output result is that the number of objects is 2, including animals and children, and according to the above two conditions, it is determined that the current scene is a qualified scene.
[0083] In some other embodiments, in addition to the unchanged output of the number of subjects, the output result of the deep learning network may not directly indicate whether the subject includes an animal or a child. Instead, when the subject is a person, the subject type may be the age characteristic of the subject, indirectly indicating whether the subject is a child. What can be finally output are the number of subjects and the age characteristic of the subject. Then, the mobile phone can match the age characteristic of the subject with the corresponding characteristic of the target scene to determine whether the current scene is a qualified scene. For example, taking the subject as a child. The mobile phone can match the age characteristic of the subject with the age characteristic of the subject in the parent-child scene to determine whether the current scene includes a child, so as to determine whether the current scene is the corresponding parent-child scene in the qualified scenes. Taking the above scene as an example, assume that the number of subjects detected is two subjects, and then one of the subjects is a person and the output age characteristic is 8 years old. Match the age characteristic with the characteristic of the target scene, where the age characteristic corresponding to the child in the target scene is less than or equal to 13 years old. It is found that 8 years old is within the range of the target characteristic of 13 years old, so it is determined that the scene is a qualified parent-child scene. In addition, in the case where the current scene includes an animal, an indication information of whether the current scene includes an animal can also be input to determine whether the current scene includes an animal.
[0084] In addition, in the above embodiments, it is described by taking the mobile phone as an example that when it is determined that the current scene meets the conditions of the number of subjects and the subject type, the current scene is determined to be a qualified scene, and then the automatic capture can be triggered. In some other embodiments of the present application, in addition to taking the above two conditions as the conditions for triggering the automatic capture, the condition of whether the subject in the current scene is looking at the camera module can also be further used as the condition for triggering the automatic capture, so as to be able to capture an image including the face of the subject. For example Figure 10 As shown, as an example, before saving the target image in the preview stream of the currently captured scene, it can be further determined whether the face angle of each subject included in the current scene is within a preset range according to the images included in the preview stream, that is, to determine whether the subjects in the images included in the preview stream are facing forward. Only when it is determined that the face angle of each subject among the multiple subjects is within the preset range, will the automatic capture be triggered. Among them, the preset range can be [0°, 45°]. Among them, the setting of 0° means that the line of sight of the subject's eyes is directly facing the center point of the camera module; 45° is the included angle between the line of sight of the subject's eyes and the camera module being 45 degrees. It should be noted that the deviation angle of 45° includes the included angles of 45 degrees up, down, left, and right with the center point of the camera module. It should be noted that the above example is described by taking the judgment of whether the current scene is a qualified scene first, and then judging whether the face angle of each subject among the multiple subjects is within the preset range. However, the judgment of the above conditions has no sequence relationship, or can be carried out simultaneously, and the embodiments of the present application do not limit this here.
[0085] In some examples, taking the current scene as a parent-child scene as an example, assume that the parent-child scene includes an adult and a child, and the face angles of both the adult and the child are 30°. As shown in Figure 10, after the mobile phone captures the preview stream through the camera module, it can perform object detection on the images included in the preview stream to obtain the number of objects 2 and that the current scene includes people. Then, the mobile phone continues to perform face detection based on the images in the preview stream, further determines that there is a child among the objects included in the current scene, and further determines whether the face angle of each object among the multiple objects included in the current scene is within the preset range (for example, the face angle of each object is 30°, within the preset range). Thus, it can be determined that the current scene is a parent-child scene, that is, it can be determined that the current scene is a qualified scene, and the mobile phone can automatically save the images in the preview stream of the currently captured scene to achieve automatic capture.
[0086] In other examples, taking the current scene as a human-pet scene as an example, assume that the human-pet scene includes an adult and an animal, and the face angles of both the adult and the animal are 40°. As shown in Figure 10, after the mobile phone captures the preview stream through the camera module, it can perform object detection on the images included in the preview stream to obtain the number of objects 2 and that the current scene includes people. Then, the mobile phone continues to perform face detection based on the images in the preview stream, determines that there is no child among the objects included in the current scene, further determines that there is an animal among the objects included in the current scene, and then continues to determine whether the face angle of each object among the multiple objects included in the current scene is within the preset range (for example, the face angle of each object is 40°, within the preset range). It can be determined that the current scene is a human-pet scene, that is, it can be determined that the current scene is a qualified scene, and the mobile phone can automatically save the images in the preview stream of the currently captured scene to achieve automatic capture.
[0087] It should be noted that when determining whether the current scene is a qualified scene based on the images included in the preview stream, each frame of the preview stream can be detected, or some frames of the preview stream can be detected. Among them, when detecting each frame of the image, the multi-frame is processed into one frame according to the fusion algorithm for subject recognition. The detection of some frames of the image is to select some frames of the image with rules for detection. For example, when determining the number of subjects in the current scene, assuming there are sixteen frames of images, four frames are processed through the fusion algorithm model, and four frames of images are obtained, and the four frames of images are used for subject recognition; it is also possible to select regularly from these sixteen frames of images, assuming that the fourth frame, the eighth frame, the twelfth frame, and the sixteenth frame are selected for subject recognition. After determining the number of subjects in the current scene, when further determining the subject type through face detection, similarly, detection can be performed every five frames, so as to save power without affecting the scene determination. In addition, the number of frames detected for subject detection and face detection can be the same or different. For example, subject detection and face detection can also be performed every four frames.
[0088] Continue to combine Figure 7 As shown in the figure, after the terminal device senses that the current scene is a qualified scene, the photo-taking decision logic can be triggered, or in other words, automatic capture can be triggered, that is, the target image in the preview stream of the current scene is automatically saved. Further, the timing of triggering the photo-taking decision logic is determined according to at least one of the image clarity and the exposure time. Because when the camera module takes an image of the current scene, its internal processing process is mainly at the moment of shooting. The camera module receives the light reflected from the object and focuses it on the film to form an inverted and reduced real image. And in this process, aspects such as the light-receiving effect will affect the clarity of the image. That is to say, image clarity is an important indicator for determining the quality of the image. If the image is not clear enough, even if the shooting image obtained by triggering the photo-taking logic is obtained, the resulting image still does not meet the user's expectations and is still a failed image. Therefore, image clarity has become a condition for measuring whether to trigger the photo-taking decision logic. In addition, since the exposure time is related to the image clarity, that is, the longer the exposure time, the greater the possibility of motion blur and jitter blur, and it is more difficult to capture an image with better clarity. Therefore, after sensing that the current scene is a qualified scene, that is to say, triggering the photo-taking decision logic also requires a certain exposure time. Therefore, in order to ensure normal triggering of the photo-taking, it is necessary to try to control a shorter exposure time. In summary, the photo-taking decision logic can be triggered when at least one of the above conditions is met.
[0089] Among them, for the image clarity, it is necessary to meet the condition that it is greater than the third threshold before triggering the photo-taking decision logic. For the exposure time of the image, it is necessary that the exposure time of the image is less than the second threshold, that is to say, when ensuring that the exposure time meets the second threshold, the photo-taking decision logic is triggered for taking pictures.
[0090] In some instances, after the mobile phone determines that the current scene is a qualified scene, it further determines whether to trigger the photographing decision logic for the image in the current scene, that is, to determine whether the image clarity of the image included in the preview stream is greater than a third threshold (the third threshold can be represented by L, and L can be set as the probability value that the clarity is greater than 0.75). If it is greater than the third threshold, the photographing decision logic can be triggered. Among them, the third threshold is determined according to multiple parameter values, such as parameter information such as edge texture intensity.
[0091] In other instances, after the mobile phone determines that the current scene is a qualified scene, it further determines whether to trigger the photographing decision logic for the image in the current scene, that is, to determine whether the exposure time of the image included in the preview stream is greater than a second threshold (the second threshold can be represented by H, and H can be set to be less than 1 / 60 ms). If the exposure time is greater than the second threshold, the photographing decision logic can be triggered. Among them, the second threshold is also determined according to multiple parameter values, such as the moving speed of the subject, the distance between the subject and the camera module, etc.
[0092] In still other embodiments, the exposure time and the image clarity can be restricted simultaneously to ensure that the final output target image can present a better effect. Combining Figure 11 As shown, after the mobile phone determines that the current scene is a qualified scene, when further determining whether to trigger the photographing decision logic for the image in the current scene, it considers that the image clarity is greater than the second threshold and the exposure time is less than the second threshold. That is to say, when determining that the current scene is a qualified scene, it determines whether the exposure time of the image is less than the second threshold (such as 1 / 60 ms). If the exposure time of the image is less than the second threshold, it further determines whether the image clarity of the image is greater than the second threshold (such as N). If the image clarity is greater than the second threshold, the photographing decision logic can be triggered; if the exposure time is greater than the second threshold or the image clarity is less than the second threshold, it is necessary to re-determine whether to trigger the photographing decision logic.
[0093] It should be noted that in the above example, the exposure time is judged first and then the image clarity. It is also possible to judge the image clarity first and then further judge the exposure time. The order is not uniquely limited. Moreover, the specific values of the above exposure time and image clarity are only examples, and the exposure time and image clarity can also be other values, which are not uniquely limited here.
[0094] Further, continuing to combine Figure 7As shown, after triggering the photo-taking decision logic, the current scene can be automatically captured. Since wonderful moments are often instantaneous, it is difficult to ensure that the frame when the decision logic is triggered is the most wonderful image. Therefore, to further improve the effect of the finally captured image, the highest-quality image can be selected as the target image within a certain range. The most wonderful frame can be selected from the frames, and the most wonderful frame can be used as the target image.
[0095] In some embodiments, before the mobile phone triggers the photo-taking decision logic, the cached image frames can be each frame in the cached preview stream as the image frames. In other embodiments, before triggering the photo-taking decision logic, the mobile phone will automatically cache a certain number of image frames, and continuously update the cached image frames without changing the quantity. That is to say, the number of image frames cached by the mobile phone before waiting to trigger the photo-taking decision logic is fixed and will not keep accumulating. For example, assume that the number of image frames that the mobile phone is set to cache is 10 frames, that is, the cached image frames are the 1st frame image to the 10th frame image. When the 11th frame image appears and the photo-taking decision logic has not been triggered, the cached image frames will be updated at this time, that is, it is the 2nd frame to the 11th frame at this time. Specifically, when triggering the photo-taking decision logic, the mobile phone can determine a certain number of image frames within a range before and after the target frame in the cached image frames. Among them, the target frame can be the frame corresponding to when the photo-taking decision logic is triggered. For example, assume that there are 10 frames of images within 5s (which are frame A image, frame B image, frame C image, frame D image, frame E image, frame F image, frame G image, frame H image, frame I image, and frame J image respectively), and the photo-taking decision logic is triggered at the 5th frame image, that is, the E frame image is the target frame. After that, the image frames within 1.5s before and after the target frame (i.e., the E frame image) when the photo-taking decision logic is triggered can be cached. That is, the finally cached image frames are the image frames corresponding to 1s to 4s, that is, the B frame image to the H frame image. After that, the mobile phone can select the most wonderful frame from the cached image frames. The so-called selection of the most wonderful frame means that when the image parameters meet any one or both of the following second conditions, it is output as the final target image. The second conditions include: the face angle of the image of the subject in the image is within a preset range and the image of the subject in the image is located at a preset position in the image. The restriction on the image of the subject in the image being located at a preset position in the image is because when an image is output, it is necessary to ensure the position of the image of the subject in the image. That is, to determine the relative position of the size of the image and the size of the image of the subject in the image. If the subject size is too large or too small or the position is too off-center, it is not beautiful and cannot be defined as the target image.
[0096] In some examples, after the mobile phone collects the preview stream, the images included in the preview stream are recognized to determine that the current scene is a qualified scene. After that, the mobile phone can determine the photographing decision logic triggered when the conditions are met based on the images in the preview stream in the current scene, and determine a certain number of image frames within the range before and after the target frame corresponding to the triggered decision logic. That is, it can cache the image frames within 1.5 seconds before and after the corresponding frame when the photographing decision logic is triggered. That is, the finally cached image frames are the target frames corresponding to 1s to 4s, that is, B-frame image, C-frame image, D-frame image, E-frame image, F-frame image, G-frame image, and H-frame image. Then, select the wonderful frames from these 7 target frames as the final target images. For example, determine whether the face angle of the subject's image in these 7 target frames is within the preset range, that is, whether the face angle is within [0°, 45°], that is, whether it meets the selection requirements of the wonderful frames. If not, end the process and re-execute the operation of the mobile phone to collect the preview stream for the current scene.
[0097] In other examples, after the mobile phone collects the preview stream, the images included in the preview stream are recognized to determine that the current scene is a qualified scene. After that, the mobile phone can determine the photographing decision logic triggered when the conditions are met based on the images in the preview stream in the current scene, and determine a certain number of target frames within the range before and after the target frame corresponding to the triggered decision logic. Then, select the wonderful frames from these target frames as the final target images. For example, determine whether the image of the subject in these target frames is located at the preset position of the image, that is, determine whether it is at the center position of the image. If it meets the requirement, it is determined as a wonderful frame and output as the target image. If not, end the process and re-execute the operation of the mobile phone to collect the preview stream for the current scene. For example, assume that a 6-inch image is to be output, and the aspect ratio of the 6-inch image is 6 inches * 4 inches. At this time, the preset range can be set so that the ratio of the subject image to the image is 2:3, that is, within the range of 3 inches * 2 inches, the subject image is located in the center position of the image, so as to ensure that the image corresponding to this frame is presented as a wonderful image, that is, the target image.
[0098] In some other examples, the face angle of the image of the subject in the image can be restricted within a preset range and the image of the subject in the image can be located at a preset position in the image simultaneously, so as to ensure that the final output target image can present a better effect. For example, after the mobile phone performs preview stream acquisition, the images included in the preview stream are recognized to determine that the current scene is a qualified scene. Then, the mobile phone can determine, based on the images in the preview stream in the current scene, to trigger the photographing decision logic when it meets the conditions, and determine a certain number of image frames within a range before and after the target frame corresponding to the triggered decision logic, that is, the image frames within 1.5 s before and after the corresponding frame when the photographing decision logic is triggered can be cached. That is, the finally cached image frames are the target frames corresponding to 1 s to 4 s, that is, B frame image, C frame image, D frame image, E frame image, F frame image, G frame image, and H frame image. Then, a wonderful frame is selected from these 7 target frames as the final target image. First, it is judged whether the face angle of the image of the subject in these 7 target frames is within the preset range, that is, whether the face angle is within [0°, 45°]. Suppose that both the A frame image and the B frame image meet the preset range at this time. Further, it is judged whether the subject image is located at the preset position in the image. Suppose that only the A frame image meets the condition among the A frame image and the B frame image, then the A frame image is output as the final target image. If the face angle of the image of the subject in the A frame image and the B frame image is not within the preset range or the subject image in the A frame image and the B frame image is not located at the preset position in the image, it means that no target image is obtained, the selection of the wonderful frame ends, and the mobile phone re-performs preview stream acquisition.
[0099] It should be noted that in the above example, it is first determined that the face angle of the image of the subject is within the preset range and then it is further judged whether the subject image is located at the preset position in the image. It can also be to first judge whether the subject image is located at the preset position in the image and then judge whether the face angle of the image of the subject is within the preset range. The specific order is not limited here.
[0100] In summary, after the terminal device receives the operation of the user to start shooting, if it is determined that the current scene is a human-pet scene or a parent-child scene, the images in the preview stream in the currently collected scene that meet the judgment conditions can be used as the final output target images. Whether shooting is performed using the front camera or the rear camera of the terminal device can meet the usage requirements of the user in the actual scene. Especially when taking a group photo and taking a self-portrait with the front camera, it can also free the user's hands, provide convenience for the user, and improve the user experience.
[0101] Some other embodiments of the present application provide a terminal device, which may include: the above-mentioned display screen, camera, memory, and one or more processors. The display screen, camera, memory, and processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can execute each function or step that the mobile phone executes in the above method embodiments. The structure of the electronic device may refer to Figure 4 the structure of the mobile phone shown.
[0102] Embodiments of the present application also provide a chip system, as Figure 12 shown, the chip system 1200 includes at least one processor 1201 and at least one interface circuit 1202. The processor 1201 and the interface circuit 1202 can be interconnected by a line. For example, the interface circuit 1202 can be used to receive signals from other devices (such as the memory of an electronic device). For another example, the interface circuit 1202 can be used to send signals to other devices (such as the processor 1201). Exemplarily, the interface circuit 1202 can read the instructions stored in the memory and send the instructions to the processor 1201. When the instructions are executed by the processor 1201, the electronic device can execute each step in the above embodiments. Of course, the chip system may also include other discrete devices, and the embodiments of the present application do not make specific limitations on this.
[0103] Embodiments of the present application also provide a computer storage medium, which includes computer instructions. When the computer instructions run on the above-mentioned electronic device, the electronic device is enabled to execute each function or step that the mobile phone executes in the above method embodiments.
[0104] Embodiments of the present application also provide a computer program product. When the computer program product runs on a computer, the computer is enabled to execute each function or step that the mobile phone executes in the above method embodiments.
[0105] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0106] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0107] The units described as separate components may or may not be physically separated. The components displayed as units can be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0108] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0109] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.
[0110] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A shooting method, characterized in that, applied to a terminal device including a camera module, the method comprising: receiving an operation by a user to open a camera application; in response to the operation, collecting a preview stream of the current scene through the camera module; in response to the current scene satisfying a first condition, saving a target image in the preview stream of the collected current scene; wherein, the first condition includes: the current scene includes multiple subjects, and there is a child among the multiple subjects; or the first condition includes: the current scene includes multiple subjects, and there is a person and an animal among the multiple subjects.
2. The method according to claim 1, characterized in that, the method further comprises: determining the number of subjects included in the current scene and whether the current scene includes a person; when the number of subjects is greater than a first threshold and the current scene includes a person, determining whether the current scene includes a child or an animal, the first threshold being an integer greater than 1.
3. The method according to claim 1 or 2, characterized in that, before saving the target image in the preview stream of the collected current scene, the method further comprises: based on the images included in the preview stream, determining that the face angles of each subject among the multiple subjects are within a preset range.
4. The method according to any one of claims 1-3, characterized in that, before saving the target image in the preview stream of the collected current scene, the method further comprises: determining that the exposure time of the images included in the preview stream is less than a second threshold; and / or, determining that the image clarity of the images included in the preview stream is greater than a third threshold.
5. The method according to any one of claims 1-4, characterized in that, the target image is an image whose image parameters in the preview stream satisfy a second condition; the second condition includes one or more of the following conditions: the face angle of the image of the subject in the image is within a preset range; the image of the subject in the image is located at a preset position in the image.
6. The method according to any one of claims 2-5, characterized in that, the method further comprises: processing the preview stream through a deep learning network to output image feature information; the image feature information is used to indicate the number of subjects and the subject types included in the current scene, and the subject types are used to indicate whether the subjects included in the current scene are children or animals; wherein, whether the current scene includes a person is determined based on the subject types, and whether the current scene includes a child or an animal is determined based on the subject types.
7. The method according to any one of claims 1-6, characterized in that, the camera module includes a front camera or a rear camera of the terminal device.
8. A terminal device, characterized in that, the terminal device includes: a memory, a camera module and one or more processors; the memory, the camera module are coupled to the processor; Among them, the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the processor, the terminal device is caused to execute the method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that it includes computer instructions; when the computer instructions run on a terminal device, the terminal device is caused to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Shooting method and device
CN105635567A
Image shooting method and device
CN105744169A
Method for attracting children to take selfies and mobile terminal
CN107024990A
Snapshot method and mobile terminal
CN107566736A
Picture processing method, picture processing device, and terminal device
CN108898587A