Shooting method and electronic device based on wonderful moment recognition
By using a multimodal large language model on electronic devices combined with image and audio to recognize exciting moments, it solves the problem that users find it difficult to capture exciting moments in dynamic scenes, and improves the shooting experience and browsing experience.
Patent Information
- Application Number
- CN202411857644.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-17
AI Technical Summary
In scenes where moving subjects are photographed, it is difficult for users to capture photos of exciting moments, resulting in poor shooting experience.
By applying a multimodal large language model (MLLM) on electronic devices, combining images collected by the camera and audio collected by the microphone, identifying and capturing exciting moments in dynamic shooting scenes.
It improves the success rate of users to capture exciting moments in dynamic scenes, improves shooting experience and browsing experience, and avoids redundant image saving.
Smart Images

Figure CN119421042B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminals, and in particular to a shooting method and electronic equipment based on wonderful moment recognition. Background Art
[0002] When shooting sports objects, such as high jump, long jump, tennis, fireworks, etc., users often have difficulty taking wonderful photos of wonderful moments due to human reaction delay, which greatly affects the user's shooting experience. Summary of the invention
[0003] The present application provides a shooting method and electronic device based on wonderful moment recognition.
[0004] In the first aspect, the present application provides a method for shooting. The method is applied to electronic devices with shooting capabilities such as mobile phones and cameras. The method includes: displaying a preview interface, in which an image captured by a camera is displayed; identifying a first key image from the image captured by the camera, the first key image is an image containing a starting posture of a first action; observing M frames of images captured by the camera backward from the moment when the first key image is identified as the starting point; in the case where a first shooting operation is detected and a first wonderful image is identified, in response to the first shooting operation, saving the first wonderful image, the first shooting operation is a shooting operation that occurs during the observation period, and the first wonderful image is an image containing the best posture of the first action in the M frames. In the case where the first shooting operation is detected but the first wonderful image is not identified, in response to the first shooting operation, saving the first image, the first image is the image captured by the camera at the moment when the first shooting operation occurs.
[0005] By implementing the method provided in the first aspect, when a user's shooting operation occurs near a wonderful image recognized by an electronic device, the electronic device can directly save the above-mentioned wonderful image as the shooting result of the above-mentioned user's shooting operation, instead of separately saving the image at the moment of the user's shooting operation and the recognized wonderful image, thereby avoiding redundancy and improving the user's shooting experience and browsing experience.
[0006] In some embodiments, after saving the first wonderful image, the method further includes: displaying a thumbnail of the first wonderful image on a preview interface; after saving the first image, the method further includes: displaying a thumbnail of the first image on a preview interface. In this way, the user can understand from the thumbnail that the electronic device has responded to the user's shooting operation, and the user can instantly view the shooting result through the thumbnail.
[0007] In some embodiments, the method further includes: detecting a second shooting operation after the observation is completed, and saving a second image captured by the camera at the time when the second shooting operation occurs in response to the second shooting operation. In this way, the time interval between the user's shooting operation and the saved image will not be too long, which can reduce the risk of deviating from the user's shooting intention.
[0008] In some embodiments, the method further includes: when the first wonderful image is identified but the first shooting operation is not detected, saving the first wonderful image. In this way, when there is no user shooting operation, the electronic device can automatically capture the wonderful image for the user to browse.
[0009] It is understandable that in other embodiments, the method further includes: when the first wonderful image is identified but the first shooting operation is not detected, the first wonderful image is discarded. In this way, when there is no user shooting operation, the electronic device will not automatically capture wonderful images, thereby avoiding automatically capturing too many images, occupying memory, and affecting the user's shooting experience.
[0010] In some embodiments, identifying the first key image from the image captured by the camera specifically includes: identifying the first key image from the image captured by the camera based on the image captured by the camera and the audio captured by the microphone. That is, the electronic device can identify the preset wonderful moments by combining the image and the audio, improve the recognition accuracy, and thus improve the user's shooting experience.
[0011] Preferably, the electronic device is pre-installed with a multimodal large language model MLLM. The electronic device processes the image captured by the camera and the audio collected by the microphone through the MLLM to determine the first key image and / or the first wonderful image.
[0012] In a specific implementation, MLLM includes a video encoder, an audio encoder and a cross encoder. The video encoder is used to process images captured by a camera to obtain image features, the audio encoder is used to process audio captured by a microphone to obtain audio features, and the cross encoder is used to process image features and audio features to obtain comprehensive features. MLLM determines whether the image corresponding to the comprehensive feature is a key image containing a starting posture of a preset action or a wonderful image containing an optimal posture of a preset action through the degree of matching between the comprehensive feature and preset description information.
[0013] In the method provided in the first aspect, the first action is a preset typical action with a wonderful moment, such as high jump, long jump, ball sports (such as tennis, badminton, etc.), fireworks, etc., and the embodiments of the present application do not specifically limit this. Among them, when the first action is ball sports and / or fireworks, the wonderful recognition effect based on images and audio provided by MLLM is particularly prominent.
[0014] In a second aspect, the present application provides an electronic device, comprising one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer programs, and when the one or more processors execute the computer programs, the electronic device executes the method described in the first aspect and any possible implementation method of the first aspect.
[0015] In a third aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, which are used to call computer instructions so that the electronic device executes the method described in the first aspect and any possible implementation method of the first aspect.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium, including a computer program. When the above-mentioned computer program runs on an electronic device, the above-mentioned electronic device executes the method described in the first aspect and any possible implementation method of the first aspect.
[0017] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect and any possible implementation of the first aspect.
[0018] It can be understood that the electronic device provided in the second aspect, the chip system provided in the third aspect, the computer storage medium provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to execute the method provided in the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of long jump provided in an embodiment of the present application;
[0020] Figure 2A It is a schematic diagram of the structure of the MLLM provided in the embodiment of the present application;
[0021] Figure 2B is a schematic diagram of image sequence segmentation provided by an embodiment of the present application;
[0022] Figure 2C is another image sequence segmentation schematic diagram provided in an embodiment of the present application;
[0023] Figure 3 is a flow chart of a shooting method based on wonderful moment recognition provided by an embodiment of the present application;
[0024] Figure 4A-4GA user interface of a shooting method based on wonderful moment recognition provided by an embodiment of the present application;
[0025] Figure 5 is a flow chart of another shooting method based on wonderful moment recognition provided by an embodiment of the present application;
[0026] Figure 6A-6B is a schematic diagram of a group of wonderful images saved during the observation period provided by an embodiment of the present application;
[0027] Figure 7 is a schematic diagram of saving an image at the time when a shooting operation occurs provided by an embodiment of the present application;
[0028] Figure 8 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0030] Figure 1 This is a schematic diagram of long jump provided in an embodiment of the present application.
[0031] like Figure 1 As shown, during the long jump process, the athlete will present the following five states in sequence: preparation, take-off, rise, landing, and landing. In general, users most want to capture the athlete in the rising state, that is, to capture the image captured by the camera at time T3. The above-mentioned rising state moment is the wonderful moment, and the image at the wonderful moment is the wonderful image. However, due to the extremely short duration of the long jump action and the problem of delayed human reaction, it is often difficult for users to capture the wonderful images of the wonderful moments. For example, it is more likely that the electronic device will detect the shooting operation made by the user before time T3, or detect the shooting operation made by the user after time T3. Before time T3 or after time T3, the user cannot capture the wonderful image of time T3.
[0032] It is understandable that the above-mentioned "rising state moment is a wonderful moment" is only an example. In other examples, the wonderful moment of the long jump action can also be the take-off state moment, the descending state moment or other state moment, and the embodiments of the present application are not limited to this. In other dynamic shooting scenes, such as high jump, tennis, fireworks and other scenes, due to the short duration of the action and the delay of human reaction, it is also difficult for users to capture wonderful images of wonderful moments. Among them, for example, the wonderful moment of the high jump scene is such as the moment of jumping to the highest point, the wonderful moment of the tennis scene is such as the moment of tennis hitting the racket, and the wonderful moment of the fireworks scene is such as the moment when the fireworks bloom to the maximum size.
[0033] In view of this, an embodiment of the present application provides a shooting method based on wonderful moment recognition.
[0034] The method can be applied to electronic devices with shooting capabilities, such as mobile phones and cameras. Among them, cameras include but are not limited to SLR cameras, mirrorless cameras, and sports cameras. Not limited to mobile phones and cameras, the electronic devices implementing the above method can also be electronic devices with shooting capabilities, such as laptops, smart wearable devices (such as smart watches), vehicle-mounted devices, smart home devices and / or smart city devices. The embodiments of this application do not impose special restrictions on this.
[0035] A multimodal large language model (MLLM) can be set in electronic devices with shooting capabilities such as mobile phones and cameras. Electronic devices can use the above MLLM to identify preset dynamic shooting scenes and determine the starting time, highlight time and end time of the dynamic shooting scene.
[0036] Figure 2A It is a schematic diagram of the structure of the MLLM provided in the embodiment of the present application.
[0037] like Figure 2A As shown, MLLM includes a video encoder and an audio encoder. The video encoder is used to process images captured by the camera and output corresponding image features. The audio encoder is used to process audio captured by the microphone and output corresponding audio features. Among them, N frames of images are an image sequence, N≥1, preferably, N=3. The video encoder processes one image sequence at a time. Correspondingly, the audio encoder processes the audio clips corresponding to one image sequence at a time.
[0038] The camera will periodically collect and generate images. Taking the frame rate of 30fps as an example, the camera will generate a frame of image every 33.3ms. Figure 2B Schematic diagram of image sequence segmentation provided by the embodiment of the present application. Figure 2B As shown, after receiving the i-th frame image F(i) captured by the camera, the electronic device can determine the image sequence corresponding to the current frame F(i): F(i-2), F(i-1), F(i), and the audio clip corresponding to the above image sequence. The electronic device can input the above image sequence and the corresponding audio clip into MLLM, and use the video encoder and audio encoder provided by MLLM to process the above image sequence and audio clip respectively to obtain image features and audio features.
[0039] Figure 2C FIG. 1 is another image sequence segmentation schematic diagram provided in an embodiment of the present application. Figure 2CAs shown, after receiving the i-th frame image F(i) captured by the camera, the electronic device can determine the image sequence corresponding to the current frame F(i): F(i-1), F(i), F(i+1), and the audio segment corresponding to the above image sequence. Among them, F(i+1) is the next frame image that the camera will output and is not yet available currently. Therefore, in Figure 2C In the image sequence segmentation method shown, the electronic device needs to wait for the camera to output F(i+1), and then input the image sequence and audio segment corresponding to F(i) into the MLLM to obtain the corresponding image features and audio features.
[0040] It can be understood that in a shooting scenario with a frame rate of 30fps or higher, the similarity between two consecutive frames is extremely high. Therefore, preferably, the electronic device can determine the image sequence corresponding to the current frame F(i) by frame sampling to reduce the algorithm running frequency and save power consumption. Taking the frame interval = 3 as an example, exemplarily, when F(i-1) is the 16th frame image captured by the camera, F(i) is the 19th frame image captured by the camera, and F(i+1) is the 22nd frame image captured by the camera.
[0041] The MLLM also includes a cross-encoder and a decoder. Based on the image features and audio features, the MLLM can fuse the image features and audio features through the cross-encoder, and then restore them through the decoder to obtain the comprehensive features describing the current image sequence and audio segment.
[0042] The trained MLLM records various reference information, such as event description information and wonderful moment description information.
[0043] Among them, the event description information is used to describe preset dynamic shooting scenarios, such as the long jump, high jump, tennis, fireworks blooming and other scenarios exemplified above. The event description information includes the image features, audio features and / or text features of the preset dynamic shooting scenario. The MLLM can determine whether the current shooting scenario is a preset dynamic shooting scenario based on the matching degree between the comprehensive features and the event description information, and determine whether the current frame is the starting frame or the ending frame of the above dynamic shooting scenario. The above starting frame is also called the starting image, and the ending frame is also called the ending image.
[0044] The wonderful moment description information is used to describe the wonderful moments of the preset dynamic shooting scenario. Similarly, the wonderful moment description information also includes image features, audio features and / or text features. The MLLM can determine whether the current frame is a wonderful frame (i.e., wonderful image) in the preset dynamic shooting scenario based on the matching degree between the comprehensive features and the wonderful moment description information.
[0045] Figure 3 It is a flowchart of a shooting method based on wonderful moment recognition provided by an embodiment of the present application.
[0046] S301, start the camera to capture images, and display the images captured by the camera in the photo preview interface.
[0047] The electronic device can detect a user operation of starting the camera, such as clicking a camera application icon to enter the camera application. In response to the above operation, the electronic device can start the camera to capture images and display a photo preview interface in which the images captured by the camera are displayed.
[0048] S302. Enable MLLM.
[0049] The electronic device may display specific controls in the photo preview interface. The above controls are also called MLLM controls. After detecting a user operation acting on the MLLM control, the electronic device may enable MLLM, use the scene recognition function provided by MLLM, identify the preset dynamic shooting scene, and determine the wonderful moments in the dynamic shooting scene. After enabling MLLM, the electronic device may execute S303-S305.
[0050] Optionally, the electronic device may not set the MLLM control. By default, after starting the camera, the electronic device may start the MLLM and execute S303-S305.
[0051] S303: Input the image captured by the camera into the MLLM.
[0052] S304, start the microphone to collect audio, and input the audio collected by the microphone into the MLLM.
[0053] refer to Figure 2A According to the introduction, after receiving the images captured by the camera and the audio collected by the microphone, based on the above images and audio, MLLM can identify the preset dynamic shooting scenes and determine the wonderful moments of the dynamic shooting scenes, that is, determine the wonderful images containing wonderful moments in the dynamic shooting scenes.
[0054] S305, capturing the wonderful images output by MLLM.
[0055] After receiving the wonderful image output by the MLLM, the electronic device can automatically capture the wonderful image. The electronic device can obtain the original (RAW) image of the frame image and write it into the memory for storage.
[0056] After the automatic capture, the electronic device can display a thumbnail of the captured image in the photo preview interface. The user can understand that the electronic device has performed an automatic capture operation based on the thumbnail. Therefore, the user can view the image captured by the electronic device through the thumbnail.
[0057] Figure 4A-4GA user interface of a shooting method based on wonderful moment recognition provided in an embodiment of the present application.
[0058] Figure 4A Schematic diagram of the photo preview interface provided by the embodiment of the present application. Figure 4A As shown, the electronic device can display an MLM open control 411 in the photo preview interface and display the image captured by the camera in the preview window 412.
[0059] refer to Figure 4A , the electronic device can detect the user operation acting on the MLLM opening control 411. In response to the above operation, the electronic device can enable MLLM, use the scene recognition function provided by MLLM, identify the preset dynamic shooting scene, and determine the wonderful moments in the dynamic shooting scene. Figure 4B After MLLM is enabled, the electronic device may display an MLLM off control 413 to replace the original MLLM on control 411. The MLLM off control 413 may be used to turn off the MLLM.
[0060] refer to Figure 4A-4F In the long jump scene, the preview window 412 can sequentially display the athlete images in different states (such as preparation, take-off, rise, landing, etc.) captured by the camera. Through the above image sequence and the audio clips at the corresponding time, MLLM can determine that the long jump scene is recognized and determine Figure 4D The image M1 shown in the preview window 412 is a wonderful photo in the current scene. Therefore, the electronic device can obtain the RAW image of M1 and write it into the memory for storage, that is, automatically capture M1. Figure 4F After automatically capturing the image M1, the electronic device may display a thumbnail of the image M1 on the review control 414 to prompt the user that the electronic device has automatically captured the image. Figure 4F-4G After detecting a user operation on the review control 414, the electronic device may display Figure 4G The gallery preview interface shown displays images automatically captured by the electronic device for users to browse.
[0061] Similarly, in other dynamic shooting scenes (such as high jump, tennis, fireworks, etc.), based on the scene recognition function and wonderful moment positioning function provided by MLLM, electronic devices can identify the above dynamic shooting scenes, determine the wonderful moments in the scenes, and then automatically capture wonderful photos of the wonderful moments, satisfying the user's demand for shooting wonderful moments and giving users a better shooting experience. Compared with the existing method of simply determining wonderful moments through image recognition, Figure 3 The method for determining wonderful moments through images and audio is more accurate, and users can get a better shooting experience.
[0062] Figure 5 This is a flowchart of another shooting method based on wonderful moment recognition provided in an embodiment of the present application.
[0063] S501, start the camera to collect images, and display the images collected by the camera in the preview interface.
[0064] S502: Input the image captured by the camera and the audio captured by the microphone into the MLLM.
[0065] S503: Identify the starting image.
[0066] refer to Figure 2A According to the introduction, after receiving the image captured by the camera and the audio captured by the microphone, based on the above image and audio, MLLM can identify the preset dynamic shooting scene and determine the starting moment in the dynamic shooting scene. When the comprehensive features corresponding to the current frame match the starting event description information in the event description information, MLLM can determine that the current frame is the starting frame of the preset dynamic shooting scene, that is, it recognizes the starting image.
[0067] S504: Observe backward the images captured by the M-frame camera.
[0068] After identifying the starting image, the electronic device may execute S504 to observe the images captured by the M-frame camera backward. An observation time that is too long or too short will affect the implementation effect of the present method. Therefore, preferably, 0.5s≤observation time≤2s. In a shooting scene with a frame rate of 30fps, 0.5s corresponds to an image captured by a 15-frame camera, and 2s corresponds to an image captured by a 60-frame camera. Therefore, preferably, 15≤M≤60.
[0069] During the observation, the electronic device can determine the wonderful images in the M frames according to the output of the MLLM. On the other hand, during the observation, the electronic device can also detect the user's shooting operation, such as the user clicking the shutter control.
[0070] S505: A user shooting operation is detected during the observation period.
[0071] After detecting the user's shooting operation, the electronic device does not respond to the shooting operation before the observation ends. After the observation ends, the electronic device executes S506 to respond to the shooting operation. If the user's shooting operation is not detected during the observation, the starting image identified in S503 is discarded, and the next starting image is waited for, and S504-S505 are repeated.
[0072] S506. The images during the observation period include wonderful images.
[0073] After detecting the user's shooting operation and the observation is finished, the electronic device can determine whether the images during the observation include wonderful images. When a wonderful image is identified, the electronic device executes S507; otherwise, the electronic device executes S508.
[0074] S507 . In response to the user's shooting operation, save the wonderful images during the observation period.
[0075] Figure 6A-6B It is a schematic diagram of a group of wonderful images saved during the observation period provided in an embodiment of the present application.
[0076] refer to Fig. 6A For example, at time T11, the electronic device can determine that F(i) is the starting frame (first key image) of a preset dynamic shooting scene based on the image sequence corresponding to F(i). Then, the electronic device can start a timer or a frame counter to observe M frames of images backward. Among them, T11~T14 is the observation interval, and time T14 is the end of observation. At time T12, the electronic device can detect the user's shooting operation Capture1 (the first shooting operation). Before ending the observation (ie, time T14), the electronic device can determine that the image F(j) at time T13 is a wonderful image (the first wonderful image). Fig. 6A As shown, T12 may be before T13. In some embodiments, as Figure 6B As shown, T12 may also be after T13. Of course, T12 may also be equal to T13. At this time, the image captured by the camera at the time of the shooting operation and the wonderful image recognized by the MLLM are the same frame image.
[0077] When the observation ends, in response to Capture 1, the electronic device can save F(j) as the shooting result of Capture 1. Similarly, after saving F(j) as the shooting result of Capture 1, the electronic device can display a thumbnail of F(j) in the photo preview interface for the user to browse the shooting results.
[0078] S508: In response to the user's shooting operation, save the image captured by the camera when the shooting operation occurs.
[0079] Figure 7 It is a schematic diagram of saving an image at the moment when a shooting operation occurs provided by an embodiment of the present application.
[0080] like Figure 7As shown, at time T12, the electronic device can detect the user's shooting operation Capture1. However, in this scenario, when the observation ends, the electronic device does not recognize any wonderful images. Therefore, in response to Capture1, the electronic device can save the image F(k) (first image) captured by the camera at the time Capture1 occurs as the shooting result of Capture1. Similarly, after saving F(k) as the shooting result of Capture1, the electronic device can display a thumbnail of F(k) in the photo preview interface for the user to browse the shooting results.
[0081] After the observation is finished and before the next observation begins, the electronic device may detect the user's shooting operation Capture2 (second shooting operation). At this time, preferably, in response to Capture2, the electronic device saves the image (second image) captured by the camera at the time when Capture2 occurs without performing any waiting operation.
[0082] In some embodiments, during the observation interval, the electronic device may not detect any user shooting operation, but the MLLM recognizes the wonderful image. At this time, in some embodiments, after the observation is completed, the electronic device can perform an automatic capture operation to save the above-mentioned wonderful image. In other embodiments, after the observation is completed, the electronic device can also discard the above-mentioned wonderful image and not perform automatic capture to avoid occupying the memory and affecting the user's shooting experience.
[0083] Implementation Figure 5 According to the method shown, when a user's shooting operation occurs near a wonderful image recognized by MLLM, the electronic device can directly save the wonderful image as the shooting result of the user's shooting operation, instead of separately saving the image at the moment of the user's shooting operation and the recognized wonderful image, thereby avoiding redundancy and improving the user's shooting experience and browsing experience.
[0084] Figure 8 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0085] The electronic device may include a processor 81, an internal memory 82, an external memory interface 83, a display screen 84, a camera 85, a communication module 86, an audio module 87, a sensor module 88, and the like.
[0086] The electronic device may include one or more processors 81. The processor 81 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated into one or more processors. The controller may generate an operation control signal according to the instruction opcode and the timing signal to complete the control of fetching and executing instructions.
[0087] The processor 81 is connected to one or more internal memories 82. Among them, the one or more internal memories 82 include random access memory (RAM) and non-volatile memory (NVM). RAM can be directly read and written by the processor 81, and can be used to store executable programs (such as machine instructions) of the operating system or other running programs, and can also be used to store user and application data, etc. NVM can also store executable programs and store user and application data, etc. The executable programs and data stored in NVM can be loaded into RAM in advance for direct reading and writing by the processor 81. A storage unit can also be set in the processor 81, and the storage unit can be a high-speed cache storage unit, which can be used to save instructions or data that the processor 81 has just used or circulated.
[0088] In the embodiment of the present application, the computer program code of the shooting method based on the wonderful moment recognition can be stored in the NVM. After starting the camera, the electronic device can load the computer program code of the shooting method stored in the NVM into the processor 81 for execution, thereby realizing Figure 3 or Figure 5 The shooting method based on wonderful moment recognition shown provides users with a wonderful moment capture function for dynamic shooting scenes.
[0089] Optionally, in some electronic devices, the electronic device may also be provided with an external memory interface 83. The external memory interface 83 may be used to connect to an external NVM to expand the storage capacity of the electronic device. The computer program code of the above-mentioned shooting method based on wonderful moment recognition may also be stored in an external NVM connected through the external memory interface 83.
[0090] The display screen 84 is used for display. The display screen 84 includes a display panel. The display panel can be a liquid crystal display (LCD). The display panel can also be made of an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), miniled, microled, micro-oled, a quantum dot light emitting diode (QLED), etc. The electronic device may include one or more display screens 84.
[0091] The camera 85 is used to capture images. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. The electronic device may include one or more cameras 85.
[0092] The electronic device can display the image capture and display capabilities provided by the ISP, the camera 85, the video codec, the GPU, the display screen 84, and the application processor. Figure 4A-4G The user interface shown.
[0093] The audio module 87 includes a microphone 87A and a speaker 87B. The microphone 87A is also called a "microphone" or a "microphone" and is used to convert a sound signal into an electrical signal. The speaker 87B is also called a "speaker" and is used to convert an audio electrical signal into a sound signal.
[0094] In an embodiment of the present application, after MLLM is enabled, the electronic device can collect sound signals through microphone 87A to obtain audio clips for identifying preset dynamic shooting scenes.
[0095] Optionally, the electronic device may include a communication module 86. The communication module 86 includes a wireless communication module and a mobile communication module. The wireless communication module can be used to provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), Bluetooth low energy (BLE), global navigation satellite system (GNSS), frequency modulation (FM), nearfield communication technology (NFC), infrared technology (IR) for use in electronic devices. The mobile communication module is used to provide mobile communication solutions including 2G / 3G / 4G / 5G, etc. for use in electronic devices. The electronic device can implement wireless communication functions through the communication module 86, such as sharing the wonderful images captured with other electronic devices.
[0096] The sensor module 88 includes a touch sensor 88A. The touch sensor 88A is also called a "touch control device". The touch sensor 88A can be set on the display screen 84, and the touch sensor 88A and the display screen 84 form a touch screen, also called a "touch control screen". The touch sensor 88A is used to detect a touch operation on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the type of touch event, such as Figure 4A-4G Based on the detected touch operation, the electronic device can provide visual output related to the touch operation through the display screen 84.
[0097] Not limited to the touch sensor 88A, the sensor module 88 may also include other sensors, such as pressure sensors, gyroscope sensors, air pressure sensors, magnetic sensors, acceleration sensors, distance sensors, proximity light sensors, fingerprint sensors, temperature sensors, ambient light sensors, bone conduction sensors, etc., so that the electronic device can achieve richer perception capabilities.
[0098] The processor 81, the internal memory 82, the external memory interface 83, the display screen 84, the camera 85, the communication module 86, the audio module 87 and the sensor module 88 are connected through a bus, and communicate and exchange data based on the above bus. The above bus includes but is not limited to an inter-integrated circuit (I2C) bus, an inter-integrated circuit sound (I2S) bus, a pulse code modulation (PCM) bus, a universal asynchronous receiver / transmitter (UART) bus, a mobile industry processor interface (MIPI) bus, a general-purpose input / output (GPIO) interface bus and / or a universal serial bus (USB) interface bus.
[0099] It is understandable that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device. Optionally, the electronic device may also include more components, such as buttons, motors, indicators, and a subscriber identification module (SIM) card interface. The embodiment of the present application does not limit this.
[0100] The term "user interface (UI)" in the specification, claims and drawings of this application refers to the medium interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface of an application is source code written in a specific computer language such as Java and extensible markup language (XML). The interface source code is parsed and rendered on the terminal device, and finally presented as content that the user can recognize, such as pictures, text, buttons and other controls. A control, also called a widget, is a basic element of a user interface. Typical controls include a toolbar, menu bar, text box, button, scrollbar, picture and text. The properties and contents of controls in the interface are defined by tags or nodes, such as XML through <textview> 、 <imgview> 、 <videoview>The controls contained in the interface are specified by nodes such as HTML, CSS, and JavaScript. A node corresponds to a control or attribute in the interface. After parsing and rendering, the node is presented as user-visible content. In addition, many applications, such as hybrid applications, usually also contain web pages in their interfaces. A web page, also known as a page, can be understood as a special control embedded in the application interface. A web page is source code written in a specific computer language, such as hypertext markup language (HTML), cascading style sheets (CSS), JavaScript (JS), etc. The web page source code can be loaded and displayed as user-recognizable content by a browser or a web page display component with similar functions to a browser. The specific content contained in a web page is also defined by tags or nodes in the web page source code. For example, HTML is defined by 、 、 <video> 、 <canvas>To define the elements and attributes of a web page.
[0101] The most common form of user interface is graphical user interface (GUI), which refers to a user interface related to computer operation that is displayed in a graphical manner. It can be an icon, window, control or other interface element displayed on the display screen of an electronic device, where a control can include icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets and other visual interface elements.
[0102] As used in the specification and appended claims of the present application, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items. As used in the above embodiments, the term "when..." may be interpreted to mean "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrase "when determining..." or "if (stated condition or event) is detected" may be interpreted to mean "if determining..." or "in response to determining..." or "when (stated condition or event) is detected" or "in response to detecting (stated condition or event)", depending on the context.
[0103] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk), etc.
[0104] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.< / canvas> < / video> < / videoview> < / imgview> < / textview>
Claims
1. A shooting method, applied to an electronic device, characterized in that: The method comprises: Display a preview interface, wherein the preview interface displays an image captured by the camera; Identifying a first key image from an image captured by a camera, wherein the first key image is an image including a starting posture of a first action; Taking the moment when the first key image is identified as the starting point, observe M frames of images captured by the camera backwards; In a case where a first shooting operation is detected and a first wonderful image is identified, in response to the first shooting operation, the first wonderful image is saved, the first shooting operation is a shooting operation occurring during an observation period, and the first wonderful image is an image in the M frames of images that includes the best posture of the first action; In a case where a first shooting operation is detected but a first wonderful image is not identified, in response to the first shooting operation, a first image is saved, where the first image is an image captured by the camera when the first shooting operation occurs.
2. The method according to claim 1, characterized in that After saving the first wonderful image, the method further includes: displaying a thumbnail of the first wonderful image on the preview interface; after saving the first image, the method further includes: displaying a thumbnail of the first image on the preview interface.
3. The method according to claim 1, characterized in that The method further comprises: A second shooting operation is detected after the observation is completed, and in response to the second shooting operation, a second image captured by the camera when the second shooting operation occurs is saved.
4. The method according to claim 1, characterized in that: The method further comprises: In a case where the first wonderful image is identified but the first shooting operation is not detected, the first wonderful image is saved.
5. The method according to claim 1, characterized in that The identifying the first key image from the image captured by the camera specifically includes: identifying the first key image from the image captured by the camera according to the image captured by the camera and the audio captured by the microphone.
6. The method according to claim 5, characterized in that The electronic device is pre-installed with a multimodal large language model MLLM, and the MLLM is used to process the image captured by the camera and the audio captured by the microphone to determine the first key image.
7. The method according to claim 6, characterized in that The electronic device recognizes the first wonderful image through the MLLM.
8. The method according to claim 7, characterized in that The MLLM includes a video encoder, an audio encoder and a cross encoder, wherein the video encoder is used to process images collected by a camera to obtain image features, the audio encoder is used to process audio collected by a microphone to obtain audio features, and the cross encoder is used to process the image features and the audio features to obtain comprehensive features; The MLLM determines whether the image corresponding to the comprehensive feature is a key image containing a starting posture of a preset action or a wonderful image containing an optimal posture of a preset action through the matching degree between the comprehensive feature and the preset description information.
9. The method according to claim 1, characterized in that: The first action includes ball games and / or fireworks.
10. An electronic device, characterized in that: The method comprises one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store a computer program, and when the one or more processors execute the computer program, the method according to any one of claims 1 to 9 is executed.
11. A chip system, the chip system is applied to electronic equipment, the chip system comprises one or more processors, characterized in that: The processor is configured to call computer instructions so that the electronic device executes the method according to any one of claims 1 to 9.
12. A computer program product comprising instructions, characterized in that When the computer program product runs on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 9.
13. A computer-readable storage medium comprising a computer program, characterized in that: When the computer program is executed on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Photographing method and electronic equipment
CN118555470A