Photographing method and device, electronic equipment and medium

By automatically identifying and capturing photos containing target information and objects during video recording, and integrating text information into the photos, the system solves the tedious problem of taking photos while recording videos, achieving a convenient photo-taking experience and efficient photo processing.

CN121509799APending Publication Date: 2026-02-10VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511931456.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

When recording videos, taking photos is cumbersome and complicated for users, making it difficult to capture wonderful moments.

Method used

Once the multimedia data meets the preset conditions, the system automatically takes photos during the recording process, uses AI technology to identify target information and objects, and integrates relevant text information into the photos.

Benefits of technology

It reduces the complexity of user operations, ensures automatic capture of wonderful moments during video recording, simplifies photo sharing and editing, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509799A_ABST
    Figure CN121509799A_ABST
Patent Text Reader

Abstract

The invention discloses a photographing method and device, electronic equipment and a medium, and belongs to the field of electronic equipment. The photographing method comprises the following steps: when it is detected that multimedia data currently recorded by the electronic equipment meets a preset condition, photographing while keeping video recording to obtain a first photo; wherein the preset condition comprises at least one of the following conditions: audio data in the multimedia data is related to target information; a video image in the multimedia data comprises a target object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of electronic devices, and particularly relates to a photographing method and device, electronic device and medium. BACKGROUND

[0002] With the continuous development of electronic devices, simple functions can no longer meet the needs of consumers. Consumers need more interesting, customized, innovative and easy-to-use new functions of smart mobile image devices. In some scenarios, such as concert stage activities, users not only need to record videos, but also often have the need to take photos during video recording in order to capture exciting moments. SUMMARY

[0003] The purpose of the embodiments of the present application is to provide a photographing method, device, electronic device and medium, which can solve the problem that the current photographing process is complicated when recording a video.

[0004] In a first aspect, the embodiments of the present application provide a photographing method, which comprises:

[0005] In the case where it is detected that multimedia data currently recorded by the electronic device meets a preset condition, a photograph is taken while the video recording is maintained to obtain a first photo; wherein the preset condition comprises at least one of the following:

[0006] The audio data in the multimedia data is related to target information;

[0007] The video image in the multimedia data contains a target object.

[0008] In a second aspect, the embodiments of the present application provide a photographing device, which comprises:

[0009] A photographing module, configured to take a photograph while the video recording is maintained to obtain a first photo in the case where it is detected that multimedia data currently recorded by the electronic device meets a preset condition;

[0010] Wherein the preset condition comprises at least one of the following:

[0011] The audio data in the multimedia data is related to target information;

[0012] The video image in the multimedia data contains a target object.

[0013] In a third aspect, the embodiments of the present application provide an electronic device, which comprises a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the method according to the first aspect.

[0014] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0015] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0016] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0017] In this embodiment of the application, during the recording process of the electronic device, if it is detected that the audio data in the multimedia data currently being recorded by the electronic device is related to the target information, and / or the time-frequency image in the multimedia data contains the target object, a first photo can be automatically taken while recording is being continued. In this way, the user can take photos without any operation during the recording process, reducing the complexity of user operation. Attached Figure Description

[0018] Figure 1 This is one of the flowcharts illustrating the photographing method according to an embodiment of this application;

[0019] Figure 2 This is a schematic diagram illustrating the activation of the intelligent photography function in an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of the first photograph and the first text information fused together in an embodiment of this application;

[0021] Figure 4 This is a schematic diagram illustrating the viewing of photos during video recording in an embodiment of this application;

[0022] Figure 5 This is one of the schematic diagrams illustrating the transition between viewing audio / video and photos in the embodiments of this application;

[0023] Figure 6 This is the second schematic diagram illustrating the transition between viewing audio / video and photos in the embodiments of this application;

[0024] Figure 7 This is a schematic diagram of the overall architecture for implementing the photographing method of the embodiments of this application;

[0025] Figure 8 It is to utilize Figure 7 The diagram shown illustrates the data processing flow of the overall architecture for implementing the image capture method.

[0026] Figure 9 This is a second schematic flowchart of the photographing method according to an embodiment of this application;

[0027] Figure 10 This is a schematic diagram of the structure of the photographing device according to an embodiment of this application;

[0028] Figure 11 This is one of the structural schematic diagrams of the electronic device according to an embodiment of this application;

[0029] Figure 12 This is a second schematic diagram of the structure of the electronic device according to an embodiment of this application. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0031] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0032] The photographing method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0033] This application provides a method for taking a picture, which can be executed by an electronic device such as a mobile phone, tablet computer, laptop computer, PDA, or in-vehicle electronic device. Some embodiments of this application use an electronic device as the execution subject to illustrate the picture-taking method provided by this application. Figure 1 As shown, the method includes:

[0034] Step 101: If the multimedia data currently being recorded by the electronic device meets preset conditions, a photo is taken while the video recording continues, resulting in a first photo; wherein the preset conditions include at least one of the following:

[0035] The audio data in the multimedia data is related to the target information;

[0036] The video images in the multimedia data contain the target object.

[0037] It should be noted here that step 101 can be performed when the smart photo-taking function is activated; in other words, the electronic device needs to enable the smart photo-taking function when starting video recording. Figure 2 For example, an electronic device's recording interface has a smart photo-taking switch. Before or during recording, the user can activate / turn on the smart photo-taking function by clicking this switch. The smart photo-taking function can be understood as an automatic photo-taking function that doesn't stop recording.

[0038] For example, the multimedia data recorded by the aforementioned electronic device can be detected by invoking an Artificial Intelligence (AI) audio-visual detection unit. For instance, during recording, the camera application transmits the real-time recorded multimedia data to the AI ​​audio-visual detection unit, which then performs real-time detection on the received multimedia data to determine whether it meets preset conditions. The AI ​​audio-visual detection unit utilizes AI technologies such as deep learning / reinforcement learning to detect the received multimedia data in real time.

[0039] For example, the above-mentioned preset conditions are conditions that characterize wonderful moments; in other words, when a wonderful moment is detected, the electronic device, specifically the camera application of the electronic device, is triggered to continuously record and take pictures so that only wonderful moments can be captured while continuously recording video.

[0040] Here, taking a concert scene as an example, the aforementioned target information may refer to the chorus information of a song; the aforementioned target object may include at least one of the following: the singer's facial expression, the singer's waving gesture, the singer's head tilting back to sing a high note, and stage effects; among them, the singer's facial expression may be a facial expression with a specific level of clarity, and the stage effects may be fireworks, ribbons, etc.

[0041] For example, in the above steps, the first photo can be obtained by performing image optimization processing on the photo after taking a picture while recording video.

[0042] In the photo-taking method of this application embodiment, when the intelligent photo-taking function is enabled, the multimedia data currently being recorded by the electronic device can be detected in real time during the recording process. If the currently recorded multimedia data is detected to be related to the target information and / or includes the target object, a photo can be automatically taken while recording to obtain a first photo. In this way, the user can take photos without any operation during the recording process, reducing the complexity of user operation.

[0043] Furthermore, as an optional implementation, the method also includes:

[0044] If the multimedia data currently recorded by the electronic device meets the preset conditions, first text information matching the audio data in the currently recorded multimedia data is obtained.

[0045] Specifically, the above steps may include: obtaining first text information associated with the audio data, and selecting a font and effects that match the emotional tone of the audio data, such as cheerful, cute, lively, or melancholic. In other words, the first text information includes text content that matches the audio data, and may also include one or more of the following: font, effects, and unique symbols.

[0046] Here, taking a concert scene as an example, the above steps can be achieved by calling the song library, performing intelligent matching, recording the chorus lyrics corresponding to the audio data at this time, and intelligently selecting appropriate fonts and effects based on the emotional tone of the music, such as light, lively or melancholic, to enhance the visual expressiveness of the lyrics.

[0047] Based on the above steps, in step 101, while maintaining the video recording, a photo is taken to obtain the first photo, including:

[0048] Sub-step 1: Perform image recognition on the second photo taken while the video is being recorded, and obtain the image recognition result; for example, the image recognition result may include the image recognition and segmentation result of the second photo.

[0049] Sub-step two: Based on the image recognition result, the first text information is superimposed onto the first target area of ​​the second photo to obtain the first photo. For example, this sub-step may specifically involve finding a first target area in the second photo where the first text information can be superimposed based on the image recognition and segmentation results of the second photo, as well as the image color and brightness of the second photo, to superimpose the first text information onto the first target area, thereby obtaining the first photo. This ensures the harmonious unity of the first text information and the first photo. Note that this sub-step involves superimposing the first text information onto the first target area, rather than rendering the first text information onto the first target area, which facilitates subsequent independent editing of the first text information.

[0050] Taking a concert scene as an example, the first sub-step described above can specifically be: using image recognition technology to identify and segment faces, bodies, stage, etc., in the first image. The second sub-step can specifically be: based on information such as image color and brightness, finding the most suitable target area for placing lyrics related to the first image (i.e., the aforementioned first text information) in the first image; then, superimposing the lyrics onto the first target area to ensure harmony between the lyrics and the image. An example of the fused first image is shown below. Figure 3 As shown.

[0051] It should be noted that the above-mentioned optional implementation method can overlay the first text information onto the second photo by calling the AI ​​image fusion unit. Specifically, for example, the camera application transmits the second photo and the first text information to the AI ​​image fusion unit, which then processes the image to obtain the first photo and sends it back to the camera application. The camera application then stores the second photo in the photo album of the electronic device, making it convenient for the user to quickly share the photo with others later.

[0052] In the above-mentioned optional implementation methods, by superimposing the first text information corresponding to the audio data obtained by intelligent matching onto the second photo to obtain the first photo, on the one hand, it can reduce the editing operations of the photo when the user shares the photo later, realizing quick photo sharing. On the other hand, it adds an explanation or description of the photo, so that when the user views the photo after the recording ends, he can quickly determine the wonderful moment of the first photo based on the first text information. This avoids the need for the user to play back the video frame by frame to match the shooting time of the first photo, simplifies the user operation, and solves the current problems of complex and inconvenient photo sharing operations.

[0053] Furthermore, as an optional implementation, the method also includes:

[0054] The image recognition result and the corresponding first text information are stored in the target location.

[0055] The above-mentioned optional implementation method is to store the image recognition results with matching / correspondence relationships and the first text information in a specific target location. This makes it easier for users to subsequently edit the first text information in the first photo to meet user needs and allow users to create more personalized commemorative photos.

[0056] As another optional implementation, the method further includes:

[0057] The system receives editing operations from the user on a first interface of the electronic device, where the first interface displays the first photo.

[0058] For example, the above steps can specifically be performed when the user views photos stored in the album after the recording ends; that is, when the user views photos taken during the recording, they can perform secondary editing on the currently viewed photo. The editing operation can be an operation on the first text information in the first photo.

[0059] Update the first text information in the first photo.

[0060] For example, the above editing operation may be: manually adjusting at least one of the following of the first text information: text position, text content, font color, font effects; or adding exclusive symbols, etc.

[0061] In the above-mentioned optional implementation methods, when users view the first photo in the album, they can also edit the first text information in the first photo to meet user needs and allow users to create more personalized photos.

[0062] After obtaining the first photograph, as another optional implementation, the method further includes:

[0063] Step A1: Receive the user's first input on the second interface of the electronic device. The main display area of ​​the second interface displays the video recording. In other words, during the recording process, i.e., when the second interface of the electronic device displays the video recording, the user can perform the first input on the second interface. This first input can be understood as an input to view photos taken during the recording process. For example, the first input can be swiping up the bottom area of ​​the second interface or swiping down the top area of ​​the second interface. Alternatively, it can be an input performed on a virtual operation object on the second interface to view photos taken during the recording process, such as a single click, double click, or long press.

[0064] Step A2: In response to the first input, display a thumbnail of at least one third photo in the second target area of ​​the second interface, the third photo including the first photo and / or a photo taken in response to the user's photo-taking operation during video recording; wherein, as Figure 4 As shown, at least one thumbnail of a third photo can be displayed sequentially from left to right within the second target area, according to the chronological order in which the third photos were taken. The second target area can be the area near the bottom of the main display area of ​​the second interface, i.e. Figure 4 The area corresponding to the dashed box.

[0065] Step A3: Receive a second input from the user regarding a first target thumbnail in the second target area; exemplarily, the first target thumbnail is a thumbnail of a third photo that the user expects to view in detail; the second input is, for example, a double-click or swipe up.

[0066] Step A4: In response to the second input, the video recording is zoomed to the third target area of ​​the second interface, and the third photo corresponding to the first target thumbnail is displayed in the main display area of ​​the second interface.

[0067] Here, with Figure 4 For example, the first target thumbnail is the second thumbnail among multiple thumbnails, and the third target area is the upper right corner of the main display area of ​​the second interface.

[0068] It should be noted that the third photo displayed in the main display area supports zooming in, zooming out, and moving in all directions, so that users can view the details of the third photo in detail.

[0069] In other words, the above-mentioned optional implementation allows users to take photos while recording video without needing to access the photo album. Instead, they can swipe up to display a pop-up bar showing thumbnails of the photos taken during the recording, arranged from left to right in chronological order of capture time. To view photo details, users can double-click the first target thumbnail to display a larger version of the photo in the main display area, or swipe up on the first target thumbnail to drag it to the main display area to display the larger version of the photo. During this process, recording will not be interrupted, and the video feed will automatically zoom to the upper right corner. This solves the current problem of having to interrupt recording to view photo details, resulting in missing key moments in photos or videos. An example of viewing photos during video recording is shown below. Figure 4 As shown.

[0070] Based on the above optional implementation methods, as a further optional implementation method, the method also includes:

[0071] Receive a third input from the user on the first target thumbnail; for example, the third input is an input for deleting the first target thumbnail and its corresponding third photo, such as clicking, swiping up, long pressing, etc., but not limited to this.

[0072] In response to the third input, the first target thumbnail and the third photo corresponding to the first target thumbnail are deleted.

[0073] It's important to note that the third input and the second input can be two different inputs. Of course, they can also be the same input. If they are the same input, for example, both the second and third inputs are swiping up, then the input can be determined based on the content currently displayed in the main display area. For instance, if the main display area does not display the third photo corresponding to the first target thumbnail, then the swipe up is the second input, and the third photo corresponding to the first target thumbnail can be displayed in the main display area. If the main display area already displays the third photo corresponding to the first target thumbnail, then the swipe up is the third input, and the first target thumbnail and its corresponding third photo can be deleted.

[0074] In the above-mentioned optional implementation methods, when viewing the photos taken during the recording process, if the user is not satisfied with the photos, they can directly delete the unsatisfactory photos through a simple third-party input, thus avoiding the need for the user to check whether each photo in the massive number of photos stored in the album is retained after the recording ends, simplifying the user's operation complexity.

[0075] As another optional implementation, when the user views the captured video and / or photos in the album after the recording ends, the method further includes:

[0076] Step B1: Receive the user's fourth input on the operation object, which is the operation object in the third interface of the electronic device. The third interface is used to display a third photo or play multimedia data.

[0077] The third photo includes the first photo and / or a photo taken in response to a user's photo-taking operation during the recording process. In other words, the third photo is a photo taken during the recording process, specifically including: a second photo intelligently captured using the method described in the embodiments of this application during the recording process, and / or a photo taken based on a user's photo-taking operation during the recording process.

[0078] For example, the above-mentioned operation object can be Figure 5 The jump label shown can be categorized as follows: when playing audio / video / multimedia data on the third interface, this jump label can also be called a quick view photo label / switch; when displaying a third photo on the third interface, this jump label can also be called a jump to video label / switch. Of course, the above operations can also be performed on labels / text boxes with drop-down menus.

[0079] For example, the fourth input mentioned above could be a single click, double click, long press, or selection.

[0080] Step B2-1, when the multimedia data is played on the third interface, in response to the fourth input, perform any of the following:

[0081] Step B2-11: Display thumbnails of multiple third photos in the fourth target area of ​​the third interface, and switch the multimedia data played on the third interface to the third photo corresponding to the second target thumbnail based on the fifth input received from the user on the second target thumbnail; wherein, the multiple third photos are related to the screen of the multimedia data currently played on the third interface.

[0082] Specifically, step B2-11 above can be implemented as follows:

[0083] First, in the fourth target area of ​​the third interface, thumbnails of multiple third photos related to the multimedia data currently being played on the third interface are displayed.

[0084] For example, such as Figure 5 As shown, the fourth target area is the area above the progress bar in the third interface. Figure 5 (The area marked by the dashed box in the middle) In other words, after the user performs a fourth input on the operation object in the third interface, the electronic device responds to the fourth input by displaying multiple thumbnails of third photos above the progress bar of the multimedia data played in the third interface. Here, the thumbnails displayed are thumbnails of third photos related to the current screen displayed in the third interface, for example, third photos taken within a preset time period before and after the current screen displayed in the third interface.

[0085] Furthermore, while displaying thumbnails of multiple third photos in the fourth target area, such as Figure 5 As shown, an identifier label can also be displayed on the progress bar of the multimedia data playback. This identifier label is used to identify the video moment corresponding to the capture of each thumbnail.

[0086] Secondly, the system receives a fifth input from the user regarding the second target thumbnail in the fourth target area; here, similar to the aforementioned description of the first target thumbnail, the second target thumbnail is a thumbnail of the third photo that the user currently wants to view, and also... Figure 5 For example, the second target thumbnail is the second thumbnail in the fourth target area.

[0087] Furthermore, in response to the fifth input, the multimedia data displayed on the third interface is switched to the third photo corresponding to the second target thumbnail.

[0088] In step B2-11 above, when the user is viewing the recorded multimedia data, the user can perform a fourth input on the operation object in the third interface playing the multimedia data. The electronic device will display a thumbnail of a third photo related to the currently displayed screen, so that the user can perform a fifth input on the thumbnail of the third photo they want to view, and display the third photo corresponding to the thumbnail in the main display area of ​​the third interface. In this way, when the user is viewing the audio or video or photos saved in the album, they can quickly switch from the audio or video to the photo they want to view, reducing the operation of the user manually exiting the audio or video playback interface and then searching for and displaying the corresponding photo. This allows the user to quickly view the audio, video and photos of the wonderful moments captured, thereby improving the user experience.

[0089] Step B2-12: The second text information is displayed in the fifth target area of ​​the third interface, and a thumbnail of a third photo related to the second text information is displayed in the sixth target area of ​​the third interface. Based on the user's input of the second text information and the thumbnail, the multimedia data played on the third interface is switched to the third photo corresponding to the thumbnail displayed in the sixth target area. The second text information is related to the currently displayed screen on the third interface. This achieves the linkage between multimedia data, text information, and thumbnails.

[0090] Specifically, step B2-12 above can be implemented as follows:

[0091] First, second text information related to the currently displayed screen on the third interface is displayed in the fifth target area of ​​the third interface, and a thumbnail of a third photo related to the second text information is displayed in the sixth target area of ​​the third interface. In other words, in this step, the fourth input is used to display the second text information and the thumbnail of the third photo corresponding to the currently playing multimedia screen on the third interface. Thus, subsequent interaction between multimedia data, photos, and text information can be achieved; for example, by sliding the text information, the currently displayed screen and / or thumbnail can be updated. Taking a concert scene as an example, the second text information might be the lyrics corresponding to the currently playing multimedia screen, and the thumbnail of the third photo might be the time corresponding to the currently playing screen or the last photo taken before that time.

[0092] For example, such as Figure 6As shown, the fifth and sixth target areas are the areas above the progress bar for multimedia data playback. Specifically, the fifth and sixth target areas can be two areas arranged in a left-right order on the third interface. The text information in the fifth target area can scroll and update along with the multimedia data displayed on the third interface. Similarly, the thumbnail of the third photo in the sixth target area will also update along with the multimedia data displayed on the third interface or the second text information displayed in the fifth target area. If no photo was taken at the time corresponding to the currently displayed multimedia screen, the thumbnail displayed in the sixth target area may not be updated, that is, the last photo taken before the time corresponding to the current multimedia screen is displayed.

[0093] Furthermore, while displaying the second text information in the fifth target area, such as Figure 6 As shown, an identifier label can also be displayed on the progress bar of the multimedia data playback. This identifier label is used to identify the video moment corresponding to the second text information mentioned above.

[0094] It should be noted here that the fourth input corresponding to the above steps and the fourth input corresponding to step B2-11 should be different types of input. For example, the fourth input corresponding to this step is a double-click operation on the object being operated on, while the fourth input corresponding to step B2-11 is a single-click operation on the object being operated on. When the object being operated on is a label / text box with a drop-down menu, the fourth input in both specific implementation methods can be an operation to select the corresponding jump method.

[0095] Secondly, based on the user's input of the second text information and the thumbnail, the multimedia data played on the third interface is switched to the third photo corresponding to the thumbnail displayed in the sixth target area. This step specifically includes:

[0096] 1) Receive a sixth input from the user on the second text information; for example, the sixth input is a swipe operation, such as swiping the second text information up and down or swiping the second text information left and right.

[0097] 2) In response to the sixth input, the second text information displayed in the fifth target area is updated to the third text information, the screen of the currently displayed multimedia data is switched to the screen corresponding to the third text information, and the thumbnail displayed in the sixth target area is updated according to the third text information; in this way, the linkage update of multimedia data, text information and photos is realized, so that when the user is watching the recorded multimedia data, he / she can quickly and accurately switch to the wonderful moment he / she wants to watch by sliding the second text information.

[0098] For example, the size of the fifth target area can be adjusted according to user needs.

[0099] 3) Receive a seventh input from the user on the thumbnail displayed in the sixth target area; for example, the seventh input may be a double-click operation, a swipe-up operation, etc.

[0100] 4) In response to the seventh input, the screen displayed on the third interface is switched to the third photo corresponding to the thumbnail displayed in the sixth target area. In this way, when the user sees a wonderful moment in the multimedia data playback, the user can quickly view the photo corresponding to that wonderful moment, reducing the complexity of user operation.

[0101] In steps B2-12 above, the user can perform a fourth input on the object being operated on to retrieve the second text information and the thumbnail of the third photo corresponding to the currently playing multimedia data screen. Then, by performing a sixth input, such as swiping, on the second text information, the displayed screen can be quickly switched to the screen of the moment the user wants to view. Then, by performing a seventh input on the thumbnail, the displayed multimedia data screen can be switched to the third photo the user wants to view. In this way, the linkage between text information, multimedia data, and photos is realized, so that when the user is viewing multimedia data or photos saved in the album, they can quickly locate the multimedia data or photos of the wonderful moment they want to view with simple operations.

[0102] Step B2-2, corresponding to step B2-1 above, is as follows: when the third photo is displayed on the third interface, in response to the fourth input, the multimedia data is played on the third interface starting from a target time, the target time corresponding to the shooting time of the third photo displayed on the third interface.

[0103] In other words, when the third photo is displayed on the third interface, if the user performs a fourth input on the object of operation, the electronic device will switch from the third photo to multimedia data. Specifically, the currently displayed photo will be switched to the screen in the multimedia data that corresponds to the shooting time of the third photo, and the multimedia data will start playing from that screen. In this way, the displayed third photo can be quickly switched to the playback of multimedia data, reducing the complexity of user operation.

[0104] It should be noted that in step B2-2 above, when switching from the third photo to multimedia data, the switched screen may or may not display the multiple thumbnails corresponding to the playing screen, the second text information corresponding to the playing screen, and the thumbnails. For example, when the third photo is a photo displayed in response to the user's selection operation in the album, after switching from displaying the third photo to playing multimedia data, the third interface may not display the multiple thumbnails corresponding to the playing screen, or it may not display the second text information and thumbnails corresponding to the playing screen; when the third photo is a photo displayed by switching from multimedia data to a photo, the third interface may display multiple thumbnails, or it may display the second text information and thumbnails, depending on the switching method, or it may switch directly to multimedia data without considering the switching method. Specifically, when the switching method is the switching method corresponding to step B2-11 above, after switching from displaying the third photo to playing multimedia data, multiple thumbnails related to the currently displayed screen are displayed when playing multimedia data on the third interface; when the switching method is the switching method corresponding to step B2-12 above, after switching from displaying the third photo to playing multimedia data, second text information and thumbnails related to the currently displayed screen are displayed when playing multimedia data on the third interface.

[0105] In the above-mentioned optional implementation methods, by enabling bidirectional navigation between photos and multimedia data, users can quickly jump to the corresponding time point in the multimedia data by clicking the operation object in the upper right corner of the photo when viewing it in the album. Similarly, when playing multimedia data, clicking the operation object in the upper right corner will quickly jump to the corresponding photo. This allows for quick viewing of multimedia data and photos capturing memorable moments, simplifying the user experience.

[0106] Below, taking a concert scene as an example, combined with Figure 7 and Figure 8 The implementation process of the above-described photographing method in the embodiments of this application will be described.

[0107] Here, it should first be noted that, as Figure 7 As shown, the overall architecture for implementing the above-mentioned photo-taking method is as follows: Figure 7 As shown, the system includes a camera application, a video stream processing unit, an audio and video acquisition device, an AI audio and video detection unit, an AI image fusion unit, and a photo album that work together in an electronic device.

[0108] like Figure 8 As shown, the implementation example of the above-mentioned photo-taking method includes the following steps:

[0109] Step 801: The camera application sends a recording request; specifically, the camera application sends a command request to the video stream processing unit to request a recording command and start recording.

[0110] Step 802, audio and video data acquisition; specifically: the video stream processing unit sends a command request to the audio and video acquisition device, the audio and video acquisition device acquires audio and video data, and the video stream processing unit transmits the audio and video data to the camera application, so that the camera application can further analyze and process the audio and video data;

[0111] Step 803: AI detects audio and video data in real time. Specifically, the audio and video data is transmitted from the camera application to the AI ​​audio and video detection unit, which uses AI technologies such as deep learning / reinforcement learning to detect when a singer is performing the chorus in real time. Once the climax of the chorus is detected, and a clear frontal expression or special gesture (such as waving or tilting the head back to sing a high note) or stage effect (such as fireworks or ribbons) is recognized, the relevant detection results are returned to the camera application. The camera application calls up the song library to intelligently match and record the lyrics of the chorus at this moment, corresponding to the aforementioned first text information. At the same time, based on the emotional tone of the music, such as upbeat, lively, or melancholic, it intelligently selects appropriate fonts and effects to enhance the visual expressiveness of the lyrics.

[0112] Step 804: The camera application sends a photo-taking request; specifically, the camera application can automatically send a photo-taking request to capture a fleeting image through the video stream processing unit. Then, the video stream processing unit optimizes the captured image and returns it to the application. Here, the image may refer to the first photo.

[0113] Step 805, AI Image Fusion; specifically, the camera application sends the chorus lyrics and the captured image to the AI ​​image fusion unit. The AI ​​image fusion unit uses image recognition technology to perform face / human / stage recognition and segmentation on the captured image. Based on information such as image color and brightness, it finds the most suitable position of the lyrics in the captured image, and then merges the lyrics and image together to ensure harmony and unity. Afterwards, relevant scene detection and lyric segmentation information are recorded in the image data, and the fused image is returned to the application for subsequent secondary editing.

[0114] Step 806, Image Storage in Album; specifically, the camera application receives the merged image, i.e., the aforementioned second photo, and stores it in the album. Users can quickly share this photo with family and friends or upload it to social media, improving the user experience of instant sharing. This allows for secondary editing of the text in the photo during sharing to meet user needs, enabling users to create more personalized concert commemorative photos. For example, within the album, users can manually change the font position, edit lyrics, change font color, change font effects, and add singer-specific symbols, etc.

[0115] Taking a concert scene as an example, the overall process of the above-described photographing method in this application embodiment is as follows: Figure 9 As shown, the steps include: Step 901, opening the smart photo-taking function; for example, this step can be in a scenario where video recording is required, by entering the phone's camera interface, selecting the menu and entering the video recording interface, and opening the smart photo-taking function, for example... Figure 2 As shown, the smart photo-taking function can be turned on by clicking the smart photo-taking switch.

[0116] Step 902: Start video recording and take a smart photo; this step can be implemented according to the aforementioned steps 801 to 804, and will not be repeated here.

[0117] Step 903, Intelligent Video Recording and Photo Taking Process; this step can be implemented according to the aforementioned step 805. Additionally, this step can also include allowing users to view photos taken during recording, enabling them to view photos without accessing the photo album or interrupting recording. Specifically, swiping up brings up a horizontal bar displaying thumbnails of photos taken during recording. The thumbnails are stored from left to right in chronological order of capture time. To view photo details, double-clicking or swiping up moves the image to the main display area. The video recording will not be interrupted and will automatically zoom to the upper right corner. Photos in the main display area can be zoomed in and out in various directions to view details. If users are not satisfied with a photo, they can swipe up on the thumbnail display to delete the thumbnail, which will also delete the full-size image.

[0118] Step 904: Shooting completed and sharing; this step can be performed according to step 806.

[0119] Step 905, quickly view videos and photos. Since users may have taken a large number of videos and photos, this application provides a method to facilitate quick viewing of the corresponding highlights. Embodiments of this application support bidirectional navigation, one such method being as follows: Figure 5 As shown, when viewing photos in the album, you can interact with the jump tabs in the upper right corner of the photos. For example, clicking one will quickly jump to the corresponding time point in the video file. While playing a video, clicking the jump tab in the upper right corner will also work. During video playback, photo thumbnails will pop up above the progress bar, and the progress bar can also display inverted triangle labels indicating the corresponding video moment. Clicking the thumbnail will then quickly jump to the larger photo image for viewing. Another bidirectional jumping method involves simultaneous quick jumps between text, video footage, and the larger photo image, such as... Figure 6As shown, when playing a video, actions such as double-clicking the jump tab in the upper right corner will bring up a lyrics bar and photo thumbnails above the progress bar. An inverted triangle label can also be displayed on the progress bar, indicating the corresponding video moment for each lyric. Users can then quickly switch to the video frame corresponding to the lyrics by swiping up / down or left / right. The lyrics bar can also be enlarged or reduced according to user needs. When a user wants to view a larger photo, clicking the thumbnail to the right of the lyrics bar will jump to the larger photo. When a user is viewing photos in their album and wants to watch a video at a specific time point, clicking the jump tab in the upper right corner of the photo will quickly jump to the corresponding time point in the video file.

[0120] The purpose of the above embodiments of this application is to provide users with a more convenient and intelligent photo-taking experience during video recording. Taking a concert scene as an example, when using this intelligent video recording and photo-taking function, users only need to turn on the switch during recording to intelligently identify the chorus, and then combine this with the detection of the singer's clear frontal expression or special actions (such as waving, tilting their head back to sing a high note) or stage effects (such as fireworks, ribbons) to automatically trigger a photo capture, while simultaneously integrating the corresponding lyrics into the captured photo. It also supports continuous recording, allowing users to quickly view photos during recording and easily edit them. Videos and photos also support bidirectional switching, allowing for quick viewing of recorded videos and photos of exciting moments. Furthermore, it supports secondary editing of captured photos. This not only improves the ability for users to capture exciting moments of concert singers but also effectively adds lyrics to the photos, thereby enhancing the user experience and facilitating timely sharing on social media.

[0121] In other words, the above-described photo-taking method in this application provides users with a convenient and intelligent concert recording and photo-taking experience through intelligent video recording and photo-taking functions. In a concert setting, it can automatically trigger photo-taking by recognizing the chorus in real time and combining it with highlights from the stage or singer, intelligently merging lyrics with captured photos to enhance visual impact. Photos can be quickly viewed without stopping recording. This solves the problem that users need to focus on taking photos to capture highlights while watching a concert, thus affecting their emotional experience. It also addresses the issue that capturing highlights is difficult and may result in missing them. Furthermore, it solves the problem that taking photos during recording requires interrupting the recording and then accessing the album, potentially missing highlights. In addition, the photo-taking method in this application supports bidirectional navigation, linking lyrics, photos, and videos for quick viewing of highlights. This solves the problem that after a concert, facing a large number of photos in the album, users need to review the video frame by frame and match the photo's capture time to determine the specific moment the photo was taken. Adding lyrics to photos later is a complex and inconvenient process for users. Furthermore, the embodiments of this application support secondary editing and quick sharing of photos, enhancing users' ability to capture memorable moments, simplifying the social media sharing process, and greatly enriching the user's interactive experience.

[0122] It should be noted that the photo-taking method in this application embodiment is illustrated using a concert scene as an example. However, the method is not limited to concert scenes, but can also be used for dramas, stage plays, or other scenes. Any scene that can be recognized by AI technology can use the photo-taking method in this application embodiment for intelligent photo-taking.

[0123] The photographing method provided in this application can be executed by a photographing device. This application uses a photographing device executing the photographing method as an example to illustrate the photographing device provided in this application.

[0124] like Figure 10 As shown, the photographing device 1000 includes:

[0125] The camera module 1001 is used to take a picture while continuing to record video, and obtain a first photo when it detects that the multimedia data currently being recorded by the electronic device meets preset conditions.

[0126] The preset conditions include at least one of the following:

[0127] The audio data in the multimedia data is related to the target information;

[0128] The video images in the multimedia data contain the target object.

[0129] The device further includes:

[0130] The acquisition module is used to acquire first text information that matches the audio data in the currently recorded multimedia data when the multimedia data currently being recorded by the electronic device meets the preset conditions.

[0131] The camera module 1001 includes:

[0132] The image recognition submodule is used to perform image recognition on the second photo taken while the video is being recorded, and to obtain the image recognition result.

[0133] The image processing submodule is used to overlay the first text information onto the first target area of ​​the second photo based on the image recognition result, so as to obtain the first photo.

[0134] The device further includes:

[0135] The first receiving module is used to receive the user's editing operation on the first interface of the electronic device, the first interface being the interface that displays the first photo;

[0136] The update module is used to update the first text information in the first photo.

[0137] The device further includes:

[0138] The second receiving module is used to receive the first input from the user on the second interface of the electronic device, and the main display area of ​​the second interface displays the video recording.

[0139] A first display module is configured to, in response to the first input, display a thumbnail of at least one third photo in a second target area of ​​the second interface, the third photo including the first photo and / or a photo taken in response to a user's photo-taking operation during video recording;

[0140] The second receiving module is used to receive a second input from the user regarding a first target thumbnail in the second target area;

[0141] The first processing module is configured to respond to the second input by scaling the video recording to a third target area of ​​the second interface and displaying the third photo corresponding to the first target thumbnail in the main display area of ​​the second interface.

[0142] The device further includes:

[0143] The third receiving module is used to receive the user's third input on the first target thumbnail;

[0144] The second processing module is configured to, in response to the third input, delete the first target thumbnail and the third photo corresponding to the first target thumbnail.

[0145] The device further includes:

[0146] The fourth receiving module is used to receive a fourth input from the user to the operation object, which is the operation object in the third interface of the electronic device. The third interface is used to display a third photo or play multimedia data; wherein, the third photo includes the first photo and / or a photo taken in response to the user's photo-taking operation during the recording process.

[0147] The third processing module is used to, in response to the fourth input, play the multimedia data on the third interface starting from a target time when the third photo is displayed on the third interface, wherein the target time corresponds to the shooting time of the third photo displayed on the third interface;

[0148] The fourth processing module is configured to, in response to the fourth input, perform any of the following when the multimedia data is played on the third interface:

[0149] In the fourth target area of ​​the third interface, thumbnails of multiple third photos are displayed, and based on the fifth input received from the user regarding the second target thumbnail, the multimedia data played on the third interface is switched to the third photo corresponding to the second target thumbnail; wherein, the multiple third photos are related to the screen of the multimedia data currently played on the third interface;

[0150] The second text information is displayed in the fifth target area of ​​the third interface, and a thumbnail of a third photo related to the second text information is displayed in the sixth target area of ​​the third interface. Based on the user's input of the second text information and the thumbnail, the multimedia data played on the third interface is switched to the third photo corresponding to the thumbnail displayed in the sixth target area.

[0151] Specifically, when the fourth processing module switches the multimedia data played on the third interface to the third photo corresponding to the thumbnail displayed in the sixth target area based on the user's input of the second text information and the thumbnail, it is used to:

[0152] Receive the user's sixth input regarding the second text information;

[0153] In response to the sixth input, the second text information displayed in the fifth target area is updated to the third text information, the screen of the currently displayed multimedia data is switched to the screen corresponding to the third text information, and the thumbnail displayed in the sixth target area is updated according to the third text information;

[0154] Receive a seventh input from the user regarding the thumbnail displayed in the sixth target area;

[0155] In response to the seventh input, the screen displayed on the third interface is switched to the third photo corresponding to the thumbnail displayed in the sixth target area.

[0156] The photographing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., and this application embodiment does not specifically limit the scope.

[0157] The camera device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0158] The photographic device provided in this application embodiment can achieve... Figures 1 to 9 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0159] Optionally, such as Figure 11 As shown, this application embodiment also provides an electronic device 1100, including a processor 1101 and a memory 1102. The memory 1102 stores a program or instructions that can run on the processor 1101. When the program or instructions are executed by the processor 1101, they implement the various steps of the above-described photographing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0160] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0161] Figure 12 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0162] The electronic device 1200 includes, but is not limited to, components such as: radio frequency unit 1201, network module 1202, audio output unit 1203, input unit 1204, sensor 1205, display unit 1206, user input unit 1207, interface unit 1208, memory 1209, and processor 1210.

[0163] Those skilled in the art will understand that the electronic device 1200 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1210 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 12 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0164] The processor 1210 is configured to take a picture while continuing video recording, and obtain a first photo, when it detects that the multimedia data currently being recorded by the electronic device meets preset conditions; wherein the preset conditions include at least one of the following:

[0165] The audio data in the multimedia data is related to the target information;

[0166] The video images in the multimedia data contain the target object.

[0167] In some embodiments, the processor 1210 is further configured to, when detecting that the multimedia data currently recorded by the electronic device meets the preset conditions, acquire first text information that matches the audio data in the currently recorded multimedia data;

[0168] Specifically, when the processor 1210 is used to take a picture while recording video to obtain a first picture, it is used to: perform image recognition on the second picture taken while recording video to obtain an image recognition result; and, based on the image recognition result, overlay the first text information onto the first target area of ​​the second picture to obtain the first picture.

[0169] The input unit 1204 is used to receive editing operations from the user on the first interface of the electronic device, the first interface being the interface that displays the first photo.

[0170] The processor 1210 is also used to update the first text information in the first photograph.

[0171] In some embodiments, the input unit 1204 is further configured to receive a first input from a user on a second interface of the electronic device, wherein the main display area of ​​the second interface displays a video recording.

[0172] The processor 1210 is also configured to, in response to the first input, control the display unit 1206 to display a thumbnail of at least one third photo in a second target area of ​​the second interface, the third photo including the first photo and / or a photo taken in response to a user's photo-taking operation during video recording;

[0173] The input unit 1204 is also configured to receive a second input from the user regarding a first target thumbnail in the second target area;

[0174] The processor 1210 is also configured to respond to the second input by controlling the display unit 1206 to zoom the video recording to the third target area of ​​the second interface and display the third photo corresponding to the first target thumbnail in the main display area of ​​the second interface.

[0175] In some embodiments, the input unit 1204 is further configured to receive a third input from the user on the first target thumbnail;

[0176] The processor 1210 is further configured to, in response to the third input, delete the first target thumbnail and the third photo corresponding to the first target thumbnail.

[0177] In some embodiments, the input unit 1204 is further configured to: receive a fourth input from a user to an operation object, the operation object being an operation object in a third interface of the electronic device, the third interface being configured to display a third photo or play multimedia data; wherein, the third photo includes the first photo and / or a photo taken in response to a user's photo-taking operation during video recording;

[0178] The processor 1210 is also configured to, in response to the fourth input, control the display unit 1206 to start playing the multimedia data on the third interface from a target time when the third photo is displayed on the third interface, wherein the target time corresponds to the shooting time of the third photo displayed on the third interface;

[0179] The processor 1210 is further configured to, in response to the fourth input, perform any of the following when the multimedia data is played on the third interface:

[0180] The display unit 1206 is controlled to display thumbnails of multiple third photos in the fourth target area of ​​the third interface, and according to the fifth input received from the user on the second target thumbnail, the multimedia data played on the third interface is switched to the third photo corresponding to the second target thumbnail; wherein, the multiple third photos are related to the screen of the multimedia data currently played on the third interface;

[0181] The display unit 1206 is controlled to display second text information in the fifth target area of ​​the third interface, and to display a thumbnail of a third photo related to the second text information in the sixth target area of ​​the third interface. Based on the user's input of the second text information and the thumbnail, the multimedia data played on the third interface is switched to the third photo corresponding to the thumbnail displayed in the sixth target area.

[0182] In some embodiments, the input unit 1204 is further configured to: receive a sixth input from the user regarding the second text information;

[0183] The processor 1210 is specifically configured to: respond to the sixth input, control the display unit 1206 to update the second text information displayed in the fifth target area to the third text information, switch the currently displayed multimedia data screen to the screen corresponding to the third text information, and update the thumbnail displayed in the sixth target area according to the third text information.

[0184] The input unit 1204 is further configured to: receive a seventh input from the user regarding the thumbnail displayed in the sixth target area;

[0185] The processor 1210 is specifically configured to: in response to the seventh input, control the display unit 1206 to switch the screen displayed on the third interface to the third photo corresponding to the thumbnail displayed in the sixth target area.

[0186] It should be understood that, in this embodiment, the input unit 1204 may include a graphics processing unit (GPU) 12041 and a microphone 12042. The GPU 12041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1206 may include a display panel 12061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1207 includes a touch panel 12071 and at least one of other input devices 12072. The touch panel 12071 is also called a touch screen. The touch panel 12071 may include a touch detection device and a touch controller. Other input devices 12072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0187] The memory 1209 can be used to store software programs and various data. The memory 1209 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1209 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1209 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0188] Processor 1210 may include one or more processing units; optionally, processor 1210 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1210.

[0189] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described photographing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0190] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0191] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described photographing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0192] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0193] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described photographing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0194] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0195] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0196] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for taking photos, characterized in that, The method includes: If the multimedia data currently being recorded by the electronic device meets preset conditions, a first photo is taken while the video recording continues; wherein the preset conditions include at least one of the following: The audio data in the multimedia data is related to the target information; The video images in the multimedia data contain the target object.

2. The method according to claim 1, characterized in that, The method further includes: If the multimedia data currently recorded by the electronic device meets the preset conditions, first text information matching the audio data in the currently recorded multimedia data is obtained; The step of taking a photo while recording video to obtain a first photo includes: Image recognition is performed on the second photo taken while the video recording is being maintained, and the image recognition result is obtained; Based on the image recognition result, the first text information is superimposed onto the first target area of ​​the second photo to obtain the first photo.

3. The method according to claim 1 or 2, characterized in that, The method further includes: The system receives editing operations from the user on a first interface of the electronic device, where the first interface is an interface that displays the first photo. Update the first text information in the first photo.

4. The method according to any one of claims 1 or 2, characterized in that, The method further includes: The system receives a first input from a user on the second interface of the electronic device, and the main display area of ​​the second interface displays the video recording. In response to the first input, at least one thumbnail of a third photo is displayed in a second target area of ​​the second interface, the third photo including the first photo and / or a photo taken in response to the user's photo-taking operation during the recording process; Receive a second input from the user regarding a thumbnail of a first target within the second target area; In response to the second input, the video recording is zoomed to the third target area of ​​the second interface, and the third photo corresponding to the first target thumbnail is displayed in the main display area of ​​the second interface.

5. The method according to claim 4, characterized in that, The method further includes: Receive a third input from the user regarding the first target thumbnail; In response to the third input, the first target thumbnail and the third photo corresponding to the first target thumbnail are deleted.

6. The method according to claim 1 or 2, characterized in that, The method further includes: The system receives a fourth input from the user regarding an operation object, which is an operation object in the third interface of the electronic device. The third interface is used to display a third photo or play multimedia data. The third photo includes the first photo and / or a photo taken in response to the user's photo-taking operation during the recording process. When the third photo is displayed on the third interface, in response to the fourth input, the multimedia data is played on the third interface starting from a target time, the target time corresponding to the shooting time of the third photo displayed on the third interface; When the multimedia data is played on the third interface, in response to the fourth input, perform any of the following: In the fourth target area of ​​the third interface, thumbnails of multiple third photos are displayed, and based on the fifth input received from the user regarding the second target thumbnail, the multimedia data played on the third interface is switched to the third photo corresponding to the second target thumbnail; wherein, the multiple third photos are related to the screen of the multimedia data currently played on the third interface; The second text information is displayed in the fifth target area of ​​the third interface, and a thumbnail of a third photo related to the second text information is displayed in the sixth target area of ​​the third interface. Based on the user's input of the second text information and the thumbnail, the multimedia data played on the third interface is switched to the third photo corresponding to the thumbnail displayed in the sixth target area.

7. The method according to claim 6, characterized in that, The step of switching the multimedia data played on the third interface to the third photo corresponding to the thumbnail displayed in the sixth target area based on the user's input of the second text information and the thumbnail includes: Receive the user's sixth input regarding the second text information; In response to the sixth input, the second text information displayed in the fifth target area is updated to the third text information, the screen of the currently displayed multimedia data is switched to the screen corresponding to the third text information, and the thumbnail displayed in the sixth target area is updated according to the third text information; Receive a seventh input from the user regarding the thumbnail displayed in the sixth target area; In response to the seventh input, the screen displayed on the third interface is switched to the third photo corresponding to the thumbnail displayed in the sixth target area.

8. A photographing device, characterized in that, The device includes: The camera module is used to take a picture while recording video and obtain a first photo when the multimedia data currently being recorded by the electronic device meets preset conditions. The preset conditions include at least one of the following: The audio data in the multimedia data is related to the target information; The video images in the multimedia data contain the target object.

9. The apparatus according to claim 8, characterized in that, The device further includes: The acquisition module is used to acquire first text information that matches the audio data in the currently recorded multimedia data when the multimedia data currently being recorded by the electronic device meets the preset conditions. The camera module includes: The image recognition submodule is used to perform image recognition on the second photo taken while the video is being recorded, and to obtain the image recognition result. The image processing submodule is used to overlay the first text information onto the first target area of ​​the second photo based on the image recognition result, so as to obtain the first photo.

10. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the photographing method as described in any one of claims 1-7.

11. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the photographing method as described in any one of claims 1 to 7.