Voice photo shooting method, display method, device, equipment and storage medium
By combining image and voice acquisition modules in smart home devices to generate target voice photos, the problems of limited content and high memory usage are solved, achieving rich recording content and low memory usage, thus improving the user experience.
Patent Information
- Application Number
- CN202411644415.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing smart home devices suffer from limited content and high memory usage when taking photos and videos.
By using the image acquisition module and the voice acquisition module in the voice photo shooting mode, the image and voice information of the subject within the field of view are acquired, and the relationship is established to generate the target voice photo.
It improves the richness of the recorded content in photos, reduces memory usage, and enhances user experience and device usability.
Smart Images

Figure CN119697507B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of information processing, and particularly relates to a photographing method and a display method of a voice photo, a device, an apparatus and a storage medium. BACKGROUND
[0002] With the continuous development of science and technology, various high technologies are applied to daily life, and various smart home devices also appear, so that smart home devices such as smart phones, smart televisions, smart refrigerators and smart sound boxes have the functions of photographing photos and videos.
[0003] In the related art, the photographing mode of the smart home device with the functions of photographing photos and videos can only perform photographing and storage of one of a photo, a video and a live photo each time. However, the photographing and storage of the photo can only record the image, resulting in single recording content; although the video and the live photo can record the image and the voice at the same time, there is a problem of large memory space occupation. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a photographing method and a display method of a voice photo, a device, an apparatus and a storage medium, which can solve the problems of single content and large memory space occupation of the smart home device in the related art during photographing.
[0005] In a first aspect, the embodiments of the present application provide a photographing method of a voice photo, applied to a first electronic device, the first electronic device comprising an image acquisition module and a voice acquisition module; the method comprising:
[0006] In a case where a voice photo photographing mode is enabled, determining a photographing object in a viewfinder range of the image acquisition module;
[0007] In a case where it is monitored that a photographing button of the first electronic device is in a pressed state, acquiring voice information corresponding to the photographing object by using the voice acquisition module;
[0008] Acquiring a target image corresponding to the viewfinder range by using the image acquisition module; the target image comprises the photographing object;
[0009] Establishing an association relationship between the photographing object in the target image and the voice information to obtain a target voice photo; the target voice photo comprises the target image and the voice information.
[0010] Optionally, in a case where the number of the photographing objects is greater than 1, the acquiring, by using the voice acquisition module, of the voice information corresponding to the photographing object in the case where it is monitored that the photographing button of the first electronic device is in the pressed state comprises:
[0011] in a case where it is monitored that the shooting button of the first electronic device is in a pressed state, receiving a first focusing instruction;
[0012] focusing on a first object indicated by the first focusing instruction in the shooting object;
[0013] acquiring first audio track information corresponding to the first object by using the voice acquisition module until a second focusing instruction is received;
[0014] focusing on a second object indicated by the second focusing instruction in the shooting object;
[0015] acquiring second audio track information corresponding to the second object by using the voice acquisition module;
[0016] determining voice information corresponding to the shooting object according to the first audio track information and the second audio track information.
[0017] Optionally, the method further comprises:
[0018] creating a first voice link identifier in a first target area corresponding to the first object in the target image, and creating a second voice link identifier in a second target area corresponding to the second object;
[0019] establishing a link relationship between the first voice link identifier and the first audio track information, and establishing a link relationship between the second voice link identifier and the second audio track information, to obtain a target voice photo.
[0020] Optionally, in a case where the number of the shooting objects is greater than 1, the method further comprises:
[0021] in a case where it is monitored that the shooting button of the first electronic device is in a pressed state, acquiring image information in the framing range by using the image acquisition module.
[0022] Optionally, the method further comprises:
[0023] performing lip analysis on the shooting objects based on the image information, and separating audio track information corresponding to each shooting object in the image information from the voice information;
[0024] determining a corresponding relationship between a shooting object in the image information and a shooting object in the target image according to the image information and the target image;
[0025] According to the correspondence, an association relationship between the shooting object in the target image and the audio track information is established, and a target voice photo is obtained.
[0026] Optionally, in a case where the number of the shooting objects is 1, in a case where it is monitored that the shooting button of the first electronic device is in a pressed state, the voice information corresponding to the shooting object is acquired by using the voice acquisition module, and the voice information corresponding to the shooting object is obtained.
[0027] In a case where it is monitored that the shooting button of the first electronic device is in a pressed state, voice information is acquired by using the voice acquisition module, and the voice information corresponding to the shooting object is obtained.
[0028] Optionally, the association relationship between the shooting object in the target image and the voice information is established, and the target voice photo is obtained, including:
[0029] A third voice link identifier is created in a third target region in the target image.
[0030] A link relationship between the third voice link identifier and the voice information corresponding to the shooting object is established, and the target voice photo is obtained.
[0031] In a second aspect, an embodiment of the present application provides a display method of a voice photo, applied to a second electronic device, the second electronic device including a display and a voice player; the method includes:
[0032] In a case where a display instruction is received, a target voice photo indicated by the display instruction is acquired; the target voice photo is obtained by using the voice photo shooting method as described above; the target voice photo includes a target image and voice information;
[0033] The target image is displayed by using the display.
[0034] The voice information is played by using the voice player.
[0035] In a third aspect, an embodiment of the present application provides a shooting device of a voice photo, applied to a first electronic device, the first electronic device including an image acquisition module and a voice acquisition module; the device includes:
[0036] A first determination module is configured to determine a shooting object in a range of view of the image acquisition module in a case where a voice photo shooting mode is enabled.
[0037] A first acquisition module is configured to acquire, in a case where it is monitored that a shooting button of the first electronic device is in a pressed state, voice information corresponding to the shooting object by using the voice acquisition module.
[0038] a second obtaining module, configured to obtain a target image corresponding to the framing range by using the image acquisition module; the target image includes the photographed object;
[0039] an association module, configured to establish an association between the photographed object in the target image and the voice information, to obtain a target voice photo; the target voice photo includes the target image and the voice information.
[0040] In a fourth aspect, an embodiment of the present application provides a display device of a voice photo, applied to a second electronic device, the second electronic device including a display and a voice player; the device includes:
[0041] a third obtaining module, configured to, in a case where a display instruction is received, obtain a target voice photo indicated by the display instruction; the target voice photo is obtained by using the voice photo photographing method described above; the target voice photo includes a target image and voice information;
[0042] a display module, configured to display the target image by using the display;
[0043] a playing module, configured to play the voice information by using the voice player.
[0044] In a fifth aspect, an embodiment of the present application provides an electronic device, including an image acquisition module, a voice acquisition module, a display, a voice player, a processor, a memory, a communication interface and a communication bus; the processor, the memory and the communication interface complete communication with each other through the communication bus; the memory is used to store executable instructions; the executable instructions make the processor execute the voice photo photographing method according to any one of the above or execute the voice photo display method according to the above.
[0045] In a sixth aspect, an embodiment of the present application provides a readable storage medium; when instructions in the readable storage medium are executed by a processor of an electronic device, the processor can execute the voice photo photographing method according to any one of the above or execute the voice photo display method according to the above.
[0046] Embodiments of the present application have the following advantages:
[0047] The photographing method of the voice photo provided in the embodiment of the present application, the first electronic device first determines the shooting object in the image capturing module's viewfinder range in the case that the voice photo shooting mode is enabled; in the case that the shooting button of the first electronic device is in the pressed state, the voice information corresponding to the shooting object is acquired by using the voice capturing module, and the target image including the shooting object corresponding to the viewfinder range is acquired by using the image capturing module; finally, the association between the shooting object in the target image and the voice information is established, and the target voice photo is obtained, so that the target voice photo obtained by shooting can record the target image in the viewfinder range and the voice information corresponding to the shooting object in the target image at the same time, the richness of the content recorded by the first electronic device in the process of photographing the photo is improved, and the occupied amount of the memory space is smaller than that of the video and the live photo. The embodiment of the present application improves the richness of the content recorded by the target voice photo, and reduces the occupied amount of the memory space of the first electronic device, and improves the accommodation and ease of use of the memory space of the first electronic device in the embodiment of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 is a step flow chart of a photographing method of a voice photo provided in the embodiment of the present application;
[0049] Figure 2 is a schematic view of a shooting object in the viewfinder range of an image capturing module provided in the embodiment of the present application;
[0050] Figure 3 is a schematic view of a target voice photo provided in the embodiment of the present application;
[0051] Figure 4 is a schematic view of a shooting object in the viewfinder range of an image capturing module provided in the embodiment of the present application;
[0052] Figure 5 is a schematic view of a target voice photo provided in the embodiment of the present application;
[0053] Figure 6 is a schematic view of a target voice photo provided in the embodiment of the present application;
[0054] Figure 7 is a schematic view of a target voice photo provided in the embodiment of the present application;
[0055] Figure 8 is a logic block diagram of a photographing method of a voice photo provided in the embodiment of the present application;
[0056] Figure 9 is a logic block diagram of another photographing method of a voice photo provided in the embodiment of the present application;
[0057] Figure 10 is a step flow chart of a voice photo display method provided by an embodiment of the present application;
[0058] Figure 11 is a logic block diagram of a voice photo shooting device provided by an embodiment of the present application;
[0059] Figure 12 is a logic block diagram of a voice photo display device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0060] The technical solutions in the embodiments of the present application will be described clearly below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.
[0061] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a category, and are not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.
[0062] Method embodiments
[0063] The voice photo shooting method provided by the embodiments of the present application will be described in detail below with reference to the drawings, specific embodiments and application scenarios.
[0064] Reference Figure 1 , a step flow chart of a voice photo shooting method provided by an embodiment of the present application is shown, which specifically includes:
[0065] Step S101, in the case where a voice photo shooting mode is enabled, determining a shooting object in a viewfinder range of an image acquisition module.
[0066] Step S102, in the case where it is monitored that a shooting button of a first electronic device is in a pressed state, acquiring voice information corresponding to the shooting object by using a voice acquisition module.
[0067] Step S103, acquiring a target image corresponding to the viewfinder range by using the image acquisition module; the target image includes the shooting object.
[0068] Step S104, establishing an association between the photographed object in the target image and the voice information, to obtain a target voice photo; the target voice photo includes the target image and the voice information.
[0069] The photographing method of the voice photo provided in the embodiments of the present application can be applied to a first electronic device, which includes an image acquisition module and a voice acquisition module. It can be understood that the first electronic device can be any electronic device including an image acquisition module and a voice acquisition module and having a data processing function. The first electronic device can include, but is not limited to, a smart terminal, a computer, a personal digital assistant (PDA), a tablet computer, an electronic book reader, a laptop computer, a vehicle-mounted device, a smart television, a wearable device, etc. Specifically, the first electronic device can be a smart home device, for example, a smart phone, a smart television, a smart refrigerator, a smart speaker, etc.
[0070] In the embodiments of the present application, the first electronic device can realize the photographing function of the voice photo based on the image acquisition module and the voice acquisition module. When the first electronic device monitors that the voice photo photographing mode is enabled, the first electronic device can determine a photographed object in the viewfinder range of the image acquisition module.
[0071] The image acquisition module includes a viewfinder, which is used to capture and provide image information to be photographed in the viewfinder range in real time, so that the user can preview the image information to be photographed. It can be understood that the viewfinder range of the image acquisition module is specifically the viewfinder range of the viewfinder, and the image information to be photographed in the viewfinder range includes the photographed object. In some embodiments, the first electronic device can further include a display. In step S101, the first electronic device can display the photographed object in the viewfinder range of the image acquisition module to the user in real time through the display.
[0072] The photographed object is a to-be-photographed object in the viewfinder range, which can include, but is not limited to, a person, an animal, a plant, etc. In the embodiments of the present application, the number of the photographed objects in the viewfinder range is at least 1.
[0073] In the embodiments of the present application, when the first electronic device monitors that the photographing button of the first electronic device is in a pressed state, the first electronic device can acquire the voice information corresponding to the photographed object by using the voice acquisition module. The photographing button of the first electronic device can be a hardware button arranged in the first electronic device, or a function button in an interface in which the first electronic device displays the photographed object in the viewfinder range of the image acquisition module to the user.
[0074] Reference Figure 2, a schematic diagram of a shooting object in a viewfinder range of an image acquisition module is shown, in Figure 2 The number of the shooting objects is 3 in the embodiment, and the shooting button of the first electronic device is a function button in an interface in which the first electronic device displays the shooting objects in the viewfinder range of the image acquisition module to the user.
[0075] Specifically, in the case of needing to acquire voice information, the user can press the shooting button of the first electronic device, so that the first electronic device performs the operation corresponding to step S102 when it is monitored that the shooting button is in the pressed state.
[0076] It can be understood that the first electronic device can continuously acquire the voice information corresponding to the shooting object by using the voice acquisition module when it is monitored that the shooting button is in the pressed state, and the first electronic device controls the voice acquisition module to end the acquisition of the voice information when the state of the shooting button is converted from the pressed state to the released state, and determines the voice information continuously acquired by the voice acquisition module between the first time when the state of the shooting button is converted from the released state to the pressed state and the second time when the state of the shooting button is converted from the pressed state to the released state as the voice information corresponding to the shooting object.
[0077] The voice information corresponding to the shooting object refers to the voice information that can be acquired by the voice acquisition module when the shooting button of the first electronic device is in the pressed state, which can be emitted by the shooting object or by other objects other than the shooting object, and the embodiment of the present application does not make a specific limitation.
[0078] In the embodiment of the present application, the first electronic device can also acquire the target image corresponding to the viewfinder range by using the image acquisition module through step S103 as the cover photo of the voice information acquired by the first electronic device through step S102.
[0079] The first electronic device can execute step S103 before step S102 or after step S102, specifically, the first electronic device can acquire the target image corresponding to the viewfinder range by using the image acquisition module through step S103 at the first time when the state of the shooting button is converted from the released state to the pressed state, and the first electronic device can also acquire the target image corresponding to the viewfinder range by using the image acquisition module through step S103 at the second time when the state of the shooting button is converted from the pressed state to the released state, and the embodiment of the present application does not make a specific limitation on the execution order of step S102 and step S103.
[0080] It should be noted that the contents and the number of the photographed objects in the range of view in step S101 and the photographed objects in the target image in step S103 are the same, but the relative positions between the photographed objects can be the same or different.
[0081] In the embodiment of the present application, after the first electronic device obtains the voice information corresponding to the photographed objects through step S102 and obtains the target image including the photographed objects through step S103, the first electronic device can perform step S104 to establish an association relationship between the photographed objects in the target image and the voice information, and obtain a target voice photo.
[0082] In the embodiment of the present application, after the first electronic device obtains the voice information corresponding to the photographed objects through step S102 and obtains the target image including the photographed objects through step S103, the first electronic device can perform step S104 to establish an association relationship between the photographed objects in the target image and the voice information, and obtain a target voice photo.
[0083] In step S104, the first electronic device can store the voice information obtained through step S102 and the target image obtained through step S103 in a one-to-one corresponding relationship in a first preset storage area of the first electronic device, and obtain a target voice photo, where the first preset storage area can be a storage area corresponding to an album in the first electronic device.
[0084] When a user views the target voice photo through the first electronic device, the first electronic device can display the target image in the target voice photo through the display, and obtain the voice information corresponding to the target image from the first preset storage area and play the voice information, thereby providing the user with a visual and auditory experience and increasing the interest in the voice photo shooting process of the first electronic device.
[0085] As an example, referring to Figure 3 , a schematic diagram of a target voice photo provided by an embodiment of the present application is shown, as shown in Figure 3 , the target voice photo includes a target image and an overall voice link identifier corresponding to voice information with a time length of 8s, the first electronic device can obtain the voice information corresponding to the target image from the first preset cache area based on the overall voice link identifier, and when a user views the target voice photo through the album of the first electronic device, the first electronic device can display the target image in the target voice photo through the display, and obtain the voice information corresponding to the target image from the first preset storage area and play the voice information; where the number of photographed objects in the target image is 3.
[0086] In the related art, the shooting mode of a smart home device with a photo and video shooting function can only perform shooting and storage of one of a photo, a video, and a live photo each time. However, the shooting and storage of a photo can only record an image, resulting in single recording content; and although a video and a live photo can simultaneously record an image and voice, there is a problem of large memory space occupation. The shooting method of a voice photo provided in the embodiments of the present application, under the condition that a voice photo shooting mode of a first electronic device is enabled, first determines a shooting object in a viewfinder range of an image acquisition module; under the condition that a shooting button of the first electronic device is in a pressed state, acquires voice information corresponding to the shooting object by using a voice acquisition module, and acquires a target image including the shooting object corresponding to the viewfinder range by using the image acquisition module; and finally establishes an association between the shooting object in the target image and the voice information, to obtain a target voice photo. The target voice photo obtained by shooting can simultaneously record a target image in the viewfinder range and voice information corresponding to the shooting object in the target image, thereby improving the richness of recording content of the first electronic device in the photo shooting process, and the memory space occupation is less than that of a video and a live photo. The embodiments of the present application improve the richness of recording content of the target voice photo, and also reduce the memory space occupation of the first electronic device, thereby improving the accommodation and ease of use of the memory space of the first electronic device in the embodiments of the present application, and further improving the interest of a user in the voice photo shooting process by using the first electronic device, and improving the life quality and use experience of the user.
[0087] As an optional implementation, referring to Figure 4 , another schematic diagram of a shooting object in a viewfinder range of an image acquisition module is shown, in Figure 4 , the shooting button of the first electronic device is a function button in an interface in which the first electronic device displays the shooting object in the viewfinder range of the image acquisition module to the user.
[0088] The photographing button of the first electronic device comprises a first button and a second button, the first button is configured to control the first electronic device to acquire the voice information corresponding to the photographing object by using the voice acquisition module, and the second button is configured to control the first electronic device to acquire the target image corresponding to the range of view by using the image acquisition module. Specifically, the first electronic device acquires the voice information corresponding to the photographing object by using the voice acquisition module in a case where it is monitored that the first button is in a pressed state. The first electronic device acquires the target image corresponding to the range of view by using the image acquisition module in a case where it is monitored that the second button is in a pressed state. It should be noted that, in a case where the photographing button comprises the first button and the second button, the first electronic device can independently realize the acquisition of the voice information and the acquisition of the target image according to the pressed state of the first button and the pressed state of the second button, respectively. It can be understood that, in a case where the first button and the second button are simultaneously in a pressed state, the first electronic device can simultaneously perform the operations corresponding to steps S102 and S103.
[0089] Optionally, in a case where the number of photographing objects is greater than 1, the step S102 of acquiring the voice information corresponding to the photographing object by using the voice acquisition module in a case where it is monitored that the photographing button of the first electronic device is in a pressed state comprises:
[0090] The step S1021 comprises: receiving a first focusing instruction in a case where it is monitored that the photographing button of the first electronic device is in a pressed state.
[0091] The step S1022 comprises: focusing on a first object indicated by the first focusing instruction in the photographing object.
[0092] The step S1023 comprises: acquiring first audio track information corresponding to the first object by using the voice acquisition module until a second focusing instruction is received.
[0093] The step S1024 comprises: focusing on a second object indicated by the second focusing instruction in the photographing object.
[0094] The step S1025 comprises: acquiring second audio track information corresponding to the second object by using the voice acquisition module.
[0095] The step S1026 comprises: determining the voice information corresponding to the photographing object according to the first audio track information and the second audio track information.
[0096] In the embodiment of the present application, in the case that the number of the shooting objects in the field of view of the image acquisition module is greater than 1, that is, the number of the shooting objects is multiple, in the process that the first electronic device monitors that the shooting button of the first electronic device is in the pressed state and acquires the voice information corresponding to the shooting objects by using the voice acquisition module, the voice information corresponding to each shooting object can be acquired in turn according to the focusing sequence of each shooting object in the field of view through steps S1021 to S1026.
[0097] The first object is any one of the multiple shooting objects, and the second shooting object is a shooting object other than the first object among the multiple shooting objects. It can be understood that the number of the first shooting objects is 1, and the number of the second shooting objects is greater than or equal to 1.
[0098] Specifically, in the case that the first electronic device monitors that the shooting button is in the pressed state and receives the first focusing instruction, the first electronic device can perform step S1022 to focus on a first object in the shooting objects indicated by the first focusing instruction by using the image acquisition module, wherein the first focusing instruction is used to instruct the first electronic device to focus on the first object among the shooting objects in the field of view; the first focusing instruction can be a focusing instruction for the first object automatically determined by the first electronic device from the shooting objects in the field of view according to a preset focusing rule; or the first focusing instruction can be a focusing instruction for the first object input by a user to the electronic device according to the shooting objects in the field of view displayed by the first electronic device.
[0099] The preset focusing rule can include but is not limited to: 1) the shooting objects in the field of view are focused in the order from left to right, and the shooting object located at the leftmost side of the field of view is determined as the first object; 2) the shooting objects in the field of view are focused in the order from right to left, and the shooting object located at the rightmost side of the field of view is determined as the first object.
[0100] After the first object indicated by the first focusing instruction in the shooting objects is focused through step S1022, the first electronic device can acquire the first audio track information corresponding to the first object by using the voice acquisition module through step S1023, and in the case that a second focusing instruction is received, control the voice acquisition module to end the acquisition of the first audio track information, and determine the voice information continuously acquired by the voice acquisition module between the third time when the first focusing instruction is received and the second time when the second focusing instruction is received as the first audio track information corresponding to the first object.
[0101] The first electronic device can perform step S1024 to focus on a second object indicated by the second focusing instruction in the shooting object after receiving the second focusing instruction, and perform step S1025 to acquire second audio track information corresponding to the second object by using the voice acquisition module. In the case where the number of second objects is greater than 1, the first electronic device can repeatedly perform steps S1024 to S1025 until all second audio track information corresponding to the second objects is acquired and / or the state of the shooting button is converted from the pressed state to the released state, and then the acquisition of the second audio track information corresponding to the second objects is completed, and step S1026 is performed to determine the voice information corresponding to the shooting object according to the first audio track information and the second audio track information. It should be noted that the process of the first electronic device performing steps S1024 and S1025 is similar to the process of performing steps S1022 and S1023, and will not be repeated here.
[0102] In step S1026, the first electronic device can combine the first audio track information and the second audio track information in chronological order to obtain the voice information corresponding to the shooting object. It can be understood that the voice information corresponding to the shooting object includes the first audio track information and the second audio track information.
[0103] As an example, the shooting object includes shooting object 1, shooting object 2 and shooting object 3, the first object is shooting object 1, the second object includes shooting object 2 and shooting object 3, and the first focusing instruction is the focusing instruction corresponding to shooting object 1, and the second focusing instruction includes a first sub-instruction corresponding to shooting object 2 and a second sub-instruction corresponding to shooting object 3. After receiving the first focusing instruction, the first electronic device performs step S1022 to focus on shooting object 1, performs step S1023 to acquire first audio track information corresponding to shooting object 1 by using the voice acquisition module, and in the case where the first sub-instruction is received, the acquisition of the first audio track information corresponding to shooting object 1 is completed; then, the first electronic device performs step S1024 to focus on shooting object 2, performs step S1025 to acquire second audio track information corresponding to shooting object 2 by using the voice acquisition module, and in the case where the second sub-instruction is received, the acquisition of the second audio track information corresponding to shooting object 2 is completed; then, the first electronic device performs step S1024 to focus on shooting object 3, performs step S1025 to acquire second audio track information corresponding to shooting object 3 by using the voice acquisition module, and in the case where the state of the shooting button of the first electronic device is converted from the pressed state to the released state, the acquisition of the second audio track information corresponding to shooting object 3 is completed, and thus the first electronic device completes the acquisition of all second audio track information corresponding to the second objects; finally, the first electronic device determines the voice information corresponding to the shooting object according to the first audio track information and the second audio track information.
[0104] The voice photograph shooting method provided in the embodiments of the present application can obtain the voice information corresponding to each shooting object in sequence according to the focusing sequence of each shooting object in the viewfinder range when the number of shooting objects is greater than 1, thereby improving the completeness and richness of the voice information and improving the realizability of the embodiments of the present application.
[0105] Optionally, the step S104 of establishing the association between the shooting object in the target image and the voice information comprises:
[0106] The step S1041 comprises: creating a first voice link identifier in a first target region corresponding to the first object in the target image, and creating a second voice link identifier in a second target region corresponding to the second object.
[0107] The step S1042 comprises: establishing a link relationship between the first voice link identifier and the first audio track information, and establishing a link relationship between the second voice link identifier and the second audio track information, to obtain a target voice photograph.
[0108] In the embodiments of the present application, when the number of shooting objects in the viewfinder range of the image acquisition module is greater than 1, i.e., the number of shooting objects is multiple, after the voice information corresponding to the shooting object is obtained through the steps S1021 to S1026, and the target image is obtained through the step S103, the first electronic device can establish the association between each shooting object and the audio track information corresponding to the shooting object through the steps S1041 to S1042, so that when the target voice photograph is viewed, the first electronic device can play the audio track information corresponding to different shooting objects according to the association between the shooting object and the audio track information corresponding to the shooting object, thereby improving the flexibility and interest of the embodiments of the present application.
[0109] The target image comprises a first target region corresponding to the first object and a second target region corresponding to the second object. It can be understood that when the number of second objects is greater than 1, the number of second target regions is also greater than 1.
[0110] The second target region can be a region in the target image where the first object is located, and the first target region can be a region in the target image close to the first object. Correspondingly, the second target region can be a region in the target image where the second object is located, and the second target region can be a region in the target image close to the second object.
[0111] Specifically, the first electronic device can first create a first voice link identifier in the first target area and a second voice link identifier in the second target area through step S1041; the first voice link identifier is an identifier in the target image for establishing a link relationship between the first object and the first audio track information; the second voice link identifier is an identifier in the target image for establishing a link relationship between the second object and the second audio track information; then, the first electronic device can establish the link relationship between the first voice link identifier and the first audio track information and the link relationship between the second voice link identifier and the second audio track information in turn according to the focusing order of the photographed objects and the acquisition order of the voice information through step S1042, to obtain the target voice photo. It can be understood that the first preset storage area at least stores the target image and the first audio track information and the second audio track information corresponding to the target image.
[0112] In the target voice photo, each photographed object can establish a link relationship with the audio track information corresponding to the photographed object. In the process of viewing the target voice photo, the first electronic device can display the target image in the target voice photo through the display; in the embodiment of the present application, the target image can also include the first voice link identifier and the second voice link identifier; the first electronic device can obtain and play the audio track information corresponding to the voice link identifier clicked by the user based on the link relationship between the first voice link identifier and the first audio track information and the link relationship between the second voice link identifier and the second audio track information in the case that the user clicks any one of the first voice link identifier and the second voice link identifier.
[0113] As an example, referring to Figure 5 , another schematic diagram of a target voice photo provided by an embodiment of the present application is shown, the photographed objects include photographed object 1, photographed object 2 and photographed object 3, the first object is photographed object 1, and the second object includes photographed object 2 and photographed object 3; as shown in Figure 5 , the target image includes photographed object 1, photographed object 2 and photographed object 3, and the first voice link identifier for establishing a link relationship with the first audio track information corresponding to photographed object 1 and the second voice link identifier for establishing a link relationship with the second audio track information corresponding to photographed object 2 and photographed object 3. In the case that the user clicks the first voice link identifier in the process of viewing the target voice photo, the first electronic device can obtain and play the first audio track information corresponding to photographed object 1 from the first preset storage area based on the link relationship between the first voice link identifier and the first audio track information.
[0114] As an optional implementation, after the first electronic device creates the first voice link identifier and the second voice link identifier through step S1041, the first electronic device can further create an overall voice link identifier in a region outside the first target region and the second target region in the target image; after the first electronic device establishes the link relationship between the first voice link identifier and the first audio track information and the link relationship between the second voice link identifier and the second audio track information through step S1042, the first electronic device can further establish a link relationship between voice information corresponding to the shooting object determined according to the first audio track information and the second audio track information and the overall voice link identifier, so that the target image further includes the overall voice link identifier in addition to the shooting object, the first voice link identifier, and the second voice link identifier. It can be understood that the first preset storage region at least stores the target image and the first audio track information and the second audio track information corresponding to the target image, and the voice information determined according to the first audio track information and the second audio track information.
[0115] As an example, referring to Figure 6 , another schematic diagram of a target voice photo provided by an embodiment of the present application is shown, the shooting object includes a shooting object 1, a shooting object 2, and a shooting object 3, the first object is the shooting object 1, and the second object includes the shooting object 2 and the shooting object 3; as shown in Figure 6 , the target image further includes an overall voice link identifier that establishes a link relationship with voice information, in addition to the shooting object 1, the shooting object 2, and the shooting object 3, and the first voice link identifier and the second voice link identifier. In the case that a user clicks the overall voice link identifier in the process of viewing the target voice photo, the first electronic device can acquire and play the voice information from the first preset storage region based on the link relationship between the overall voice link identifier and the voice information.
[0116] Optionally, in the case that the number of the shooting objects is greater than 1, the voice photo shooting method provided by an embodiment of the present application further includes:
[0117] Step A11, in the case that the shooting button of the first electronic device is in a pressed state, the image information in the viewfinder range is acquired by using the image acquisition module.
[0118] The image information refers to continuous image information in the viewfinder range acquired by using the image acquisition module in the process that the first electronic device acquires voice information by using the voice acquisition module. It can be understood that the image information is image information in the viewfinder range continuously acquired by the image acquisition module between a first moment when the state of the shooting button is converted from a released state to a pressed state and a second moment when the state of the shooting button is converted from the pressed state to the released state.
[0119] In the embodiments of the present application, in the case that the number of the photographed objects in the field of view of the image acquisition module is greater than 1, that is, the number of the photographed objects is multiple, in order to assign the audio track information corresponding to different photographed objects in the voice information obtained in step S102 to the corresponding photographed object, the first electronic device can obtain the image information in the field of view by using the image acquisition module in the case that the photographed button of the first electronic device is in the pressed state through step A11 before step S104 is performed.
[0120] Optionally, the step S104 of establishing the association relationship between the photographed object in the target image and the voice information to obtain the target voice photo comprises:
[0121] The step S1043 comprises: performing lip shape analysis on the photographed object based on the image information, and separating the audio track information corresponding to each photographed object in the image information from the voice information.
[0122] The step S1044 comprises: determining the correspondence relationship between the photographed object in the image information and the photographed object in the target image according to the image information and the target image.
[0123] The step S1045 comprises: establishing the association relationship between the photographed object in the target image and the audio track information according to the correspondence relationship to obtain the target voice photo.
[0124] In the embodiments of the present application, in the case that the number of the photographed objects in the field of view of the image acquisition module is greater than 1, that is, the number of the photographed objects is multiple, in order to assign the audio track information corresponding to different photographed objects in the voice information obtained in step S102 to the corresponding photographed object, the first electronic device can obtain the image information in the field of view by using the image acquisition module through step A11, and in the case that the voice information corresponding to the photographed object is obtained through step S102, the association relationship between the photographed object in the target image and the voice information is established through the corresponding operations of step S1043 to step S1045 to obtain the target voice photo before step S104 is performed.
[0125] Specifically, in step S1043, the first electronic device can use a voice segmentation clustering method of mixed features of mel frequency cepstral coefficients and gamma frequency cepstral coefficients based on the image information to realize separation of the human voice, and use AI recognition technology to perform lip shape analysis on the photographed object in the image information, so that the audio track information corresponding to each photographed object in the image information can be separated from the voice information obtained in step S102, so as to match different audio track information to different photographed objects in the image information.
[0126] It should be noted that the content and quantity of the photographed objects in the image information are the same as the content and quantity of the photographed objects in the target image, but the relative positions between the photographed objects in the image information can be the same as or different from the relative positions between the photographed objects in the target image. In order to improve the accuracy of the audio track information distribution, in step S1044, the first electronic device can determine the correspondence between the photographed objects in the image information and the photographed objects in the target image according to the image information and the target image.
[0127] In step S1045, the first electronic device can associate the audio track information corresponding to the photographed object in the image information separated from the voice information in step S1043 to the corresponding photographed object in the target image according to the correspondence between the photographed objects in the image information and the photographed objects in the target image, establish the association relationship between the photographed objects in the target image and the audio track information, and obtain the target voice photo.
[0128] The specific implementation process of associating the audio track information corresponding to the photographed object in the image information separated from the voice information in step S1043 to the corresponding photographed object in the target image, establishing the association relationship between the photographed objects in the target image and the audio track information can refer to the detailed description of steps S1041 to S1042. To avoid repetition, details are not repeated here.
[0129] It should be noted that after the first electronic device establishes the association relationship between the photographed objects in the target image and the audio track information to obtain the target voice photo in step S1045, the first electronic device can delete the image information obtained in step A11 to release the storage space of the first electronic device.
[0130] In the embodiment of the present application, the voice information obtained by step S102 can be voice information composed of audio track information obtained in the focusing order by the method described in steps S1021 to S1026. The voice information obtained by step S102 can be voice information emitted by each photographed object simultaneously between the first moment when the state of the shooting button is converted from the released state to the pressed state and the second moment when the state of the shooting button is converted from the pressed state to the released state.
[0131] The voice photo shooting method provided by the embodiment of the present application is suitable for the application scenarios in which multiple photographed objects in the framing range output voice simultaneously, and the application scenarios in which multiple photographed objects in the framing range output voice in the focusing order. The voice photo shooting method is not limited by the order of the voice output between the photographed objects, ensures the accuracy of the distribution of the audio track information in the voice information to each photographed object, improves the realizability of the embodiment of the present application, and expands the application range of the embodiment of the present application.
[0132] Optionally, in the case that the number of the photographing objects is 1, the step S102 further comprises:
[0133] The step S1027 comprises: in the case that the photographing button of the first electronic device is in the pressed state, acquiring the voice information by using the voice acquisition module, to obtain the voice information corresponding to the photographing object.
[0134] In the embodiment of the present application, in the case that the number of the photographing objects in the field of view of the image acquisition module is 1, the first electronic device can acquire the voice information corresponding to the photographing object by using the voice acquisition module in the case that the photographing button of the first electronic device is in the pressed state, and the voice information corresponding to the photographing object can be acquired by the step S1027.
[0135] Specifically, in the case that there is only one photographing object in the field of view, the first electronic device can acquire the voice information by using the voice acquisition module in the case that the photographing button of the first electronic device is in the pressed state, and the voice information acquired by the voice acquisition module between the first time when the state of the photographing button is converted from the released state to the pressed state and the second time when the state of the photographing button is converted from the pressed state to the released state is determined as the voice information corresponding to the photographing object.
[0136] Optionally, the step S104 further comprises:
[0137] The step S1046 comprises: creating a third voice link identifier in the third target region in the target image.
[0138] The step S1047 comprises: establishing a link relationship between the third voice link identifier and the voice information corresponding to the photographing object, to obtain the target voice photo.
[0139] In the embodiment of the present application, in the case that the number of the photographing objects in the field of view of the image acquisition module is 1, after the voice information corresponding to the photographing object is acquired by the step S1027 and the target image is acquired by the step S103, the first electronic device can establish the association relationship between the photographing object and the voice information by the steps S1046 to S1047, so that the first electronic device can play the voice information corresponding to the photographing object according to the association relationship between the photographing object and the voice information in the case that the user views the target voice photo, and the flexibility and the interest of the embodiment of the present application are improved.
[0140] The third target region can be any region in the target image.
[0141] Specifically, first, the first electronic device creates a third voice link identifier in the third target region through step S1046; wherein the third voice link identifier is an identifier in the target image for establishing a link relationship between the photographed object and the voice information.
[0142] Then, the first electronic device establishes a link relationship between the third voice link identifier and the voice information obtained in step S1027 through step S1047, to obtain the target voice photo. Referring to Figure 7 , another schematic diagram of a target voice photo provided by an embodiment of the present application is shown, as shown in Figure 7 , the target image includes a photographed object and a third voice link identifier.
[0143] The photographing method of the voice photo provided by the embodiment of the present application provides an implementation manner for obtaining voice information corresponding to a photographed object and establishing an association relationship between the photographed object and the voice information in an application scenario where the number of photographed objects is 1, thereby expanding the application range of the embodiment of the present application.
[0144] As an example, referring to Figure 8, a logic block diagram of a voice photo shooting method provided by an embodiment of the present application is shown, specifically: the first electronic device determines a shooting object in the range of view of the image acquisition module in the case that the voice photo shooting mode is enabled; the first electronic device receives an automatic focusing instruction or a manual focusing instruction in the case that it is monitored that the shooting button of the first electronic device is in a pressed state, and focuses on the shooting object in the range of view according to the automatic focusing instruction or the manual focusing instruction; the first electronic device stores the image information in the range of view obtained by the image acquisition module to a first preset storage area in the case that the shooting button is in the pressed state; the first electronic device stores the voice information corresponding to the shooting object obtained by the voice acquisition module to the first preset storage area in the case that the shooting button is in the pressed state; the first electronic device stores the target image corresponding to the range of view obtained by the image acquisition module to the first preset storage area in the case that it is monitored that the state of the shooting button is converted from the pressed state to the released state; the first electronic device performs AI analysis according to the image information, the target image and the voice information in the first preset storage area, establishes the association between the shooting object in the target image and the audio track information, obtains a target voice photo, and stores the target voice photo to the first preset storage area. Wherein, the specific implementation process of the first electronic device performing AI analysis according to the image information, the target image and the voice information in the first preset storage area, and establishing the association between the shooting object in the target image and the audio track information can refer to the detailed description of steps S1043 to S1045, which will not be repeated here.
[0145] As another example, with reference to Figure 9, shows the logic block diagram of another voice photo shooting method provided by the embodiment of the application, specifically: the first electronic device determines the shooting object in the viewfinder range of the image acquisition module in the case that the voice photo shooting mode is enabled; the first electronic device receives the automatic focusing instruction or the manual focusing instruction in the case that the shooting button of the first electronic device is monitored to be in the pressed state, and focuses on the shooting object in the viewfinder range according to the automatic focusing instruction or the manual focusing instruction; the first electronic device acquires the target image corresponding to the viewfinder range by using the image acquisition module in the case that the state of the shooting button is monitored to be converted from the released state to the pressed state, and stores the target image to the first preset storage area; the first electronic device acquires the image information in the viewfinder range by using the image acquisition module in the case that the shooting button is in the pressed state, and stores the image information to the first preset storage area; the first electronic device acquires the voice information corresponding to the shooting object by using the voice acquisition module in the case that the shooting button is in the pressed state, and stores the voice information to the first preset storage area; the first electronic device performs AI analysis according to the image information, the target image and the voice information in the first preset storage area, establishes the association relationship between the shooting object in the target image and the audio track information, obtains the target voice photo, and stores the target voice photo to the first preset storage area. Wherein, the specific implementation process that the first electronic device performs AI analysis according to the image information, the target image and the voice information in the first preset storage area, and establishes the association relationship between the shooting object in the target image and the audio track information can refer to the detailed description of steps S1043 to S1045, which will not be repeated here.
[0146] With reference to Figure 10 , shows the step flow chart of a voice photo display method provided by the embodiment of the application, which specifically includes:
[0147] Step S201, in the case that the display instruction is received, the target voice photo indicated by the display instruction is acquired; the target voice photo includes a target image and voice information.
[0148] Step S202, the target image is displayed by using the display.
[0149] Step S203, the voice information is played by using the voice player.
[0150] The display method of the voice photo provided in the embodiments of the present application can be applied to a second electronic device, and the second electronic device includes a display and a voice player. It can be understood that the second electronic device is any electronic device including a display and a voice player. The second electronic device can include, but is not limited to, a smart terminal, a computer, a personal digital assistant, a tablet computer, an electronic book reader, a laptop computer, a vehicle-mounted device, a smart television, a wearable device, and the like. Specifically, the second electronic device can be a smart home device, for example, a smart phone, a smart television, a smart refrigerator, a smart speaker, and the like.
[0151] In the embodiments of the present application, the second electronic device and the first electronic device can be the same device or different devices.
[0152] Specifically, the second electronic device can obtain a target voice photo indicated by the display instruction in the case where the display instruction is received, and the target voice photo is obtained by the first electronic device through the voice photo shooting method as described above.
[0153] The target voice photo includes a target image and voice information, the target image is a target image corresponding to a shooting range of the image acquisition module obtained by the first electronic device, and the target image includes a shooting object in the shooting range. In the case where the voice photo shooting mode is enabled, the first electronic device can determine the shooting object in the shooting range of the image acquisition module by using the image acquisition module. The voice information is voice information corresponding to the shooting object obtained by the first electronic device by using the voice acquisition module in the case where it is monitored that the shooting button of the first electronic device is in a pressed state. After the voice information and the target image are obtained, the first electronic device establishes an association relationship between the shooting object in the target image and the voice information to obtain the target voice photo obtained by the second electronic device in step S201.
[0154] In the case where the second electronic device and the first electronic device are different electronic devices, the second electronic device can obtain the target voice photo from the first preset storage area in the first electronic device and store the target voice photo in the second preset storage area of the second electronic device before step S201. In step S201, the second electronic device can obtain the target voice photo from the second preset storage area in the case where the display instruction is received.
[0155] The manner in which the second electronic device obtains the target voice photo from the first preset storage area in the first electronic device can include, but is not limited to, obtaining through a communication connection with the first electronic device, obtaining by using a mobile storage medium, and the like.
[0156] The display instruction is an instruction sent by the user to the second electronic device by clicking an icon or a photo name corresponding to the target voice photo in the photo album of the second electronic device, and the display instruction is used to instruct the second electronic device to obtain and display the target voice photo.
[0157] In the process of displaying the target voice photo, the second electronic device can display the target image in the target voice photo by using the display in step S202, and play the voice information in the target voice photo by using the voice player in step S203.
[0158] In the embodiments of the present application, the second electronic device can execute step S203 at the same time as executing step S202, or the second electronic device can first execute step S202, and then execute step S203 after receiving the voice playing instruction.
[0159] Optionally, in some embodiments, the target image of the target voice photo includes a shooting object, a voice link identifier establishing a link relationship with the audio track information corresponding to the shooting object, and an overall voice link identifier establishing a link relationship with the voice information. The voice information includes the audio track information corresponding to the shooting object.
[0160] In the process of the user viewing the target voice photo by using the second electronic device, the user can send a voice playing instruction to the second electronic device by clicking the voice link identifier establishing a link relationship with the audio track information corresponding to the shooting object in the target image, or clicking the overall voice link identifier establishing a link relationship with the voice information in the target image. In the case that the user sends a voice playing instruction to the second electronic device by clicking the voice link identifier establishing a link relationship with the audio track information corresponding to the shooting object, the second electronic device can obtain the audio track information establishing a link relationship with the voice link identifier from the second preset cache area, and play the audio track information. In the case that the user sends a voice playing instruction to the second electronic device by clicking the overall voice link identifier establishing a link relationship with the voice information, the second electronic device can obtain the voice information establishing a link relationship with the overall voice link identifier from the second preset cache area, and play the voice information.
[0161] It can be understood that after the second electronic device stores the target voice photo into the second preset storage area of the second electronic device, the second preset storage area stores the target image in the target voice photo, the audio track information establishing a link relationship with the voice link identifier corresponding to the shooting object in the target image, and the voice information establishing a link relationship with the overall voice link identifier in the target image.
[0162] Optionally, in the process that the user views the target voice photo through the second electronic device, the user can long press the voice link identifier in the target image corresponding to the audio track information establishing a link relationship or the whole voice link identifier establishing a link relationship with the voice information to adjust the sound size of the voice player playing the voice information or the audio track information or directly mute the voice bar.
[0163] Optionally, in the process that the user views the target voice photo through the second electronic device, the user can long press the voice link identifier in the target image corresponding to the audio track information establishing a link relationship or the whole voice link identifier establishing a link relationship with the voice information to select to convert the voice information or the audio track information into text.
[0164] The display method of the voice photo provided by the embodiment of the application, the second electronic device acquires the target voice photo indicated by the display instruction in the case that the display instruction is received, and displays the target image in the target voice photo by using the display and plays the voice information in the target voice photo by using the voice player. The embodiment of the application improves the richness of the recording content of the target voice photo, reduces the occupied amount of the memory space of the second electronic device, improves the accommodation and ease of use of the memory space of the second electronic device in the embodiment of the application, and further improves the interest of the user in displaying the voice photo through the second electronic device, and improves the life quality and use experience of the user.
[0165] Device embodiment
[0166] As shown in Figure 11 , Figure 11 A logic block diagram of a voice photo shooting device provided by the embodiment of the application is shown, which is applied to a first electronic device, and the first electronic device includes an image acquisition module and a voice acquisition module. The device can include:
[0167] A determination module 1101 is configured to determine a shooting object in a viewfinder range of the image acquisition module in the case that a voice photo shooting mode is enabled.
[0168] A first acquisition module 1102 is configured to acquire voice information corresponding to the shooting object by using the voice acquisition module in the case that it is monitored that a shooting button of the first electronic device is in a pressed state.
[0169] A second acquisition module 1103 is configured to acquire a target image corresponding to the viewfinder range by using the image acquisition module. The target image includes the shooting object.
[0170] The association module 1104 is configured to establish an association between the shooting object in the target image and the voice information, and obtain a target voice photo; the target voice photo includes the target image and the voice information.
[0171] Optionally, in a case where the number of the shooting objects is greater than 1, the first acquisition module comprises:
[0172] The receiving sub-module is configured to receive a first focusing instruction in a case where it is monitored that a shooting button of the first electronic device is in a pressed state.
[0173] The first focusing sub-module is configured to focus on a first object indicated by the first focusing instruction in the shooting object.
[0174] The first acquisition sub-module is configured to acquire first audio track information corresponding to the first object by using the voice collection module until a second focusing instruction is received.
[0175] The second focusing sub-module is configured to focus on a second object indicated by the second focusing instruction in the shooting object.
[0176] The second acquisition sub-module is configured to acquire second audio track information corresponding to the second object by using the voice collection module.
[0177] The first determination sub-module is configured to determine voice information corresponding to the shooting object according to the first audio track information and the second audio track information.
[0178] Optionally, the association module comprises:
[0179] The first creation sub-module is configured to create a first voice link identifier in a first target region corresponding to the first object in the target image, and create a second voice link identifier in a second target region corresponding to the second object.
[0180] The first establishment sub-module is configured to establish a link relationship between the first voice link identifier and the first audio track information, and establish a link relationship between the second voice link identifier and the second audio track information, to obtain a target voice photo.
[0181] Optionally, in a case where the number of the shooting objects is greater than 1, the device further comprises:
[0182] The fourth acquisition module is configured to acquire image information in the framing range by using the image collection module in a case where it is monitored that a shooting button of the first electronic device is in a pressed state.
[0183] Optionally, the association module comprises:
[0184] The audio track separation submodule is used to perform lip-sync analysis on the shooting object based on the image information, and separate the audio track information corresponding to each shooting object in the image information from the voice information;
[0185] The second determining submodule is used to determine the correspondence between the shooting object in the image information and the shooting object in the target image based on the image information and the target image;
[0186] The second submodule is used to establish the association between the photographed object in the target image and the audio track information according to the correspondence, so as to obtain the target audio photograph.
[0187] Optionally, when the number of subjects to be photographed is 1, the first acquisition module includes:
[0188] The third acquisition submodule is used to acquire voice information using the voice acquisition module when the camera button of the first electronic device is detected to be pressed, thereby obtaining the voice information corresponding to the subject being photographed.
[0189] Optionally, the associated module includes:
[0190] The second creation submodule is used to create a third voice link identifier in the third target region of the target image;
[0191] The third submodule is used to establish a link between the third voice link identifier and the voice information corresponding to the shooting object, so as to obtain the target voice photo.
[0192] like Figure 12 As shown, Figure 12 This illustration shows a logic block diagram of a voice-activated photo display device according to an embodiment of this application, applied to a second electronic device, the second electronic device including a display and a voice player; the device may include:
[0193] The third acquisition module 1201 is used to acquire a target voice photograph indicated by a display instruction upon receiving a display instruction; the target voice photograph is obtained by the voice photograph shooting method described above; the target voice photograph includes a target image and voice information;
[0194] Display module 1202 is used to display the target image using the display screen;
[0195] The playback module 1203 is used to play the voice information using the voice player.
[0196] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0197] The various embodiments in the specification are described in progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between various embodiments can be referred to each other.
[0198] The electronic device provided by the embodiment of the present application comprises an image acquisition module, a voice acquisition module, a display, a voice player, a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface complete the communication between each other through the communication bus; the memory is used for storing executable instructions, and the executable instructions make the processor execute the photographing method of the voice photo as described above, or execute the display method of the voice photo as described above.
[0199] The embodiment of the present application further provides a non-transitory computer readable storage medium, when the instructions in the readable storage medium are executed by the processor of the electronic device, the processor can execute the photographing method of the voice photo as described above, or execute the display method of the voice photo as described above.
[0200] The various embodiments in the specification are described in progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between various embodiments can be referred to each other.
[0201] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device or computer program product. Therefore, the embodiments of the present application can adopt a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can adopt a computer program product in the form of one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes.
[0202] The embodiments of the present application are described with reference to flowcharts and / or block diagrams according to the method, terminal device (system) and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal device produce a device for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The device for implementing the functions specified in one flow or multiple flows and / or blocks Figure 1 The device for implementing the functions specified in one flow or multiple flows and / or blocks
[0203] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flow Figure 1 one or more flow or block Figure 1 one or more blocks or blocks specified in the flow.
[0204] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 one or more flow or block Figure 1 one or more blocks or blocks specified in the flow.
[0205] While preferred embodiments of the application have been described, those skilled in the art will appreciate that other modifications and variations are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the application, equivalents can be substituted for elements recited herein. Further, those skilled in the art will appreciate that not all combinations of components recited herein are necessarily the most preferred combinations. Accordingly, the appended claims are intended to embrace all such alternatives, modifications and variations as falling within the scope of the present application.
[0206] Finally, it is to be understood that the phraseology or terminology employed herein, such as "first" and "second", etc., are for descriptive purposes only and should not be construed to be indicative of a necessary order or sequence. Also, the use of "including", "containing", or "comprising" and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Further, the terms "a" and "an" herein do not denote a limitation of quantity, but rather denote the presence of at least one of the referenced items. Further, the term "plurality" herein does not denote a limitation of quantity, but rather denotes a quantity of at least two.
[0207] The voice photo shooting method, display method, device, equipment and storage medium provided by the present application are described in detail above, and the principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present application; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges can be changed, and the above description of the present application should not be understood as a limitation.
Claims
1. A photographing method of a voice photo, characterized by, The method is applied to a first electronic device, and the first electronic device comprises an image acquisition module and a voice acquisition module. In a case where a voice photo shooting mode is enabled, a shooting object in a viewfinder range of the image acquisition module is determined. In a case where a shooting button of the first electronic device is detected to be in a pressed state, voice information corresponding to the shooting object is acquired by using the voice acquisition module. A target image corresponding to the viewfinder range is acquired by using the image acquisition module, and the target image comprises the shooting object. An association between the shooting object in the target image and the voice information is established, and a target voice photo is obtained. The target voice photo comprises the target image and the voice information. In a case where the number of the shooting objects is greater than one, the method further comprises: In a case where the shooting button of the first electronic device is detected to be in the pressed state, a first focusing instruction is received. A first object indicated by the first focusing instruction in the shooting object is focused. First audio track information corresponding to the first object is acquired by using the voice acquisition module until a second focusing instruction is received. A second object indicated by the second focusing instruction in the shooting object is focused. Second audio track information corresponding to the second object is acquired by using the voice acquisition module. The voice information corresponding to the shooting object is determined according to the first audio track information and the second audio track information.
2. The method of claim 1, wherein, The association between the shooting object in the target image and the voice information is established, and the target voice photo is obtained, comprising: A first voice link identifier is created in a first target area corresponding to the first object in the target image, and a second voice link identifier is created in a second target area corresponding to the second object. A link relationship between the first voice link identifier and the first audio track information is established, and a link relationship between the second voice link identifier and the second audio track information is established, and a target voice photo is obtained.
3. The method of claim 1, wherein, In a case where the number of the shooting objects is greater than one, the method further comprises: In a case where the shooting button of the first electronic device is detected to be in the pressed state, image information in the viewfinder range is acquired by using the image acquisition module.
4. The method of claim 3, wherein, The association between the shooting object in the target image and the voice information is established, and the target voice photo is obtained, comprising: Mouth shape analysis is performed on the shooting object based on the image information, and audio track information corresponding to each shooting object in the image information is separated from the voice information. According to the image information and the target image, a corresponding relationship between the shooting object in the image information and the shooting object in the target image is determined. According to the corresponding relationship, an association between the shooting object in the target image and the audio track information is established, and a target voice photo is obtained.
5. A display method of a voice photo, characterized by, The method is applied to a second electronic device, and the second electronic device comprises a display and a voice player. In a case where a display instruction is received, a target voice photo indicated by the display instruction is acquired; the target voice photo is obtained by using the voice photo shooting method in any one of claims 1 to 4; and the target voice photo comprises a target image and voice information. The target image is displayed by using the display. The voice information is played by using the voice player.
6. A voice photo shooting apparatus, characterized by comprising: The device is applied to a first electronic device, and the first electronic device comprises an image acquisition module and a voice acquisition module. A determination module is configured to determine a shooting object in a viewfinder range of the image acquisition module in a case where a voice photo shooting mode is enabled. A first acquisition module is configured to acquire voice information corresponding to the shooting object by using the voice acquisition module in a case where it is monitored that a shooting button of the first electronic device is in a pressed state. A second acquisition module is configured to acquire a target image corresponding to the viewfinder range by using the image acquisition module; and the target image comprises the shooting object. An association module is configured to establish an association between the shooting object in the target image and the voice information, and obtain a target voice photo; and the target voice photo comprises the target image and the voice information. In a case where the number of the shooting objects is greater than 1, the first acquisition module comprises: A receiving sub-module is configured to receive a first focusing instruction in a case where it is monitored that the shooting button of the first electronic device is in the pressed state. A first focusing sub-module is configured to focus on a first object indicated by the first focusing instruction in the shooting object. A first acquisition sub-module is configured to acquire first audio track information corresponding to the first object by using the voice acquisition module until a second focusing instruction is received. A second focusing sub-module is configured to focus on a second object indicated by the second focusing instruction in the shooting object. A second acquisition sub-module is configured to acquire second audio track information corresponding to the second object by using the voice acquisition module. A first determination sub-module is configured to determine the voice information corresponding to the shooting object according to the first audio track information and the second audio track information.
7. A display device of a voice photo, characterized by, The device is applied to a second electronic device, and the second electronic device comprises a display and a voice player. A third acquisition module is configured to acquire a target voice photo indicated by a display instruction in a case where the display instruction is received; the target voice photo is obtained by using the voice photo shooting method in any one of claims 1 to 4; and the target voice photo comprises a target image and voice information. A display module is configured to display the target image by using the display. A playing module is configured to play the voice information by using the voice player.
8. An electronic device, comprising: The electronic device comprises an image acquisition module, a voice acquisition module, a display, a voice player, a processor, a memory, a communication interface and a communication bus; the processor, the memory and the communication interface complete communication with each other through the communication bus. The memory is configured to store executable instructions, and the executable instructions cause the processor to perform the photographing method of the voice photo according to any one of claims 1 to 4, or perform the display method of the voice photo according to claim 5.
9. A readable storage medium, characterized by, When the instructions in the readable storage medium are executed by the processor of the electronic device, the processor can perform the photographing method of the voice photo according to any one of claims 1 to 4, or perform the display method of the voice photo according to claim 5.
Citation Information
Patent Citations
Photographing method and mobile terminal
CN106791442A
Method, device, and system for sharing photographs on basis of voice recognition
WO2019112145A1