Electronic device, its control method, program, and recording medium

The electronic device facilitates high-quality content generation by allowing users to select and combine video and audio from multiple sources based on shooting positions and quality, addressing the limitations of existing methods.

JP7721303B2Active Publication Date: 2025-08-12CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021060114
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-31
Publication Date
2025-08-12
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

Existing methods for combining audio and video content do not facilitate easy generation of high-quality content, requiring re-capture or re-recording if the quality is found to be low.

Method used

An electronic device that allows users to select and synthesize video or audio content from multiple sources, displaying shooting positions and quality information, enabling high-quality content generation through user-controlled combination.

Benefits of technology

Enables easy creation of high-quality content by allowing users to select and combine high-quality video and audio components from multiple sources, improving the overall media production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007721303000001
    Figure 0007721303000001
  • Figure 0007721303000002
    Figure 0007721303000002
  • Figure 0007721303000003
    Figure 0007721303000003
Patent Text Reader

Abstract

To provide a technology with which it is possible to easily generate a high-quality content.SOLUTION: An electronic device according to the present invention includes: selection means configured to select a moving image or a sound included in at least any of a plurality of second contents respectively photographed by a plurality of second cameras different from a first camera; and acquisition means configured to, in a case where a moving image included in at least any of the plurality of second contents is selected, acquire a third content including the selected moving image and a sound included in a first content photographed by the first camera, and in a case where a sound included in at least any of the plurality of second contents is selected, acquire a fourth content including the selected sound and a moving image included in the first content.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for synthesizing moving images and audio. [Background technology]

[0002] Conventionally, media content has been widely shared using social networking sites (SNS) and cloud services. Content is also being created by combining multiple elements. For example, Patent Document 1 discloses a technique for combining audio captured from a microphone with an image captured from a camera. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-220848 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the method disclosed in Patent Document 1, audio captured from a predetermined microphone is combined with an image captured from a predetermined camera. Therefore, if the user realizes after capturing an image that the image is of low quality, the user must re-capture the image. Similarly, if the user realizes after recording that the audio is of low quality, the user must re-record the audio. In other words, the method disclosed in Patent Document 1 does not easily allow for the generation of high-quality content.

[0005] An object of the present invention is to provide a technique that enables high-quality content to be easily generated. [Means for solving the problem]

[0006] The electronic device of the present invention comprises: a first selection means for selecting a first content acquired by a first camera in response to a user operation; a second selection means for selecting whether to synthesize a video or an audio with the first content in response to a user operation, separately from the selection of the content to be synthesized; A plurality of second contents each captured by a plurality of second cameras different from the first camera. a control means for controlling the display of a plurality of items corresponding to the plurality of second contents, respectively; a third selection means for selecting a second content to be combined with the first content from the plurality of second contents in response to a user operation; When a video is selected, Included in the second content selected by the third selection means Video and The selected by the first selection means The audio included in the first content Synthesize Third content Generate death, The second selection means When Audio is selected, Included in the second content selected by the third selection means Audio and The selected by the first selection means The video included in the first content Synthesize The fourth content Generate do Generate means and 、 With The control means controls the display device to further display items indicating the subject, the shooting position of the first content, and the shooting positions of each of the plurality of second contents. It is characterized by: [Effects of the Invention]

[0007] According to the present invention, high-quality content can be easily generated. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is an external view of a digital camera. [Figure 2] FIG. 1 is a block diagram of a digital camera. [Figure 3] FIG. 1 is a configuration diagram of a cloud system. [Figure 4] FIG. 1 is a diagram showing a shooting scene and video content. [Figure 5] FIG. 10 is a diagram illustrating an application screen. [Figure 6] FIG. 10 is a diagram illustrating an application screen. [Figure 7] 10 is a flowchart of the overall synthesis process and the synthesis process. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, a preferred embodiment of the present invention will be described with reference to the drawings. 1 is an external view of a digital camera 100 (imaging device). A display unit 28 displays images and various information. A mode switch 60 is an operation unit for switching between various modes. A connector 112 is a connector between the digital camera 100 and a connection cable 111 for connecting to an external device such as a personal computer or printer. An operation unit 70 is an operation unit including various switches, buttons, a touch panel, and other operation members that accept various operations from the user. A controller wheel 73 is a rotatable operation member included in the operation unit 70. A power switch 72 is a push button for switching the power on and off. A recording medium 200 is a recording medium such as a memory card. A recording medium slot 201 is a slot for storing the recording medium 200. The recording medium 200 stored in the recording medium slot 201 can communicate with the digital camera 100. When the recording medium 200 is stored in the recording medium slot 201, it becomes possible to record images to the recording medium 200 and play back images recorded on the recording medium 200. A lid 202 is a lid for the recording medium slot 201. FIG. 1 shows a state in which the cover 202 is opened and a part of the recording medium 200 is removed from the slot 201 and exposed.

[0010] FIG. 2 is a block diagram showing an example configuration of a digital camera 100. In FIG. 2, the photographing lens 103 is a lens group including a zoom lens and a focus lens. The shutter 101 is a shutter with an aperture function. The imaging unit 22 is an imaging element (image sensor) composed of a CCD, CMOS element, or the like that converts an optical image into an electrical signal. The A / D converter 23 converts an analog signal into a digital signal. The A / D converter 23 is used to convert the analog signal output from the imaging unit 22 into a digital signal. The barrier 102 covers the imaging system of the digital camera 100, including the photographing lens 103, to prevent the imaging system, including the photographing lens 103, shutter 101, and imaging unit 22, from getting dirty or being damaged.

[0011] The image processing unit 24 performs predetermined pixel interpolation, resizing such as reduction, and color conversion processing on the data from the A / D converter 23 or the data from the memory control unit 15. The image processing unit 24 also performs predetermined arithmetic processing using the captured image data. The system control unit 50 performs exposure control and distance measurement control based on the arithmetic results obtained by the image processing unit 24. This allows TTL (through-the-lens) type AF (autofocus) processing, AE (autoexposure) processing, and EF (flash pre-flash) processing to be performed. The image processing unit 24 also performs predetermined arithmetic processing using the captured image data, and performs TTL type AWB (auto white balance) processing based on the arithmetic results obtained.

[0012] The output data from the A / D converter 23 is written into the memory 32 via the image processing unit 24 and the memory control unit 15, or directly via the memory control unit 15. The memory 32 stores image data obtained by the imaging unit 22 and converted into digital data by the A / D converter 23, as well as image data to be displayed on the display unit 28. The memory 32 has a storage capacity sufficient to store a predetermined number of still images and a predetermined period of moving images and audio.

[0013] The memory 32 also serves as a memory (video memory) for image display. The D / A converter 13 converts the image display data stored in the memory 32 into an analog signal and supplies it to the display unit 28. In this way, the display image data written to the memory 32 is displayed by the display unit 28 via the D / A converter 13. The display unit 28 displays an image on a display device such as an LCD in accordance with the analog signal from the D / A converter 13. The digital signal that has been A / D converted by the A / D converter 23 and stored in the memory 32 is D / A converted by the D / A converter 13 into an analog signal, which is then transferred to and displayed on the display unit 28. By performing this sequentially, a function as an electronic viewfinder can be realized, and a through-image display (live view display (LV display)) can be performed. Hereinafter, the image displayed in live view display will be referred to as a live view image (LV image).

[0014] The nonvolatile memory 56 is a memory serving as an electrically erasable and recordable recording medium, and may be, for example, an EEPROM. The nonvolatile memory 56 stores constants, programs, etc. for the operation of the system control unit 50. The programs referred to here are computer programs for executing various flowcharts described later in this embodiment.

[0015] The system control unit 50 is a control unit made up of at least one processor and / or at least one circuit, and controls the entire digital camera 100. The system control unit 50 executes programs recorded in the nonvolatile memory 56 described above to realize each process of this embodiment, which will be described later. The system memory 52 may be, for example, a RAM. Constants and variables for the operation of the system control unit 50, programs read from the nonvolatile memory 56, and the like are loaded into the system memory 52. The system control unit 50 also performs display control by controlling the memory 32, D / A converter 13, display unit 28, etc.

[0016] The system timer 53 is a timekeeping unit that measures the time used for various controls and the time of a built-in clock.

[0017] The sound collection unit 61 collects sound and inputs the obtained sound data to the sound processing unit 62. The sound collection unit 61 has a microphone, a conversion unit that converts the sound received by the microphone into sound data, etc. The sound processing unit 62 performs noise reduction processing, amplification processing, etc. on the sound data input from the sound collection unit 61.

[0018] The geomagnetic sensor 63 detects the vertical and horizontal components of the geomagnetism and detects the angle between a reference direction based on north and the optical axis of the imaging unit 22 as an azimuth angle, thereby detecting the imaging direction (photographing direction) of the digital camera 100. The geomagnetic sensor 63 is configured using, for example, an acceleration sensor, a gyro sensor, etc.

[0019] The GPS receiver 64 measures geographic information using artificial satellites. For example, the GPS receiver 64 transmits a signal to an artificial satellite and receives a response. The GPS receiver 64 then determines the geographical position (for example, latitude or longitude The GPS receiver 64 can identify the shooting location, etc.

[0020] The communication unit 65 is a wireless or wired The communication unit 65 transmits and receives video signals and audio signals to and from external devices connected via a cable. The communication unit 65 can also connect to a wireless LAN (Local Area Network) or the Internet. The communication unit 65 can also communicate with external devices via Bluetooth (registered trademark) or Bluetooth Low Energy. The communication unit 65 can transmit images (including live images) captured by the imaging unit 22 and images recorded on the recording medium 200 to external devices such as cloud data storage 250, and can also receive image data and various other information from external devices.

[0021] The mode selector switch 60 and the operation unit 70 are operation means for inputting various operational instructions to the system control unit 50. The mode selector switch 60 switches the operation mode of the system control unit 50 to one of still image recording mode, video shooting mode, playback mode, etc. Modes included in the still image recording mode include auto shooting mode, auto scene determination mode, manual mode, aperture priority mode (Av mode), shutter speed priority mode (Tv mode), and program AE mode. There are also various scene modes and custom modes that provide shooting settings for each shooting scene. The mode selector switch 60 allows the user to directly switch to one of these modes. Alternatively, after switching to a list screen of shooting modes with the mode selector switch 60, the user may selectively switch to one of the multiple displayed modes using other operation members. Similarly, the video shooting mode also has multiple Modes may also be included.

[0022] Each operating member of the operating unit 70 is assigned a function appropriate for each situation by selecting and operating various function icons displayed on the display unit 28, and acts as various function buttons. The function buttons include, for example, an end button, a back button, an image forward button, a jump button, a filter button, and an attribute change button. For example, when the menu button is pressed, a menu screen on the display unit 28 on which various settings can be made is displayed. The user can intuitively make various settings using the menu screen displayed on the display unit 28, the four directional buttons (up, down, left, and right), and the SET button.

[0023] The controller wheel 73 is a rotatable operating member included in the operation unit 70, and is used together with the directional buttons to indicate a selection item. When the controller wheel 73 is rotated, an electrical pulse signal is generated according to the amount of rotation, and the system control unit 50 controls each unit of the digital camera 100 based on this pulse signal. This pulse signal can determine the angle at which the controller wheel 73 is rotated and the number of rotations made. The controller wheel 73 may be any operating member that can detect a rotation operation. For example, the controller wheel 73 may be a dial operating member that rotates itself in response to a user's rotation operation to generate a pulse signal. Alternatively, the controller wheel 73 may be an operating member made of a touch sensor that does not rotate itself but detects the rotation of the user's finger on the controller wheel 73 (a so-called touch wheel).

[0024] The power supply control unit 80 is composed of a battery detection circuit, a DC-DC converter, a switch circuit for switching between powered blocks, etc., and detects whether a battery is installed, the battery type, and the remaining battery charge. The power supply control unit 80 also controls the DC-DC converter based on the detection results and instructions from the system control unit 50, and supplies the required voltage for the required period to each unit, including the recording medium 200. The power supply unit 30 is composed of primary batteries such as alkaline batteries or lithium batteries, secondary batteries such as NiCd batteries, NiMH batteries, or Li batteries, an AC adapter, etc.

[0025] The recording medium I / F 18 is an interface with a recording medium 200 such as a memory card. The recording medium 200 is a recording medium such as a memory card for recording captured images, and is composed of a semiconductor memory, an optical disk, a magnetic disk, or the like.

[0026] The digital camera 100 has a touch panel 70a, which is one of the operation members included in the operation unit 70, that can detect touch operations on the display unit 28. The touch panel 70a and the display unit 28 can be configured as an integrated unit. For example, the touch panel 70a is configured so that its light transmittance does not interfere with the display on the display unit 28, and is attached to the upper layer of the display surface of the display unit 28. Input coordinates on the touch panel 70a are then associated with display coordinates on the display screen of the display unit 28. This makes it possible to provide a GUI (Graphical User Interface) that allows the user to directly operate the screen displayed on the display unit 28. The system control unit 50 can detect the following operations or states on the touch panel 70a.

[0027] A finger or pen that has not been touching the touch panel 70a touches the touch panel 70a again. In other words, the start of touching (hereinafter referred to as touch-down). A state in which the touch panel 70a is touched with a finger or a pen (hereinafter referred to as Touch-On). A finger or a pen is moved while touching the touch panel 70a (hereinafter referred to as Touch-Move). The finger or pen that has been touching the touch panel 70a is released from the touch panel 70a. In other words, the touch ends (hereinafter referred to as "touch-up"). A state in which nothing is touching the touch panel 70a (hereinafter referred to as Touch-Off)

[0028] When a touch down is detected, a touch on is also detected at the same time. After a touch down, a touch on is usually continued to be detected unless a touch up is detected. A touch move is also detected when a touch on is detected. Even if a touch on is detected, a touch move will not be detected unless the touch position has moved. Once it is detected that all fingers or pens that were touching have touched up, a touch off occurs.

[0029] These operations and states, as well as the coordinates of the position where the finger or pen touches the touch panel 70a, are notified to the system control unit 50 via the internal bus. The system control unit 50 then determines what kind of operation (touch operation) was performed on the touch panel 70a based on the notified information. Regarding touch-move, the direction of movement of the finger or pen moving on the touch panel 70a can also be determined for each vertical and horizontal component on the touch panel 70a based on changes in the position coordinates. A slide operation is determined to have been performed when a touch-move of a predetermined distance or more is detected. An operation in which a finger is touched on the touch panel 70a, quickly moved a certain distance, and then released is called a flick. A flick is, in other words, an operation in which the finger quickly traces the touch panel 70a as if flicking it. When a touch-move of a predetermined distance or more at a predetermined speed or faster is detected and a touch-up is then detected, a flick is determined to have been performed (a flick can be determined to have followed a slide operation). Furthermore, a touch operation in which multiple points (for example, two points) are touched together and the touch positions are brought closer together is called a pinch in, and a touch operation in which the touch positions are moved farther apart is called a pinch out. Pinch out and pinch in are collectively called a pinch operation (or simply a pinch). The touch panel 70a may be of any of various touch panel types, such as a resistive film type, a capacitance type, a surface acoustic wave type, an infrared type, an electromagnetic induction type, an image recognition type, or an optical sensor type. There are types that detect a touch by contact with the touch panel, and types that detect a touch by the approach of a finger or pen to the touch panel, and either type is acceptable.

[0030] The cloud data storage 250 is capable of storing information such as image data, and is capable of sending and receiving information to and from the communication unit 65 of the digital camera 100 .

[0031] Fig. 3 is a diagram showing an example of the configuration of a cloud system according to this embodiment. In the cloud system of Fig. 3, a cloud server 300 includes a cloud control unit 301, a cloud storage unit 302, a cloud display control unit 303, a cloud communication unit 304, and the like. The cloud server 300 is connected to a global network 320. In Fig. 3, a camera (digital camera) 100, a smartphone 330, and a PC (personal computer) 360 are also connected to the global network 320. The global network 320 is connected to an NTP (Network Time Protocol) server 321, and devices connected to the global network 320 can synchronize their time with the NTP server 321.

[0032] The present invention is applicable to the camera 100, the cloud server 300, the smartphone 330, the PC 360, etc. The present invention can also be understood as a cloud system including the cloud server 300 and electronic devices (terminals) connected to the cloud server 300.

[0033] Cloud communication unit 304 receives signals from each device connected to cloud server 300, such as camera 100, smartphone 330, and PC 360, via global network 320, and converts the signals into control signals for cloud server 300. It also transmits information about cloud server 300 to each device. Cloud storage unit 302 is composed of ROM 302A and RAM 302B, and includes programs that operate cloud server 300 and work memory for storing and processing audio data and image data. Cloud display control unit 303 controls information displayed on a video output device (not shown) of cloud server 300 and each device connected to cloud server 300. One example of a method for displaying information on each device is a method using a browser. Cloud control unit 301 controls the entire cloud server 300 based on signals transmitted and received from an input device (not shown) of cloud server 300, cloud storage unit 302, cloud display control unit 303, and cloud communication unit 304.

[0034] Cloud server 300 is connected to cloud data storage 250. Cloud server 300 can store data from each terminal connected to cloud server 300 in cloud data storage 250, and can transmit data stored in cloud data storage 250 to each terminal.

[0035] Cloud server 300 is connected to image / audio separation unit 371 and image / audio synthesis unit 372. Image / audio separation unit 371 separates data stored in cloud data storage 250 or cloud storage unit 302 into image data and audio data. Each separated data is stored in cloud data storage 250 or cloud storage unit 302. Image / audio synthesis unit 372 synthesizes image data and audio data stored in cloud data storage 250 or cloud storage unit 302 into synthesized data consisting of audio and image. The synthesized data is stored in cloud data storage 250 or cloud storage unit 302. Note that image / audio separation unit 371 and image / audio synthesis unit 372 may be part of cloud server 300.

[0036] FIG. 4(A) shows a fireworks venue, which is an example of a location where photography according to this embodiment is performed. Subject 401 (fireworks) is photographed as video content by photographer 411 at point A, photographer 412 at point B, photographer 413 at point C, and photographer 414 at point D. Video content can include video and audio. In this embodiment, photographer 411 photographs with camera 100, and photographer 412 photographs with smartphone 330. Each photographer accesses cloud server 300 and saves the video content in cloud data storage 250. In addition, the video content photographed by each photographer includes information on the shooting position and shooting direction (posture) as meta information.

[0037] Fig. 4(B) shows video content 421 captured by photographer 411, and Fig. 4(C) shows video content 431 captured by photographer 412. Video content 421 is made up of video 422 and audio 423. Video content 431 is made up of video 432 and audio 433. In video 432, the fireworks extend beyond the maximum angle of view due to incorrect settings on smartphone 330, resulting in low video quality. On the other hand, video 422 captured by photographer 411 has the fireworks captured within the entire angle of view, resulting in good video quality.

[0038] FIG. 5(A) shows an example of an application screen 500 of the smartphone 330. The application screen 500 is, for example, a browser screen or a dedicated application screen. Here, the smartphone 330 corresponds to the first camera of the present invention, and the other device (terminal) used for taking the image corresponds to the second camera of the present invention. The application screen 500 is displayed based on a control signal from the cloud display control unit 303. The smartphone 33 An input device such as a touch panel (not shown) of the smartphone 330 can be used to operate the application screen 500. The smartphone 330 can transmit a control signal to the cloud control unit 301 in response to an operation on the application screen 500.

[0039] The application screen 500 is composed of multiple components. The file selection button 501 is a button for selecting video content shot by the smartphone 330. The location display window 510 is a window that displays the shooting locations of multiple video content stored in the cloud data storage 250. The location display window 510 also displays the location of the subject 401 (sound source). The shooting location of the video content 502 selected by the file selection button 501 is also displayed in an identifiable manner as the shooting location of the video content shot by the smartphone 330. When video content is selected by the file selection button 501, the cloud control unit 301 acquires (data of) the selected video content from the cloud data storage 250. The acquired data (information) includes not only video data and audio data, but also information on the shooting time and shooting location of the video content 502. The acquired data (information) may also include information on the shooting direction of the video content 502. Based on the acquired data (information on the shooting position of video content 502), cloud display control unit 303 displays the shooting position of video content 502 in position display window 510. Similarly, data of other video content is acquired, and the shooting positions of the other video content are displayed in position display window 510. The position of subject 401 is set in advance.

[0040] Data of video content 502 is also displayed in user content information window 540 under the control of cloud display control unit 303. In user content information window 540, reduced images (thumbnails) of the video are displayed as items (image information) related to the video of video content 502. Specifically, thumbnails of multiple frames within specified period 561, which is the period specified by seek bar 560, are displayed. Specified period 561 can be changed by a touch operation by the user. In addition, a waveform of the audio is also displayed as an item (audio information) related to the audio of video content 502. The video data and audio data of video content 502 are obtained by image / audio separation unit 371 and stored in cloud storage unit 302.

[0041] Audio / image selection button 520 is a button for selecting whether to synthesize audio or video for video content 502. Video content 502 has fireworks that extend beyond the maximum angle of view, resulting in low video quality. In such a case, the user selects "Image (synthesize video)" with audio / image selection button 520. Therefore, the case of synthesizing video will be explained here. The case of synthesizing audio will be explained later.

[0042] The selected content information window 530 is a window that displays data of the video content other than video content 502 from among the video content displayed in the position display window 510. As with the user content information window 540, thumbnails of multiple frames in the specified period 561 are displayed. Furthermore, the selected content information window 530 displays a composite selection cursor 531 and display switch buttons 532 and 533. The composite selection cursor 531 is a cursor for selecting video content to be composited with the video content 502, and the selection result window 534 displays the result of the selection made using the composite selection cursor 531 (for example, the identifier of the selected video content). The display switch buttons 532 and 533 are buttons that are used when only some of the video content can be displayed in the selected content information window 530, and are buttons for switching the video content displayed in the selected content information window 530.

[0043] The content confirmation button group 550 includes a play / pause button 551, a stop button 552, a fast-forward button 553, and a fast-rewind button 554. In the user content information window 540, the video content 502 (thumbnail and waveform) is displayed so that the video content 502 can be played back. If the play / pause button 551 is touched when the video content 502 is not being played back, playback of the video content 502 begins. If the play / pause button 551 is touched while the video content 502 is being played back, playback of the video content 502 is paused. If the fast-forward button 553 is touched, the video content 502 is fast-forwarded, and if the fast-rewind button 554 is touched, the video content 502 is fast-rewinded. Playback of the video content selected by the composite selection cursor 531 may be controlled instead of the video content 502. Playback of all video content displayed in the selected content information window 530 may be controlled simultaneously. For example, the video content (part or all) displayed in the selected content information window 530 may be played in conjunction with the playback of the video content 502. If the playback of multiple video contents is controlled simultaneously, the user can view multiple video contents at the same time. The playback and display of the video contents are controlled by the cloud control unit 301 and the cloud display control unit 303.

[0044] Composition start button 503 is a button for generating (obtaining) new video content (composite content) by combining video and audio. When composition start button 503 is touched, smartphone 330 transmits a control signal for starting composition to cloud server 300. The control signal is input to cloud control unit 301 via cloud communication unit 304. Cloud control unit 301 performs control to generate composite content in response to receiving the control signal.

[0045] Here, audio / image selection button 520 has been used to select combining video with video content 502. Therefore, cloud control unit 301 performs control to combine the audio (audio data) of video content 502 with the video (video data) of the video content displayed in selection result window 534. Specifically, cloud control unit 301 instructs image / audio separation unit 371 to separate the video data from the data of the video content displayed in selection result window 534. Image / audio separation unit 371 stores the separated video data in cloud storage unit 302 and notifies cloud control unit 301 that separation is complete. Upon receiving the notification of separation completion, cloud control unit 301 instructs image / audio synthesis unit 372 to combine the audio data of video content 502 with the video data of the video content displayed in selection result window 534. All of this data is stored in cloud storage unit 302. The image / audio synthesis unit 372 stores the video content (composite content) obtained by synthesis in the cloud data storage 250, the cloud storage unit 302, or the smartphone 330 (user terminal). The user can check the composite content using the smartphone 330.

[0046] During compositing, a progressive bar (not shown) may be displayed to notify the user (photographer 412) of the progress of compositing. The file name of the composite content may be, for example, a file name obtained by adding a prefix or postfix to the file name of the video content 502, or it may not be. A file name input box (not shown) for the user to input a file name may be arranged on the application screen 500, and an arbitrary file name input in the file name input box may be set as the file name of the composite content.

[0047] FIG. 5(B) shows a modified example of the seek bar 560. In FIG. 5(B), a synthesis start pointer 562 for specifying the synthesis start time and a synthesis end pointer 563 for specifying the synthesis end time are superimposed on the seek bar 560. The positions of the synthesis start pointer 562 and the synthesis end pointer 563 can be changed by dragging or touching. In this case, the image / audio synthesis unit 372 replaces the video or audio from the synthesis start time to the synthesis end time of the video content 502 with that of another video content to generate the composite content. This type of configuration is suitable when it is desired to combine other video content only during periods when the quality of video content 502 is poor.

[0048] FIG. 5(C) shows a modified example of application screen 500. In selected content information window 530 in FIG. 5(C), image quality information of the video (specifically, resolution 535) is displayed as image information. Displaying the image quality information makes it possible to select a video with high image quality. Also displayed as image information is a user rating (user rating) 536. The user rating is, for example, the number of "likes" managed on a social networking site (SNS). Displaying user rating 536 allows the user (photographer 412) to select a video based on the ratings of other users. As a result, composite content that is more likely to receive ratings from third parties can be generated.

[0049] The image information is not limited to these, and may include, for example, information on the shooting direction. Doing so makes it possible to select a video shot in a direction close to the direction from the cameraman 412 to the subject 401 (sound source). A video shot in a direction close to the direction from the cameraman 412 to the subject 401 (sound source) is a video that closely resembles the view from the shooting position of the video content 502, or a video that matches the sound of the video content 502. This makes it possible to generate highly realistic composite content. The image quality information may include information on color depth, dark noise, etc. Doing so makes it possible to select a clear video from among videos shot at night.

[0050] FIG. 5(D) shows a modified example of the application screen 500. In the application screen 500 of FIG. 5(D), a plurality of video contents can be selected in a selected content information window 530. In the selected content information window 530, seek bars 571, 573, and 575 for specifying a compositing period are displayed. Pointers 564 and 565 linked to a compositing start pointer 562 and a compositing end pointer 563 are displayed in the seek bars 571, 573, and 575, respectively. When a plurality of video contents are selected, a pointer 566 for specifying a plurality of compositing periods is displayed. For example, when N video contents (N is 2 or more) are selected, N-1 pointers 566 for specifying N compositing periods are displayed. The cloud control unit 301 sets a plurality of compositing periods according to the positions of the pointers 566. The time of the compositing start pointer 562 (compositing start time) is the start time of the first compositing period, and the time of the compositing end pointer 563 (compositing end time) is the end time of the last compositing period. Pointer 566 can also be said to be a pointer for specifying the switching time between multiple synthesis periods. In this case, image / audio synthesis unit 372 generates synthesized content by replacing the video or audio of multiple periods (multiple synthesis periods) of video content 502 with that of multiple other video contents. In this way, high-quality synthesized content can be obtained over a long period of time.

[0051] In the example of FIG. 5(D), video content A (video content shot at point A) and video content C (video content shot at point C) are selected, and the seek bar 57 of video content C is 3A single pointer 566 is displayed in the image. The period from the synthesis start time to the time of pointer 566 is set as synthesis period 572 for synthesizing video content A, and the period from the time of pointer 566 to the synthesis end time is set as synthesis period 574 for synthesizing video content C. Image and audio synthesis unit 372 replaces the video or audio of video content 502 during synthesis period 572 with that of video content A. Furthermore, image and audio synthesis unit 372 replaces the video or audio of video content 502 during synthesis period 574 with that of video content C. In this way, the synthesized content is generated.

[0052] The cloud control unit 301 may narrow down candidates for video content to be combined with the video content 502. Here, if the video content 502 is a video uploaded to an SNS, Consider a case where a user (photographer 412) has added tag information to video content 502. In this case, among multiple video contents uploaded to the SNS, video contents to which tag information related to the tag information of video content 502 has been added may be selected as candidates for video contents to be composited with video content 502. This makes it possible to generate composite content that matches the values of photographer 412.

[0053] Although an example has been shown in which the cloud server 300 separates video content into video and audio, or synthesizes video and audio to generate composite content, the present invention is not limited to this. For example, at least one of the processes of separation and synthesis may be incorporated into an application of the smartphone 330. In other words, the smartphone 330 may perform the separation and synthesis.

[0054] Although an example has been shown in which video content from each terminal is stored in the cloud data storage 250, the present invention is not limited to this. For example, video content data may be directly transferred from another terminal to the data storage (not shown) of the smartphone 330 via the global network 320. In this case, only data for the period to be included in the composite content (data for the composite period) may be transferred. This makes it possible to reduce the amount of data communication. The data to be transferred may be data before separation or data after separation. The data to be transferred may be data to be included in the composite content (i.e., data for either video or audio), or it may not be. If only the data to be included in the composite content is transferred, the amount of data communication can be reduced. communication The amount can be further reduced.

[0055] Video content may have a denial period during which video and audio cannot be included in other video content. In FIG. 5(E), the photographer 414 has set a denial period 586. The selected content information window 530 displays the denial period 586, and the denial period 586 cannot be used for video composition. This prevents the sharing of video and audio during a specific period. Data during the denial period is not transferred to the smartphone 330.

[0056] A case will be described where audio synthesis is selected with audio / image selection button 520. FIG. 6(A) shows an example of an application screen 600 of PC 360. Application screen 600 is, for example, a browser screen or a dedicated application screen. In this case, the user of PC 360 is a user of camera 100 (photographer 411), who is using PC 360 to edit video content 602 captured by camera 100. In this case, camera 100 corresponds to the first camera of the present invention, and the other device (terminal) used for capturing the video corresponds to the second camera of the present invention.

[0057] The data of the video content 602 selected by the user (photographer 411) is displayed in the user content information window 640. In the user content information window 640, a reduced image (thumbnail) of the video is displayed as an item (image information) related to the video of the video content 602. In addition, as an item (audio information) related to the audio of the video content 602, the waveform of the audio is also displayed. The audio of the video content 602 shot by the photographer 411 is silent, and the quality of the audio is low. In such a case, the user selects "Audio (synthesize audio)" using the audio / image selection button 520.

[0058] The selected content information window 630 is a window that displays data of video content other than the video content 602 among the video content displayed in the position display window 510. In FIG. 5(A), in order to composite video with the video content 502 selected by the user, the selected content information window 530 displays image information as data of video content other than the video content 502. On the other hand, in FIG. 6(A), In order to synthesize audio for the video content 602 selected by the user, audio information is displayed in the selected content information window 630 as data for video content other than the video content 602. Of course, both image information and audio information may be displayed. In FIG. 6(A), an audio waveform 635 is displayed as the audio information. The user (photographer 411) can check the audio waveform 635 and identify a preferred audio.

[0059] When the composition start button 503 is touched, the PC 360 transmits a composition start control signal to the cloud server 300. The control signal is input to the cloud control unit 301 via the cloud communication unit 304. The cloud control unit 301 performs control to generate composite content in response to the reception of the control signal.

[0060] In this example, audio / image selection button 520 is used to select audio synthesis for video content 602. Therefore, cloud control unit 301 performs control to synthesize the video (video data) of video content 602 with the audio (audio data) of the video content displayed in selection result window 534. Specifically, cloud control unit 301 instructs image / audio separation unit 371 to separate audio data from the data of the video content displayed in selection result window 534. Image / audio separation unit 371 stores the separated video data in cloud storage unit 302 and notifies cloud control unit 301 that separation is complete. Upon receiving the notification of separation completion, cloud control unit 301 instructs image / audio synthesis unit 372 to synthesize the video data of video content 602 with the audio data of the video content displayed in selection result window 534. All of this data is stored in cloud storage unit 302. The image and audio synthesis unit 372 stores the video content (synthesized content) obtained by synthesis in the cloud data storage 250, the cloud storage unit 302, or the PC 360 (user terminal). The user can check the synthesized content using the PC 360.

[0061] FIG. 6(B) shows a modified example of the application screen 600. In the selected content information window 630 of FIG. 6(B), volume information 636 indicating the volume is displayed as audio information. By displaying the volume information 636, audio at an appropriate volume can be selected. Furthermore, S / N ratio information 637 indicating the S / N ratio is also displayed as audio information. This is suitable when normalizing and synthesizing audio. Furthermore, a user rating value 638 is also displayed as audio information. By displaying the user rating value 638, the user (photographer 411) can select a video based on the ratings of other users. As a result, it becomes possible to generate composite content that is more likely to receive ratings from third parties.

[0062] The audio information is not limited to these, and may include, for example, information on the distance from the subject 401 (sound source) to the shooting position. Because sound quality depends on the distance from the sound source, it becomes possible to select audio with high sound quality. The audio information may also include information on the shooting direction. This makes it possible to select audio from a shooting direction close to the direction from the cameraman 411 to the subject 401 (sound source). Audio from a shooting direction close to the direction from the cameraman 411 to the subject 401 is audio close to the audio heard at the shooting position of the video content 602, or audio that matches the video of the video content 602. This makes it possible to generate composite content with a high sense of realism.

[0063] Furthermore, if the audio to be synthesized is stereo audio, the L (left) and R (right) channels of the audio to be synthesized may be adjusted depending on the shooting position (recording position) of the audio to be synthesized, the shooting position of the video content 602, and the position of the subject 401 (sound source). For example, if the shooting position (recording position) of the audio to be synthesized is exactly opposite the shooting position of the video content 602 with the subject 401 (sound source) in between, the L and R channels of the audio to be synthesized may be swapped. This makes it possible to optimize the relationship between the video and audio in the synthesized content and reduce any sense of incongruity.

[0064] Fig. 7(A) is a flowchart showing an example of the overall synthesis process in cloud server 300. This process is realized by cloud control unit 301 expanding a program recorded in ROM 302A into RAM 302B and executing it. For example, when a terminal accesses cloud server 300, the process in Fig. 7(A) starts.

[0065] In step S701, cloud control unit 301 acquires data of the video content (first content) selected by file selection button 501 from cloud data storage 250. This enables display in user content information window 540. It also makes it possible to display the shooting location of the first content in location display window 510. Information on the shooting location (location information) is stored, for example, in the meta information of the video content.

[0066] In step S702, the cloud control unit 301 determines whether to synthesize audio or video for the first content. The user selects whether to synthesize audio or video for the first content using the audio / video selection button 520. If video is to be synthesized, the process proceeds to step S710, and if audio is to be synthesized, the process proceeds to step S720.

[0067] In step S710, the cloud control unit 301 uses the image / audio separation unit 371 to separate (extract) audio data from the data (first content data) acquired in step S701.

[0068] In step S711, the cloud control unit 301 searches for multiple video contents (multiple second contents) related to the first content from the cloud data storage 250. For example, multiple video contents whose shooting locations are close to the shooting location of the first content are searched for from the cloud data storage 250. If the first content is video content uploaded to an SNS, the multiple video contents may be searched for based on tag information.

[0069] In step S712, the cloud control unit 301 acquires image information (video thumbnail, resolution, color depth, dark noise, shooting direction, user rating, etc.) and shooting location information for each of the multiple second contents. This allows display in the selected content information window 530. It also makes it possible to display the shooting location for each of the multiple second contents in the location display window 510.

[0070] In step S720, the cloud control unit 301 uses the image / audio separation unit 371 to separate (extract) video data from the data (first content data) acquired in step S701.

[0071] In step S721, the cloud control unit 301 searches the cloud data storage 250 for a plurality of video contents (a plurality of second contents) related to the first content, similar to step S711.

[0072] In step S722, the cloud control unit 301 acquires audio information (audio waveform, volume, S / N ratio, distance from the subject (sound source) to the shooting position, shooting direction, user evaluation value, etc.) of each of the multiple second contents and information on the shooting position. This makes it possible to display in the selected content information window 630. It also makes it possible to display the shooting position of each of the multiple second contents in the position display window 510.

[0073] In step S703, the cloud control unit 301 selects a composite image from a plurality of second contents. The user can specify (select) the second content to be combined using the combining selection cursor 531. The cloud control unit 301 selects the second content specified by the user using the combining selection cursor 531. At this time, the cloud control unit 301 may further determine the combining period.

[0074] In step S704, the cloud control unit 301 generates composite content using the image / audio separation unit 371 and the image / audio synthesis unit 372 (synthesis process). Specifically, when the processes of steps S710 to S712 have been performed, the cloud control unit 301 uses the image / audio separation unit 371 to extract video data from the data of the second content selected in step S703. Then, the cloud control unit 301 uses the image / audio synthesis unit 372 to synthesize the audio data extracted in step S710 (audio data of the first content) with the video data of the second content selected in step S703. When the processes of steps S720 to S722 have been performed, the cloud control unit 301 uses the image / audio separation unit 371 to extract audio data from the data of the second content selected in step S703. Then, cloud control unit 301 synthesizes the video data extracted in step S720 (video data of the first content) with the audio data of the second content selected in step S703 using image / audio synthesis unit 372. Cloud control unit 301 stores the data of the synthesized content in cloud data storage 250, cloud storage unit 302, or the user terminal.

[0075] In the synthesis process of step S704, the audio data and video data are synthesized based on, for example, the time of the NTP server 321. In other words, the audio data and video data are synthesized so that audio data and video data recorded (audio or video) at the same time are played back at the same time. However, if the distance from the subject (sound source) to the shooting position differs greatly between the audio data and video data to be synthesized, the video and audio will be out of sync, resulting in synthesized content with a low sense of realism (unnatural synthesized content).

[0076] Therefore, it is preferable that the cloud control unit 301 adjusts the time position of at least one of the first content and the second content based on the distance from the subject to the shooting position of the first content and the distance from the subject to the shooting position of the second content. In this case, the cloud control unit 301 combines the adjusted first content and second content.

[0077] This makes it possible to generate composite content with a more realistic feel. The position of the subject may be set in advance, or may be determined by triangulation based on information about the shooting position and shooting direction (posture) added to the video content.

[0078] For example, in FIG. 4A, there is a large difference between the distance SA from subject position S to shooting position A and the distance SB from subject position S to shooting position B. Therefore, if the video content (either video or audio) from shooting position A and the video content (the other of video and audio) from shooting position B are combined without adjusting the time positions, a time difference Δt (shift) will occur between the video and audio according to the difference between the distance SA and the distance SB. The time difference Δt can be calculated using the following equation 1. In equation 1, V is the speed of sound. Δt = (SA - SB) / V (Equation 1)

[0079] Therefore, the cloud control unit 301 adjusts the time position of at least one of the video contents A and B so that there is a time difference Δt between the video content at the shooting position A (video content A) and the video content at the shooting position B (video content B). The time position is adjusted so that the difference between the video and the audio in the composite content is reduced. For example, the audio of video content A is delayed by the time difference Δt from the audio of video content B, so that the cloud control unit 301 adjusts the time position of at least one of the video contents A and B so that there is a time difference Δt between the video content A and the audio of video content B. The control unit 301 delays the time position of the video content B by the time difference Δt. Note that the amount of adjustment may be greater or smaller than Δt as long as the difference between the video and audio in the composite content is reduced.

[0080] The above example does not take into account the shooting magnification (when the shooting magnification is 1x). As the shooting magnification increases, the shooting position is brought closer to the subject. Therefore, to generate more realistic composite content, it is preferable to consider the difference in shooting magnification (difference in angle of view). Here, let f1 be the lens focal length of the camera that shot video content A, and D1 be the image circle diameter of that camera. Let f2 be the lens focal length of the camera that shot video content B, and D2 be the image circle diameter of that camera. Let D0 be the image circle diameter of a sensor size of 35mm, and f0 be the reference focal length. For example, image circle diameter D0 is the image circle diameter of a full-frame sensor, and reference focal length f0 is 50mm. Now, consider a case where distances SA and SB are sufficiently greater than focal lengths f1 and f2. In this case, the actual time difference Δt is determined by the difference between the distance SA multiplied by the ratio of the focal length f1 to the image circle diameter D1 and the distance SB multiplied by the ratio of the focal length f2 to the image circle diameter D2, as shown in the following equation 2. Δt=(SA·D1 / f1-SB·D 2 / f 2 ) / (V D0 / f0) ...(Formula 2)

[0081] By adjusting the time position of at least one of the video contents A and B so that there is a time difference Δt between them, as calculated by Equation 2, it is possible to generate highly realistic composite content even when combining video contents that have been zoomed in.

[0082] In addition, when replacing the entire audio or video of a video content (first content) with that of another video content (second content), it is preferable to acquire data of a period including the period of the first content as data of the second content and use it for synthesis. startData is acquired from a time before and close to the end time of the first content to a time after and close to the end time of the shooting of the first content. This makes it possible to prevent video or audio periods from occurring in the composite content when the time position is adjusted and the composite content is performed.

[0083] Fig. 7B is a flowchart showing an example of the synthesis process in step S704 in Fig. 7A. This process is realized by the cloud control unit 301 expanding a program recorded in ROM 302A into RAM 302B and executing it.

[0084] In step S731, the cloud control unit 301 acquires information on the position of the subject 401, which is the sound source, and the speed of sound. The position of the subject 401 may be searched for via the global network 320. As information on the speed of sound, an approximate value of 340 m / s may be acquired, or the speed of sound may be calculated more precisely from the humidity and temperature at the time of shooting.

[0085] In step S732, the cloud control unit 301 acquires information about the shooting location of the first content. In step S733, the cloud control unit 301 acquires information about the shooting location of the second content. The shooting location information is acquired from, for example, meta information of the video content.

[0086] In step S734, the cloud control unit 301 calculates the distance from the position of the subject 401 to the shooting position of the first content and the distance from the position of the subject 401 to the shooting position of the second content based on the processing results of steps S731 to S733.

[0087] In step S735, the cloud control unit 301 calculates the following equation based on the above-described equation 1 or 2: , and calculate the time difference Δt.

[0088] In step S736, the cloud control unit 301 adjusts the time position of at least one of the first content and the second content based on the time difference Δt calculated in step S735, and combines the first content and the second content.

[0089] According to the present embodiment described above, the video and audio to be recorded are not limited to those captured by the device the user is using for shooting, but can be selected from video and audio captured at a location distant from the user, allowing the user to record the desired video and audio. This allows the user to easily generate high-quality video content without having to reshoot the video.

[0090] The various controls described above as being performed by the cloud control unit 301 may be performed by a single piece of hardware, or the entire device may be controlled by multiple pieces of hardware (e.g., multiple processors or circuits) sharing the processing.

[0091] Furthermore, although the present invention has been described in detail based on preferred embodiments thereof, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Furthermore, each of the above-described embodiments merely represents one embodiment of the present invention, and each embodiment can be combined as appropriate.

[0092] Furthermore, in the above-described embodiment, the present invention has been described with reference to an example in which it is applied to a cloud server, but this is not limited to this example. The present invention can also be applied to electronic devices capable of editing video content, and control devices that communicate with electronic devices (including network cameras) via wired or wireless communication and remotely control the electronic devices, as well as to electronic devices themselves. For example, the present invention can be applied to personal computers, PDAs, mobile phone terminals, portable image viewers, printers, digital photo frames, music players, game consoles, e-book readers, imaging devices, and the like. The present invention can also be applied to video players, display devices (including projection devices), tablet terminals, smartphones, AI speakers, home appliances, in-vehicle devices, and the like.

[0093] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]

[0094] 300: Cloud server 301: Cloud control unit

Claims

1. A first selection means for selecting a first content acquired by a first camera in response to a user operation; a second selection means for selecting whether to synthesize a video or an audio with the first content in response to a user operation, separately from the selection of the content to be synthesized; a control means for controlling the display of a plurality of items corresponding to a plurality of second contents photographed by a plurality of second cameras different from the first camera; a third selection means for selecting a second content to be combined with the first content from the plurality of second contents in response to a user operation; a generating means for generating a third content by synthesizing the video included in the second content selected by the third selecting means and the audio included in the first content selected by the first selecting means when a video is selected by the second selecting means, and for generating a fourth content by synthesizing the audio included in the second content selected by the third selecting means and the video included in the first content selected by the first selecting means when audio is selected by the second selecting means; and The control means controls the display so as to further display items indicating the subject, the shooting position of the first content, and the shooting positions of each of the plurality of second contents. An electronic device characterized by:

2. further comprising the first camera; The plurality of second cameras are provided in a device different from the electronic device.

2. The electronic device according to claim 1, wherein the electronic device is a semiconductor device.

3. The control means When a video is selected by the second selection means, a plurality of items related to the video included in the plurality of second contents are displayed as the plurality of items; 3. The electronic device according to claim 1, wherein when audio is selected by the second selection means, the electronic device is controlled so that a plurality of items related to the audio included in the plurality of second contents are displayed as the plurality of items.

4. the items related to the video include thumbnails of the video; The thumbnail images included in the items related to the video are displayed so that the video can be played.

4. The electronic device according to claim 3.

5. The electronic device according to claim 4 , wherein the plurality of thumbnail images included in the plurality of items can be simultaneously played back as moving images.

6. The item related to the video indicates at least one of image quality information, a user rating, and a shooting direction.

6. The electronic device according to claim 3, wherein the first and second electrodes are electrically connected to the first and second electrodes.

7. The items related to the audio indicate at least one of volume, S / N ratio, waveform, user evaluation value, distance from the subject to the shooting position, and shooting direction.

7. The electronic device according to claim 3, wherein the first and second electrodes are electrically connected to the first and second electrodes.

8. The generating means generates content in which video or audio for a part of the first content is replaced with that of the second content selected by the third selecting means, as the third content or the fourth content.

8. The electronic device according to claim 1, wherein the electronic device is a semiconductor device.

9. The generating means generates the third content or the fourth content by replacing the video or audio of the first content for a plurality of periods with the video or audio of the plurality of second content selected by the third selecting means.

9. The electronic device according to claim 8.

10. The generating means acquires data of a period including a period of the first content as data of the second content selected by the third selecting means in order to generate the third content or the fourth content.

8. The electronic device according to claim 1, wherein the electronic device is a semiconductor device.

11. The generating means acquires data of a period to be included in the third content or the fourth content as data of the second content selected by the third selecting means in order to generate the third content or the fourth content.

10. The electronic device according to claim 8 or 9.

12. At least one of the plurality of second contents has a denial period set therein during which the video and audio cannot be combined with other content.

12. The electronic device according to claim 8, 9 or 11.

13. The generating means adjusting a time position of at least one of the first content selected by the first selection means and the second content selected by the third selection means based on a first distance, which is a distance from a shooting position of the first content selected by the first selection means to a subject, and a second distance, which is a distance from a shooting position of the second content selected by the third selection means to the subject; The adjusted first content and the adjusted second content are combined to generate the third content or the fourth content.

13. The electronic device according to claim 1.

14. The generating means adjusts a time position of at least one of the first content selected by the first selecting means and the second content selected by the third selecting means so that a time difference according to a difference between the first distance and the second distance is created between the first content selected by the first selecting means and the second content selected by the third selecting means.

14. The electronic device according to claim 13.

15. The generating means adjusts a time position of at least one of the first content selected by the first selecting means and the second content selected by the third selecting means so that a time difference corresponding to a difference between a distance obtained by multiplying the first distance by a ratio of an image circle diameter to a focal length of the first camera and a distance obtained by multiplying the second distance by a ratio of an image circle diameter to a focal length of the second camera is created between the first content selected by the first selecting means and the second content selected by the third selecting means.

14. The electronic device according to claim 13.

16. the first content is content uploaded to a social networking site (SNS), The second content is a content to which tag information related to tag information of the first content is added, among a plurality of contents uploaded to the SNS.

16. The electronic device according to claim 1, wherein the electronic device is a semiconductor device.

17. A first selection step of selecting first content acquired by a first camera in response to a user operation; a second selection step of selecting whether to synthesize a video or an audio with the first content in response to a user operation, separately from the selection of the content to be synthesized; a control step of controlling the display device to display a plurality of items corresponding to a plurality of second contents respectively captured by a plurality of second cameras different from the first camera; a third selection step of selecting, from the plurality of second contents, a second content to be combined with the first content in response to a user operation; a generation step of generating a third content by combining the video included in the second content selected in the third selection step with the audio included in the first content captured by the first camera selected in the first selection step when a video is selected in the second selection step, and generating a fourth content by combining the audio included in the second content selected in the third selection step with the video included in the first content selected in the first selection step when audio is selected in the second selection step; and In the control step, control is performed so that items indicating the subject, the shooting position of the first content, and the shooting positions of each of the plurality of second contents are further displayed. A method for controlling an electronic device.

18. A program for causing a computer to function as each of the means of the electronic device according to any one of claims 1 to 16.

19. A computer-readable storage medium storing a program for causing a computer to function as each of the means of the electronic device according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Digital platform for synchronized editing of user-generated videos

    JP2016508309A

  • Data processing apparatus, data processing method and program

    JP2019220848A