Content Creation by Sound Control

The system addresses unprofessional segments in content creation by using voice control to start and stop recording and automatically trim segments with voice commands, improving content quality and reducing editing complexity.

JP2025521229AActive Publication Date: 2025-07-08LEMON CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024572400
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-10
Filing Date
2023-06-01
Publication Date
2025-07-08
Estimated Expiration
2043-06-01

AI Technical Summary

Technical Problem

Existing content creation methods, such as video recording, often include unprofessional segments due to button or touch screen operations, which can degrade content quality and increase data size, and trimming these segments requires additional editing steps.

Method used

A system for content creation by voice control, where voice commands are used to start and stop recording, and segments with voice commands are automatically trimmed, improving content quality and reducing editing complexity.

Benefits of technology

Enables hands-free content creation with improved professional appearance and reduced data size by using voice commands to control recording and automatically removing segments with voice cues, enhancing user experience and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521229000001_ABST
    Figure 2025521229000001_ABST
Patent Text Reader

Abstract

This disclosure describes techniques for content creation by voice control. The techniques include monitoring voice commands spoken by a creator. In response to recognizing a first voice command spoken by the creator, recording of the content may be started. In response to recognizing a second voice command spoken by the creator, recording of the content may be stopped. A timestamp associated with the second voice command may be created. Based on the timestamp, segments may be automatically deleted from the content. The segments may include the recording of the second voice command.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Patent No. 17 / 838,022, filed on June 10, 2022, with the title "Content Creation by Voice Control", the disclosure of which is hereby incorporated by reference in its entirety.

Background Art

[0002] Communication using Internet - based tools is increasing. The Internet - based tools can be any software or platform. Existing social media platforms enable communication by allowing users to share information such as images and videos via static applications or web pages. As communication devices such as mobile phones are becoming increasingly powerful, people continue to seek new forms of entertainment, social networking, and communication methods.

Brief Description of the Drawings

[0003] The following detailed description is better understood when read in conjunction with the accompanying drawings. For purposes of illustration, exemplary embodiments of various aspects of the present disclosure are shown in the accompanying drawings, but the invention is not limited to the specific methods and means disclosed.

[0004]

Figure 1

[0005]

Figure 2

[0006]

Figure 3

[0007]

Figure 4

[0008]

Figure 5

[0009]

Figure 6

[0010]

Figure 7

[0011]

Figure 8

[0012]

Figure 9

[0013]

Figure 10

[0014]

Figure 11

[0015]

Figure 12

[0016]

Figure 13

[0017]

Figure 14

DETAILED DESCRIPTION OF THE INVENTION

[0018] Users of the content creation platform can create content that includes "start" and / or "stop" operation segments. For example, a user of the content creation platform may create a video indicating that the user starts and / or stops a camera recording, as the only way to start and / or stop a camera recording may be button selection or touch screen operation. However, such segments do not add value to the content, or add little value. Furthermore, such segments may lead to an unprofessional look of the published content, and / or such segments may increase the content data size, which may lead to an increase in the server upload delay.

[0019] To address the drawbacks of including these segments in the published content, these segments may be trimmed before the content is published. However, trimming these segments in this way may involve additional content editing steps. Furthermore, the content quality may be degraded by the re-encoding process after editing.

[0020] Therefore, improvement of content creation technology is required. In particular, technology for content creation by voice control is required. A technology that enables creation of content such as video by voice control will be described. It is also possible to monitor voice commands spoken by a content creator. In response to recognizing a first voice command spoken by the creator, recording of content such as video may be started. In response to recognizing a second voice command spoken by the creator, recording of the content may be stopped. A time stamp associated with the second voice command may be created. Based on the time stamp, segments may be deleted from the content. The deleted segments may include recording of the second voice command.

[0021] The technology for content creation by voice control described in this specification may be used by a system for delivering content. FIG. 1 shows an exemplary system 100 for delivering content. System 100 may include a server 102 and a plurality of client devices 104. The server 102 and the plurality of client devices 104a - 104n may communicate with each other via one or more networks 132.

[0022] The server 102 may be located in a data center such as a single building, or may be distributed at different geographical locations (e.g., several buildings). The server 102 may provide services via the one or more networks 120. Network 132 includes various network devices such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or similar devices. Network 132 may include physical links such as coaxial cable links, twisted pair cable links, fiber optic links, combinations thereof, etc. Network 132 may include wireless links such as cellular links, satellite links, Wi-Fi links, etc.

[0023] Server 102 may include a plurality of computing nodes that host various services. In one embodiment, the node hosts a video service 112. The video service 112 may include a content streaming service such as an Internet Protocol video streaming service. The video service 112 may be configured to deliver content 123 via various transmission technologies. The video service 112 is configured to provide content 123 such as video, audio, text data, combinations thereof, and the like. The content 123 may include a content stream (e.g., a video stream, an audio stream, an information stream), a content file (e.g., a video file, an audio file, a text file), and / or other data. The content 123 can be stored in the database 122. For example, the video service 112 may include a video sharing service, a video hosting platform, a content delivery platform, a collaborative game platform, and the like.

[0024] In one embodiment, the content 123 distributed or provided by the video service 112 includes video. The video may have a time interval of less than a predetermined time limit, such as 1 minute, 5 minutes, or other predetermined minutes. By way of example and not limitation, the video may include at least one and up to four 15-second segments combined with each other. The short video time interval can provide viewers with rapid and continuous entertainment that enables users to view a large number of videos within a short time frame. Such rapid and continuous entertainment may be popular on social media platforms.

[0025] The video may include a pre-recorded audio overlay, such as a pre-recorded song or audio clip from a television program or movie. If the short video includes a pre-recorded audio overlay, the short video may be characterized by one or more people lip-syncing, dancing, or moving their bodies in other ways along with the pre-recorded audio. For example, the short video may be characterized by an individual completing a "dance challenge" to a hit song, or the short video may be characterized by two people participating in a lip-sync or a dance for two. As another example, the short video may be characterized by an individual achieving a challenge that requires moving their body in response to a pre-recorded audio overlay, such as in correspondence with the beat or rhythm of a pre-recorded song characterized by the pre-recorded audio overlay. Other videos may not include a pre-recorded audio overlay. For example, these videos may be characterized by an individual doing sports, playing pranks, or giving advice such as beauty or fashion advice, cooking tips, or home interior tips.

[0026] In one embodiment, the content 123 may be output to different client devices 104 via the network 132. The content 123 may be streamed to the client device 104. The content stream may be a stream of video received from the video service 112. The plurality of client devices 104 may be configured to access the content 123 from the video service 112. In one embodiment, the client device 104 may include a content application 106. The content application 106 outputs (e.g., displays, renders, presents) the content 123 to a user associated with the client device 104. The content may include video, audio, comments, text data, and the like.

[0027] The plurality of client devices 104 may include any type of computing device, such as a mobile device, a tablet device, a laptop computer, a desktop computer, a smart TV, or other smart devices (e.g., smart watches, smart speakers, smart glasses, smart helmets), a gaming device, a set-top box, a digital streaming device, a robot, etc. The plurality of client devices 104 may be associated with one or more users. A single user may access the server 102 using one or more of the plurality of client devices 104. The plurality of client devices 104 may move to various locations and access the server 102 using different networks.

[0028] The video service 112 may be configured to receive input from a user. The user may be registered as a user of the video service 112 and may also be a user of the content application 106 operating on the client device 104. The user input may include a video created by the user, a user comment associated with the video, or a "like" associated with the video. The user input may include a connection request and user input data such as text data, digital image data, or user content. The connection request may include a request to connect to the video service 112 from the client devices 104a - d. The user input data may include information that a user connected to the video service 112, such as a video and / or a user comment, desires to share with other connected users of the video service 112.

[0029] Video service 112 may be able to receive different types of inputs from users using different types of client devices 104. For example, a user using content application 106 on a first user device such as a mobile phone or a tablet may be able to create and upload a video using content application 106. A user using content application 106 on a different mobile phone or tablet may be able to view, comment on, or "like" a video or a comment written by another user. In another example, a user using content application 106 on a smart TV, laptop, desktop, or gaming device may not be able to create and upload a video or comment on a video using content application 106. Instead, a user using content application 106 on a smart TV, laptop, desktop, or gaming device may only be able to use content application 106 to view a video, view comments left by other users, or "like" a video.

[0030] In one embodiment, a user may use content application 106 on client device 104 to create a video, such as a short video, and upload the video to server 102. Client device 104 may be able to access interface 108 of content application 106. Interface 108 may include input elements. For example, the input elements may be configured to enable the user to create a video. To create a short video, the user may give content application 106 permission to access an image acquisition device such as a camera of client device 104 or a microphone. Using content application 106, the user may select a time interval of the video or set the speed of the video, such as "slow motion" or "speed up".

[0031] The user may edit the video using the content application 106. The user may add effects such as one or more texts, filters, sounds, or beauty effects to the video. To add a pre-recorded audio overlay to the video, the user may select a song or a sound clip from the sound library of the content application 106. The sound library may include different songs, sound effects, or audio clips from movies, albums, or TV shows. In addition to, or instead of, adding a pre-recorded audio overlay to the video, the user may use the content application 106 to add a narration to the video. The narration may be the sound recorded by the user using the microphone of the client device 104. The user can add a text overlay to the short video and may specify, using the content application 106, when the text overlay appears in the video. The user may assign captions, location tags, and one or more hashtags to the video to indicate the theme of the video. The content application 106 may prompt the user to select a frame of the video for use as a "cover image" for the video.

[0032] After the user creates the video, the user may use the content application 106 to upload the video to the server 102 and / or save the video locally on the user device 104. When the user uploads the video to the server 102, the user may select whether to make the video viewable by all other users of the content application 106 or only by a subset of the users of the content application 106. The video service 112 may store the uploaded video and any metadata associated with the video in one or more databases 122.

[0033] In one embodiment, the user may provide an input on a video using the content application 106 on the client device 104. The client device 104 may access an interface 108 of the content application 106 that enables the user to provide an input associated with the video. The interface 106 may include input elements. For example, the input elements may be configured to receive inputs from the user such as comments or "likes" associated with a particular video. If the input is a comment, the content application 106 may allow the user to set an emoji associated with their input. The content application 106 may be able to determine time information about the input, such as when the user wrote the comment. The content application 106 may be able to send the input and the associated metadata to the server 102. For example, the content application 106 may send the comment, the identifier of the user who wrote the comment, and the time information about the comment to the server 102. The video service 112 may store the input and the associated metadata in the database 122.

[0034] The video service 112 may be configured to output the uploaded video and user input to other users. The user may be registered as a user of the video service 112 and may view videos created by other users. The user may also be a user of the content application 106 operating on the client device 104. The content application 106 may output (display, render, present) the video and user comments to the user associated with the client device 104. The client device 104 can access the interface 108 of the content application 106. The interface 108 may include output elements. The output elements may be configured to display information about different videos so that the user can select and view the videos. For example, the output elements may be configured to display a plurality of cover images, subtitles, or hashtags associated with the video. The output elements may also be configured to arrange the videos according to the categories associated with each video.

[0035] In one embodiment, the user comments associated with the video may be output to other users who are viewing the same video. For example, all users accessing the video may view the comments associated with the video. The video service 112 may output the comments associated with the video simultaneously. The comments may be output by the video service 112 in real time or near real time. The content application 106 may display the video and comments on the client device 104 in various ways. For example, the comments may be displayed as an overlay on the content, or as an overlay next to the content. As another example, a user who wants to view the comments of other users associated with the video may need to select a button to view the comments. The comments may be displayed with animation when shown. For example, the comments may be scrolled across the video or across the overlay.

[0036] In an embodiment, a user may use the content application 106 on the client device 104 to create a video, such as a short video, using voice control and upload the video to the server 102. In this way, a user of the video service 112 can perform "hands-free" video recording. The user does not need to press a recording button, move away from the client device 104 to take a video, and then return to the client device 104 to press a recording stop button, and can control video creation using voice control. For example, after moving away from the client device 104, the user can instruct the content application 106 to start recording a video. Similarly, without returning to the client device 104, the user can instruct the content application 106 to pause and / or end the video recording. To create a video using voice control, the user may give the content application 106 permission to access an image acquisition device, such as a camera of the client device 104, and a microphone.

[0037] In an embodiment, the video service 112 and / or the content application 106 includes a voice recognition unit 117. The voice recognition unit 117 may be set to listen for or monitor keywords associated with video creation. For example, the voice recognition unit 117 may be set to listen for or monitor voice commands spoken by a content creator (i.e., a user of the video service 112). The voice recognition unit 117 may be started in response to a user indicating that they want to create content such as a video. For example, the user may select a button on the interface 108 of the content application 106 indicating that the user wants to create and / or upload a new video to the server 102. When the user selects a button indicating that the user wants to create and / or upload a new video to the server 102, the voice recognition unit 117 may be activated. When the voice recognition unit 117 is activated, the voice recognition unit 117 may start listening for keywords in voice commands spoken by the user.

[0038] In an embodiment, the voice recognition unit 117 may start executing a corresponding operation in response to hearing a specific keyword in a voice command spoken by a user. For example, the video service 112 and / or the content application 106 may include a recording unit 118. When the voice recognition unit 117 recognizes a keyword associated with starting video recording, the voice recognition unit 117 may trigger the recording unit 118 to start video recording. Examples of keywords associated with starting video recording include, but are not limited to, "start", "action", "record", "go", etc. When the voice recognition unit 117 recognizes a keyword associated with pausing video recording, the voice recognition unit 117 may trigger the recording unit 118 to pause (not end) video recording. Examples of keywords associated with pausing video recording include, but are not limited to, "pause", "break", etc. When the voice recognition unit 117 recognizes a keyword associated with resuming paused video recording, the voice recognition unit 117 may trigger the recording unit 118 to resume (unpause) video recording. Examples of keywords associated with resuming paused video recording include, but are not limited to, "continue", "restart", "action", etc. When the voice recognition unit 117 recognizes a keyword associated with stopping video recording, the voice recognition unit 117 may trigger the recording unit 118 to stop or end video recording. Examples of keywords associated with stopping video recording include, but are not limited to, "stop", "cut", "end", etc.

[0039] In an embodiment, the recording unit 118 may generate a time stamp associated with the execution of various operations. For example, when the recording unit 118 starts recording a video at time A in response to being triggered by the voice recognition unit 117 or the like, the recording unit 118 may generate a time stamp associated with time A. As another example, when the recording unit 118 temporarily stops recording a video at time B in response to being triggered by the voice recognition unit 117 or the like, the recording unit 118 may generate a time stamp associated with time B. When the recording unit 118 resumes recording a video at time C in response to being triggered by the voice recognition unit 117 or the like, the recording unit 118 may generate a time stamp associated with time C. As yet another example, when the recording unit 118 stops recording a video at time D in response to being triggered by the voice recognition unit 117 or the like, the recording unit 118 may generate a time stamp associated with time D. The recording unit 118 may store these time stamps in the database 124 as time stamp data 125. The time stamp data 125 may indicate, for each time stamp, the time associated with the time stamp and the corresponding operation executed at that time.

[0040] In an embodiment, the recording unit 118 may store an unedited copy or version of video recording as unedited content 127 in a database 126 or the like. The unedited content 127 may be edited using the timestamp data 125. For example, the video service 112 and / or the content application 106 may include an editing unit 119. The editing unit 119 may be set to edit the unedited content 127 using the timestamp data 125. For example, the editing unit 119 may be set to delete one or more segments from the unedited video based on the timestamp data 125. The one or more segments to be deleted may include one or more unedited videos in which a voice command appears. For example, the segment to be deleted may include a segment in which there is a scene where the user speaks a voice command such as "pause" or "stop". In this way, the final, edited video does not have a scene where the user speaks a voice command. The final, edited video may be stored in the database 122 as content 123 for distribution to the client devices 104a - n.

[0041] In an embodiment, the voice recognition unit 117 may start executing other corresponding operations in response to hearing other keywords in the voice command spoken by the user. For example, the voice recognition unit 117 may be set to start executing an operation while the video is being recorded. If the user can use voice commands while recording a video, it may improve the quality of the created video. For example, using voice commands while recording a video makes accessible camera operations that are inaccessible unless the user stops recording, approaches the user device, and physically presses a button or touch screen.

[0042] For example, when the voice recognition unit 117 recognizes a keyword associated with zooming in or out during the video recording process, the voice recognition unit 117 may trigger the camera operation unit 210 to zoom in or out while recording the video. Examples of keywords associated with zooming in or out during video recording include, but are not limited to, "zoom in" and "zoom out". When the voice recognition unit 117 recognizes a keyword associated with focusing the camera during video recording, the voice recognition unit 117 may trigger the camera operation unit 210 to focus (e.g., in the near field or far field). Examples of keywords associated with focusing the camera during video recording include, but are not limited to, "focus". When the voice recognition unit 117 recognizes a keyword associated with adding or removing a flash during video recording, the voice recognition unit 117 may trigger the camera operation unit 210 to turn the flash on or off. Examples of keywords associated with adding or removing a flash during video recording include, but are not limited to, "flash on" and "flash off". When the voice recognition unit 117 recognizes a keyword associated with using a wide-angle camera or a telephoto camera during video recording, the voice recognition unit 117 may trigger the camera operation unit 210 to use the wide-angle camera or the telephoto camera. Examples of keywords associated with using a wide-angle camera or a telephoto camera during video recording include, but are not limited to, "wide lens" and "telephoto". When the voice recognition unit 117 recognizes a keyword associated with adding a filter or an effect during video recording, the voice recognition unit 117 may trigger the camera operation unit 210 to add the filter or the effect while recording the video. Examples of keywords associated with adding a filter or an effect during the video recording process include, but are not limited to, "add filter" and "add effect".

[0043] As another example, the voice recognition unit 117 may be set to search for specific content, such as one or more videos, in response to having heard one or more keywords indicating a content search within the voice command spoken by the user. For example, the user may speak a command such as "Search for cat videos." The voice recognition unit 117 may recognize one or more keywords within this command and, in response, present the search results for cat videos on the interface 108 of the client device 104.

[0044] FIG. 2 is a schematic diagram 200 showing a process for content creation by voice control according to the present disclosure. The user may indicate that they want to create content such as a video. For example, the user may select a button on the interface 108 of the content application 106 indicating that the user wants to create and / or upload a new video to the server 102. When the user selects a button indicating that the user wants to create and / or upload a new video to the server 102, the voice recognition unit 117 may be activated. When the voice recognition unit 117 is activated, the voice recognition unit 117 may start listening for keywords within the voice command spoken by the user. For example, the voice recognition unit 117 may listen for keywords obtained by the voice input 206 of the client device 104. The voice input 206 may include, for example, one or more microphone inputs.

[0045] The voice recognition unit 117 may recognize the first keyword 202 obtained by the voice input 206 at the first time. The first keyword 202 may be a keyword associated with the start of video recording. For example, the first keyword 202 may include "start", "action", "record", "go", etc. When the voice recognition unit 117 recognizes the first keyword 202, the voice recognition unit 117 may trigger the video recording unit 118 to start the camera operation 210. The camera operation 210 may be, for example, video recording.

[0046] The voice recognition unit 117 may recognize the second keyword 204 obtained by the voice input 206 at a second time after the first time. The second keyword 204 may be a keyword associated with the stop of the video recording. The second keyword 204 may be obtained by the voice input 206, for example, during the video recording. The second keyword 204 may include "stop", "cut", "end", etc. When the voice recognition unit 117 recognizes the second keyword 204, the voice recognition unit 117 may trigger the video recording unit 118 to end the camera operation 210. The camera operation 210 may be, for example, video recording.

[0047] The editing unit 119 may be set to edit the recorded video using, for example, the timestamp data 125. For example, the editing unit 119 may be set to delete one or more segments from the unedited video based on the timestamp data 125. The one or more segments to be deleted may include one or more unedited videos in which voice commands appear. For example, the segment to be deleted may include a segment in which there is a scene where the user speaks the second keyword 204. In this way, the final edited video does not have a scene where the user speaks a voice command. The final edited video may be stored in the database 122 as the content 123 for distribution to the client devices 104a~n.

[0048] Figure 3 shows an exemplary process 300 for content creation by voice control. Although described as a series of operations in FIG. 3, those skilled in the art should understand that in various embodiments, the described operations may be added, removed, rearranged, or modified. At 302, voice commands spoken by the creator may be monitored. For example, the voice recognition unit (i.e., voice recognition unit 117) may be set to listen for or monitor keywords associated with video creation. The voice commands may be obtained, for example, by one or more microphones of the client device. The voice recognition unit may be started in response to the user indicating that they want to create content such as a video (i.e., the monitoring of voice commands may be started). For example, the user may select a button on the interface of the content application indicating that the user wants to create and / or upload a new video to the server. When the user selects a button indicating that the user wants to create and / or upload a new video to the server, the voice recognition unit may be activated. When the voice recognition unit is activated, the voice recognition unit may start listening for keywords in the voice commands spoken by the user.

[0049] At 304, in response to recognizing a first voice command spoken by the creator, recording of the content may be started. For example, when the voice recognition unit recognizes the first voice command, the voice recognition unit may start recording the video. The first voice command may include keywords associated with starting the recording of the video. Examples of keywords associated with starting the recording of a video include, but are not limited to, "start", "action", "record", "go", etc.

[0050] The content may continue to be recorded until the voice recognition unit recognizes the second voice command. In 306, in response to recognizing the second voice command spoken by the creator, the recording of the content may be stopped. For example, when the voice recognition unit recognizes the second voice command, the voice recognition unit may stop the recording of the video. The second voice command may include keywords associated with stopping or ending the video recording. Examples of keywords associated with stopping the video recording include, but are not limited to, "stop", "cut", "end", etc.

[0051] In 308, a time stamp associated with the second voice command may be created. For example, when the recording of the content is stopped at time C in response to recognizing the second voice command or the like, a time stamp associated with time C may be generated. The time stamp may be stored in the database as time stamp data. The time stamp data may indicate the time associated with the time stamp and the corresponding operation (i.e., the recording is stopped) performed at that time.

[0052] The content may be edited using the time stamp data. In 310, based on the time stamp, segments may be automatically deleted from the content. The segments may include the recording of the second voice command. For example, the editing unit (i.e., editing unit 119) may be set to automatically delete one or more segments from the unedited video based on the time stamp data. The one or more segments to be deleted may include one or more unedited videos in which a voice command, such as the second voice command, appears. For example, the segments to be deleted may include segments in which there is a scene where the user speaks a voice command such as "pause" or "stop". In this way, the final, edited video does not have a scene where the user speaks a voice command.

[0053] FIG. 4 shows an exemplary process 400 for content creation by voice control. Although described as a series of operations in FIG. 4, those skilled in the art should understand that in various embodiments, the described operations may be added, removed, rearranged, or modified. At 402, voice commands spoken by the creator may be monitored. For example, the voice recognition unit (i.e., voice recognition unit 117) may be set to listen for or monitor keywords associated with video creation. The voice commands may be obtained, for example, by one or more microphones of the client device. The voice recognition unit may be started in response to the user indicating that they want to create content such as a video (i.e., the monitoring of voice commands may be started). For example, the user may select a button on the interface of the content application indicating that the user wants to create and / or upload a new video to the server. When the user selects a button indicating that the user wants to create and / or upload a new video to the server, the voice recognition unit may be activated. When the voice recognition unit is activated, the voice recognition unit may start listening for keywords in the voice commands spoken by the user.

[0054] At 404, in response to recognizing a first voice command spoken by the creator, recording of the content may be started. For example, when the voice recognition unit recognizes the first voice command, the voice recognition unit may start recording a video. The first voice command may include keywords associated with starting the recording of the video. Examples of keywords associated with starting the recording of a video include, but are not limited to, "start", "action", "record", "go", etc.

[0055] In 406, the voice input may be divided into a first voice stream for recognizing voice commands and a second voice stream for recording voice. Dividing the voice input into a first stream for recognizing voice commands and a second voice stream for recording voice can facilitate the recognition of voice commands, such as a second voice command indicating that recording should be stopped.

[0056] Content may continue to be recorded until the voice recognition unit recognizes the second voice command. In 408, in response to recognizing the second voice command spoken by the creator, the recording of the content may be stopped. For example, when the voice recognition unit recognizes the second voice command, the voice recognition unit may stop the recording of the video. The second voice command may include keywords associated with stopping or ending the video recording. Examples of keywords associated with stopping video recording include, but are not limited to, "stop", "cut", "end", etc.

[0057] In 410, a time stamp associated with the second voice command may be created. For example, if the recording of the content is stopped at time C in response to recognizing the second voice command or the like, a time stamp associated with time C may be generated. The time stamp may be stored in a database as time stamp data. The time stamp data may indicate the time associated with the time stamp and the corresponding operation (i.e., the recording is stopped) performed at that time.

[0058] Content may be edited using timestamp data. In 412, based on the timestamp, segments may be automatically deleted from the content. The segment may include a recording of a second voice command. For example, an editing department (i.e., editing department 119) may be set to delete one or more segments from the unedited video based on the timestamp data. The one or more segments to be deleted may include an unedited video in which a voice command, for example a second voice command, appears. For example, the segment to be deleted may include a segment in which there is a scene where the user speaks a voice command such as "pause" or "stop". Thus, in the final, edited video, there is no scene where the user speaks a voice command.

[0059] FIG. 5 shows an exemplary process 500 for content creation by voice control. Although described in FIG. 5 as a series of operations, those skilled in the art should understand that in various embodiments, the described operations may be added, removed, rearranged, or modified. In 502, the voice command spoken by the creator may be monitored using a voice recognition process (i.e., voice recognition unit). For example, the voice recognition unit (i.e., voice recognition unit 117) may be set to listen for or monitor keywords associated with video creation. The voice command may be obtained, for example, by one or more microphones of a client device. The voice recognition unit may be started in response to the user indicating a desire to create content such as a video (i.e., start monitoring voice commands). For example, the user may select a button on the interface of the content application indicating that the user wishes to create and / or upload a new video to the server. When the user selects a button indicating that the user wishes to create and / or upload a new video to the server, the voice recognition unit may be activated. When the voice recognition unit is activated, the voice recognition unit may start listening for keywords in the voice command spoken by the user.

[0060] In 504, in response to recognizing a first voice command spoken by the creator, recording of content may be started. For example, when the voice recognition unit recognizes the first voice command, the voice recognition unit may start recording a video. The first voice command may include keywords associated with starting the recording of the video. Examples of keywords associated with starting the recording of the video include, but are not limited to, "start", "action", "record", "go", etc.

[0061] In 506, while recording audio by the voice recording process, the voice command may be continuously monitored by the voice recognition process. Content including video and / or audio may be continuously recorded until the voice recognition unit recognizes a second voice command. In 508, in response to recognizing a second voice command spoken by the creator, recording of the content may be stopped. For example, when the voice recognition unit recognizes the second voice command, the voice recognition unit may stop recording the video. The second voice command may include keywords associated with stopping or ending the recording of the video. Examples of keywords associated with stopping the recording of the video include, but are not limited to, "stop", "cut", "end", etc.

[0062] In 510, a time stamp associated with the second voice command may be created. For example, when recording of the content is stopped at time C in response to recognizing the second voice command or the like, a time stamp associated with time C may be generated. The time stamp may be stored in the database as time stamp data. The time stamp data may indicate the time associated with the time stamp and the corresponding action (i.e., the recording is stopped) executed at that time.

[0063] The content may be edited using timestamp data. At 512, based on the timestamp, segments may be automatically deleted from the content. The segment may include a recording of a second voice command. For example, the editing department (i.e., editing department 119) may be set to delete one or more segments from the unedited video based on the timestamp data. The one or more segments to be deleted may include one or more unedited videos in which a voice command, for example a second voice command, appears. For example, the segments to be deleted may include segments where there is a scene where the user speaks a voice command such as "pause" or "stop". Thus, in the final, edited video, there is no scene where the user speaks a voice command.

[0064] FIG. 6 shows an exemplary process 600 for content creation by voice control. Although described as a series of operations in FIG. 6, those skilled in the art should understand that in various embodiments, the described operations may be added, removed, rearranged, or modified. At 602, voice commands spoken by the creator may be monitored. For example, the voice recognition unit (i.e., voice recognition unit 117) may be set to listen for or monitor keywords associated with video creation. The voice commands may be obtained, for example, by one or more microphones of the client device. The voice recognition unit may be started in response to the user indicating that they want to create content such as a video (i.e., the monitoring of voice commands may be started). For example, the user may select a button on the interface of the content application indicating that the user wants to create and / or upload a new video to the server. When the user selects a button indicating that the user wants to create and / or upload a new video to the server, the voice recognition unit may be activated. When the voice recognition unit is activated, the voice recognition unit may start listening for keywords in the voice commands spoken by the user.

[0065] In 604, in response to recognizing the first voice command spoken by the creator, recording of the content may be started. For example, when the voice recognition unit recognizes the first voice command, the voice recognition unit may start recording a video. The first voice command may include keywords associated with starting the recording of the video. Examples of keywords associated with starting the recording of the video include, but are not limited to, "start", "action", "record", "go", etc.

[0066] Recording of the content may continue until the voice recognition unit recognizes the second voice command. In 606, in response to recognizing the second voice command spoken by the creator, recording of the content may be stopped. For example, when the voice recognition unit recognizes the second voice command, the voice recognition unit may stop recording the video. The second voice command may include keywords associated with stopping or ending the recording of the video. Examples of keywords associated with stopping the recording of the video include, but are not limited to, "stop", "cut", "end", etc.

[0067] In 608, a time stamp associated with the second voice command may be created. For example, if recording of the content is stopped at time C in response to recognizing the second voice command or the like, a time stamp associated with time C may be generated. The time stamp may be stored in the database as time stamp data. The time stamp data may indicate the time associated with the time stamp and the corresponding action (i.e., the recording is stopped) performed at that time.

[0068] The content may be edited using timestamp data. At 610, based on the timestamp, the final content may be created by automatically deleting segments from the content. The segments include recordings of second voice commands. For example, the editing department (i.e., editing department 119) may be set to delete one or more segments from the unedited video based on the timestamp data. The one or more segments to be deleted may include one or more unedited videos where a voice command, such as a second voice command, appears. For example, the segments to be deleted may include segments where there are scenes where the user speaks voice commands such as "pause" or "stop". Thus, in the final, edited video, there are no scenes where the user speaks voice commands. At 612, the final content may be output. For example, the final edited video may be uploaded to a video service and distributed to other users of the video service for consumption.

[0069] FIG. 7 is a diagram showing an exemplary UI 700 according to the present disclosure. The user may indicate a desire to create content such as a video. For example, the user may select a button 702 on the interface of the content application to indicate that the user wants to create and / or upload a new video to the server. When the user selects the button 702 indicating that the user wants to create and / or upload a new video to the server, a preview mode may be started as an option. FIG. 8 is a diagram showing an exemplary UI 800 according to the present disclosure. The UI 800 displays a camera feed in the preview mode. In the preview mode, the user may be able to see the background or scenery present in the video. If the user does not like the background or scenery displayed in the preview mode, the user may move the client device and / or direct the camera towards a different background or scenery. If the user is satisfied with the background or scenery, the user may want to start shooting the video.

[0070] When the user wants to start shooting a video, the user may want to activate the listening mode. When the listening mode is activated, the voice recognition unit may be activated. When the voice recognition unit is activated, the voice recognition unit may start listening for keywords in the voice command spoken by the user. For example, the voice recognition unit may listen for keywords obtained by voice input of the client device. The voice input may include, for example, one or more microphone inputs.

[0071] To start the listening mode, the user may, for example, select button 802. Alternatively, when the user selects button 702 indicating that the user wants to create and / or upload a new video to the server, the client device may automatically enter the listening mode. For example, the client device may automatically enter the listening mode when it is in the preview mode.

[0072] When the client device is in the listening mode, as shown in FIG. 9, an instruction 902 for the listening mode may appear on the UI 900. The voice recognition unit may recognize the first keyword obtained by voice input at the first time. The first keyword may be a keyword associated with the start of video recording. For example, the first keyword may include "start", "action", "record", "go", etc. When the voice recognition unit recognizes the first keyword, the voice recognition unit may start video recording. For example, as shown in the UI 1000 of FIG. 10, the client device may enter the recording mode. The client device may remain in the recording mode (i.e., continue to record the video) until the voice recognition unit recognizes the second keyword obtained by voice input at the second time after the first time. The second keyword may be a keyword associated with the stop of the video recording. The second keyword may be obtained by voice input, for example, during the video recording. The second keyword may include "stop", "cut", "end", etc. When the voice recognition unit recognizes the second keyword, the voice recognition unit may trigger the video recording unit to end the video recording.

[0073] In some embodiments, the voice recognition unit (e.g., the voice recognition unit 117) may be set to start the execution of a specific operation during the content (e.g., video) editing process. If the user can use voice commands while editing the video, it may improve the editing process. For example, due to limited screen space, there may be many effects or filters that cannot be all displayed on a single screen. As a result, the user may have to scroll multiple times to search for a specific effect or filter. If the user can use voice commands to search for and / or apply filters or effects during the editing process instead, the editing process may be more efficient.

[0074] FIG. 11 is a schematic diagram 1100 showing a process for video editing by voice control according to the present disclosure. The user may indicate that the user wants to edit content (e.g., a previously recorded video) in any suitable manner. For example, the user may select a button indicating that the user wants to edit the video on the interface 108 of the content application 106. When the user selects a button indicating that the user wants to edit the video, the voice recognition unit 117 may be activated. When the voice recognition unit 117 is activated, the voice recognition unit 117 may start listening for keywords in the voice command spoken by the user. For example, the voice recognition unit 117 may listen for keywords obtained by voice input (e.g., voice input 1106 of the client device 104). The voice input 1106 may include, for example, one or more microphone inputs.

[0075] The voice recognition unit 117 may recognize a first keyword 1102 obtained by the voice input 1106 at a first time. For example, the first keyword 1102 may be a keyword associated with adding a visual effect during the video editing process. When the voice recognition unit 117 recognizes the first keyword 1102 associated with adding a visual effect during the video editing process, the video editing unit 119 may be triggered to add the visual effect to the video. Examples of keywords associated with adding a visual effect during the video editing process include, but are not limited to, "[visual effect name] add", "[visual effect name] apply", etc. As shown in the UI 1200 of FIG. 1200, the "[visual effect name]" may be the title of the visual effect, such as "Flash", "Smog", or "Flower".

[0076] The voice recognition unit 117 may recognize the second keyword 1104 acquired by the voice input 1106 at a second time after the first time. The second keyword 1104 may be a keyword associated with adding a motion effect during the video editing process. When the voice recognition unit 117 recognizes the second keyword 1104 associated with adding a motion effect during the video editing process, the video editing unit 119 may be triggered to add the motion effect to the video. Examples of keywords associated with adding a motion effect during the video editing process include, but are not limited to, "[Motion effect name] add", "[Motion effect name] apply", etc. As shown in the UI 1300 of FIG. 13, the "[Motion effect name]" may be the title of the motion effect, such as "wing", "counter", "bean", or "handwriting".

[0077] Additionally or alternatively, when the voice recognition unit 117 recognizes a keyword associated with adding a filter during the video editing process, the voice recognition unit 117 may trigger the video editing unit 119 to add the filter to the video. Examples of keywords associated with adding a filter during the video editing process include, but are not limited to, "[Filter name] add", "[Filter name] apply", etc.

[0078] FIG. 14 shows a computing device that may be used in various ways, such as the services, networks, modules, and / or devices shown in FIG. 1. With respect to the architecture example of FIG. 1, the message service, interface service, processing service, content service, cloud network, and client may each be implemented by one or more instances of the computing device 1400 of FIG. 14. The computer architecture shown in FIG. 14 represents a conventional server computer, workstation, desktop computer, laptop computer, tablet, network device, PDA, electronic reader, digital cellular phone, or other computing node, and may be used to execute any aspect of the computers described herein, such as implementing the methods described herein.

[0079] The computing device 1400 may include a substrate or "motherboard" which is a printed circuit board that can be connected to a plurality of components or devices via a system bus or other electrical communication path. One or more central processing units (CPUs) 1404 may operate in conjunction with a chipset 1406. The CPU 1404 may be a standard programmable processor that performs arithmetic and logical operations necessary for the operation of the computing device 1400.

[0080] The CPU 1404 may be executed by operating switching elements that distinguish and change these states to perform operations necessary to transition from one discrete physical state to the next. The switching elements may typically include an electronic circuit that maintains one of two binary states, such as a flip-flop, and an electronic circuit that provides an output state based on a logical combination of the states of one or more other switching elements, such as a logic gate. By combining these basic switching elements, more complex logic circuits including registers, adders / subtractors, arithmetic logic units, floating point units, etc. may be created.

[0081] The CPU 1404 may be augmented or replaced by other processing units such as a GPU. The GPU may include processing units specialized for, but not necessarily limited to, high - level parallel computing such as graphics and other visualization - related processing.

[0082] The chipset 1406 may provide an interface between the CPU 1404 and the remaining components and devices on the substrate. The chipset 1406 may provide an interface to the random - access memory (RAM) 1408 used as the main memory within the computing device 1400. The chipset 1406 may also provide an interface to a computer - readable storage medium, such as a read - only memory (ROM) 1420 or non - volatile RAM (NVRAM) (not shown), for storing basic routines that can start up the computing device 1400 and facilitate the transmission of information between various components and devices. According to the aspects described herein, the ROM 1420 or NVRAM may store other software components necessary for the operation of the computing device 1400.

[0083] The computing device 1400 may operate in a network environment using logical connections to remote computing nodes and computer systems via a local - area network (LAN). The chipset 1406 may include functionality for providing a network connection via a network - interface controller (NIC) 1422, such as a gigabit Ethernet adapter. The NIC 1422 may be capable of connecting the computing device 1400 to other computing nodes via the network 1416. It should be understood that multiple NICs 1422 may be present within the computing device 1400 to connect the computing device to other types of networks and remote computer systems.

[0084] Computing device 1400 may be connected to a mass storage device 1428 that provides a non-volatile memory device for a computer. The mass storage device 1428 may store system programs, application programs, other program modules, and data described in more detail herein. The mass storage device 1428 may be connected to the computing device 1400 via a memory controller 1424 connected to the chipset 1406. The mass storage device 1428 may be composed of one or more physical storage units. The mass storage device 1428 may include a management unit 1410. The memory controller 1424 may interface with the physical storage unit via a serial attached SCSI (SAS) interface, a serial advanced technology attachment (SATA) interface, a fiber channel (FC) interface, or other types of interfaces for physically connecting and transmitting data between the computer and the physical storage unit.

[0085] The computing device 1400 may store data on the mass storage device 1428 by converting the physical state of the physical storage unit to reflect the stored information. The specific conversion of the physical state may depend on various factors and different embodiments of this specification. Examples of such factors include, but are not limited to, the technology for implementing the physical storage device and whether the mass storage device 1428 is characterized as a primary storage device or a secondary storage device, etc.

[0086] For example, computing device 1400 may store information in mass storage device 1428 by issuing instructions via memory controller 1424 to change the magnetic properties at specific locations within a magnetic disk drive unit, the reflective or refractive properties at specific locations within an optical storage unit, or the electrical properties of specific capacitors, transistors, or other discrete components within a solid-state memory unit. Without departing from the scope and spirit of this specification, other conversions of physical media are possible, and the foregoing examples are provided merely to facilitate its explanation. Computing device 1400 may further read information from mass storage device 1428 by detecting the physical state or characteristics of one or more specific locations within a physical memory unit.

[0087] In addition to the aforementioned mass storage device 1428, computing device 1400 may be able to access other computer-readable storage media for storing and retrieving information such as program modules, data structures, or other data. One of ordinary skill in the art should understand that a computer-readable storage media can be any available media that provides non-transitory data storage and is accessible by computing device 1400.

[0088] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, temporary and non-temporary computer-readable storage media, as well as removable and non-removable media implemented in any method or technology. Computer-readable storage media includes, but is not limited to, RAM, ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory or other solid-state memory technologies, compact disc ROM (CD-ROM), digital versatile disc (DVD), high-definition DVD (“HD-DVD”), BLU-RAY or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, other magnetic storage devices, or any other media that can be used to store desired information in a non-transitory manner.

[0089] A mass storage device such as the mass storage device 1428 shown in FIG. 14 may store an operating system for controlling the operation of the computing device 1400. The operating system may include one version of the LINUX operating system. The operating system may include one version of the MICROSOFT WINDOWS® SERVER operating system. According to another aspect, the operating system may include one version of the UNIX® operating system. Also, various mobile phone operating systems such as IOS and ANDROID® may be used. It should be understood that other operating systems may also be used. The mass storage device 1428 may store other systems or applications and data used by the computing device 1400.

[0090] The mass storage device 1428 or other computer-readable storage media may also be encoded with computer-executable instructions that, when loaded into the computing device 1400, convert the computing device from a general-purpose computing system to a dedicated computer capable of implementing the aspects described herein. As described above, these computer-executable instructions convert the computing device 1400 by defining how the CPU 1404 transitions between states. The computing device 1400 may be able to access a computer-readable storage medium that stores computer-executable instructions that can execute the methods described herein when executed by the computing device 1400.

[0091] A computing device, such as the computing device 1400 shown in FIG. 14, may further include an input / output controller 1432 for receiving and processing inputs from a plurality of input devices such as a keyboard, a mouse, a touchpad, a touch screen, an electronic stylus pen, or other types of input devices. Similarly, the input / output controller 1432 may provide outputs to a display such as a computer monitor, a flat panel display, a digital projector, a printer, a plotter, or other types of output devices. It should be understood that the computing device 1400 may not include all of the components shown in FIG. 14, may include other components not explicitly shown in FIG. 14, or may utilize an architecture that is completely different from the architecture shown in FIG. 14.

[0092] As described herein, the computing device may be a physical computing device such as the computing device 1400 of FIG. 14. The computing node may also include a virtual machine host process and one or more virtual machine instances. The computer-executable instructions may be indirectly executed by the physical hardware of the computing device by interpreting and / or executing instructions stored and executed within the context of the virtual machine.

[0093] It should be understood that the methods and systems are not limited to a particular method, a particular component, or a particular embodiment. It should also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting.

[0094] As used in the specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed using the antecedent "about," it should be understood that the particular value forms another embodiment. Further, it should be understood that each end point of a range is significant both in relation to the other end point and independently of the other end point.

[0095] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and the specification includes instances where said event or circumstance occurs and instances where it does not.

[0096] Throughout the description and claims of this specification, the word "comprising" and variations of the word such as "comprises" and "comprising" are meant to mean "including but not limited to," and are not intended to exclude, for example, other components, integers, or steps. "Exemplary" is used to mean "an example of" and is not intended to convey an indication of a preferred or desired embodiment. "Such as" is not used in a limiting sense and is used for interpretive purposes.

[0097] Components that can be used to implement the described methods and systems are described. When describing combinations, subsets, interactions, groups, etc. of these components, specific references to each of the various individual and collective combinations and permutations of these components may not be explicitly described, and it should be understood that each of them is specifically contemplated and described herein for all methods and systems. This applies to all aspects of the present application, including but not limited to operations in the described methods. Thus, if there are various additional operations that can be performed, it should be understood that each of these additional operations can be performed in any particular embodiment or combination of embodiments of the described method.

[0098] The present method and system can be more easily understood by referring to the following detailed description of the preferred embodiments and examples included therein, as well as the accompanying drawings and their descriptions.

[0099] As those skilled in the art will understand, the method and system may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Further, the present method and system may take the form of a computer program product on a computer-readable storage medium having computer-readable program instructions (e.g., computer software) embodied thereon. More specifically, the present method and system may take the form of web-implemented computer software. Any suitable computer-readable storage medium, including a hard disk, CD-ROM, optical storage device, or magnetic storage device, may be utilized.

[0100] Referring to the block diagrams and flowcharts of methods, systems, apparatuses, and computer program products, embodiments of the methods and systems will be described below. It should be understood that each block of the block diagrams and flowcharts, as well as combinations of blocks in the block diagrams and flowcharts, may be implemented by computer program instructions. These computer program instructions may be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed on the computer or other programmable data processing apparatus generate means for realizing the functions specified in one or more blocks of the flowchart.

[0101] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including computer-readable instructions for realizing the functions specified in one or more blocks of the flowchart. The computer program instructions may be loaded onto a computer or other programmable data processing apparatus to provide steps for realizing the functions specified in one or more blocks of the flowchart, and to cause a series of operation steps to be executed on the computer or other programmable data processing apparatus to generate a computer-implemented process.

[0102] The various characteristics and processes described above may be used independently of each other or may be combined in various ways. All possible combinations and sub - combinations are intended to fall within the scope of the present disclosure. Further, in some implementations, some methods or process blocks may be omitted. The methods and processes described herein are not limited to any particular order, and the associated blocks or states may be executed in other suitable orders. For example, the described blocks or states may be executed in an order other than the specifically described order, or a plurality of blocks or states may be combined within a single block or state. The exemplary blocks or states may be executed sequentially, in parallel, or in some other way. Blocks or states may be added to or removed from the described exemplary embodiments. The exemplary systems and components described herein may be configured differently from those described. For example, elements may be added, removed, or rearranged compared to the described exemplary embodiments.

[0103] Also, various items are shown to be stored in memory or on a storage device during use, and it should be understood that these items or portions thereof may be transferred between memory and other storage devices for memory management and data integrity purposes. Alternatively, in other embodiments, software modules and / or portions or all of the system may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. Further, in some embodiments, portions or all of the system and / or module may be implemented or provided in other ways, such as at least partially in firmware and / or hardware, where the hardware includes, but is not limited to, one or more application specific integrated circuits (ASICs), standard integrated circuits, controllers (e.g., including microcontrollers and / or embedded controllers by executing appropriate instructions), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), etc. Portions or all of the modules, systems, and data structures may be stored (e.g., as software instructions or structured data) on a computer-readable medium such as a hard disk, memory, network, or portable media product for reading by an appropriate device or via an appropriate connection. The systems, modules, and data structures may be transmitted as generated data signals (e.g., as part of a carrier wave or other analog or digital propagation signal) on various computer-readable transmission media including wireless-based media and wired / cable-based media, and the computer-readable transmission media may take various forms (e.g., as part of a single or multiplexed analog signal or as multiple discrete digital packets or frames). In other embodiments, such computer program products may take other forms. Thus, the present invention can be implemented with other computer system configurations.

[0104] While methods and systems have been described in connection with preferred embodiments and specific examples, the embodiments herein are intended to be exemplary rather than limiting in all respects, so it is not intended to limit the scope to specific embodiments.

[0105] Unless otherwise specified, the methods described herein do not require that their operations be performed in a particular order. Thus, where a method claim does not actually recite an order for the operations to follow, or where the operations are not specifically recited in the claims or the specification as being limited to a particular order, no inference of order is intended in any manner. This applies to any possible non-expressive basis for interpretation, including logical issues regarding the placement of steps or the operation flow, the plain meaning derived from grammatical construction and punctuation, the number or type of embodiments described in the specification.

[0106] As will be apparent to those skilled in the art, various modifications and changes are possible without departing from the scope or spirit of the disclosure. Considering this specification and the practice described herein, other embodiments will be apparent to those skilled in the art. This specification and the exemplary drawings are only intended to be regarded as exemplary, and the true scope and spirit are indicated by the following claims.

Claims

1. A method for content creation by voice control, comprising: monitoring voice commands spoken by a creator; starting recording of content in response to recognizing a first voice command spoken by the creator; stopping recording of the content in response to recognizing a second voice command spoken by the creator; creating a time stamp associated with the second voice command; automatically deleting, based on the time stamp, a segment including the recording of the second voice command from the content; A method comprising the above steps.

2. Recording the content includes recording video and audio. The method according to claim 1.

3. The method according to claim 2, further comprising splitting an audio input into a first audio stream for recognizing the voice command and a second audio stream for recording the audio. The method according to claim 2.

4. monitoring the voice command by a voice recognition process; recording the audio by a voice recording process; The method according to claim 2, further comprising the above steps.

5. creating final content by automatically deleting the segment from the content; outputting the final content; The method according to claim 1, further comprising the above steps.

6. The method according to claim 1, wherein the voice command further includes a voice command indicating adding at least one visual effect or auditory effect to the content. The method according to claim 1.

7. The method according to claim 1, wherein the voice command further includes a voice command indicating searching for a video associated with a specific topic. The method according to claim 1.

8. The method according to claim 1, wherein the voice command further includes a voice command indicating editing the content in a specific way during an editing process. The method according to claim 1.

9. A system for content creation by voice control, comprising: at least one computing device communicating with a computer memory; the computer memory including computer-readable instructions; the computer-readable instructions, when executed by the at least one computing device, cause the system to: monitor voice commands spoken by a creator; start recording of content in response to recognizing a first voice command spoken by the creator; In response to recognizing a second voice command spoken by the creator, stopping the recording of the content; Creating a time stamp associated with the second voice command; Automatically deleting, based on the time stamp, a segment including the recording of the second voice command from the content; Configuring to execute an operation including: A system. **Claim 10** Recording the content includes recording video and audio. The system according to claim 9. **Claim 11** The operation further includes: Splitting an audio input into a first audio stream for recognizing the voice command and a second audio stream for recording the voice. The system according to claim 10. **Claim 12** The operation further includes: Monitoring the voice command by a voice recognition process; Recording the voice by a voice recording process. The system according to claim 10. **Claim 13** The operation further includes: Creating final content by automatically deleting the segment from the content; Outputting the final content. The system according to claim 9. **Claim 14** The voice command further includes at least one of a voice command indicating adding at least one visual effect or auditory effect to the content, or a voice command indicating searching for a video associated with a specific topic. The system according to claim 9. **Claim 15** A non-transitory computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, cause the processor to: Monitor a voice command spoken by a creator; Start recording content in response to recognizing a first voice command spoken by the creator; Stop recording the content in response to recognizing a second voice command spoken by the creator; Create a time stamp associated with the second voice command; Automatically delete, based on the time stamp, a segment including the recording of the second voice command from the content. A non-transitory computer-readable storage medium for executing an operation including the above. **Claim 16** ​ ​ ​ ​ ​ ​ ​ Recording the content includes recording video and audio. The non-transitory computer-readable storage medium according to claim 15. **Claim 17** The operation further includes splitting the voice input into a first voice stream for recognizing the voice command and a second voice stream for recording the voice. The non-transitory computer-readable storage medium according to claim 16. **Claim 18** The operation further includes monitoring the voice command by a voice recognition process and recording the voice by a voice recording process. The non-transitory computer-readable storage medium according to claim 16. **Claim 19** The operation further includes creating final content by automatically deleting the segment from the content and outputting the final content. The non-transitory computer-readable storage medium according to claim 16. **Claim 20** The voice command further includes a voice command indicating adding at least one visual effect or auditory effect to the content. The non-transitory computer-readable storage medium according to claim 16.

Citation Information

Patent Citations

  • Recording / Reproducing device provided with voice control function, recording / Reproducing method and storage medium

    JP2000076737A

  • Voice control type audiovisual recording device and voice control method

    JP2001203974A

  • Imaging apparatus, control method, and program

    JP2021145256A