Video editing method and device, electronic equipment and computer readable medium
By acquiring and parsing voice data to implement video editing commands, and combining user characteristics and historical operation configuration templates, the complexity of existing video editing methods is solved, providing a more user-friendly voice editing interface and simplifying the operation process.
Patent Information
- Application Number
- CN202310110611.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-10
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2043-02-10
AI Technical Summary
Existing video editing methods are complex and unfriendly through UI and user interaction, especially for specific groups such as the elderly and children.
By acquiring and parsing voice data to obtain target video editing instructions, and modifying the video based on these instructions, the system provides both a voice editing interface and a UI control editing interface. It also combines user characteristic information and historical operation data to configure template modules, thereby enabling video editing.
It simplifies video editing operations, improves user-friendliness, and allows users to easily edit videos using voice commands, reducing learning costs and operational complexity.
Smart Images

Figure CN118488263B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video, in particular to a video editing method and device, electronic equipment and computer readable medium. BACKGROUND
[0002] With the development of communication technology, the intelligent degree of terminal devices such as mobile phones and tablet computers is continuously improved to meet the various needs of users. At present, a video editor can be installed in a terminal device, and a user can select a certain number of pictures and / or video clips, confirm the playing order of the pictures and / or video clips, additionally add background music to the video, and finally play the pictures and / or video clips selected by the user in sequence to generate a video. SUMMARY
[0003] The present application provides a video editing method, device, electronic equipment and computer readable medium to improve the above-mentioned defects.
[0004] In a first aspect, the embodiments of the present application provide a video editing method applied to an electronic device, the method comprising: obtaining voice data; parsing the voice data to obtain a target video editing instruction; and modifying a first video based on the target video editing instruction to obtain a second video.
[0005] In a second aspect, the embodiments of the present application also provide a video editing device applied to an electronic device, the device comprising: an obtaining unit, a parsing unit and a modifying unit. The obtaining unit is configured to obtain voice data. The parsing unit is configured to parse the voice data to obtain a target video editing instruction. The modifying unit is configured to modify a first video based on the target video editing instruction to obtain a second video.
[0006] In a third aspect, the embodiments of the present application also provide an electronic device, comprising: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to execute the above-mentioned method.
[0007] In a fourth aspect, the embodiments of the present application also provide a computer readable medium, wherein the readable storage medium stores a program code executable by a processor, and the program code is executed by the processor to make the processor execute the above-mentioned method.
[0008] The video editing method and device, the electronic device and the computer readable medium provided in the application obtain voice data, analyze the voice data to obtain a target video editing instruction, and modify a first video based on the target video editing instruction to obtain a second video. Therefore, a user can edit a video to obtain an edited video by inputting voice data, and the operation is more friendly compared with a traditional UI interaction operation.
[0009] Other features and advantages of the embodiments of the present application will be described in the following description, and part of them will become apparent from the description, or will be understood through implementation of the embodiments of the present application. The purposes and other advantages of the embodiments of the present application can be achieved and obtained through the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0011] Figure 1 A method flowchart of a video editing method provided by an embodiment of the present application is shown;
[0012] Figure 2 A method flowchart of a video editing method provided by another embodiment of the present application is shown;
[0013] Figure 3 A schematic diagram of an editing selection interface provided by an embodiment of the present application is shown;
[0014] Figure 4 A schematic diagram of a first video editing interface provided by an embodiment of the present application is shown;
[0015] Figure 5 A schematic diagram of a second video editing interface provided by an embodiment of the present application is shown;
[0016] Figure 6 A method flowchart of a video editing method provided by still another embodiment of the present application is shown;
[0017] Figure 7 A schematic diagram of an instruction selection interface provided by an embodiment of the present application is shown;
[0018] Figure 8 A schematic diagram of an instruction selection interface provided by another embodiment of the present application is shown;
[0019] Figure 9A schematic diagram of a video editing process provided in an embodiment of this application is shown;
[0020] Figure 10 A block diagram of a video editing apparatus provided in one embodiment of this application is shown;
[0021] Figure 11 A module block diagram of a video editing apparatus provided in another embodiment of this application is shown;
[0022] Figure 12 A block diagram of an electronic device provided in an embodiment of this application is shown;
[0023] Figure 13 The present application provides a storage unit for storing or carrying program code that implements the video editing method according to the present application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.
[0025] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0026] With the development of communication technology, the intelligence level of terminal devices such as mobile phones and tablets is constantly improving to meet various user needs. Currently, video editors can be installed on terminal devices. Users can select a certain number of pictures and / or video clips, then confirm the playback order of the pictures and / or video clips. In addition, users can add background music to the video, and finally play the selected pictures and / or video clips in sequence to generate a video.
[0027] Current video editing generally relies on UI and user operation to complete complex editing tasks. Whether on PC or mobile, there is a complete video editing interface, and users can complete video editing by clicking different buttons.
[0028] However, the inventors found in their research that editing videos through the aforementioned UI and user interaction has a high learning curve for users. They need to learn the functions of each UI control and understand the meaning of each interactive operation. Therefore, the operation is complicated and not user-friendly, especially for certain groups of people (such as the elderly and children).
[0029] Please see Figure 1 , Figure 1 An embodiment of this application provides a video editing method applied to the aforementioned electronic device. The method includes steps S101 to S103.
[0030] S101: Acquire voice data.
[0031] For example, while the first video is playing on the screen of an electronic device, the user can input voice data through the electronic device. Alternatively, the electronic device can acquire the user-inputted voice data even when the first video is not playing. Furthermore, the voice data may not be user-inputted but automatically generated by other devices or software; for example, voice software may automatically generate and play voice based on user-inputted text content. This is not a limitation.
[0032] S102: Parse the voice data to obtain the target video editing instructions.
[0033] The electronic device analyzes the voice data to obtain its corresponding semantic content, which is a natural language description of the voice data. After obtaining the voice content, the target video editing instruction can be derived from it. For example, the semantic content of the voice data might be "Please delete 1-2 seconds of video content." The electronic device parses the voice data to obtain this semantic content and uses it as the target video editing instruction. Alternatively, a natural language description can be pre-set for each video editing instruction. This description can be understood as a password or keyword corresponding to the video editing instruction. After parsing the voice data to obtain its corresponding voice content, the electronic device searches for the keywords corresponding to that voice content to obtain the video editing instruction corresponding to those keywords, which is then used as the target video editing instruction.
[0034] S103: Modify the first video based on the target video editing instructions to obtain the second video.
[0035] It should be noted that the first video can be a pre-acquired video, such as a video downloaded in advance from the network, a video captured in advance by the camera of an electronic device, a video captured in advance by video capture software of an electronic device, or a video generated from pre-acquired material data, and there is no limitation here.
[0036] In one implementation, the voice data can be collected by the electronic device while playing a first video. For example, upon detecting user input of voice data, the electronic device pauses the playback of the first video, edits the first video based on the target video editing instruction corresponding to the voice data, and then resumes playback of the first video. The resulting video after the first video finishes playback is recorded as the second video. Alternatively, during the playback of the first video, multiple voice data points can be collected, and the target video editing instruction corresponding to each voice data point can be determined. Instead of executing the editing operation immediately, the video editing instruction is recorded. After the first video finishes playback, all video editing instructions are executed uniformly to obtain the second video. In other words, in this embodiment, the first video can be modified immediately upon determining the target editing instruction during playback. After the user confirms the completion of the modification, the editing of the first video can be considered complete, thus obtaining the second video, and the modified first video can be saved as the second video.
[0037] Therefore, the video editing method provided in this application acquires voice data; parses the voice data to obtain target video editing instructions; and modifies a first video based on the target video editing instructions to obtain a second video. Thus, for videos generated from source material, users can edit the video by inputting voice data to obtain the edited video, which is more user-friendly compared to traditional UI interaction.
[0038] Please see Figure 2 , Figure 2 An embodiment of this application provides a video editing method applied to the aforementioned electronic device, which includes steps S201 to S204.
[0039] S201: Generate the first video based on the pre-acquired material data.
[0040] The source material can be an initial video recorded by the user using a video capture device, an initial image captured by an image capture device, or an initial audio captured by an audio capture device. It can also be an initial video, initial image, or initial audio downloaded from the internet, or an initial video, initial image, or initial audio sent by other devices; there are no limitations on this. After acquiring the aforementioned initial video, initial image, and / or initial audio source material, the electronic device generates a video based on that source material, which is named the first video. For example, the first video can be generated by inserting initial audio at a specific timestamp in the initial video, i.e., adding background music (BGM) to the initial video. The specific method of generating the first video from the source material is not limited here.
[0041] S202: When the first video editing interface corresponding to the first video is displayed on the screen of the electronic device, acquire voice data.
[0042] It should be noted that editing the first video can include two interfaces: a first video editing interface and a second video editing interface. The first video editing interface refers to the interface where the user edits the video using voice commands, while the second video editing interface refers to the interface where the user edits the video using UI controls. In other words, when the first video editing interface corresponding to the first video is displayed on the screen of the electronic device, an editing selection interface is displayed on the screen of the electronic device before acquiring voice data. This editing selection interface includes a first selection control. If a first operation is detected on the first selection control, the first video editing interface is displayed in response to the first operation; if a second operation is detected on the second selection control, the second video editing interface is displayed in response to the second operation.
[0043] For example, after obtaining the first video, an editing control is provided on the display interface of the first video. When the user triggers the editing control, they can enter the editing selection interface, such as... Figure 3 As shown, the editing selection interface includes a manual editing selection control 301 and a voice editing selection control 302. The first selection control is the voice editing selection control 302, and the second selection control is the manual editing selection control 301.
[0044] Users can access the voice editing selection control 302 by clicking on it. Figure 4 The first video editing interface shown can be accessed by clicking the manual editing selection control 301. Figure 5 The second video editing interface is shown. Figure 4 As shown, the first video editing interface is for users to modify the first video via voice commands, as follows: Figure 4As shown, the first video editing interface displays the playback screen and progress bar corresponding to the first video. Users can see the content of the first video in the first video editing interface and control the playback content of the first video in the interface by dragging the progress bar.
[0045] In addition, the first video editing interface also displays a voice control 401. By triggering the voice control 401, the user can perform voice editing on the video playing on the first video editing interface while it is displayed on the screen. For example, if the user presses and holds the voice control 401 in the first video editing interface while inputting voice, the electronic device can collect the user's voice input data.
[0046] Furthermore, such as Figure 5 As shown, Figure 5 The second video editing interface shown refers to the interface through which users edit videos using the UI controls on that screen, such as... Figure 5 As shown, the second video editing interface displays multiple editing controls 501, each corresponding to a video editing instruction. The user inputs different video editing instructions by triggering the editing control. That is, from the at least one editing control, a target editing control is determined; based on the video editing instruction corresponding to the target editing control, the first video is modified, and the resulting video can be named the third video. It should be noted that both the aforementioned first and second operations can be click operations.
[0047] Therefore, after the electronic device generates the first video based on the pre-acquired material, the user displays the first video editing interface corresponding to the first video on the screen of the electronic device through the aforementioned operations. Then, the user inputs voice in the first video editing interface, and the electronic device collects the voice to obtain voice data.
[0048] S203: Parse the voice data to obtain the target video editing instructions.
[0049] S204: Modify the first video based on the target video editing instructions to obtain the second video.
[0050] One implementation method is to pause the playback of the first video within the first video editing interface when user-input voice data is detected. After editing the first video based on the target video editing instruction corresponding to the voice data, playback resumes. Once the first video finishes playing and the user confirms completion of the editing within the first video editing interface, the resulting video is recorded as the second video. Alternatively, during the playback of the first video within the first video editing interface, multiple voice data points are collected, and the target video editing instruction corresponding to each voice data point is determined. The editing operation is not executed immediately; instead, the video editing instruction is recorded. After the first video finishes playing, all video editing instructions are executed uniformly to obtain the second video. In other words, in this embodiment, the first video can be modified immediately upon determining the target editing instruction during playback. After the user confirms completion of the modification (e.g., exiting the first video editing interface), the editing of the first video can be considered complete, thus obtaining the second video, and the modified first video is saved as the second video.
[0051] Therefore, the video editing method provided in this application generates a first video based on pre-acquired source material; while displaying a first video editing interface corresponding to the first video on the screen of the electronic device, it acquires voice data; it parses the voice data to obtain a target video editing instruction; and it modifies the first video based on the target video editing instruction to obtain a second video. Thus, for videos generated from source material, users can edit the video by inputting voice data to obtain an edited video, which is more user-friendly compared to traditional UI interaction.
[0052] Please see Figure 6 , Figure 6 An embodiment of this application provides a video editing method applied to the aforementioned electronic device, which includes steps S601 to S604.
[0053] S601: Generate a first video based on pre-acquired material data through a template module. The template module includes multiple functional modules, each corresponding to a video editing operation.
[0054] The template module is used to perform a series of editing operations on the video based on the media resources given by the user. The template module includes multiple functional modules, each of which corresponds to a video editing operation. In other words, the template module can be understood as performing the video editing operations corresponding to each functional module on the input material based on the functional modules configured in the template module, and outputting a video, which is denoted as the first video.
[0055] It is understandable that this template module can use multiple functional modules, and the configured functional modules can be at least one of the aforementioned available functional modules; that is, it can be all or some of them. Considering that users' needs for similar functions in the generated video are not all-encompassing—for example, when choosing video filters, they usually choose one from several available filters, rather than using all filters in the same video—it is necessary to select certain functional modules from among these multiple modules as the template module's functional modules. The parameters of each functional module should also be adjusted accordingly. For example, if a functional module's function is to adjust video brightness, setting its parameter to +5 corresponds to increasing the video brightness by 5%, and setting the parameter to -10 corresponds to decreasing the video brightness by 10%. Therefore, the functional modules of the template module need to be configured before use.
[0056] For example, the template module can be configured in two ways. The first is by user configuration. Specifically, the template module can have a configuration interface that displays the identifiers of different functional modules and a description of each module's function. Users can view the function of each module on this interface and select the desired module as the template module's functional module. The template module can then edit the pre-acquired materials to obtain the first video based on the user-configured functional modules.
[0057] It should be noted that users can also configure the template module based on input voice commands. Specifically, during the template module configuration phase, the user inputs a configuration voice command. The electronic device recognizes and parses this voice command, obtaining multiple different video editing requirements. Then, it determines the functional module corresponding to each requirement as the functional module of the template module. For example, if the input voice command 1 is "Add a transition between videos and change the music to music I've been listening to recently," then the electronic device, recognizing this voice command 1, determines the video editing requirements as transition selection, adding transition effects, setting transition duration, selecting audio files, and replacing audio tracks. Then, it determines the corresponding functional modules for each of these requirements, thus completing the configuration operation of the template module's functional modules. Furthermore, considering that after recognizing the various video editing requirements obtained from the aforementioned voice command 1, a certain video editing requirement may correspond to multiple functional modules, the electronic device can display the identifiers and function descriptions of all functional modules on the screen. The user can then select the desired functional module from these multiple modules as the functional module of the template module. Of course, it's also possible to arbitrarily select a functional module for each video editing need, combine these modules to create multiple initial modules, and then allow the user to choose the desired module from these initial modules as the final template module. Specific details are not limited here.
[0058] The second approach is to determine template modules based on user characteristic information. For example, the video type of the first video can be determined, and then the corresponding functional modules can be identified as included in the template module. For instance, if the video type is determined to be a game, the corresponding video effects can be identified as Effect 1 and Effect 2 based on the user's characteristic information. Effect 1 and Effect 2 can be video display effects, special effects, music, etc., without limitation. Then, Effect 1 is identified as corresponding to Function Module 1, meaning that the editing operation of Function Module 1 can bring the effect of Effect 1 to the video. Effect 2 is identified as corresponding to Function Module 2, and Function Module 1 and Function Module 2 can then be used as the template module. Specifically, the user's interest in different types of videos with different effects can be determined based on the user's characteristic information. Different effects can be the video editing functions corresponding to the functional modules supported by the template module, such as audio, transition types, transition effects, and transition duration, without limitation. This determines the video editing effects that the user is interested in when processing this type of video, thereby identifying the interested functional modules and configuring the interested template modules.
[0059] Specifically, the feature information includes at least one of the following: user identity information, user interest information, attribute information of the product used by the user, and operational information of the product used by the user. The operational information of the product used by the user may be operational information of a video editing application installed on an electronic device. The feature information includes feature identifiers and feature data. The feature identifier reflects the specific object of the product operated by the user when operating the preset product, thus reflecting the product parameters of the preset product. Specifically, the feature identifier can also be user information, in which case it can be a user identity tag, a user interest tag, an attribute tag of the product used by the user, and an operational information tag of the product used by the user. The feature data is the data corresponding to the feature identifier. For example, if the feature identifier is a user identity tag, then the data corresponding to the feature identifier is user identity information.
[0060] Specifically, the feature data can be obtained through a data acquisition component installed within a preset product. More specifically, the specific implementation for obtaining the feature information of a user using the preset product involves acquiring the user's feature information sent by the data acquisition component within the preset product. This data acquisition component can be a program module capable of collecting feature data corresponding to feature identifiers within the preset product and sending the feature representation and feature data together to a server.
[0061] Specifically, the product refers to electronic devices and software products. The software product can be an application program. Therefore, the data acquisition component is an SDK component installed within the system and application program of the electronic device. An SDK, or Software Development Kit, is a collection of development tools used to build application software for specific software packages, software frameworks, hardware platforms, operating systems, etc. Specifically, the SDK component is installed within the client or electronic device and is bound to it. The SDK component provides an access interface for other applications to access the client's data, and can also actively collect data from the client and send it to other clients or terminals.
[0062] In this embodiment of the application, a data acquisition component is set within the operating system and application of the electronic device. The data acquisition component acquires feature information based on the pre-set embedding points within the electronic device or application and sends it to the server. As one implementation, multiple embedding points are set within the application or the operating system of the electronic device. Each embedding point represents the user's operation behavior or the running state of the application. For example, the launch of the application is a embedding point, and the user opening a certain interface of the application is a embedding point. When these behaviors occur, that is, when the embedding points are triggered, the corresponding data can be stored in the local storage space corresponding to the application. As the data to be reported by the application, it is stored locally, and the data to be reported is feature information.
[0063] In this embodiment, the server can integrate user data reported by the system SDK and application SDK of the electronic device, and use statistical and data mining techniques to extract and standardize user features to construct a comprehensive and three-dimensional user profile. Specifically, the user data is feature data, and the user profile is obtained based on the feature data and feature identifiers. In this embodiment, the user profile includes user basic tags, user interest and preference tags, user device attribute and behavior tags, user application behavior tags, etc. Among them, the user basic tags correspond to user identity tags, which refer to the user's basic demographic attribute tags (including gender, age, region, etc.). The feature data corresponding to these tags is user identity data, and the acquisition methods for this data include user reporting, algorithm mining, etc. User interest and preference tags correspond to user interest tags, which correspond to the user's interest content. The acquisition methods for these tags can also be user reporting, algorithm mining, etc. User device attribute tags correspond to the attribute information of the product used by the user. The corresponding feature data is the configuration parameters of the product used by the user, such as memory capacity, battery capacity, or screen size. The acquisition methods for these tags can be user reporting or collection through SDK components within the user device. User device behavior tags correspond to user operation tags on electronic devices. The corresponding feature data is the data generated by the user's operation of the electronic device, which can be obtained by collecting data through the SDK components within the electronic device's operating system. User application behavior tags correspond to user operation tags on applications installed on electronic devices. The corresponding feature data is the data generated by the user's operation of applications installed on the electronic device, which can also be obtained by collecting data through the SDK components within the electronic device's applications.
[0064] It should be noted that the difference between user operation information on electronic devices and user operation information on applications installed on electronic devices can be that user operation information on electronic devices refers to data generated by the user operating system applications on the electronic device, such as data generated by the user operating the camera interface or Bluetooth connection function on the electronic device, while user operation information on applications installed on the electronic device refers to data generated by the user collecting data from non-system applications on the electronic device, such as data generated by applications other than the system applications that come with the electronic device.
[0065] Therefore, electronic devices can determine a user's profile through the aforementioned server, i.e., obtain the user's characteristic information, and thus determine the video editing effects that the user is interested in. For example, user characteristic information includes the user's age. Based on the user's age, the user's age distribution can be determined, and the video editing effects commonly used by users within that age distribution (e.g., commonly used filters, background music, etc.) can be identified, and then used as functional modules of interest to the user. As another example, based on the user's operations on applications with video editing functions, the video editing effects frequently used by the user when editing videos or images can be determined, and then used as functional modules of interest to the user. This method of determining the user's interest functional modules, compared to subsequently determining the functional modules of template modules based on the user's historical video editing commands, can more accurately determine the functional modules included in the user's template modules even with limited historical data.
[0066] The third approach involves statistically analyzing frequently used video editing commands by users within a preset time period to determine the template module. Specifically, before generating the first video from pre-acquired materials using the template module, specific video editing commands used by the user more than a preset threshold within the preset time period are identified; the template module is then generated based on these specified video editing commands. The preset time period can be a time period corresponding to a preset duration preceding the current moment, which can be set based on actual usage needs and is not limited here. Each user operation records the editing commands involved, and after multiple operations, a template is automatically generated based on the frequency of the same commands. For example, if a user has used the same filter and transition effects in their last ten edits, this filter and transition type, along with the relevant parameters, are stored as a template, which can be used directly in the next edit.
[0067] S602: When the first video editing interface corresponding to the first video is displayed on the screen of the electronic device, acquire voice data.
[0068] S603: Determine the target semantic content corresponding to the voice data through the semantic recognition module.
[0069] S604: Search for the target video editing instruction corresponding to the target semantic content in the semantic library, wherein the semantic library includes multiple semantic contents and video editing instructions corresponding to each semantic content.
[0070] The speech recognition module comprises two parts: a first part and a second part. The first part is the semantic recognition function, used for speech data recognition and translation. Specifically, it accepts the user's speech input, converts it into corresponding natural language, and then obtains the corresponding video editing instructions. This speech recognition module has a semantic library containing multiple semantic contents and corresponding video editing instructions for each semantic content, as shown in Table 1.
[0071] Table 1
[0072]
[0073]
[0074] Based on this semantic database, the implementation method for finding the target video editing instruction corresponding to the target semantic content within the semantic database can be as follows: A semantic recognition module determines the target semantic content corresponding to the voice data; the semantic database searches for voice content matching the target semantic content as candidate semantic content; and then, the video editing instruction corresponding to the candidate semantic content is determined as the target editing instruction. For example, if the user inputs voice data to darken by 50%, the semantic recognition module can determine from the semantic database that the target video editing instruction corresponding to this voice data is to reduce the value of the brightness filter by half.
[0075] Furthermore, for ambiguous language that doesn't specify the exact editing operation, or may correspond to multiple editing operations, such as "add a nice background," where it's unclear which filter the user specifically wants to add, a series of filter effect demonstrations will be provided for the user to choose from. The user can then click on the desired filter effect or specify which filter effect they want. In other words, the implementation method for searching for the target video editing instruction corresponding to the target semantic content in the semantic library can be as follows: searching for multiple candidate video editing instructions corresponding to the target semantic content in the semantic library; displaying at least one of the candidate video editing instructions on the screen of the electronic device; and determining the candidate video editing instruction selected by the user as the target video editing instruction.
[0076] As one implementation method, searching for multiple candidate video editing instructions corresponding to the target semantic content in the semantic library can mean searching for semantic content in the semantic library with a matching degree greater than or equal to a first threshold, and using this as the semantic content that matches the target semantic content in the semantic library. After recognizing the speech data, each semantic content corresponding to the speech data can be obtained as the target semantic content. For example, if the speech data is "Add a transition between videos and change the music to music I often listen to recently," then the target semantic content obtained from recognizing the speech data could be transition selection, adding transition effects, setting transition duration, selecting audio files, replacing audio tracks, etc. Assuming that each semantic content can find a matching semantic content in the semantic library, each target semantic content can correspond to a video editing instruction. For example, the video editing instructions corresponding to transition selection, adding transition effects, setting transition duration, selecting audio files, and replacing audio tracks are respectively transition instructions, effect instructions, setting transition duration instructions, adding audio files instructions, and replacing audio tracks instructions. The electronic device can then use these various video editing instructions as candidate video editing instructions, display them on the screen, and determine the candidate video editing instruction selected by the user as the target video editing instruction. For example... Figure 7 The instruction selection interface shown can be displayed as a pop-up window on the first video editing interface, showing multiple candidate video editing instructions. Users can select the candidate video editing instructions they are interested in as the target video editing instructions.
[0077] As another implementation method, for some ambiguous language where the specific editing operation is unclear, or where it may correspond to multiple editing operations, such as "add a nice background," and it's unclear which filter the user specifically wants to add, a series of filter effect demonstration images will be provided for the user to choose from. In other words, the aforementioned ambiguous language refers to semantic content corresponding to the language data where no semantic content with a matching degree greater than or equal to a first threshold can be found in the semantic library. Therefore, a fuzzy search method can be used to find semantic content similar to the target semantic content in the semantic library. Thus, the implementation method for searching multiple video editing instructions corresponding to the target semantic content in the semantic library is as follows: if no semantic content with a matching degree greater than or equal to the first threshold can be found in the semantic library, then candidate semantic content with a matching degree greater than a second threshold and less than the first threshold is found in the semantic library; the video editing instructions corresponding to the candidate semantic content in the semantic library are then used as candidate video editing instructions. In other words, all candidate semantic content in the semantic library that matches the target semantic content with a degree greater than the second threshold and less than the first threshold can be considered as semantic content similar to the target semantic content. For example, if the target semantic content is "adjust brightness", there is no semantic content "adjust brightness" in the semantic library, but there are two semantic contents: "brighten" and "darken". Therefore, both "brighten" and "darken" are considered as approximate semantic content of "adjust brightness", and the video editing instructions corresponding to these semantic contents are considered as candidate video editing instructions. Meanwhile, the video editing instructions corresponding to semantic content that can be completely matched can be directly used as the target video editing instructions.
[0078] like Figure 8 As shown, the selected video editing command can be displayed on the screen of the electronic device, and... Figure 7 compared to, Figure 8 The difference is that it can choose not to display the video editing instructions corresponding to the semantic content that can be completely matched, and can also display the target semantic content corresponding to the candidate semantic content on the screen, and display the candidate video editing instructions corresponding to the target semantic content in the display position of the target semantic content. Then, the user can select one candidate video editing instruction from the candidate video editing instructions corresponding to the target semantic content as the target video editing instruction for each target semantic content.
[0079] In addition, the second part of the aforementioned semantic recognition module is the updating of the semantic library. That is, based on the video editing instructions selected by the user, the semantic library can be updated. Specifically, after determining the candidate video editing instructions selected by the user as the target video editing instructions, the target semantic content and the target video editing instructions are updated to the semantic library. In other words, when the target semantic content corresponds to multiple candidate video editing instructions, the video editing instructions selected by the user from the multiple candidate video editing instructions are determined as the target video editing instructions, and the correspondence between the target semantic content and the target video editing instructions is updated to the semantic library.
[0080] It should be noted that the aforementioned implementation of determining the selected video editing instruction can be as follows: when multiple selected video editing instructions are displayed on the electronic screen, determine the user's trigger position on the electronic screen, and determine the instruction corresponding to that trigger position as the selected video editing instruction. Alternatively, when multiple selected video editing instructions are displayed on the electronic screen, determine the user's input voice as the first voice, recognize the first voice, and determine the selected video editing instruction corresponding to the first voice.
[0081] S605: Modify the first video based on the target video editing instructions to obtain the second video.
[0082] The editing engine mainly performs specific editing tasks based on the semantically translated editing instructions. There are no restrictions on the choice of editing engine here. Any engine that can edit video can be used as an editing engine module and embedded in this embodiment.
[0083] like Figure 9 As shown, users first import the necessary original materials such as videos, photos, and audio. Then, using some preset templates, a raw editable video is generated. Next, users can provide feedback on the video modifications via voice. The creation engine's voice recognition module converts the user's speech into specific editing commands, such as "increase brightness" to increase the video's exposure, "delete content between 3-5 seconds" to crop the video, and "delete the top half of the video" to resize the video. Simultaneously, the semantic library is updated based on the user's language characteristics, as everyone's description of effects may differ. Once the modifications are complete and the user is satisfied, the target video is generated. Furthermore, the template library is updated based on this user's editing actions for future use.
[0084] It should be noted that the aforementioned voice input is not limited to Chinese; it can receive multiple languages, and the corresponding semantic library can also support multiple languages. In addition, Chinese can also receive multiple regional dialects, and the system can recognize and perform escaping operations for regional dialects.
[0085] Please see Figure 10 The diagram illustrates a structural block diagram of a video editing device 1000 provided in an embodiment of this application. The device may include: an acquisition unit 1001, a parsing unit 1002, and a modification unit 1003.
[0086] Acquisition unit 1001 is used to acquire voice data;
[0087] The parsing unit 1002 is used to parse the voice data to obtain target video editing instructions;
[0088] Modification unit 1003 is used to modify the first video based on the target video editing instructions to obtain the second video.
[0089] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0090] Please see Figure 11 The diagram shows a structural block diagram of a video editing device 1100 provided in an embodiment of this application. The device may include: a video generation unit 1101, an acquisition unit 1102, a parsing unit 1103, and a modification unit 1104.
[0091] The video generation unit 1101 is used to generate a first video based on pre-acquired material data.
[0092] Furthermore, the video generation unit 1101 is used to display an editing selection interface on the screen of the electronic device, the editing selection interface including a first selection control; if a first operation is detected on the first selection control, the first video editing interface is displayed in response to the first operation.
[0093] Furthermore, the video generation unit 1101 is also configured to, if a second operation is detected on the second selection control, respond to the second operation and display a second video editing interface, wherein at least one editing control is displayed in the second video editing interface; determine the target editing control that has been touched from the at least one editing control; modify the first video based on the video editing instruction corresponding to the target editing control to obtain a third video.
[0094] Furthermore, the video generation unit 1101 is also used to generate a first video based on pre-acquired material data through a template module. The template module includes multiple functional modules, each corresponding to a video editing operation.
[0095] Furthermore, the video generation unit 1101 is also used to obtain a specified video editing instruction whose number of uses by the user within a preset time period exceeds a preset threshold; and to generate the template module based on the specified video editing instruction.
[0096] The acquisition unit 1102 is used to acquire voice data when the first video editing interface corresponding to the first video is displayed on the screen of the electronic device.
[0097] The parsing unit 1103 is used to parse the voice data to obtain target video editing instructions.
[0098] Furthermore, the parsing unit 1103 is also used to determine the target semantic content corresponding to the voice data through the semantic recognition module; and to search for the target video editing instruction corresponding to the target semantic content in the semantic library, wherein the semantic library includes multiple semantic contents and video editing instructions corresponding to each semantic content.
[0099] Furthermore, the parsing unit 1103 is also used to search for multiple candidate video editing instructions corresponding to the target semantic content in the semantic library; display each candidate video editing instruction on the screen of the electronic device; and determine the candidate video editing instruction selected by the user as the target video editing instruction.
[0100] Furthermore, the parsing unit 1103 is also configured to, if no semantic content with a matching degree greater than or equal to the first threshold can be found in the semantic library, then find all candidate semantic content in the semantic library with a matching degree greater than the second threshold and less than the first threshold; and take the video editing instruction corresponding to each candidate semantic content in the semantic library as the candidate video editing instruction.
[0101] Furthermore, the parsing unit 1103 is also used to update the target semantic content and the target video editing instructions to the semantic library.
[0102] Modification unit 1104 is used to modify the first video based on the target video editing instructions to obtain the second video.
[0103] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0104] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0105] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0106] Please refer to Figure 12 This document illustrates a structural block diagram of an electronic device according to an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications can be stored in the memory 120 and configured to be executed by one or more processors 110, and the one or more applications are configured to perform the methods described in the foregoing method embodiments.
[0107] Processor 110 may include one or more processing cores. Processor 110 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data of the electronic device 100 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 110 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.
[0108] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the electronic device 100 during use (such as phonebook data, audio and video data, chat log data, etc.).
[0109] Please refer to Figure 13 This diagram illustrates a structural block diagram of a computer-readable medium provided in an embodiment of this application. The computer-readable medium 1300 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0110] Computer-readable medium 1300 may be an electronic storage device such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, computer-readable medium 1000 includes a non-volatile computer-readable storage medium. Computer-readable medium 1300 has storage space for program code 1310 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 1310 may be compressed, for example, in a suitable form.
[0111] The video editing method, apparatus, electronic device, and computer-readable medium provided in this application generate a first video based on pre-acquired source material; while displaying a first video editing interface corresponding to the first video on the screen of the electronic device, voice data is acquired; the voice data is parsed to obtain a target video editing instruction; and the first video is modified based on the target video editing instruction to obtain a second video. Therefore, for videos generated from source material, users can edit the video by inputting voice data to obtain an edited video, which is more user-friendly compared to traditional UI interaction.
[0112] This application completes video editing operations based on the user's voice input, simplifying the video editing process and lowering the learning threshold for a wide range of users. Whether it is the elderly, children, or young people who do not want to learn complicated editing processes, they can easily edit and create the videos they want.
[0113] In addition, to help users edit videos more intelligently, a semantic library and a template library have been added. On the one hand, a mapping is used to complete the editing operation based on the user's language habits. On the other hand, the user's preferred effects are recorded to generate templates. At the beginning of the editing, an initial version of the effect is generated based on these. Then, the user may only need to say a few words to modify the video content. This not only reduces the difficulty of creation, but also reduces the time cost of editing and creation.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A video editing method characterized by, The method is applied to an electronic device, and the method comprises: Obtaining a configuration voice input by a user, parsing the configuration voice to obtain a plurality of different video editing requirements, and determining a function module corresponding to each requirement as a function module of a template module; Editing a first video by using a function module configured by the user through the template module from pre-acquired material data; Obtaining voice data in a case of playing the first video; Parsing the voice data to obtain a target video editing instruction; Modifying the first video based on the target video editing instruction to obtain a second video; The modification of the first video based on the target video editing instruction to obtain the second video comprises: Pausing a playing operation of the first video, editing the first video based on the target video editing instruction corresponding to the voice data, and then continuing to play the first video; Obtaining a video as the second video when the first video is played to the end.
2. The method of claim 1, wherein, The parsing of the voice data to obtain the target video editing instruction comprises: Determining target semantic content corresponding to the voice data through a semantic recognition module; Determining a target video editing instruction corresponding to the target semantic content based on a semantic library, wherein the semantic library comprises a plurality of semantic contents and a video editing instruction corresponding to each semantic content.
3. The method of claim 2, wherein, The determination of the target video editing instruction corresponding to the target semantic content based on the semantic library comprises: Searching for a plurality of candidate video editing instructions corresponding to the target semantic content in the semantic library; Displaying at least one candidate video editing instruction on a screen of the electronic device; Determining a candidate video editing instruction selected by the user as the target video editing instruction.
4. The method of claim 3, wherein, The searching for the plurality of video editing instructions corresponding to the target semantic content in the semantic library comprises: If a semantic content with a matching degree greater than or equal to a first threshold value cannot be found in the semantic library, searching for a candidate semantic content with a matching degree greater than a second threshold value and less than the first threshold value in the semantic library; Taking a video editing instruction corresponding to the candidate semantic content in the semantic library as a candidate video editing instruction.
5. The method of claim 3, wherein, After the determination of the candidate video editing instruction selected by the user as the target video editing instruction, the method further comprises: Updating the target semantic content and the target video editing instruction to the semantic library.
6. The method of claim 1, wherein, The obtaining of the voice data comprises: Obtaining the voice data in a case of displaying a first video editing interface corresponding to the first video on a screen of the electronic device.
7. The method of claim 6, wherein, Before the obtaining of the voice data in the case of displaying the first video editing interface corresponding to the first video on the screen of the electronic device, the method further comprises: Displaying an editing selection interface on the screen of the electronic device, wherein the editing selection interface comprises a first selection control; If a first operation directed to the first selection control is detected, displaying the first video editing interface in response to the first operation.
8. The method of claim 7, wherein, The editing selection interface further comprises a second selection control, and the method further comprises: If a second operation directed to the second selection control is detected, displaying a second video editing interface in response to the second operation, wherein at least one editing control is displayed in the second video editing interface; From the at least one editing control, a target editing control being touched is determined; Based on a video editing instruction corresponding to the target editing control, the first video is modified to obtain a third video.
9. A video editing apparatus characterized by comprising: The application is applied to an electronic device, and the device comprises: An acquisition unit is configured to acquire a configuration voice input by a user, parse the configuration voice to obtain a plurality of different video editing requirements, determine a function module corresponding to each requirement as a function module of a template module, edit a first video by using a function module configured by the user through the template module from pre-acquired material data, and acquire voice data in a case of playing the first video; An analysis unit is configured to analyze the voice data to obtain a target video editing instruction; A modification unit is configured to modify the first video based on the target video editing instruction to obtain a second video. The modification of the first video based on the target video editing instruction to obtain the second video comprises: Pausing a playing operation of the first video, editing the first video based on the target video editing instruction corresponding to the voice data, and then continuing to play the first video; In a case of playing the first video to the end, the obtained video is recorded as the second video.
10. An electronic device, comprising: The device comprises: One or more processors; A memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to execute the method according to any one of claims 1-8.
11. A computer readable medium characterized by The computer readable medium stores processor executable program codes, and the program codes are executed by the processor to make the processor execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Image adjustment method, electronic equipment and storage medium
CN113467735A
Video preview method and device, readable medium and electronic equipment
CN115022696A