Video recording processing method and apparatus, and electronic device and computer storage medium
By acquiring recording and video tags through the device-side terminal and combining them with voice control, the system-side terminal automatically selects and generates the initial video, solving the problem of low recording processing efficiency and realizing intelligent recording generation and efficient video production.
Patent Information
- Application Number
- PCT/CN2024/123769
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2026-04-16
AI Technical Summary
Recordings exported from outdoor video equipment require users to manually browse and select highlights, resulting in low video processing efficiency.
The system acquires recording data and video tags through the device-side terminal, automatically selects and generates the initial video, and combines voice control to realize recording mode selection, recording start, marking of highlight moments and ending of recording, simplifying the recording processing workflow.
It improves video recording efficiency, reduces manual editing time for users, and enables intelligent video generation and efficient video production.
Smart Images

Figure CN2024123769_16042026_PF_FP_ABST
Abstract
Description
Video recording methods, devices, electronic equipment, and computer storage media Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a video recording processing method, apparatus, electronic device, and computer storage medium. Background Technology
[0002] With the advancement of technology, video processing methods are becoming increasingly diverse. In some solutions, recording devices such as outdoor dashcams support video export functions, allowing users to edit exported videos and create vlogs or other content to record memorable moments in outdoor environments.
[0003] However, in related technologies, the videos exported by outdoor dashcams are generally in days. Users need to spend a lot of time browsing these exported videos and manually selecting the best clips from each video to make videos. This is time-consuming and laborious, resulting in low video processing efficiency.
[0004] Summary of the Invention
[0005] This application provides a video recording processing method, apparatus, electronic device, and computer storage medium, which can improve the efficiency of video recording processing.
[0006] In a first aspect, embodiments of this application provide a video recording processing method applied to a system-side terminal, the method comprising:
[0007] Obtain the video recording sent by the device-side terminal and the video tags carried in the recording;
[0008] Display the video recording selection page; the video recording selection page includes the first historical video recording taken by the device-side terminal within a preset time period, the first historical video recording includes the video recording;
[0009] In response to the user's first trigger operation based on the recording selection page, a first initial video is generated based on the first historical recording and the video tags carried by each recording in the first historical recording.
[0010] In one possible implementation, the video tag includes a highlight video tag for marking highlight moments in the recording;
[0011] In response to a user's first trigger operation based on input on the recording selection page, a first initial video is generated based on the first historical recording and the video tags carried by each recording in the first historical recording, including:
[0012] In response to the user's first trigger operation based on the recording selection page, the first historical recording is input into the video content recognition model to obtain the recognition results of each recording in the first historical recording;
[0013] Based on the identification results of each video in the first historical video and the video tags carried by each video in the first historical video, the first key video segment in the first historical video is extracted; the first key video segment includes the highlight video segment corresponding to the highlight moment;
[0014] Remove the jitter from the first key video segment to obtain the first initial video segment;
[0015] A first initial video is generated based on the first initial video segment.
[0016] In one possible implementation, after displaying the recording selection page, the method further includes:
[0017] In response to a second trigger operation input by the user on the recording selection page, a second initial video is generated based on the video tags carried by each recording in the second historical recording selected by the user on the recording selection page.
[0018] In one possible implementation, in response to a second trigger operation input by the user on the recording selection page, a second initial video is generated based on the video tags carried by each recording in the second historical recording selected by the user on the recording selection page, including:
[0019] In response to the second trigger operation input by the user based on the recording selection page, the second historical recording is input into the video content recognition model to obtain the recognition results of each recording in the second historical recording;
[0020] Based on the identification results of each video in the second historical video and the video tags carried by each video in the second historical video, the second key video segment is extracted from the second historical video;
[0021] Remove the jitter from the second key video segment to obtain the second initial video segment;
[0022] A second initial video is generated based on the second initial video segment.
[0023] In one possible implementation, after generating a first initial video based on a first historical video and the video tags carried by each video in the first historical video, in response to a first trigger operation input by the user on the video selection page, the method further includes:
[0024] Display the video editing page; the video editing page includes a video preview area, which is used to display a preview video associated with the first initial video.
[0025] In one possible implementation, the video editing page also includes a segment editing area for displaying a first initial video segment from the first initial video;
[0026] After displaying the video editing page, the methods also include:
[0027] In response to a user's clipping operation on any first initial video segment in the clip editing area, adjust the video duration of any first initial video segment;
[0028] and / or
[0029] In response to a user's speed adjustment operation for any first initial video segment in the segment editing area, adjust the playback speed of any first initial video segment;
[0030] and / or
[0031] In response to a user's segment replacement operation on any first initial video segment in the segment editing area, replace any first initial video segment with the target video segment;
[0032] and / or
[0033] In response to a user's operation to adjust the order of any first initial video segment in the segment editing area, the appearance order of any first initial video segment in the first initial video is adjusted.
[0034] In one possible implementation, the video editing page also includes a template editing area, which displays multiple video templates, each with a different video output duration and video presentation effect.
[0035] After displaying the video editing page, the methods also include:
[0036] In response to a third trigger operation by the user on any video template in the template editing area, the first initial video is adjusted based on that video template.
[0037] In one possible implementation, the method also includes:
[0038] In response to the fourth trigger action input by the user based on the video editing page, the preview video displayed in the video preview area is used as the target video.
[0039] Secondly, embodiments of this application provide a video recording processing method applied to a device-side terminal, the method comprising:
[0040] In response to the user's first control command, determine the target recording mode;
[0041] In response to the user's second control command, recording begins in the target recording mode;
[0042] In response to a third control command from the user, generate highlight video tags to mark highlight moments in the recording;
[0043] In response to the user's fourth control command, the recording is stopped, and the recording and the video tags carried by the recording are sent to the system-side terminal. The video tags include highlight video tags used to mark the highlight moments in the recording, so that the system-side terminal can obtain the recording and the video tags carried by the recording and display the recording selection page. The recording selection page includes the first historical recordings taken by the device-side terminal within a preset time period. The first historical recordings include recordings. In response to the user's first trigger operation based on the recording selection page, a first initial video is generated based on the first historical recordings and the video tags carried by each recording in the first historical recordings.
[0044] The control methods for the first, second, third, and fourth control commands include voice control.
[0045] Thirdly, embodiments of this application provide a video recording processing apparatus applied to a system-side terminal, the apparatus comprising:
[0046] The acquisition module is used to acquire the video recordings sent by the device-side terminal and the video tags carried in the recordings;
[0047] The display module is used to display the video recording selection page; the video recording selection page includes the first historical video recording taken by the device-side terminal within a preset time period, and the first historical video recording includes the video recording.
[0048] The response module is used to respond to the first trigger operation input by the user based on the recording selection page, and to generate a first initial video based on the first historical recording and the video tags carried by each recording in the first historical recording.
[0049] Fourthly, embodiments of this application provide a video recording processing apparatus applied to a device-side terminal, the apparatus comprising:
[0050] The first response module is used to determine the target recording mode in response to the user's first control command;
[0051] The second response module responds to the user's second control command and starts recording in the target recording mode;
[0052] The third response module, in response to the user's third control command, generates highlight video tags for marking highlight moments in the recording;
[0053] The fourth response module, in response to the user's fourth control command, ends the recording and sends the recording and the video tags carried by the recording to the system-side terminal. The video tags include highlight video tags used to mark the highlight moments in the recording, so that the system-side terminal can obtain the recording and the video tags carried by the recording and display the recording selection page. The recording selection page includes the first historical recording captured by the device-side terminal within a preset time period. The first historical recording includes recordings. In response to the user's first trigger operation based on the recording selection page, a first initial video is generated based on the first historical recording and the video tags carried by each recording in the first historical recording.
[0054] The control methods for the first, second, third, and fourth control commands include voice control.
[0055] Fifthly, embodiments of this application provide an electronic device, including: a processor and a memory; wherein the memory stores a computer program, the computer program being adapted to be loaded by the processor and execute the method steps provided in the first or second aspect of embodiments of this application.
[0056] In a sixth aspect, embodiments of this application provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps provided in the first or second aspect of embodiments of this application.
[0057] The aforementioned video recording processing method, apparatus, electronic device, and computer storage medium, at the device-side terminal, determine the target recording mode by responding to a first control command input by the user in the form of voice or other control methods. This helps the device-side terminal configure the recording parameters corresponding to the target recording mode, improving the recording effect. By responding to a second control command input by the user in the form of voice or other control methods, recording is started in the target recording mode, allowing the user to start recording at any time according to actual needs, improving the flexibility of recording. By responding to a third control command input by the user in the form of voice or other control methods, highlight video tags are generated to mark highlight moments in the recording, enabling real-time marking of highlight moments in the recording without affecting shooting, simplifying the subsequent video recording processing flow. By responding to a fourth control command input by the user in the form of voice or other control methods, recording is ended, and the recording and the video tags carried by the recording are sent to the system-side terminal, allowing the user to stop recording at any time according to actual needs, improving the effectiveness of the recorded content. The entire video recording processing process at the device-side terminal not only improves the recording quality but also simplifies the subsequent video recording processing flow at the system-side terminal, improving the efficiency of video recording processing.
[0058] The aforementioned video recording processing method, apparatus, electronic device, and computer storage medium, at the system-side terminal, by acquiring the video recordings and video tags carried by the device-side terminal, help the system-side terminal automatically select exciting segments from the recordings based on the recordings and their tags. By displaying a video selection page, which includes first historical recordings taken by the device-side terminal within a preset time period, and these first historical recordings include the video recordings, users can intuitively see all the recordings taken by the device-side terminal within the preset time period, improving the efficiency of video viewing. In response to a first trigger operation input by the user based on the video selection page, a first initial video is generated based on the first historical recordings and the video tags carried by each recording in the first historical recordings, greatly reducing the time and effort required for users to manually edit the recordings, effectively simplifying the video processing workflow, and realizing intelligent video generation. The entire video recording processing process at the system-side terminal eliminates the need for users to spend a lot of time and effort creating videos from the recordings exported from the device-side terminal; users only need to perform a simple trigger operation, and the system-side terminal can automatically process the recordings and intelligently generate videos, effectively improving video processing efficiency. Attached Figure Description
[0059] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 is a schematic diagram of a video recording processing system provided in an exemplary embodiment of this application;
[0061] Figure 2 is a schematic flowchart of a video recording method provided in an exemplary embodiment of this application;
[0062] Figure 3 is a schematic diagram of the main interface of a recorder APP provided in an exemplary embodiment of this application;
[0063] Figure 4 is a schematic diagram of a video recording selection page provided in an exemplary embodiment of this application;
[0064] Figure 5 is a flowchart illustrating another video recording method provided in an exemplary embodiment of this application;
[0065] Figure 6 is a schematic diagram of a video editing page provided in an exemplary embodiment of this application;
[0066] Figure 7 is a schematic diagram of a segment editing pop-up window provided in an exemplary embodiment of this application;
[0067] Figure 8 is a flowchart illustrating another video recording method provided in an exemplary embodiment of this application;
[0068] Figure 9 is a schematic diagram of the structure of a video recording processing apparatus provided in an exemplary embodiment of this application;
[0069] Figure 10 is a schematic diagram of another video recording processing apparatus provided in an exemplary embodiment of this application;
[0070] Figure 11 is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this application. Detailed Implementation
[0071] First, let's introduce an application scenario applicable to the embodiments of this application: An outdoor dashcam combines the functions of a dashcam and an action camera, allowing users to shoot from inside a car or outdoors. During recording, users can input control commands via voice or other means to select recording modes (such as normal recording mode, slow-motion recording mode, self-timer recording mode, etc.), start or stop recording, and mark highlights in the recording.
[0072] This outdoor dashcam comes with a dashcam app installed on your phone. Users can not only view the videos uploaded by the dashcam through the app, but also process the videos to quickly generate valuable videos (such as highlight video clips corresponding to user-marked highlight moments), making it easier and more efficient for users to record and share precious moments.
[0073] The video recording processing method provided in this application embodiment can be applied to the video recording processing system shown in Figure 1. This video recording processing system includes a device-side terminal 11 and a system-side terminal 12. The device-side terminal 11 can be connected to the system-side terminal 12 wirelessly (e.g., via Bluetooth, Wi-Fi) or via a wired connection. The device-side terminal 11 is equipped with at least two cameras and at least two microphones in different directions. The device-side terminal 11 can be, but is not limited to, various outdoor recorders, action cameras, monitoring terminals, etc. The system-side terminal 12 is equipped with an app compatible with the device-side terminal 12. The system-side terminal 12 can be, but is not limited to, various smartphones, tablets, laptops, personal computers, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc.
[0074] Specifically, in response to a user's first control command, the device-side terminal 11 determines the target recording mode; in response to a user's second control command, it starts recording in the target recording mode; in response to a user's third control command, it generates highlight video tags to mark highlights in the recording; in response to a user's fourth control command, it ends recording and sends the recording and the video tags carried by the recording (including the highlight video tags used to mark highlights in the recording) to the system-side terminal 12 via wireless or wired means. The control methods of the aforementioned first, second, third, and fourth control commands include voice control. The system-side terminal 12 obtains the recording and the video tags carried by the recording sent by the device-side terminal 11; displays a recording selection page, which includes first historical recordings taken by the device-side terminal within a preset time period, the first historical recordings including the recordings; in response to a user's first trigger operation input based on the recording selection page, it generates a first initial video based on the first historical recordings and the video tags carried by each recording in the first historical recordings.
[0075] In one embodiment, as shown in FIG2, a video recording processing method is provided. Taking the application of this method to the device-side terminal 11 and system-side terminal 12 shown in FIG1 as an example, the method includes the following steps:
[0076] S201: The device-side terminal 11 responds to the user's first control command and determines the target recording mode.
[0077] The first control instruction is used to specify the target recording mode selected by the user; the control form of the first control instruction includes, but is not limited to, voice control, button control, gesture control, touch screen control, etc.
[0078] Optionally, the device-side terminal 11 has built-in multiple different recording modes (including but not limited to normal recording mode, slow-motion recording mode, selfie recording mode, etc.) to meet user needs in different scenarios, and the target recording mode is one of the above-mentioned different recording modes. After the user issues a first control command to the device-side terminal 11 in the form of voice control, the device-side terminal 11 collects the user's first control command through its built-in microphone, and then uses Automatic Speech Recognition (ASR) technology to convert the first control command into text to parse the user's intent, thereby determining the target recording mode that the user wishes to use. Then, it obtains the recording parameters corresponding to the target recording mode, including but not limited to the camera to be turned on in the target recording mode (such as a front camera or a rear camera), the resolution in the target recording mode, the frame rate in the target recording mode, and the focus mode in the target recording mode.
[0079] For example, after the user inputs "selfie video mode" by voice, the device terminal 11 responds to the first control command input by the user's voice, determines that the target video mode that the user wants to use is the selfie video mode, and then obtains the recording parameters corresponding to the selfie video mode, including the camera (front camera) to be turned on in the selfie video mode, the resolution in the selfie video mode, the frame rate in the selfie video mode, the focus mode in the selfie video mode, etc.
[0080] In this embodiment, by responding to the first control command input by the user in the form of voice or other control methods, the target recording mode is determined. This helps the device-side terminal configure the recording parameters corresponding to the target recording mode, ensuring that the recording mode meets the user's needs and improving the recording effect and user experience.
[0081] S202: The device-side terminal 11 responds to the user's second control command and starts recording in the target recording mode.
[0082] The second control command is used to control the device-side terminal 11 to start recording in the target recording mode specified by the user; the control form of the second control command includes, but is not limited to, voice control, button control, gesture control, touch screen control, etc.
[0083] Optionally, the user can issue a second control command to the device terminal 11 via voice control and / or button control. After the user issues the second control command to the device terminal 11 via voice control, the device terminal 11 collects the user's second control command through its built-in microphone, uses automatic speech recognition technology to convert the second control command into text, and interprets the user's intent to start recording in the target recording mode selected by the user, that is, to start recording with the recording parameters corresponding to the target recording mode.
[0084] For example, after the user inputs "selfie video recording mode" and "start recording" by voice, the device terminal 11 responds to the second control command input by the user's voice and starts recording with the recording parameters corresponding to the selfie video recording mode.
[0085] In this embodiment, recording is started in the target recording mode in response to a second control command input by the user in the form of voice or other control, which makes it convenient for the user to start recording at any time according to actual needs and improves the flexibility of recording.
[0086] S203: The device-side terminal 11 responds to the user's third control command and generates a highlight video tag for marking highlight moments in the recording.
[0087] The third control command is used to mark the highlight moments in the video recording; the control methods of the third control command include, but are not limited to, voice control, button control, gesture control, touch screen control, etc.
[0088] Optionally, during recording on the device-side terminal 11, the user can issue a third control command to the device-side terminal 11 via voice control or other means to mark highlights in the recording, i.e., generate highlight video tags to mark highlights in the recording. This allows the system-side terminal 12 to quickly locate exciting segments in the recording and generate videos based on the highlight video tags generated by the backup terminal 11. Understandably, the same recording segment can correspond to multiple highlight video tags for marking highlights in the recording.
[0089] Specifically, after the user issues a third control command to the device terminal 11 via voice control, the device terminal 11 collects the user's third control command through its built-in microphone, uses automatic speech recognition technology to convert the third control command into text, parses the user's intent, and thus generates a highlight video tag containing the highlight moment timestamp information. It is worth noting that the highlight video tag is not limited to including the highlight moment timestamp information; it may also include descriptive information about the highlight moment, etc.
[0090] For example, after the user voice inputs "mark highlight moment" or "lock video", the device terminal 11 responds to the third control command input by the user's voice input and generates a highlight video tag containing timestamp information of the highlight moment (e.g., 10:25:32). Furthermore, after the user voice inputs "mark highlight moment: goal moment" or "lock video, lock goal moment", the device terminal responds to the third control command input by the user's voice input and generates a highlight video tag containing timestamp information of the highlight moment and descriptive information of the highlight moment (e.g., goal moment).
[0091] In this embodiment, by responding to a third control command input by the user in the form of voice or other control methods, highlight video tags are generated to mark the highlight moments in the video recording. This allows for the instant marking of highlight moments in the video recording without affecting the shooting process, thus simplifying the subsequent video recording process.
[0092] S204: In response to the user's fourth control command, the device-side terminal 11 ends the recording and sends the recording and the video tag carried in the recording to the system-side terminal 12.
[0093] The fourth control command is used to control the device-side terminal 11 to end the current recording; the control form of the fourth control command includes, but is not limited to, voice control, button control, gesture control, touch screen control, etc.
[0094] Optionally, the user can issue a fourth control command to the device-side terminal 11 via voice control and / or button control. After the user issues the fourth control command to the device-side terminal 11 via voice control, the device-side terminal 11 acquires the user's fourth control command through its built-in microphone, converts the fourth control command into text using automatic speech recognition technology, and interprets the user's intent to end the currently ongoing recording. For example, after the user voice inputs "End recording," the device-side terminal 11 responds to the user's voice input second control command and ends the currently ongoing recording.
[0095] In this embodiment, the recording is terminated in response to the fourth control command input by the user in the form of voice or other control, and the recording and the video tag carried by the recording are sent to the system-side terminal. This allows the user to stop recording at any time according to actual needs, improving the effectiveness of the recorded content and thus improving the recording processing efficiency of the subsequent system-side terminal.
[0096] S205: System-side terminal 12 obtains the video recording and video tag carried by the video recording sent by device-side terminal 11.
[0097] Optionally, the system-side terminal 12 can connect to the device-side terminal 11 wirelessly (e.g., via Bluetooth, Wi-Fi, etc.) or via wired connection to obtain the video recordings and video tags (including highlight video tags) sent by the device-side terminal 11. Furthermore, after obtaining the video recordings and video tags sent by the device-side terminal 11, the system-side terminal 12 will also push a recording generation notification message to the APP associated with the device-side terminal 11.
[0098] The following explanation uses a smartphone as an example, with device-side terminal 11 being an outdoor dashcam. Understandably, system-side terminal 12 can connect to device-side terminal 11 (outdoor dashcam) wirelessly (e.g., via Bluetooth, Wi-Fi) and / or via wired connection. System-side terminal 12 has a dashcam app installed that is compatible with the outdoor dashcam.
[0099] For example, please refer to Figure 3. After the user opens the recorder APP on the system-side terminal 12, they can click the settings control 31 in the upper right corner of the recorder APP main interface 30 to complete the connection between the system-side terminal 12 and the outdoor recorder. After the system-side terminal 12 is successfully connected to the outdoor recorder, the recorder APP main interface 30 will display a video recording generation prompt message 32, indicating that a new video has been generated, and the user can create a video clip based on the generated video with one click.
[0100] In this embodiment, the system-side terminal obtains the video recording and video tags carried by the device-side terminal and displays video generation prompts. This not only helps the system-side terminal automatically select exciting segments from the video recording based on the recording and the tags carried by the recording, but also helps the user start the one-click video creation process based on the video generation prompts, thereby improving the efficiency of video processing.
[0101] S206: The system-side terminal 12 displays the recording selection page.
[0102] The video recording selection page includes the first historical video recording taken by the device terminal 11 within a preset time period (such as the most recent day, the most recent week, the most recent month, etc.), and the first historical video recording includes the aforementioned video recording. Understandably, the video recording selection page may also include other historical video recordings taken by the device terminal 11 within other preset time periods (i.e., not the first preset time period) as well as video recordings selected for generating the video.
[0103] Optionally, the preset time period is the most recent day, and the system-side terminal 12 displays the first historical video recording taken by the device-side terminal 11 within the most recent day on the video recording selection page. In addition, the system-side terminal 12 can also display other historical videos taken by the device-side terminal 11 before the most recent day on the video recording selection page.
[0104] For example, when a user clicks the "View" button 33 on the main interface 30 of the recorder app shown in Figure 3, the system-side terminal 12 displays the recording selection page 40 as shown in Figure 4. The recording selection page 40 displays multiple sets of historical recordings 41 taken by the device-side terminal 11 in reverse chronological order by day, along with the corresponding recording dates 42 for each set of historical recordings. The first historical recording 43 consists of all recordings taken by the outdoor recorder on the most recent day (as shown in the figure, "today, February 23"). In addition, the bottom of the recording selection page 40 also displays the recordings 44 selected for video generation. It is worth noting that the recordings 44 selected for video generation can be automatically determined by the system-side terminal 12 or selected by the user. By default, the system-side terminal 12 uses the first historical recording (i.e., all recordings taken by the device-side terminal 11 within the most recent day) as the recording 44 used for video generation.
[0105] In this embodiment, the system-side terminal displays multiple sets of historical videos taken by the device-side terminal and their corresponding recording dates in reverse chronological order by day on the recording selection page. This not only makes it convenient for users to see the videos taken by the device-side terminal intuitively, improving the efficiency of viewing the videos, but also makes it easier for users to select and combine different historical videos to generate videos based on the recording selection page, speeding up the recording processing and saving users' creation time.
[0106] S207: In response to the first trigger operation input by the user based on the recording selection page, the system-side terminal 12 generates a first initial video based on the first historical recording and the video tags carried by each recording in the first historical recording.
[0107] The first triggering operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing.
[0108] For example, referring to Figure 4, the system-side terminal 12, in response to the user's click operation on the confirmation button 45 in the recording selection page 40, generates a first initial video based on the first historical recording (i.e., the recording 44 selected for video generation shown in Figure 4) and the video tags (including highlight video tags) carried by each recording in the first historical recording. It is worth noting that in this embodiment, the recording 44 selected for video generation is automatically selected by the system-side terminal 12, that is, no manual video selection is required during the generation of the first initial video; the user only needs to perform a simple click operation to generate the video with one click.
[0109] In this embodiment, by responding to the first trigger operation input by the user based on the recording selection page, a first initial video is generated based on the first historical recording automatically selected by the system-side terminal and the video tags carried by each recording in the first historical recording. This greatly reduces the time and effort required for the user to manually edit the recordings, effectively simplifies the recording processing flow, and realizes intelligent video generation.
[0110] The above-described video recording method, on the device-side terminal, determines the target recording mode by responding to a first control command input by the user in the form of voice or other control methods. This helps the device-side terminal configure the recording parameters corresponding to the target recording mode, improving the recording effect. By responding to a second control command input by the user in the form of voice or other control methods, recording is started in the target recording mode, allowing the user to start recording at any time according to actual needs, improving the flexibility of recording. By responding to a third control command input by the user in the form of voice or other control methods, highlight video tags are generated to mark the highlights in the recording. This allows for the instant marking of highlights in the recording without affecting shooting, simplifying the subsequent recording processing flow. By responding to a fourth control command input by the user in the form of voice or other control methods, recording is ended, and the recording and the video tags carried by the recording are sent to the system-side terminal, allowing the user to stop recording at any time according to actual needs, improving the effectiveness of the recorded content. On the system-side terminal, by acquiring the video recordings and their accompanying video tags sent by the device-side terminal, the system-side terminal can automatically select highlights from the recordings based on the recordings and tags. By displaying a recording selection page, which includes the first historical recordings taken by the device-side terminal within a preset time period, users can easily see all the recordings taken by the device-side terminal within that time period, improving the efficiency of viewing recordings. Responding to the user's first trigger operation based on the recording selection page, a first initial video is generated based on the first historical recordings and their accompanying video tags. This significantly reduces the time and effort required for manual recording editing, effectively simplifying the recording processing workflow and achieving intelligent video generation. Throughout the entire recording processing process, while improving recording quality, the system eliminates the need for users to spend significant time and effort creating videos from exported recordings. Users only need to perform a simple trigger operation, and the system can automatically process the recordings and intelligently generate videos, effectively improving recording processing efficiency.
[0111] In one embodiment, as shown in FIG5, another video recording method is provided. Taking the application of this method to the device-side terminal 11 and system-side terminal 12 shown in FIG1 as an example, the method includes the following steps:
[0112] S501: The device-side terminal 11 responds to the user's first control command, determines the target recording mode, and generates a video type tag corresponding to the target recording mode.
[0113] The control methods of the first control command include, but are not limited to, voice control, button control, gesture control, and touch screen control.
[0114] Optionally, the device-side terminal 11 has built-in multiple different recording modes (including but not limited to normal recording mode, slow-motion recording mode, selfie recording mode, etc.) to meet user needs in different scenarios. The target recording mode is one of the above-mentioned different recording modes. After the user issues a first control command to the device-side terminal 11 in the form of voice control, the device-side terminal 11 collects the user's first control command through its built-in microphone, and then uses automatic speech recognition technology to convert the first control command into text to parse the user's intent, thereby determining the target recording mode that the user wants to use, and generating a corresponding video type tag for the target recording mode, such as a normal recording tag, a slow-motion recording tag, a selfie recording tag, etc., so that the system-side terminal 12 can intelligently extract key video segments from the recording based on the video type tag. Then, the device-side terminal 11 obtains the recording parameters corresponding to the target recording mode, including but not limited to the camera to be turned on in the target recording mode (such as a front camera or a rear camera), the resolution in the target recording mode, the frame rate in the target recording mode, the focus mode in the target recording mode, etc.
[0115] For example, after the user inputs "selfie video mode" via voice, the device-side terminal 11 responds to the first control command input by the user's voice, determines that the target recording mode the user wishes to use is the selfie video mode, and generates a selfie video tag corresponding to the selfie video mode, so that the system-side terminal 12 can intelligently extract key video segments from the recording based on the selfie video tag. Then, the device-side terminal 11 obtains the recording parameters corresponding to the selfie video mode, including the camera to be turned on (front camera) in the selfie video mode, the resolution in the selfie video mode, the frame rate in the selfie video mode, the focus mode in the selfie video mode, etc.
[0116] In this embodiment, by responding to the first control command input by the user in the form of voice or other control methods, the target recording mode is determined. This not only helps the system-side terminal configure recording parameters according to the target recording mode, ensuring that the recording mode meets the user's needs, but also helps the device-side terminal generate video type tags based on the target recording mode selected by the user. This allows the system-side terminal to intelligently extract key video segments from the recording based on the video type tags, improving the recording effect while simplifying the subsequent recording processing flow, thereby effectively improving the recording processing efficiency.
[0117] S502: The device-side terminal 11 responds to the user's second control command and starts recording in the target recording mode.
[0118] Specifically, S502 is the same as S202, and will not be repeated here.
[0119] S503: The device-side terminal 11 responds to the user's third control command and generates a highlight video tag for marking highlight moments in the recording.
[0120] Specifically, S503 is the same as S203, and will not be repeated here.
[0121] S504: In response to the user's fourth control command, the device-side terminal 11 ends the recording and sends the recording and the video tag carried in the recording to the system-side terminal 12.
[0122] The video tags include video type tags corresponding to the target recording mode, highlight video tags used to mark highlights in the recording, and video scene tags used to mark video scenes in the recording. The video scenes in the aforementioned recordings include one or more of the following: human figures, scenery (such as natural landscapes, cityscapes, etc.), vehicles, and moving images (such as walking, cycling, skiing, etc.).
[0123] Understandably, the aforementioned video type tag and highlight video tag are generated based on the aforementioned first control command and third control command, respectively. The aforementioned video scene tag is intelligently generated based on the scene recognition model of the device-side terminal 11 and is used to identify the video scene corresponding to different recording segments. The aforementioned video scene tag can be generated either during the recording process of the device-side terminal 11 or after the recording process has ended; this embodiment of the application does not limit this.
[0124] It is worth noting that the aforementioned video recordings can carry one or more highlight video tags and video scene tags. For example, if the user issues multiple third control commands during the recording process on the device terminal 11, the device terminal 11 will respond to all the user's third control commands by generating multiple highlight video tags to mark the highlight moments in the recording. If the recording captured by the device terminal 11 contains multiple different video scenes, the device terminal 11 will generate multiple video scene tags to mark the video scenes in the recording based on these different video scenes. In addition, the aforementioned highlight video tags contain timestamp information corresponding to the highlight moments, used to mark the specific time points when the highlight moments appear in the recording; the aforementioned video scene tags contain timestamp information corresponding to the video scenes, used to mark the specific time periods when the video scenes appear in the recording.
[0125] In this embodiment, the device-side terminal sends multiple tags of different dimensions, such as video type tags, highlight video tags, and video scene tags, to the system-side terminal. This allows the system-side terminal to quickly locate valuable segments in the recording based on these tags and generate videos from them, effectively simplifying the recording process and improving recording efficiency.
[0126] S505: System-side terminal 12 obtains the video recording and video tag carried by the video recording sent by device-side terminal 11.
[0127] The video tags include video type tags corresponding to the target recording mode, highlight video tags used to mark highlights in the recording, and video scene tags used to mark video scenes in the recording. The video scenes in the aforementioned recordings include one or more of the following: human figures, scenery (such as natural landscapes, cityscapes, etc.), vehicles, and moving images (such as walking, cycling, skiing, etc.).
[0128] Specifically, S505 is the same as S205, and will not be repeated here.
[0129] S506: The system-side terminal 12 displays the recording selection page.
[0130] The video recording selection page includes a first historical video recording taken by the device-side terminal within a preset time period, and the first historical video recording includes video recordings.
[0131] Specifically, S506 is the same as S206, and will not be repeated here.
[0132] S507: In response to the first trigger operation input by the user based on the recording selection page, the system-side terminal 12 inputs the first historical recording into the video content recognition model to obtain the recognition results of each recording in the first historical recording.
[0133] The first triggering operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing. The recognition results of each video in the first historical video include, but are not limited to, the speech recognition results obtained after performing speech recognition on each video in the first historical video, and the face recognition results obtained after performing face recognition on each video in the first historical video.
[0134] Optionally, in response to a first trigger operation input by the user based on the recording selection page, the system-side terminal 12 inputs the first historical recordings into the video content recognition model to obtain the video content recognition results of each recording in the first historical recordings. These video content recognition results include speech recognition results obtained by performing speech recognition on each recording in the first historical recordings and face recognition results obtained by performing face recognition on each recording in the first historical recordings. This video content recognition model can be deployed either in the cloud or on the system-side terminal 12; this embodiment of the application does not impose any limitations on this.
[0135] Specifically, in response to the user's first trigger operation based on the recording selection page, the system-side terminal 12 inputs the first historical recording into the video content recognition model. The video content recognition model then extracts the audio and video frames of each recording in the first historical recording. Next, it performs speech recognition on the extracted audio of each recording, converting the audio into text to obtain the speech recognition result for each recording. Finally, it performs face recognition on the extracted video frames of each recording, marking video frames containing target faces as target video frames to obtain the face recognition result for each recording. The target face is a clear and unobstructed face in the recording.
[0136] S508: The system-side terminal 12 extracts the first key video segment from the first historical video based on the identification results of each video in the first historical video and the video tags carried by each video in the first historical video.
[0137] The recognition results for each video in the first historical video recording include the speech recognition results and the face recognition results corresponding to each video in the first historical video recording. The video tags carried by each video in the first historical video recording include a video type tag corresponding to the target recording mode, a highlight video tag used to mark highlight moments in the recording, and a video scene tag used to mark video scenes in the recording. The first key video segment includes the highlight video segment corresponding to the highlight moment.
[0138] Optionally, the first key video segment also includes the stage video segment corresponding to valuable audio in the first historical recording, the stage video segment corresponding to the video frame (target video frame) containing the target face in the first historical recording, and the stage video segment corresponding to the video scene tag marked in the recordings in the first historical recording with video type tags of slow-motion recording and selfie recording.
[0139] Specifically, the system-side terminal 12 performs semantic analysis on the speech recognition results corresponding to each video recording in the first historical video recording, and retains the video segments containing important information, dialogue, or other meaningful content in the audio as valuable audio-corresponding video segments. For example, if a user says something describing the scenery captured by the device-side terminal 11 as nice during the recording process, the system-side terminal 12 will retain the corresponding video segment containing that user's words. Based on the face recognition results corresponding to each video recording in the first historical video recording, the system-side terminal 12 retains a video segment before and after each target video frame in the first historical video recording, for example, retaining a video segment of a preset duration (e.g., 3 seconds) before and after the target video frame as the corresponding video segment for the target video frame.
[0140] Furthermore, for recordings carrying highlight video tags in the first historical recordings, the system-side terminal 12 uses the timestamp information corresponding to the highlight moments in the highlight video tags to extract the highlight video segments corresponding to the highlight moments in the first historical recordings. For example, it retains video segments of a preset duration (e.g., 5 seconds) before and after the highlight moment as the highlight video segments corresponding to the highlight moments. The system-side terminal 12 also filters out recordings with video type tags of slow motion and self-shot recording based on the video type tags carried by each recording in the first historical recordings, that is, it filters out slow motion recordings and self-shot recordings in the first historical recordings. Then, for the video scene tags carried by the filtered slow motion recordings and self-shot recordings, it uses the timestamp information corresponding to the video scene in the video scene tags to extract the stage video segments in the slow motion recordings and self-shot recordings.
[0141] Understandably, the system-side terminal 12 can extract the stage video segments corresponding to valuable audio, the stage video segments corresponding to the target video frames, the highlight video segments corresponding to highlight moments, and the stage video segments from slow-motion recordings and self-shot recordings from the first historical recording, as the first key video segments in the first historical recording.
[0142] S509: The system-side terminal 12 removes the jitter from the first key video segment to obtain the first initial video segment.
[0143] Optionally, the system-side terminal 12 extracts video frames frame by frame from the first key video segment, uses motion estimation techniques (such as optical flow, feature point matching, etc.) to detect motion changes between each frame, calculates the displacement of each frame relative to a reference frame (such as the first frame or an intermediate frame) based on the motion estimation results, then uses a motion compensation algorithm to correct the position of each video frame to eliminate jitter, uses a filter to smooth the corrected frames, and then reassembles the processed frames into a new video segment, namely the first initial video segment.
[0144] S510: The system-side terminal 12 generates a first initial video based on the first initial video segment.
[0145] Optionally, the system-side terminal 12 will stitch together the first initial video segment after removing the jittered image, according to the time sequence, to generate the first initial video.
[0146] In this embodiment, the system terminal can extract the first key video segment from the first historical video recording based on the first historical video recording and the video tags carried by each video recording in the first historical video recording, and remove the jitter in the first key video segment, thereby intelligently generating the first initial video. This improves the efficiency of video recording while ensuring the quality of video recording.
[0147] S511: The system-side terminal 12 displays the video editing page.
[0148] The video editing page includes a video preview area, a clip editing area, and a template editing area. The video preview area displays a preview video associated with the first initial video; the clip editing area displays the first initial video clip from the first initial video; and the template editing area displays multiple video templates, each with a different output duration and presentation effect.
[0149] For example, please refer to Figure 6. The video editing page 60 includes a video preview area 61, a segment editing area 62, and a template editing area 63. The system-side terminal 12 displays the first initial video generated above in the video preview area 61. Users can edit the first initial video based on the segment editing area 62 and the template editing area 63. Each time an editing operation is performed, the edited video will be displayed in real time in the video preview area 61 for users to browse and view the editing effect.
[0150] In this embodiment, the system-side terminal divides the video editing page into a video preview area, a segment editing area, and a template editing area. This not only facilitates users in editing the initial video based on the video editing page, but also ensures that users can preview the edited video effect in real time based on the video editing page. This makes the video editing process more personalized, efficient, and convenient, and improves the efficiency of video recording.
[0151] S512: The system-side terminal 12 responds to the user's segment trimming operation on any first initial video segment in the segment editing area and adjusts the video duration of any first initial video segment.
[0152] The clipping operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing.
[0153] For example, referring to Figure 6, the system-side terminal 12 adjusts the video duration of the first initial video segment 64 in the first initial video in response to a user's click and progress bar cropping operation on one of the first initial video segments 64 in the segment editing area 62. Understandably, the user can also perform segment cropping operations on other first initial video segments in the segment editing area 62, such as first initial video segment 65, first initial video segment 66, and first initial video segment 67, thereby adjusting the video duration accordingly.
[0154] Specifically, in response to a user's click on one of the first initial video segments 64 in the segment editing area 62, the system-side terminal 12 displays a segment editing pop-up window 70 as shown in Figure 7. The user can adjust and view the playback progress of the first initial video segment 64 based on the video progress bar 71 in the segment editing pop-up window 70, select a video segment for trimming, and click the confirmation control 72 in the upper right corner of the segment editing pop-up window 70 to save the edited first initial video segment 64, thereby adjusting the video duration of the first initial video segment 64.
[0155] It is worth noting that, in response to a user's clipping operation on any initial video segment in the clip editing area, after adjusting the video duration of any initial video segment, the complete video obtained after clipping will be displayed in the video preview area of the video editing page.
[0156] S513: The system-side terminal 12 responds to the user's speed adjustment operation for any first initial video segment in the segment editing area, and adjusts the playback speed of any first initial video segment.
[0157] The speed adjustment operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long pressing.
[0158] For example, referring to Figure 6, the system-side terminal 12 adjusts the playback speed of the first initial video segment 64 in response to a user's click and speed adjustment type selection operation on one of the first initial video segments 64 in the segment editing area 62. Understandably, the user can also perform speed adjustment operations on other first initial video segments in the segment editing area 62, such as first initial video segment 65, first initial video segment 66, first initial video segment 67, etc., thereby adjusting the playback speed accordingly.
[0159] Specifically, in response to a user's click on one of the first initial video segments 64 in the segment editing area 62, the system-side terminal 12 displays a segment editing pop-up window 70 as shown in Figure 7. The user can click on the linear speed control 73 or the curve speed control 74 in the segment editing pop-up window 70 to select the corresponding speed type. For example, after clicking on the curve speed control 74, the user can select the montage speed type, jump speed type, bullet time speed type, etc. under curve speed in the segment editing pop-up window 70, and click on the confirmation control 72 in the upper right corner of the segment editing pop-up window 70 to save the edited first initial video segment 64, thereby adjusting the playback speed of the first initial video segment.
[0160] In addition, the system-side terminal 12 can also respond to the user's background audio adjustment operation on any of the first initial video segments in the segment editing area, adjusting the background audio of any of the first initial video segments. Specifically, after the system-side terminal 12 responds to the user's click operation on one of the first initial video segments 64 in the segment editing area 62 and displays the segment editing pop-up window 70 as shown in Figure 7, the user can click the background audio adjustment control 75 in the segment editing pop-up window 70 to choose to turn off or on the original sound in the first initial video segment 64, and click the confirmation control 72 in the upper right corner of the segment editing pop-up window 70 to save the edited first initial video segment 64, thereby adjusting the background audio of the first initial video segment.
[0161] It is worth noting that, in response to the user's speed adjustment operation on any first initial video segment in the clip editing area, after adjusting the playback speed of any first initial video segment, or in response to the user's background sound adjustment operation on any first initial video segment in the clip editing area, the complete video obtained after adjusting the playback speed or the complete video obtained after adjusting the background sound will be displayed in the video preview area of the video editing page.
[0162] S514: In response to a user's segment replacement operation on any first initial video segment in the segment editing area, the system-side terminal 12 replaces any first initial video segment with the target video segment.
[0163] The segment replacement operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing.
[0164] For example, referring to Figure 6, in response to a user's long-press operation on one of the first initial video segments 64 in the segment editing area 62, the system-side terminal 12 displays a video segment replacement pop-up window. The user can select the desired target video segment from the multiple video segments displayed in the pop-up window and click on it, thereby causing the system-side terminal 12 to replace the first initial video segment 64 with the user-selected target video segment. Understandably, the user can also perform segment replacement operations on other first initial video segments in the segment editing area 62, such as first initial video segment 65, first initial video segment 66, and first initial video segment 67, thereby adjusting the playback speed accordingly.
[0165] It is worth noting that the multiple video clips displayed in the aforementioned video clip replacement pop-up can be extracted from the first historical recording and / or other historical recordings, and this application embodiment does not limit this. In response to a user's clip replacement operation on any first initial video clip in the clip editing area, after replacing any first initial video clip with the target video clip, the complete video obtained after the clip replacement will be displayed in the video preview area of the video editing page.
[0166] S515: The system-side terminal 12 responds to the user's operation to adjust the order of any first initial video segment in the segment editing area, and adjusts the appearance order of any first initial video segment in the first initial video.
[0167] The order adjustment operations include, but are not limited to, custom interactive actions such as clicking, swiping, and long-pressing. The clip editing area is used to display the initial video clips according to their order of appearance in the initial video.
[0168] For example, referring to Figure 6, the system-side terminal 12 adjusts the appearance order of the first initial video segment 64 in the first initial video in response to the user's click and swipe operations on the first initial video segment 64 in the segment editing area 62. Understandably, the first initial video segment 64 appears before the first initial video segment 65 in the first initial video, and the first initial video segment 65 appears before the first initial video segment 66 in the first initial video. When the user clicks on the first initial video segment 64 and drags it between the first initial video segments 65 and 67 through a swipe operation, the appearance order of the first initial video segment 64 in the first initial video also changes accordingly to be between the first initial video segments 65 and 67.
[0169] It is worth noting that, in response to the user's operation of adjusting the order of any first initial video segment in the segment editing area, after adjusting the appearance order of any first initial video segment in the first initial video, the complete video with the adjusted segment order will be displayed in the video preview area of the video editing page.
[0170] In this embodiment, users can perform editing operations such as clipping, speed adjustment, clip replacement, and order adjustment on each initial video clip in the clip editing area of the video editing page according to actual needs, ensuring that the video content is presented in the desired manner and improving the flexibility of recording processing. Furthermore, after each clip editing operation, the edited video is automatically displayed in the video preview area of the video editing page, allowing users to easily view the editing results and improving the convenience and intelligence of the operation. The above method improves recording processing efficiency while ensuring the quality of the recording process.
[0171] S516: The system-side terminal 12 responds to the user's third trigger operation for any video template in the template editing area and adjusts the first initial video based on any video template.
[0172] The third triggering action includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing.
[0173] Optionally, the multiple video templates displayed in the template editing area each have predefined video output duration, background music, cropping rules, speed adjustment strategies, and the timing and position of video description text. This means that each video template has a different output duration and presentation effect. The system-side terminal 12 responds to the user's click on any video template in the template editing area, adjusting the initial video according to the selected template.
[0174] Please refer to Figure 6. The template editing area 63 of the video editing page 60 displays multiple video templates. Each video template is labeled with the corresponding video output duration. Users can select a video template based on the video output duration according to their personal needs to adjust the initial video.
[0175] For example, in response to a user's click on the video template 68 in the template editing area 63, the system-side terminal 12 applies the video template 68 to the initial video and displays the complete video obtained after template editing in the video preview area 61 of the video editing page 60.
[0176] In this embodiment, users can easily adjust the initial video based on any video template in the template editing area through a simple trigger operation. Since each video template includes elements such as the output duration, background music, cropping rules, speed adjustment strategy, and the timing and position of the video description text, users do not need to manually adjust these video elements one by one, greatly simplifying the editing process and improving video generation efficiency. Furthermore, this embodiment is based on a "start with the end in mind" strategy, allowing users to select a video template based on their desired final output video duration, effectively meeting their personalized needs. This method improves recording processing efficiency while ensuring recording quality.
[0177] S517: The system-side terminal 12 responds to the fourth trigger operation input by the user based on the video editing page and uses the preview video displayed in the video preview area as the target video.
[0178] The fourth trigger action includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing. The target video is the video the user expects to obtain; it can be the unedited initial video or a video generated after the user edits the initial video through the video editing page.
[0179] Optionally, in response to a fourth trigger operation input by the user based on the video editing page, the system-side terminal 12 may use the unedited initial video in the video preview area or the video generated after the user edits the initial video through the video editing page as the target video.
[0180] For example, please refer to Figure 6. In response to the user's click operation on the save button 68 in the video editing page 60, the system-side terminal 12 saves the unedited initial video in the video preview area 61 or the video generated after the user edits the initial video through the video editing page 60 as the target video.
[0181] In this embodiment, by responding to the fourth trigger operation input by the user based on the video editing page, the video associated with the initial video in the video preview area is saved as the target video, thus avoiding data loss and improving the efficiency of recording processing.
[0182] The aforementioned video recording processing method, through a series of steps including recording mode selection, recording start / stop control, recording marking, recording transmission, recording content recognition, key content extraction, removal of jittery frames, generation of the initial video, video clip editing, video template editing, and saving the target video, makes the entire video recording process more efficient and convenient. Users can not only quickly obtain high-quality initial videos but also personalize them according to their needs and preferences, thereby generating high-quality target video content. Furthermore, during video editing, users can intuitively view the edited video in the video preview area of the video editing page, making the editing process more efficient and intuitive. After editing, users only need to simply click to save the target video, simplifying the recording processing workflow and improving the user experience. By adopting this method, recording processing efficiency is improved, recording processing quality is ensured, and personalized user needs are effectively met.
[0183] In one embodiment, as shown in FIG8, another video recording method is provided. Taking the application of this method to the device-side terminal 11 and system-side terminal 12 shown in FIG1 as an example, the method includes the following steps:
[0184] S801: The device-side terminal 11 responds to the user's first control command, determines the target recording mode, and generates a video type tag corresponding to the target recording mode.
[0185] Specifically, S801 is the same as S501, and will not be repeated here.
[0186] S802: The device-side terminal 11 responds to the user's second control command and starts recording in the target recording mode.
[0187] Specifically, S802 is the same as S202, and will not be repeated here.
[0188] S803: The device-side terminal 11 responds to the user's third control command and generates a highlight video tag for marking highlight moments in the recording.
[0189] Specifically, S803 is the same as S203, and will not be repeated here.
[0190] S804: In response to the user's fourth control command, the device-side terminal 11 ends the recording and sends the recording and the video tag carried in the recording to the system-side terminal.
[0191] Specifically, S804 is the same as S504, and will not be repeated here.
[0192] S805: System-side terminal 12 obtains the video recording and video tag carried by the video recording sent by device-side terminal 11.
[0193] The video tags include video type tags corresponding to the target recording mode, highlight video tags used to mark highlights in the recording, and video scene tags used to mark video scenes in the recording. The video scenes in the aforementioned recordings include one or more of the following: human figures, scenery (such as natural landscapes, cityscapes, etc.), vehicles, and moving images (such as walking, cycling, skiing, etc.).
[0194] Specifically, S505 is the same as S205, and will not be repeated here.
[0195] S806: The system-side terminal 12 displays the recording selection page.
[0196] The video recording selection page includes a first historical video recording taken by the device-side terminal within a preset time period, and the first historical video recording includes video recordings.
[0197] Specifically, S806 is the same as S206, and will not be repeated here.
[0198] S807: In response to the second trigger operation input by the user based on the recording selection page, the system-side terminal 12 generates a second initial video based on the video tags carried by each recording in the second historical recording selected by the user on the recording selection page.
[0199] The second triggering operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing.
[0200] Understandably, the video recording selection page includes the first historical video recording taken by the device terminal 11 within a preset time period (such as the most recent day, the most recent week, the most recent month, etc.) and other historical video recordings taken by the device terminal 11 within other preset time periods (i.e., not the first preset time period). Specifically, assuming the aforementioned preset time period is the most recent day, the video recording selection page includes both the first historical video recording taken by the device terminal 11 within the most recent day and other historical video recordings taken by the device terminal 11 before the most recent day. Users can select the video recordings to generate a video based on all the video recordings displayed on the video recording selection page.
[0201] For example, referring to Figure 4, the user can click the selection control 46 in the upper right corner of the recording selection page 40 to start recording selection. When the user clicks any one or more recordings in multiple groups of historical recordings 41, the system terminal 12 responds to the user's click operation based on the recording selection page, displays the second historical recording selected by the user at the bottom of the recording selection page 40 (corresponding to the recording 44 selected for video generation in Figure 4), and generates a second initial video based on the second historical recording selected by the user on the recording selection page and the video tags carried by each recording in the second historical recording.
[0202] In one embodiment, the system-side terminal 12 responds to a second trigger operation input by the user based on the recording selection page, inputs the second historical recording into the video content recognition model, and obtains the recognition results of each recording in the second historical recording; based on the recognition results of each recording in the second historical recording and the video tags carried by each recording in the second historical recording, it extracts the second key video segment from the second historical recording; the second key video segment includes the highlight video segment corresponding to the highlight moment; it removes the jitter in the second key video segment to obtain the second initial video segment; and it generates the second initial video based on the second initial video segment.
[0203] The recognition results for each video in the second historical video include the speech recognition results and the face recognition results for each video in the second historical video. The video tags carried by each video in the second historical video include a video type tag corresponding to the target recording mode, a highlight video tag for marking highlight moments in the video, and a video scene tag for marking video scenes in the video. The second key video segment includes the highlight video segment corresponding to the highlight moment.
[0204] Optionally, the second key video segment also includes the stage video segment corresponding to valuable audio in the second historical recording, the stage video segment corresponding to the video frame (target video frame) containing the target face in the second historical recording, and the stage video segment corresponding to the video scene tag marked in the recordings in the second historical recording with video type tags of slow-motion recording and selfie recording.
[0205] Specifically, the system-side terminal 12 performs semantic analysis on the speech recognition results corresponding to each video recording in the second historical video recording, and retains the video segments containing important information, dialogue, or other meaningful content in the audio as valuable audio-corresponding video segments. For example, if a user says something describing the scenery captured by the device-side terminal 11 as nice during the recording process, the system-side terminal 12 will retain the corresponding video segment containing that user's words. Based on the face recognition results corresponding to each video recording in the second historical video recording, the system-side terminal 12 retains a video segment before and after each target video frame in the second historical video recording, for example, retaining a video segment of a preset duration (e.g., 3 seconds) before and after the target video frame as the corresponding video segment for the target video frame.
[0206] Furthermore, for recordings in the second historical recordings carrying highlight video tags, the system-side terminal 12 uses the timestamp information corresponding to the highlight moments in the highlight video tags to extract the highlight video segments corresponding to the highlight moments in the second historical recordings. For example, it retains video segments of a preset duration (e.g., 5 seconds) before and after the highlight moment as the highlight video segments corresponding to the highlight moments. The system-side terminal 12 also filters out recordings with video type tags of slow motion and self-shot recording based on the video type tags carried by each recording in the second historical recordings, that is, it filters out slow motion recordings and self-shot recordings in the second historical recordings. Then, for the video scene tags carried by the filtered slow motion recordings and self-shot recordings, it uses the timestamp information corresponding to the video scene in the video scene tags to extract the stage video segments in the slow motion recordings and self-shot recordings.
[0207] Understandably, the system-side terminal 12 can extract the stage video segments corresponding to valuable audio, the stage video segments corresponding to the target video frames, the highlight video segments corresponding to highlight moments, and the stage video segments from slow-motion recordings and self-shot recordings from the second historical recordings, as the second key video segments in the second historical recordings.
[0208] Furthermore, the system-side terminal 12 extracts video frames frame by frame from the second key video segment, employs motion estimation technology to detect motion changes between each frame, calculates the displacement of each frame relative to the reference frame based on the motion estimation results, and then uses a motion compensation algorithm to correct the position of each video frame, eliminating jitter. After smoothing the corrected frames using a filter, the processed frames are recombined into a new video segment, namely the second initial video segment. Then, the second initial video segment, after removing jitter, is spliced together in chronological order to generate the second initial video.
[0209] In this embodiment, the system terminal can extract the second key video segment from the second historical video based on the video tags carried by each video in the second historical video and remove the jitter in the second key video segment by the user with a simple trigger operation, thereby intelligently generating the second initial video. This improves the efficiency of video processing while ensuring the quality of video processing.
[0210] S808: The system-side terminal 12 displays the video editing page.
[0211] Specifically, S808 is identical to S511, and will not be repeated here.
[0212] S809: The system-side terminal 12 responds to the user's segment trimming operation on any second initial video segment in the segment editing area and adjusts the video duration of any second initial video segment.
[0213] The clipping operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing.
[0214] For example, referring to Figure 6, the system-side terminal 12 adjusts the video duration of the second initial video segment 64 in the second initial video in response to a user's click and progress bar cropping operation on one of the second initial video segments 64 in the segment editing area 62. Understandably, the user can also perform segment cropping operations on other second initial video segments in the segment editing area 62, such as second initial video segment 65, second initial video segment 66, second initial video segment 67, etc., thereby adjusting the video duration accordingly.
[0215] Specifically, in response to a user's click on one of the second initial video segments 64 in the segment editing area 62, the system-side terminal 12 displays a segment editing pop-up window 70 as shown in Figure 7. The user can adjust and view the playback progress of the second initial video segment 64 based on the video progress bar 71 in the segment editing pop-up window 70, select a video segment for trimming, and click the confirmation control 72 in the upper right corner of the segment editing pop-up window 70 to save the edited second initial video segment 64, thereby adjusting the video duration of the second initial video segment 64.
[0216] It is worth noting that, in response to a user's clipping operation on any of the second initial video segments in the clip editing area, after adjusting the video duration of any of the second initial video segments, the complete video obtained after clipping will be displayed in the video preview area of the video editing page.
[0217] S810: The system-side terminal 12 responds to the user's speed adjustment operation for any second initial video segment in the segment editing area, and adjusts the playback speed of any second initial video segment.
[0218] The speed adjustment operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long pressing.
[0219] For example, referring to Figure 6, the system-side terminal 12 adjusts the playback speed of the second initial video segment 64 in response to a user's click and speed adjustment type selection operation on one of the second initial video segments 64 in the segment editing area 62. Understandably, the user can also perform speed adjustment operations on other second initial video segments in the segment editing area 62, such as second initial video segment 65, second initial video segment 66, second initial video segment 67, etc., thereby adjusting the playback speed accordingly.
[0220] Specifically, in response to a user's click on one of the second initial video segments 64 in the segment editing area 62, the system-side terminal 12 displays a segment editing pop-up window 70 as shown in Figure 7. The user can click on the linear speed control 73 or the curve speed control 74 in the segment editing pop-up window 70 to select the corresponding speed type. For example, after clicking on the curve speed control 74, the user can select the montage speed type, jump speed type, bullet time speed type, etc. under curve speed in the segment editing pop-up window 70, and click on the confirmation control 72 in the upper right corner of the segment editing pop-up window 70 to save the edited second initial video segment 64, thereby adjusting the playback speed of the second initial video segment.
[0221] Furthermore, the system-side terminal 12 can also respond to the user's background audio adjustment operation on any second initial video segment in the segment editing area, adjusting the background audio of any second initial video segment. Specifically, after the system-side terminal 12 responds to the user's click operation on one of the second initial video segments 64 in the segment editing area 62 and displays the segment editing pop-up window 70 as shown in Figure 7, the user can click the background audio adjustment control 75 in the segment editing pop-up window 70 to choose to turn off or on the original sound in the second initial video segment 64, and click the confirmation control 72 in the upper right corner of the segment editing pop-up window 70 to save the edited second initial video segment 64, thereby adjusting the background audio of the second initial video segment.
[0222] It is worth noting that, in response to the user's speed adjustment operation on any second initial video segment in the clip editing area, after adjusting the playback speed of any second initial video segment, or in response to the user's background sound adjustment operation on any second initial video segment in the clip editing area, the complete video obtained after adjusting the playback speed or the complete video obtained after adjusting the background sound will be displayed in the video preview area of the video editing page.
[0223] S811: In response to a user's segment replacement operation on any second initial video segment in the segment editing area, the system-side terminal 12 replaces any second initial video segment with the target video segment.
[0224] The segment replacement operation includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing.
[0225] For example, referring to Figure 6, in response to a user's long-press operation on one of the second initial video segments 64 in the segment editing area 62, the system-side terminal 12 displays a video segment replacement pop-up window. The user can select the desired target video segment from the multiple video segments displayed in the pop-up window and click on it, thereby causing the system-side terminal 12 to replace the second initial video segment 64 with the user-selected target video segment. Understandably, the user can also perform segment replacement operations on other second initial video segments in the segment editing area 62, such as second initial video segments 65, 66, and 67, thereby adjusting the playback speed accordingly.
[0226] It is worth noting that the multiple video clips displayed in the aforementioned video clip replacement pop-up can be extracted from the second historical recording and / or other historical recordings, and this application embodiment does not limit this. In response to the user's clip replacement operation on any second initial video clip in the clip editing area, after replacing any second initial video clip with the target video clip, the complete video obtained after the clip replacement will be displayed in the video preview area of the video editing page.
[0227] S812: The system-side terminal 12 responds to the user's operation to adjust the order of any second initial video segment in the segment editing area, and adjusts the appearance order of any second initial video segment in the second initial video.
[0228] The order adjustment operations include, but are not limited to, custom interactive actions such as clicking, swiping, and long-pressing. The clip editing area is used to display the initial video clips according to their order of appearance in the initial video.
[0229] For example, referring to Figure 6, the system-side terminal 12 adjusts the appearance order of the second initial video segment 64 in the second initial video in response to the user's click and swipe operations on the second initial video segment 64 in the segment editing area 62. Understandably, the second initial video segment 64 appears before the second initial video segment 65 in the second initial video, and the second initial video segment 65 appears before the second initial video segment 66 in the second initial video. When the user clicks on the second initial video segment 64 and drags it between the second initial video segment 65 and the second initial video segment 67 through a swipe operation, the appearance order of the second initial video segment 64 in the second initial video also changes accordingly to be between the second initial video segment 65 and the second initial video segment 67.
[0230] It is worth noting that, in response to the user's operation of adjusting the order of any second initial video segment in the segment editing area, after adjusting the appearance order of any second initial video segment in the second initial video, the complete video with the adjusted segment order will be displayed in the video preview area of the video editing page.
[0231] In this embodiment, users can perform editing operations such as clipping, speed adjustment, clip replacement, and order adjustment on each second initial video clip in the clip editing area of the video editing page according to actual needs, ensuring that the video content is presented in the desired manner and improving the flexibility of recording processing. Furthermore, after each clip editing operation, the edited video is automatically displayed in the video preview area of the video editing page, allowing users to easily view the editing results and improving the convenience and intelligence of the operation. The above method improves recording processing efficiency while ensuring the quality of recording processing.
[0232] S813: The system-side terminal 12 responds to the user's third trigger operation for any video template in the template editing area and adjusts the second initial video based on any video template.
[0233] The third triggering action includes, but is not limited to, custom interactive actions such as clicking, swiping, and long-pressing.
[0234] Optionally, the multiple video templates displayed in the template editing area each have predefined video output duration, background music, cropping rules, speed adjustment strategies, and the timing and position of video description text. This means that each video template has a different output duration and presentation effect. The system-side terminal 12 responds to the user's click on any video template in the template editing area, adjusting the second initial video according to the selected template.
[0235] Please refer to Figure 6. The template editing area 63 of the video editing page 60 displays multiple video templates. Each video template is labeled with the corresponding video output duration. Users can select a video template based on the video output duration according to their personal needs to adjust the second initial video.
[0236] For example, in response to a user's click on the video template 68 in the template editing area 63, the system-side terminal 12 applies the video template 68 to the initial video and displays the complete video obtained after template editing in the video preview area 61 of the video editing page 60.
[0237] In this embodiment, users can easily adjust the second initial video based on any video template in the template editing area through a simple trigger operation. Since each video template includes elements such as the video's output duration, background music, cropping rules, speed adjustment strategy, and the timing and position of the video description text, users do not need to manually adjust these video elements one by one, greatly simplifying the editing process and improving video generation efficiency. Furthermore, this embodiment is based on a "start with the end in mind" strategy, allowing users to select a video template based on their desired final output video duration, effectively meeting their personalized needs. This method improves recording processing efficiency while ensuring recording quality.
[0238] S814: The system-side terminal 12 responds to the fourth trigger operation input by the user based on the video editing page and uses the preview video displayed in the video preview area as the target video.
[0239] Specifically, S814 is the same as S517, and will not be repeated here.
[0240] The aforementioned video recording processing method, through a series of steps including recording mode selection, recording start / stop control, recording marking, recording transmission, recording content recognition, key content extraction, removal of jittery frames, generation of a second initial video, video clip editing, video template editing, and saving the target video, makes the entire video recording process more efficient and convenient. Users can not only quickly obtain high-quality second initial videos but also personalize them according to their needs and preferences, thereby generating high-quality target video content. Furthermore, during video editing, users can intuitively view the edited video in the video preview area of the video editing page, making the editing process more efficient and intuitive. After editing, users only need to simply click to save the target video, simplifying the recording processing workflow and improving the user experience. By adopting this method, video recording processing efficiency is improved, recording processing quality is ensured, and personalized user needs are effectively met.
[0241] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0242] Based on the inventive concept of the above-described video recording method, as shown in FIG9, this application embodiment also provides a video recording processing apparatus 900 for implementing the above-described video recording method applied to a system-side terminal. The video recording processing apparatus 900 includes:
[0243] The acquisition module 901 is used to acquire the video recording sent by the device-side terminal and the video tag carried in the video recording;
[0244] Display module 902 is used to display a video recording selection page; the video recording selection page includes a first historical video recording taken by the device-side terminal within a preset time period, the first historical video recording includes video recordings;
[0245] The response module 903 is used to respond to the first trigger operation input by the user based on the recording selection page, and to generate a first initial video based on the first historical recording and the video tags carried by each recording in the first historical recording.
[0246] In one possible implementation, the video tags include highlight video tags for marking highlight moments in the recording; the response module 903 is specifically configured to respond to a first trigger operation input by the user based on the recording selection page, input the first historical recording to the video content recognition model, and obtain the recognition results of each recording in the first historical recording; based on the recognition results of each recording in the first historical recording and the video tags carried by each recording in the first historical recording, extract the first key video segment in the first historical recording; the first key video segment includes the highlight video segment corresponding to the highlight moment; remove the jitter in the first key video segment to obtain the first initial video segment; and generate the first initial video based on the first initial video segment.
[0247] In one possible implementation, the response module 903 is further configured to respond to a second trigger operation input by the user based on the recording selection page, and generate a second initial video based on the video tags carried by each recording in the second historical recording selected by the user on the recording selection page.
[0248] In one possible implementation, the response module 903 is further configured to respond to a second trigger operation input by the user based on the recording selection page, input the second historical recording to the video content recognition model, obtain the recognition results of each recording in the second historical recording; extract the second key video segment from the second historical recording based on the recognition results of each recording in the second historical recording and the video tags carried by each recording in the second historical recording; remove the jitter in the second key video segment to obtain the second initial video segment; and generate the second initial video based on the second initial video segment.
[0249] In one possible implementation, the display module 902 is also used to display a video editing page; the video editing page includes a video preview area for displaying a preview video associated with the first initial video.
[0250] In one possible implementation, the second acquisition module 802 is specifically used to acquire user binding information; and generate video description information corresponding to multiple video segments based on the video capture dates and user binding information corresponding to multiple video segments.
[0251] In one possible implementation, the video editing page further includes a segment editing area for displaying first initial video segments from the first initial video; the response module 903 is also configured to adjust the video duration of any first initial video segment in response to a user's segment trimming operation on any first initial video segment in the segment editing area; and / or adjust the playback speed of any first initial video segment in response to a user's speed adjustment operation on any first initial video segment in the segment editing area; and / or replace any first initial video segment with a target video segment in response to a user's segment replacement operation on any first initial video segment in the segment editing area; and / or adjust the appearance order of any first initial video segment in the first initial video in response to a user's order adjustment operation on any first initial video segment in the segment editing area.
[0252] In one possible implementation, the video editing page also includes a template editing area, which displays multiple video templates, each with a different video output duration and presentation effect; the response module 903 is also used to respond to a third trigger operation by the user on any one of the video templates in the template editing area, and to adjust the first initial video based on any one of the video templates.
[0253] In one possible implementation, the response module 903 is also configured to respond to a fourth trigger operation based on user input on the video editing page, and use the preview video displayed in the video preview area as the target video.
[0254] Based on the inventive concept of the above-described video recording method, as shown in FIG10, this application embodiment also provides a video recording processing apparatus 1000 for implementing the above-described video recording method applied to a device-side terminal. The video recording processing apparatus 1000 includes:
[0255] The first response module 1001 is used to determine the target recording mode in response to the user's first control command;
[0256] The second response module 1002 responds to the user's second control command and starts recording in the target recording mode;
[0257] The third response module 1003, in response to the user's third control command, generates highlight video tags for marking highlight moments in the recording;
[0258] The fourth response module 1004, in response to the user's fourth control command, ends the recording and sends the recording and the video tags carried by the recording to the system-side terminal. The video tags include highlight video tags used to mark the highlight moments in the recording, so that the system-side terminal can obtain the recording and the video tags carried by the recording and display the recording selection page. The recording selection page includes the first historical recording captured by the device-side terminal within a preset time period. The first historical recording includes recordings. In response to the user's first trigger operation based on the recording selection page, a first initial video is generated based on the first historical recording and the video tags carried by each recording in the first historical recording.
[0259] The control methods for the first, second, third, and fourth control commands include voice control.
[0260] Each module in the aforementioned video recording processing device 900 and video recording processing device 1000 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0261] This application embodiment also provides an electronic device, which can be a terminal, and its internal structure diagram is shown in Figure 11. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor of the electronic device provides computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the electronic device is used for exchanging information between the processor and external devices. The communication interface of the electronic device is used for wired or wireless communication with external user terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a video recording method. The display unit of the electronic device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the electronic device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the electronic device, or external keyboards, touchpads, or mice, etc.
[0262] Those skilled in the art will understand that the structure shown in Figure 11 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.
[0263] This application also provides a computer storage medium storing instructions that, when run on a computer or processor, cause the computer or processor to perform one or more steps in the above embodiments. If the constituent modules of the above-described electronic device are implemented as software functional units and sold or used as independent products, they can be stored in the above-described computer-readable storage medium.
[0264] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer storage medium or transmitted through the computer storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).
[0265] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.
[0266] The embodiments described above are merely preferred embodiments of this application and are not intended to limit the scope of this application. Any modifications and improvements made by those skilled in the art to the technical solutions of this application without departing from the spirit of this application should fall within the protection scope defined by the claims.
[0267] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A video recording processing method, applied to a system-side terminal, characterized in that, The method includes: Obtain the video recording sent by the device-side terminal and the video tag carried by the video recording; Display a video recording selection page; the video recording selection page includes a first historical video recording taken by the device-side terminal within a preset time period, and the first historical video recording includes the video recording; In response to a first trigger operation input by the user based on the video selection page, a first initial video is generated based on the first historical video and the video tags carried by each video in the first historical video.
2. The method as described in claim 1, characterized in that, The video tags include highlight video tags used to mark highlight moments in the recording; The step of responding to a first trigger operation input by the user based on the video selection page, generating a first initial video based on the first historical video and the video tags carried by each video in the first historical video, includes: In response to a first trigger operation input by the user based on the video selection page, the first historical video is input into the video content recognition model to obtain the recognition results of each video in the first historical video. Based on the identification results of each video in the first historical video and the video tags carried by each video in the first historical video, a first key video segment is extracted from the first historical video; the first key video segment includes the highlight video segment corresponding to the highlight moment; Remove the jitter from the first key video segment to obtain the first initial video segment; A first initial video is generated based on the first initial video segment.
3. The method as described in claim 1, characterized in that, After displaying the video recording selection page, the method further includes: In response to the second trigger operation input by the user based on the video recording selection page, a second initial video is generated based on the second historical video selected by the user on the video recording selection page and the video tags carried by each video in the second historical video.
4. The method as described in claim 3, characterized in that, The response to the second trigger operation input by the user based on the video selection page, generating a second initial video based on the second historical video selected by the user on the video selection page and the video tags carried by each video in the second historical video, includes: In response to a second trigger operation input by the user based on the video selection page, the second historical video is input into the video content recognition model to obtain the recognition results of each video in the second historical video. Based on the identification results of each video in the second historical video and the video tags carried by each video in the second historical video, the second key video segment is extracted from the second historical video; Remove the jitter from the second key video segment to obtain the second initial video segment; A second initial video is generated based on the second initial video segment.
5. The method as described in claim 1, characterized in that, After generating a first initial video based on the first historical video and the video tags carried by each video in the first historical video in response to a first trigger operation input by the user on the video selection page, the method further includes: The video editing page is displayed; the video editing page includes a video preview area, which is used to display a preview video associated with the first initial video.
6. The method as described in claim 5, characterized in that, The video editing page also includes a segment editing area, which is used to display a first initial video segment from the first initial video. After displaying the video editing page, the method further includes: In response to the user's segment trimming operation on any first initial video segment in the segment editing area, the video duration of any first initial video segment is adjusted; and / or In response to the user's speed adjustment operation for any first initial video segment in the segment editing area, the playback speed of the any first initial video segment is adjusted; and / or In response to the user's segment replacement operation for any first initial video segment in the segment editing area, the any first initial video segment is replaced with the target video segment; and / or In response to the user's operation to adjust the order of any first initial video segment in the segment editing area, the order in which the any first initial video segment appears in the first initial video is adjusted.
7. The method as described in claim 5, characterized in that, The video editing page also includes a template editing area, which is used to display multiple video templates. Each of the multiple video templates has a different video output duration and video presentation effect. After displaying the video editing page, the method further includes: In response to the user's third trigger operation for any video template in the template editing area, the first initial video is adjusted based on the arbitrary video template.
8. The method as described in claim 5, characterized in that, The method further includes: In response to the fourth trigger operation input by the user based on the video editing page, the preview video displayed in the video preview area is used as the target video.
9. A video recording processing method, applied to a device-side terminal, characterized in that, The method includes: In response to the user's first control command, determine the target recording mode; In response to the user's second control command, recording is initiated in the target recording mode; In response to a third control command from the user, a highlight video tag is generated to mark the highlight moments in the recording; In response to the user's fourth control command, the recording is terminated, and the recording and the video tags carried by the recording are sent to the system-side terminal. The video tags include highlight video tags used to mark highlight moments in the recording, so that the system-side terminal can obtain the recording and the video tags carried by the recording and display a recording selection page. The recording selection page includes a first historical recording taken by the device-side terminal within a preset time period. The first historical recording includes the recording. In response to the user's first trigger operation based on the recording selection page, a first initial video is generated based on the first historical recording and the video tags carried by each recording in the first historical recording. The control methods of the first control command, the second control command, the third control command, and the fourth control command include voice control.
10. A video recording processing device, applied to a system-side terminal, characterized in that, The device includes: The acquisition module is used to acquire the video recording sent by the device-side terminal and the video tag carried by the video recording; The display module is used to display a video recording selection page; the video recording selection page includes a first historical video recording taken by the device-side terminal within a preset time period, and the first historical video recording includes the video recording. The response module is used to respond to a first trigger operation input by the user based on the recording selection page, and to generate a first initial video based on the first historical recording and the video tags carried by each recording in the first historical recording.
11. A video recording processing apparatus, applied to a device-side terminal, characterized in that, The device includes: The first response module is used to determine the target recording mode in response to the user's first control command; The second response module, in response to the user's second control command, starts recording in the target recording mode; The third response module, in response to the user's third control command, generates highlight video tags for marking highlight moments in the recording; The fourth response module, in response to the user's fourth control command, ends the recording and sends the recording and the video tags carried by the recording to the system-side terminal. The video tags include highlight video tags for marking highlight moments in the recording, so that the system-side terminal can obtain the recording and the video tags carried by the recording and display a recording selection page. The recording selection page includes a first historical recording taken by the device-side terminal within a preset time period. The first historical recording includes the recording. In response to the user's first trigger operation input based on the recording selection page, a first initial video is generated based on the first historical recording and the video tags carried by each recording in the first historical recording. The control methods of the first control command, the second control command, the third control command, and the fourth control command include voice control.
12. An electronic device, characterized in that, include: A processor and a memory; the memory stores a computer program, and the processor executes the computer program to implement the method steps of any one of claims 1-8 or 9.
13. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the method steps as claimed in any one of claims 1-8 or 9.
Citation Information
Patent Citations
Video processing method and mobile terminal
CN110557565A
Video processing method and device, readable medium and electronic equipment
CN111447489A
Video data processing method, device, equipment and medium
CN112565825A
Video processing method, electronic equipment and readable medium
CN114827342A
Video processing method, electronic equipment, chip system and storage medium
CN118474448A