Video note generation method and device and electronic equipment
By setting content generation controls in the video playback window and generating pending content in combination with the current video frame, the problem of users needing to frequently switch video players and note-taking tools is solved, and the efficiency of note generation is improved.
Patent Information
- Application Number
- CN202510280258.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, users need to frequently switch video players and note-taking tools when viewing videos, resulting in low note-generating efficiency.
A method for generating video notes is provided. By setting a content generation control in the video playback window, detecting that the user selects the control, obtains the current video frame and combines the content generation control to perform content generation processing, generates the content to be processed and displays in the note processing area, and determines the video notes corresponding to the video according to user operations.
Eliminates software isolation between video player and note-taking tool, and users do not need to switch tools frequently, improving the efficiency of note generation.
Smart Images

Figure CN120128752A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as smart network disks, video processing, deep learning, natural language processing, computer vision, and large models, and in particular to a method, device, and electronic device for generating video notes. Background Art
[0002] Currently, when users use a video player to watch a video and need to take notes, they need to use a note-taking tool to edit the notes. However, there is software isolation between the video player and the note-taking tool, which causes the user to frequently switch between the video player and the note-taking tool, resulting in low note generation efficiency. Summary of the invention
[0003] The present invention provides a method, device and electronic device for generating video notes.
[0004] According to one aspect of the present disclosure, a method for generating video notes is provided, the method comprising: displaying a video playback window; a video playback area is provided in the video playback window; at least one content generation control is displayed in the video playback area; when a selection operation is detected for a first content generation control among at least one of the content generation controls, a current video frame of the video played in the video playback window is obtained; content generation processing is performed in combination with the first content generation control and the current video frame to obtain content to be processed, and the content to be processed is displayed in an editing component in a note processing area of the video playback window; and the video note corresponding to the video is determined based on the operation on the content to be processed in the editing component and the content to be processed.
[0005] According to another aspect of the present disclosure, a device for generating video notes is provided, the device comprising: a window display module for displaying a video playback window; a video playback area is provided in the video playback window; at least one content generation control is displayed in the video playback area; a video frame acquisition module for acquiring a current video frame of a video played in the video playback window when a selection operation is detected for a first content generation control in at least one of the content generation controls; a content generation module for performing content generation processing in combination with the first content generation control and the current video frame to obtain content to be processed, and displaying the content to be processed in an editing component in a note processing area of the video playback window; a note determination module for determining the video note corresponding to the video based on the operation on the content to be processed in the editing component and the content to be processed.
[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the video note generation method proposed above in the present disclosure.
[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the video note generation method proposed above in the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program implements the steps of the video note generation method proposed above in the present disclosure when executed by a processor.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Description of the Drawings
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0012] Figure 2A is a schematic diagram according to the second embodiment of the present disclosure;
[0013] Figure 2B is a schematic diagram showing the content generation control;
[0014] Figure 3A is a schematic diagram according to the third embodiment of the present disclosure;
[0015] Figure 3B is a schematic diagram of the note assistance mode and the content to be processed;
[0016] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0017] Figure 5 is a block diagram of an electronic device for implementing the video note generation method of the embodiments of the present disclosure. Detailed Embodiments
[0018] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0019] When users watch videos using a video player such as a network disk, there are a large number of pause behaviors, mainly for doing questions, taking notes, or marking key content. However, the existing note-taking tools lack sufficient linkage with the video, resulting in users having to frequently switch tools during the learning process, with low efficiency.
[0020] In addition, existing artificial intelligence (AI) note-taking tools have pain points in extracting blackboard writing content, screenshot illustrations, and text collation. For example, verbatim transcription is time-consuming, the screenshot path is long, and the original video cannot be located. Although some note-taking tools support audio-video to text conversion, they lack linkage with the video player and can only import the generated video courseware and video manuscripts into the note-taking tool for audio-video to text processing, and then select and insert the text content into the note.
[0021] In view of the above problems, the present disclosure proposes a method, apparatus, and electronic device for generating video notes.
[0022] Figure 1 FIG. is a schematic diagram according to the first embodiment of the present disclosure. It should be noted that the method for generating video notes in the embodiments of the present disclosure can be applied to a device for generating video notes, and the device can be configured in an electronic device so that the electronic device can execute the function of generating video notes.
[0023] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, etc., which are hardware devices with various operating systems, touch screens, and / or display screens.
[0024] Among them, the device for generating video notes can also be software in the electronic device, such as software for generating video notes. In the following embodiments, the execution entity is taken as an example of an electronic device for description.
[0025] As Figure 1 shown, the method for generating video notes can include the following steps:
[0026] Step 101: Display a video playback window. A video playback area is set in the video playback window, and at least one content generation control is displayed in the video playback area.
[0027] In the embodiments of the present disclosure, an electronic device can display a video playback window, which is a graphical user interface element that can be used to display videos. The video playback area is used to play and display videos. Users can watch videos in the video playback area and perform basic playback controls, such as play, pause, etc.
[0028] Among them, the video playback area may refer to a partial area in the video playback window.
[0029] In the embodiments of the present disclosure, the content generation control is a clickable control within the video playback area, used to trigger the content generation function. After the user clicks the content generation control, the electronic device can automatically generate note content related to the video content. Among them, the content generation control may include a content generation button.
[0030] In some embodiments, the content generation control may include a text extraction control, a speech recognition control, a summary extraction control, etc.
[0031] In some embodiments, in order to meet the diverse content generation needs of users and improve the user experience, the content generation control includes at least one of the following: a screenshot control, a subtitle extraction control, and a text extraction control.
[0032] Among them, different content generation controls generate different note contents. For example, assuming that the content generation control includes a text extraction control, a speech recognition control, a summary extraction control, etc., the corresponding note content may include text, video summary / video frame summary, etc.
[0033] Step 102: When a selection operation for a first content generation control among at least one content generation control is detected, obtain the current video frame of the video being played in the video playback window.
[0034] In the embodiments of the present disclosure, the first content generation control may refer to any one of the at least one content generation control. The current video frame may refer to the video frame being played when the selection operation is detected. The electronic device can detect the selection operation for any content generation control in the video playback area, so as to obtain the current video frame of the video being played in the video playback window.
[0035] Step 103: Perform content generation processing by combining the first content generation control and the current video frame to obtain the content to be processed, and display the content to be processed in the editing component in the note processing area of the video playback window.
[0036] In the embodiments of the present disclosure, the note processing area is used to process notes related to videos. The note processing area can display editing components, and users can operate on the content to be processed within the editing components. For example, the editing components can be in the form of cards, pop-up windows, etc.
[0037] Among them, the note processing area also includes an area for displaying controls such as a font color adjustment control, a font size adjustment control, a note saving control, a note generation control, a mind map generation control, and an exercise generation control.
[0038] Among them, the video playback area and the note processing area can refer to two adjacent areas or two non-adjacent areas. It should be noted that the present disclosure does not limit the area sizes of the video playback area and the note processing area. The area sizes of the video playback area and the note processing area can be the same or different. As an example, in order not to affect the user's video viewing experience, the area size of the video playback area can be larger than that of the note processing area.
[0039] In the embodiments of the present disclosure, the content to be processed can include at least one of the current video frame, the video frame summary of the current video frame, the text in the current video frame, and / or the speech text.
[0040] Among them, a content generation large model matching the first content generation control can be used to combine the first content generation control and the current video frame for content generation processing to obtain the content to be processed.
[0041] It should be noted that the content generation large models corresponding to the respective content generation controls can be different large models or the same large model. Among them, when displaying the content to be processed in the editing component in the note processing area, the insertion display position of the editing component can be determined according to the position of the cursor in the note processing area. As an example, the editing component can be automatically inserted at the position of the cursor according to the cursor position indicated by the user; if the cursor loses focus, it can be inserted at the previous cursor position; if there is information in the line where the cursor is located, it will automatically wrap, and if there is no information in the line where the cursor is located, there is no need to wrap.
[0042] Step 104, determine the video note corresponding to the video according to the operation on the content to be processed in the editing component and the content to be processed.
[0043] In the embodiments of the present disclosure, the operations on the content to be processed can include operations such as moving the position of the content to be processed, adjusting the font size / color of the content to be processed, and marking part or all of the content in the content to be processed. The electronic device can combine the operations on the content to be processed and generate corresponding video notes according to the content to be processed.
[0044] The method for generating video notes according to the embodiments of the present disclosure includes displaying a video playback window; a video playback area is set in the video playback window; at least one content generation control is displayed in the video playback area; when a selection operation on a first content generation control among at least one content generation control is detected, obtaining the current video frame of the video being played in the video playback window; performing content generation processing by combining the first content generation control and the current video frame to obtain the content to be processed, and displaying the content to be processed in an editing component in the note processing area of the video playback window; determining the video note corresponding to the video according to the operation on the content to be processed in the editing component and the content to be processed; wherein, by setting the content generation control, the video playback and note processing functions are integrated into one, so that the video note can be automatically generated by combining the selection operation on the content generation control and the current video frame, eliminating the software isolation between the video player and the note tool, and the user does not need to import the relevant content of the current video frame into the note tool for note processing, improving the note generation efficiency.
[0045] Among them, in order to provide a better user interaction experience for users, different video playback windows can be displayed in different situations. As Figure 2A shown, Figure 2A is a schematic diagram according to the second embodiment of the present disclosure, Figure 2A The embodiment shown may include the following steps:
[0046] Step 201, when the video in the video playback area is in a playing state and no video playback related controls are displayed in the video playback area, display the first video playback window; a screenshot control in the content generation controls is displayed in the video playback area of the first video playback window; the screenshot control is partially transparent.
[0047] In the embodiments of the present disclosure, the video playback related controls may include controls such as video play / pause controls, progress bar adjustment controls, volume adjustment controls, and brightness adjustment controls.
[0048] In the embodiments of the present disclosure, a video playback area is set in the first video playback window; a screenshot control is displayed in the video playback area.
[0049] In order to avoid the screenshot control affecting the user's video viewing experience, when the video is in a playing state and the video player is in an immersive state (no video playback related controls are displayed in the video playback area), the screenshot control can be set to have a certain transparency.
[0050] Step 202, when the video in the video playback area is in a playing state and video playback related controls are displayed in the video playback area, display the second video playback window; a screenshot control is displayed in the video playback area of the second video playback window; the screenshot control is opaque.
[0051] In an embodiment of the present disclosure, a video playback area is provided in the second video playback window; a screenshot control is displayed in the video playback area.
[0052] The transparent screenshot control is not easy for the user to click. To ensure the convenience of the user's operation on the screenshot control, when the video is in the playback state and the video player is in the non-immersive state (video playback-related controls are displayed in the video playback area), the screenshot control can be set to have no transparency.
[0053] According to steps 201 and 202, when the video in the video playback area is in the playback state, the screenshot control in the content generation control is displayed in the video playback areas of the first video playback window and the second video playback window. This design only displays the screenshot control during video playback, achieving function convergence (convergence of the content generation control), making the video playback area more concise and clear, and at the same time facilitating the user to quickly take screenshots while watching the video, improving the user experience.
[0054] It should be noted that when the video in the video playback area is in the paused state, all content generation controls can be displayed in the video playback area of the video playback window, such as Figure 2B shown, Figure 2B is a schematic diagram of the display of the content generation control. The video screenshot in the figure corresponds to the screenshot control, the extracted subtitle corresponds to the subtitle extraction control, and the intercepted text corresponds to the text extraction control.
[0055] In addition, the content generation control also corresponds to a close control. After clicking the close control, the screenshot control, the subtitle extraction control, and the text extraction control all disappear. Among them, when the user closes the pop-up window or floating layer for displaying the content generation control through the close control in the video player, the display state of the content generation control will remain closed within the current video player session opened this time. That is, even if the user switches videos within the video selection, the content generation control will not appear. However, when the user reopens the video player next time, the content generation control will resume display, which not only ensures the convenience of the user's operation within the current session but also ensures that the function is still available when the user uses it next time.
[0056] Step 203, when the video in the video playback area is in the playback state and the mouse pointer hovers within the first sub-region of the video playback area, display the third video playback window; the screenshot control, the subtitle extraction control, and the text extraction control in the content generation control are displayed in the video playback area of the third video playback window.
[0057] In the embodiments of the present disclosure, the first sub-region may refer to the entire region of the video player or the video stream region, and the video stream region does not include the intelligent screenshot entry region. Among them, when the video playing region is in the narrow screen state, the video stream region does not include the sidebar page region where the content generation control is located. The sidebar page region refers to the region where the pop-up window or floating layer for displaying the content generation control is located as described above.
[0058] When the mouse pointer hovers within the first sub-region, the functions are automatically expanded, and a screenshot control, a subtitle extraction control, and a text extraction control are displayed in the video playing region, facilitating the user to capture video frames, as well as extract subtitle and text information at any time.
[0059] In the embodiments of the present disclosure, when the mouse pointer hovers over the region where the content generation control is located, the pop-up window or floating layer displays the function description text of the content generation control, and / or the pop-up window or floating layer displays the shortcut keys corresponding to the content generation control.
[0060] Among them, the function description text is used to describe the functions of the content generation control, and the shortcut keys corresponding to the content generation control are used to quickly call the functions of the content generation control.
[0061] For example, the shortcut key corresponding to the screenshot control can be command+shift+d, the shortcut key corresponding to the subtitle extraction control can be command+shift+z, and the shortcut key corresponding to the text extraction control can be command+shift+x.
[0062] Among them, displaying the function description text of the content generation control in the pop-up window or floating layer helps the user intuitively and clearly understand the functions of the content generation control; displaying the shortcut keys corresponding to the content generation control in the pop-up window or floating layer facilitates the user to quickly memorize and operate, greatly improving the operation convenience.
[0063] It should be noted that steps 201-step 203 respectively correspond to the display methods of the video playing window in three cases.
[0064] Step 204, when a selection operation on the first content generation control among at least one content generation control is detected, obtain the current video frame of the video being played in the video playing window.
[0065] In the embodiments of the present disclosure, when a selection operation on the first content generation control among at least one content generation control is detected, determine whether a note processing region is set in the video playing window; when the note processing region is not set in the video playing window, display the note processing region in the video playing window.
[0066] Among them, the note processing area can be displayed in the video playback window according to the preset area size of the note processing area.
[0067] Among them, when the first content generation control is selected and there is no note processing area set in the video playback window, automatically displaying the note processing area in the video playback window helps the user to immediately record or process notes related to the video content.
[0068] Step 205: Perform content generation processing on the basis of the first content generation control and the current video frame to obtain the content to be processed, and display the content to be processed in the editing component in the note processing area of the video playback window.
[0069] In the embodiments of the present disclosure, the first content generation control is a screenshot control; the generation process of the content to be processed includes: taking a screenshot of the current video frame by combining the screenshot control to obtain a screenshot image; combining the screenshot image and the timestamp of the current video frame to generate the content to be processed.
[0070] Among them, the timestamp can be used to locate the position of the current video frame in the video being played in the video playback window. It should be noted that there is no binding relationship between the screenshot image and the timestamp, and either party can be deleted separately.
[0071] Among them, automatically generating the content to be processed in combination with the timestamp of the current video frame provides an accurate time positioning for the screenshot image, facilitating the user to quickly trace the specific time point of the screenshot image in the video.
[0072] In the embodiments of the present disclosure, the first content generation control is a subtitle extraction control; the generation process of the content to be processed includes: obtaining the current timestamp of the current video frame and the voice information of the video within the first time period before the current timestamp; performing speech recognition processing on the voice information to obtain subtitle text; combining the subtitle text and the timestamp of the current video frame to generate the content to be processed.
[0073] Among them, the first time period can be preset. For example, the first time period can be 2 minutes.
[0074] Among them, the subtitle extraction control can automatically recognize the voice information associated with the current video frame and automatically generate the content to be processed in combination with the timestamp of the current video frame, with high efficiency. In addition, performing speech recognition on the voice information of the video within the first time period before the current timestamp to obtain subtitle text can obtain more text content, facilitating the user to quickly understand the notes when reviewing; in addition, the timestamp can be used to accurately locate the current video frame when reviewing.
[0075] In an embodiment of the present disclosure, the first content generation control is a text extraction control; the generation process of the content to be processed includes: taking a screenshot of the current video frame to obtain a screenshot image; performing text recognition processing on the screenshot image to obtain recognized text; and generating the content to be processed by combining the recognized text and the timestamp of the current video frame.
[0076] Among them, optical character recognition (OCR) technology can be used to perform text recognition processing on the screenshot image.
[0077] Among them, through the text extraction control, the text in the screenshot image can be automatically recognized, and the content to be processed can be automatically generated in combination with the timestamp of the current video frame, with high efficiency, and the accurate positioning of the current video frame can be achieved during review through the timestamp.
[0078] In an embodiment of the present disclosure, the video being played in the video playback window is paused, and the video playback related controls in the video playback window are stopped from being displayed; a screenshot of the current video frame is taken according to the detected screenshot frame to obtain a screenshot image.
[0079] Among them, after the text extraction control is selected, the video playback is automatically paused, avoiding interference from the dynamic clutter of the picture during the screenshot operation. At the same time, hiding the video playback related controls can provide a clean interface environment for the screenshot operation, and thus a clear and high-quality screenshot image can be obtained.
[0080] It should be noted that the generated content to be processed may not be displayed in the editing component either, but directly inserted into the note processing area.
[0081] Step 206: Determine the video note corresponding to the video according to the operation on the content to be processed in the editing component and the content to be processed.
[0082] Among them, it should be noted that for the detailed content of steps 204 to 206, reference can be made to Figure 1 Steps 102 to 104 in the shown embodiment, and details will not be described here again.
[0083] The method for generating video notes according to the embodiments of the present disclosure includes: when the video in the video playback area is in the playback state and no video playback-related controls are displayed in the video playback area, a first video playback window is displayed; a screenshot control in the content generation controls is displayed in the video playback area of the first video playback window; the screenshot control is partially transparent; when the video in the video playback area is in the playback state and video playback-related controls are displayed in the video playback area, a second video playback window is displayed; the screenshot control is displayed in the video playback area of the second video playback window; the screenshot control is opaque; when the video in the video playback area is in the playback state and the mouse pointer hovers within a first sub-region in the video playback area, a third video playback window is displayed; the screenshot control, subtitle extraction control, and text extraction control in the content generation controls are displayed in the video playback area of the third video playback window; when a selection operation for a first content generation control among at least one content generation control is detected, the current video frame of the video being played in the video playback window is obtained; content generation processing is performed in combination with the first content generation control and the current video frame to obtain the content to be processed, and the content to be processed is displayed in an editing component in the note processing area of the video playback window; according to the operation on the content to be processed in the editing component and the content to be processed, a video note corresponding to the video is determined; wherein, the screenshot control in the first video playback window is partially transparent, which can avoid the screenshot control affecting the user's video viewing experience; the screenshot control in the second video playback window is opaque, which can ensure the convenience of the user's operation on the screenshot control; the screenshot control, subtitle extraction control, and text extraction control are displayed in the third video playback window, which is convenient for the user to capture video frames at any time and extract subtitle and text information.
[0084] Among them, in order to assist the user in intelligent organization of notes to help the user efficiently record key content, in the note assistance mode, when the video is paused, the extraction content (note material) can be automatically obtained, and then the video note is generated. As Figure 3A shown, Figure 3A is a schematic diagram according to the third embodiment of the present disclosure, Figure 3A The embodiment shown may include the following steps:
[0085] Step 301, display a video playback window; a video playback area is set in the video playback window; a pause control is displayed in the video playback area.
[0086] In the embodiments of the present disclosure, in addition to the content generation controls, a pause control is also displayed in the video playback area, where the pause control belongs to the video playback-related controls.
[0087] Step 302, when a selection operation for the pause control is detected, determine whether the note assistance mode in the note processing area of the video playback window is in the on state.
[0088] In an embodiment of the present disclosure, the note assistance mode has a switch control. If the note assistance mode is in the on state, the electronic device can, when the pause control is selected, assist in generating content to be processed. Specifically, it automatically pauses the video playback to obtain a paused video frame, and performs at least two of screenshotting, subtitle extraction, and text extraction on the paused video frame to obtain extraction content, and thus generates content to be processed in combination with the extraction content.
[0089] Step 303: When the note assistance mode is in the on state, obtain the paused video frame of the video.
[0090] In an embodiment of the present disclosure, the paused video frame may refer to the video frame displayed when the video is paused.
[0091] In some embodiments, if the quality of the video frame displayed when the video is paused does not meet the set requirements (for example, the content is blurred), the electronic device may select a video frame with quality meeting the set requirements near the pause time point as the paused video frame.
[0092] Step 304: Perform at least two of screenshotting, subtitle extraction, and text extraction on the paused video frame to obtain extraction content.
[0093] In an embodiment of the present disclosure, the extraction content may include at least two of the screenshot image of the paused video frame, the subtitle text obtained from the voice information of the video within the first time period before the time stamp corresponding to the paused video frame, and the recognized text within the screenshot image of the paused video frame.
[0094] Among them, through the note assistance mode, intelligent sorting can be performed when the video is paused, and the extraction content (note materials) can be automatically obtained, without the user having to manually select the content generation control after pausing the video, helping the user record efficiently and saving the user's time for organizing notes.
[0095] Step 305: Generate content to be processed in combination with the extraction content.
[0096] In an embodiment of the present disclosure, the electronic device can generate content to be processed according to the extraction content. As Figure 3B shown, Figure 3B is a schematic diagram of the note assistance mode and the content to be processed. In the figure, the AI assistance mode corresponds to the note assistance mode, and the content to be processed shown in the lower right rectangular box is presented in the form of a card.
[0097] In some embodiments, when the extraction content includes the screenshot image obtained by screenshotting and the subtitle text obtained by subtitle extraction, obtain the time stamp of the paused video frame; generate content to be processed in combination with the screenshot image, the time stamp of the paused video frame, and the subtitle text.
[0098] Among them, by skillfully integrating the screenshot image, the timestamp of the accurately obtained paused video frame, and the subtitle text, content to be processed with rich content and accurate information can be obtained, which helps users review and share.
[0099] In an embodiment of the present disclosure, the content to be processed generated can be displayed in the note processing area. Among them, in the case where the content to be processed is not displayed in the note processing area, during the generation process of the content to be processed, a generating-in progress prompt text is displayed within the editing component; and / or, in the case where the content generation fails, a generation-failure prompt text is displayed within the editing component. Among them, during the process of performing content generation processing, displaying the generating-in progress prompt text or the generation-failure prompt text within the editing component enables the user to more clearly and intuitively grasp the content generation status and improve the user experience.
[0100] Among them, the display method of the generating-in progress prompt text and the generation-failure prompt text is the same as the display method of the content to be processed, that is, the insertion position of the prompt text is determined according to the position of the cursor.
[0101] Among them, at least one of the following processing controls is set for the generating-in progress prompt text in the note processing area: a termination control (after clicking, the generating-in progress prompt text directly disappears and the task is cancelled); at least one of the following processing controls is set for the generation-failure prompt text in the note processing area: a close control (after clicking, the generation-failure prompt text disappears), a retry control (changes back to the generating-in progress prompt text and regenerates).
[0102] Step 306, determine the video note corresponding to the video according to the operation on the content to be processed within the editing component and the content to be processed.
[0103] In an embodiment of the present disclosure, in the case where no operation on the content to be processed within the editing component is detected and new content to be processed is generated, the content to be processed is discarded; the new content to be processed is displayed in the note processing area.
[0104] Among them, for example, assume that the unoperated content to be processed is the content to be processed obtained after selecting the screenshot control, or the subtitle extraction control, or the text extraction control, and a new screenshot control selection event or pause control selection event, etc. occurs later, generating new content to be processed. In this case, the previous content to be processed is discarded. Another example, assume that the unoperated content to be processed is the content to be processed obtained after selecting the pause control; a new pause control selection event or screenshot control selection event, etc. occurs later, generating new content to be processed, then the previous content to be processed is discarded.
[0105] Among them, when generating new content to be processed, the old content to be processed is automatically discarded, ensuring that the content to be processed in the note processing area can be updated in a timely manner following the user's operations. At the same time, it also avoids redundancy of the content to be processed and improves the neatness of the note processing area.
[0106] In an embodiment of the present disclosure, when no operation on the content to be processed within the editing component is detected and the content of the video note changes, the content to be processed is discarded.
[0107] Among them, the change in the content of the video note may refer to the change or disappearance of the cursor position (such as when the user inputs content); to ensure the neatness of the content of the video note, after discarding the content to be processed, the line break behavior that occurs is restored to the state without cards (i.e., automatic backspace processing).
[0108] Among them, when no operation is performed on the content to be processed within the editing component, once the content of the video note changes, the current content to be processed is automatically discarded, ensuring the timeliness and accuracy of the content to be processed and avoiding the retention of outdated or incorrect information.
[0109] In an embodiment of the present disclosure, determining the video note corresponding to the video according to the operation on the content to be processed within the editing component and the content to be processed may include the following situations:
[0110] (1) When the operation on the content to be processed is an insertion operation, the content to be processed is determined as the content in the video note;
[0111] (2) When the operation on the content to be processed is a discard operation, the determination of the content to be processed as the content in the video note is discarded;
[0112] (3) When the operation on the content to be processed is a regeneration operation, the determination of the content to be processed as the content in the video note is discarded, and the content generation process of the content to be processed is regenerated by combining the first content generation control and the current video frame again.
[0113] Among them, at least one of the following processing controls is set for the content to be processed in the note processing area: discard control, copy control, insert control, and regeneration control.
[0114] Among them, when the discard control is selected, it is determined that a discard operation on the content to be processed is detected; when the copy control is selected, it is determined that an insertion operation on the content to be processed is detected, and an operation of copying the content in the content to be processed to the clipboard is detected; when the insert control is selected, it is determined that an insertion operation on the content to be processed is detected; when the regeneration control is selected, it is determined that a regeneration operation on the content to be processed is detected.
[0115] Among them, the user can perform flexible operations such as inserting, abandoning, or changing templates on the content to be processed according to actual needs. This design enables the user to easily manage the content to be processed, thereby improving the user's operation experience.
[0116] It should be noted that if the content to be processed is displayed in the form of a card, a pop-up window, or a floating layer, when the operation on the content to be processed is an insertion operation, after inserting the content to be processed into the note, the corresponding card, pop-up window, or floating layer can disappear; when the operation on the content to be processed is an abandonment operation, abandoning the content to be processed can mean deleting the corresponding card, pop-up window, or floating layer.
[0117] In the embodiments of the present disclosure, determine the number of consecutive abandonment operations on the content to be processed within the editing component in the video playback window; when the number is greater than or equal to the first number threshold, switch the note assistance mode in the note processing area to the closed state, and display a closing prompt text in the pop-up window or floating layer.
[0118] Among them, consecutive abandonment operations can include performing abandonment processing on multiple consecutively generated contents to be processed; the first number threshold can be preset. For example, the first number threshold can be 3; the closing prompt text is used to prompt the user that the note assistance mode has been closed.
[0119] Among them, when the number of consecutive abandonments of the content to be processed within the editing component in the video playback window reaches or exceeds the set threshold, it is recognized that the user may not want to turn on the note assistance mode. Therefore, automatically switching the note assistance mode in the note processing area to the closed state can avoid the inconvenience and interference brought to the user by the content to be processed generated by the note assistance mode. In addition, by displaying the closing prompt text in the pop-up window or floating layer, it is ensured that the user can timely understand the change in the on / off state of the note assistance mode, thereby further optimizing the user experience.
[0120] Among them, it should be noted that for the detailed content of step 306, reference can be made to Figure 1 step 104 in the shown embodiment, and details will not be described here again.
[0121] The method for generating video notes according to the embodiments of the present disclosure includes displaying a video playback window; a video playback area is set in the video playback window; a pause control is displayed in the video playback area; when a selection operation on the pause control is detected, determining whether the note assistance mode in the note processing area of the video playback window is in an enabled state; when the note assistance mode is in an enabled state, obtaining a paused video frame of the video; performing at least two of screenshot, subtitle extraction, and text extraction on the paused video frame to obtain extraction content; combining the extraction content to generate content to be processed; determining a video note corresponding to the video according to the operation on the content to be processed in the editing component and the content to be processed; wherein, through the note assistance mode, the extraction content can be automatically obtained when the video is paused, without the user manually selecting a content generation control, which helps the user record efficiently and saves the user's time for organizing notes.
[0122] To clearly illustrate the solution of the present disclosure, the following will be described in conjunction with the display logic, function description text, and shortcut keys of the screenshot control, subtitle extraction control, and text extraction control:
[0123] 1. For learning videos, display the screenshot control, subtitle extraction control, and text extraction control in the video playback area.
[0124] 2. Display logic:
[0125] a. During video playback: The functions are collapsed, and only the screenshot control is displayed.
[0126] When the video player is in the immersive state, the screenshot control has a certain transparency;
[0127] When the video player is in the non-immersive state (when there are video playback-related controls), the screenshot control is displayed normally without transparency.
[0128] b. When the video is paused: The functions are automatically expanded, and all three function controls are revealed.
[0129] c. When the user manually hovers: The functions are automatically expanded, all three function controls are revealed, and the close control corresponding to the display content generation control is displayed.
[0130] 3. Hot zone
[0131] When the mouse hovers over the video stream, player controls appear
[0132] Original hot zone: The entire area of the video player (network disk).
[0133] Current hot zone: The video stream area (excluding the sidebar page area where the content generation control is located in the narrow screen state), and the video stream area does not include the intelligent screenshot entry area (dynamically adjusted based on the current entry area height).
[0134] 4. Sidebar Adjustment
[0135] Wide Screen State: The appearance / disappearance of the sidebar page area is consistent with the video playback related controls.
[0136] Narrow Screen State: When the video player is in the immersive state, increase the transparency display.
[0137] 5. Function Description Text and Shortcut Keys
[0138] Function description text and shortcut keys when hovering over the screenshot control: Quick screenshot command+shift+d
[0139] Function description text and shortcut keys when hovering over the subtitle extraction control: Extract nearby subtitles command+shift+z
[0140] Function description text and shortcut keys when hovering over the text extraction control: Extract text within the screenshot command+shift+x
[0141] Among them, the shortcut keys corresponding to the content generation control should avoid conflicts with the existing shortcut keys of the user.
[0142] 6. Close Control
[0143] The content generation control corresponds to a close control. After clicking the close control, the screenshot control, subtitle extraction control, and text extraction control all disappear. Among them, the close logic remains closed within the video player session opened this time, that is, even if the user cuts the video within the selected episode, the content generation control will not appear; when the user reopens the video player next time, the function controls are normally displayed.
[0144] The following will separately explain the screenshot control, subtitle extraction control, text extraction control, and the pause state AI function linkage guidance of the AI note assistance mode based on AI:
[0145] 1. Screenshot Control:
[0146] After clicking the screenshot control or selecting the screenshot control through the shortcut key, open the sidebar and automatically position the note tab (Tabulation, tab). Among them, if the size of the video player does not meet the sidebar opening condition, the player size will be automatically adjusted.
[0147] Operation Negative Feedback: When the user continuously operates the screenshot control within a short period of time (for example, the number of operations on the screenshot control is greater than once within 0.5s), a toast prompt "Do not operate frequently" appears.
[0148] Insertion position of the content to be processed / editing component: Insert at the note cursor position (if there is content after the cursor, it will automatically wrap and be placed after the insertion content loading area; if there is no content, it will be inserted from this line). If the user's cursor loses focus, it will be inserted at the cursor position recorded by the user last time.
[0149] After clicking the screenshot control, the following situations may occur:
[0150] Loading: Prompt "Screenshot image uploading..."
[0151] Insertion successful: The screenshot image directly enters the note, or the editing component including the screenshot image directly enters the note
[0152] Insertion failed: Prompt "Upload failed"; if it is due to space limit, prompt "Insufficient space to upload".
[0153] Among them, in the case of insertion failure, the following processing controls may appear: Discard (clear after clicking), Retry (restore the loading state and re-upload the image), Go to expand capacity.
[0154] After clicking the screenshot control, the content inserted in the note processing area is: timestamp + screenshot image, or the editing component including timestamp + screenshot image. Among them, the timestamp and the screenshot image are not in a binding relationship, and either party can be deleted separately.
[0155] 2. Subtitle extraction control:
[0156] The subtitle extraction control is used to extract the voice information or subtitles of the previous 2 minutes corresponding to the current video frame to form a subtitle text; among them, if the current video does not support subtitle extraction, the subtitle extraction control will be dimmed, and a toast prompt will appear after clicking: The current video does not support subtitle extraction.
[0157] After clicking the subtitle extraction control or selecting the subtitle extraction control through the shortcut key, the sidebar will be opened and the note tab will be automatically positioned. Among them, if the video player size does not meet the sidebar opening condition, the player size will be automatically adjusted.
[0158] When the user operates the subtitle extraction control continuously within a short period of time (for example, the number of operations on the subtitle extraction control is greater than once within 0.5s), a toast prompt "Do not operate frequently" will appear.
[0159] Insertion position of the content to be processed / editing component: Insert at the note cursor position (if there is content after the cursor, it will automatically wrap and be placed after the insertion content loading area; if there is no content, it will be inserted from this line). If the user's cursor loses focus, it will be inserted at the cursor position recorded by the user last time.
[0160] After clicking the subtitle extraction control, the following situations may occur:
[0161] Inserting: Display a loading bar during extraction, with a prompt "Extracting subtitles, please wait a moment."
[0162] Insertion successful: The subtitle text directly enters the note, or the editing component including the subtitle text directly enters the note
[0163] Insertion failed: Prompt "Subtitle extraction failed, please try again."
[0164] Among them, in the case of insertion failure, the following processing controls can appear: Abandon (clears after clicking), Retry (restores the loading state and extracts subtitles again).
[0165] Among them, the output method of the subtitle text is: After all content is generated, it is output uniformly, in a non-streaming output manner.
[0166] 3. Text extraction control:
[0167] After clicking the text extraction control, the video pauses, the video player enters an immersive state, and the all-content generation control (tool entry) is hidden. A screenshot of the current video frame is taken to obtain a screenshot image, and text recognition processing is performed on the screenshot image to obtain the recognized text. Among them, the video playback pause caused by taking the screenshot automatically resumes after the screenshot is completed.
[0168] It should be noted that after clicking the text extraction control, a screenshot of the current video frame can be automatically taken to obtain a screenshot image, or after clicking the text extraction control, the video pauses, the video player enters an immersive state, and the sidebar is hidden. At the same time, the text extraction control and the screenshot control are displayed in a set area in the video playback area. After the user operates the screenshot control, a screenshot of the current video frame is taken to obtain a screenshot image. Among them, the set area also displays the close controls corresponding to the text extraction control and the screenshot control.
[0169] When the user continuously operates the text extraction control within a short period of time (for example, the number of operations on the text extraction control is greater than once within 0.5s), a toast prompt "Do not operate frequently" appears.
[0170] Insertion position of the content to be processed / editing component: Insert at the note cursor position (if there is content after the cursor, it will automatically wrap and be placed after the insertion content loading area; if there is no content, it will be inserted from this line). If the user's cursor loses focus, it will be inserted at the cursor position recorded by the user last time.
[0171] After clicking the screenshot control, the following situations can occur:
[0172] Inserting: Display a loading bar during extraction, with a prompt "Extracting text, please wait a moment."
[0173] Insertion successful: The recognized text directly enters the note, or the editing component including the recognized text directly enters the note. Among them, handling controls such as a discard control, a copy control, and an insertion control are set for the editing component.
[0174] Insertion failed: Prompt "Failed to extract text, please try again".
[0175] Among them, in the case of insertion failure, the following handling controls can appear: Discard (clears after clicking), Retry (restores the loading state and extracts the text again).
[0176] 4. Linkage guidance of the paused AI function based on the AI note assistance mode
[0177] AI note assistance mode:
[0178] The AI note assistance mode has a switch control. One of the prerequisite conditions for realizing the linkage guidance of the paused AI function is that the switch control is on. If the switch control is off, the linkage guidance of the paused AI function cannot be realized.
[0179] The switch control of the AI note assistance mode is displayed only when there is an associated video. Among them, if the user has intervened in the switch state, the switch control is preferentially controlled according to the user's intervention; if the user has not intervened in the switch state, it is judged according to whether the note file in the note processing area is empty. If it is empty, the switch state is on, and if it is not empty, the switch state is off.
[0180] Linkage guidance of the paused AI function:
[0181] When a selection operation on the pause control is detected and the note assistance mode is in the on state, a screenshot and subtitle extraction are performed on the paused video frame of the video to obtain the extracted content.
[0182] Assume that the editing component is in the form of a card. In this case, the editing component can be called a pause guidance card. Among them, in response to the completion of the screenshot, the editing component, that is, the pause guidance card (at this time, the text enters the loading state), can be displayed in the note processing area. At this time, the pause guidance card only includes the screenshot image and the corresponding timestamp. After the subtitle extraction is completed, the pause guidance card includes the complete content to be processed.
[0183] Among them, after the screenshot is completed, the pause guidance card can be automatically inserted at the position of the cursor. If the cursor loses focus, it appears at the previous cursor position. Among them, if there is information in the line where the cursor is located, it will automatically wrap; if there is no information in the line where the cursor is located, there is no need to wrap.
[0184] For the pause guidance card shown after the screenshot is completed, the card disappears when any of the following situations occur: the user exits the video player, when the user changes or disappears the cursor position, such as when the user enters content, uses the AI function of any note (triggering any insertable card), such as triggering the mind map generation control or the exercise generation control. Among them, when the card disappears, the line break behavior that occurs returns to the state without the card (i.e., automatic backspace processing).
[0185] The pause guidance card is unique: only 1 pause guidance card can be retained at the same time. If a new pause guidance card is triggered, replacement occurs (for example, when the pause guidance card is triggered after pausing, and the user pauses again after resuming playback, a new pause guidance card is triggered). Among them, if the cursor position does not change, the content of the current pause guidance card is replaced (the transition effect can be confirmed through the user interface (UI)); if the cursor position changes, the original pause guidance card disappears, and a new pause guidance card is inserted at the new cursor position.
[0186] Among them, a discard control (delete the card after clicking), a copy control (after clicking, the card content is inserted into the note and the card disappears, and at the same time the card content is copied to the clipboard, and a toast prompt: the content has been copied to the clipboard), and an insert control (after clicking, the card content is inserted into the note and the card disappears) are set for the pause guidance card. It should be noted that if the insert control is clicked when the text is in the loading state (subtitle extraction is not completed), the content inserted in the note processing area is the screenshot image and the corresponding timestamp, and the subtitle extraction task is interrupted.
[0187] It should be noted that when the text loading is completed, text content is displayed in the pause guidance card. If the text content exceeds the set number of lines, the excess content is omitted.
[0188] In addition, if the user clicks the discard control set for the pause guidance card multiple times in a row (such as three times), the AI note assistance mode is turned off, and a pop-up window or floating layer is triggered to display a closing prompt text. Among them, each user can trigger the closing logic of the AI note assistance mode at most once; the pop-up window or floating layer can be closed in any of the following ways: click the "I know" control, automatically close after a set duration (such as 10 seconds), trigger other floating layers in the top function bar, or operate any function in the top function bar.
[0189] To implement the above embodiments, the present disclosure also provides a video note generation device. As Figure 4 shown, Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure. The video note generation device 40 may include: a window display module 401, a video frame acquisition module 402, a content generation module 403, and a note determination module 404.
[0190] Among them, the window display module 401 is used to display a video playback window; a video playback area is set in the video playback window; at least one content generation control is displayed in the video playback area; the video frame acquisition module 402 is used to obtain the current video frame of the video being played in the video playback window when a selection operation on the first content generation control among at least one content generation control is detected; the content generation module 403 is used to perform content generation processing by combining the first content generation control and the current video frame to obtain the content to be processed, and display the content to be processed in the editing component in the note processing area of the video playback window; the note determination module 404 is used to determine the video note corresponding to the video according to the operation on the content to be processed in the editing component and the content to be processed.
[0191] As a possible implementation manner of the embodiment of the present disclosure, the content generation control includes at least one of the following: a screenshot control, a subtitle extraction control, and a text extraction control.
[0192] As a possible implementation manner of the embodiment of the present disclosure, the window display module 401 is specifically used to display a first video playback window when the video in the video playback area is in a playback state and no video playback related controls are displayed in the video playback area; the screenshot control in the content generation control is displayed in the video playback area of the first video playback window; the screenshot control is partially transparent; or, display a second video playback window when the video in the video playback area is in a playback state and video playback related controls are displayed in the video playback area; the screenshot control is displayed in the video playback area of the second video playback window; the screenshot control is opaque.
[0193] As a possible implementation manner of the embodiment of the present disclosure, the window display module 401 is specifically used to display a third video playback window when the video in the video playback area is in a playback state and the mouse pointer hovers in the first sub - area of the video playback area; the screenshot control, the subtitle extraction control, and the text extraction control in the content generation control are displayed in the video playback area of the third video playback window.
[0194] As a possible implementation manner of the embodiment of the present disclosure, the device further includes: a display module, used to pop up a window or display a floating layer with the function description text of the content generation control, and / or pop up a window or display a floating layer with the shortcut key corresponding to the content generation control when the mouse pointer hovers in the area where the content generation control is located.
[0195] As a possible implementation manner of the embodiments of the present disclosure, the device further includes: a region determination module, configured to determine whether a note processing region is set in the video playback window when a selection operation on a first content generation control among at least one content generation control is detected; a region display module, configured to display a note processing region in the video playback window when the note processing region is not set in the video playback window.
[0196] As a possible implementation manner of the embodiments of the present disclosure, the first content generation control is a screenshot control; the content generation module 403 is specifically configured to perform a screenshot process on the current video frame in combination with the screenshot control to obtain a screenshot image; and generate content to be processed in combination with the screenshot image and the timestamp of the current video frame.
[0197] As a possible implementation manner of the embodiments of the present disclosure, the first content generation control is a subtitle extraction control; the content generation module 403 is specifically configured to obtain the current timestamp of the current video frame and the voice information of the video within a first time period before the current timestamp; perform a voice recognition process on the voice information to obtain subtitle text; and generate content to be processed in combination with the subtitle text and the timestamp of the current video frame.
[0198] As a possible implementation manner of the embodiments of the present disclosure, the first content generation control is a text extraction control; the content generation module 403 is specifically configured to perform a screenshot process on the current video frame to obtain a screenshot image; perform a text recognition process on the screenshot image to obtain recognized text; and generate content to be processed in combination with the recognized text and the timestamp of the current video frame.
[0199] As a possible implementation manner of the embodiments of the present disclosure, the content generation module 403 is specifically configured to perform a pause playback process on the video played in the video playback window and stop displaying the video playback related controls in the video playback window; and perform a screenshot process on the current video frame according to the detected screenshot frame to obtain a screenshot image.
[0200] As a possible implementation manner of the embodiments of the present disclosure, a pause control is displayed in the video playback area; the device further includes: a state determination module, configured to determine whether the note assistance mode in the note processing region of the video playback window is in an on state when a selection operation on the pause control is detected; a paused video frame acquisition module, configured to acquire a paused video frame of the video when the note assistance mode is in an on state; a processing module, configured to perform at least two of screenshot, subtitle extraction, and text extraction on the paused video frame to obtain extracted content; and a card generation module, configured to generate content to be processed in combination with the extracted content.
[0201] As a possible implementation manner of an embodiment of the present disclosure, the extracted content includes a screenshot image obtained by taking a screenshot and subtitle text obtained by subtitle extraction; the card generation module is specifically configured to obtain the timestamp of the paused video frame; and generate the content to be processed by combining the screenshot image, the timestamp of the paused video frame, and the subtitle text.
[0202] As a possible implementation manner of an embodiment of the present disclosure, the device further includes: a card abandonment module, configured to abandon the content to be processed when no operation is detected on the content to be processed in the editing component and new content to be processed is generated; a first card display module, configured to display the new content to be processed in the editing component in the note processing area.
[0203] As a possible implementation manner of an embodiment of the present disclosure, the device further includes: a card abandonment module, configured to abandon the content to be processed when no operation is detected on the content to be processed in the editing component and the content of the video note changes.
[0204] As a possible implementation manner of an embodiment of the present disclosure, the note determination module 404 is specifically configured to, when the operation on the content to be processed is an insertion operation, determine the content to be processed as the content in the video note; when the operation on the content to be processed is an abandonment operation, abandon determining the content to be processed as the content in the video note; when the operation on the content to be processed is a regeneration operation, abandon determining the content to be processed as the content in the video note, and regenerate the content to be processed by combining the first content generation control and the current video frame again.
[0205] As a possible implementation manner of an embodiment of the present disclosure, the device further includes: a number determination module, configured to determine the number of consecutive abandonment operations on the content to be processed in the editing component in the video playback window; a state switching module, configured to switch the note assistance mode in the note processing area to the closed state and pop up a window or display a floating layer with a closing prompt text when the number is greater than or equal to the first number threshold.
[0206] As a possible implementation manner of an embodiment of the present disclosure, the device further includes: a second card display module, configured to display a generating prompt text in the editing component during the generation of the content to be processed when the content to be processed is not displayed in the note processing area; and / or display a generation failure prompt text in the editing component when the content generation fails.
[0207] The video note generation device according to the embodiments of the present disclosure displays a video playback window; a video playback area is set in the video playback window; at least one content generation control is displayed in the video playback area; when a selection operation for a first content generation control among the at least one content generation control is detected, the current video frame of the video being played in the video playback window is obtained; content generation processing is performed by combining the first content generation control and the current video frame to obtain content to be processed, and the content to be processed is displayed in an editing component in the note processing area of the video playback window; a video note corresponding to the video is determined according to the operation on the content to be processed in the editing component and the content to be processed; wherein, by setting the content generation control, the video playback and note processing functions are integrated into one, so that video notes can be automatically generated by combining the selection operation for the content generation control and the current video frame, eliminating the software isolation between the video player and the note tool, and the user does not need to import the relevant content of the current video frame into the note tool for note processing, improving the note generation efficiency.
[0208] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved are all carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0209] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0210] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0211] As Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0212] Multiple components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0213] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the method for generating video notes. For example, in some embodiments, the method for generating video notes can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method for generating video notes described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the method for generating video notes in any other appropriate manner (e.g., by means of firmware).
[0214] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0215] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0216] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0217] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0218] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0219] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server combined with a blockchain.
[0220] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0221] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating a video note, the method comprising: Display the video playback window; The video play window is provided with a video play area; At least one content generation control is displayed in the video playback area; When a selection operation is detected for a first content generation control in at least one of the content generation controls, obtaining a current video frame of the video played in the video playback window; Performing content generation processing in combination with the first content generation control and the current video frame to obtain content to be processed, and displaying the content to be processed in an editing component in a note processing area of the video playback window; The video notes corresponding to the video are determined according to the operation on the content to be processed in the editing component and the content to be processed.
2. The method according to claim 1, wherein: The content generation control includes at least one of the following: a screenshot control, a subtitle extraction control, and a text extraction control.
3. The method according to claim 1 or 2, wherein: The video playback window includes: When the video in the video play area is in a play state and no video play related controls are displayed in the video play area, a first video play window is displayed; a screenshot control in the content generation control is displayed in the video play area of the first video play window; the screenshot control is partially transparent; or, When the video in the video playback area is in playback state and video playback related controls are displayed in the video playback area, a second video playback window is displayed; the screenshot control is displayed in the video playback area of the second video playback window; and the screenshot control is opaque.
4. The method according to claim 3, wherein: The video playback window further includes: When the video in the video playback area is in playback state and the mouse pointer is suspended in the first sub-area of the video playback area, a third video playback window is displayed; and the screenshot control, subtitle extraction control and text extraction control in the content generation control are displayed in the video playback area of the third video playback window.
5. The method according to claim 1, wherein: The method further comprises: When the mouse pointer hovers over the area where the content generation control is located, a pop-up window or a floating layer displays a function description text of the content generation control, and / or a pop-up window or a floating layer displays a shortcut key corresponding to the content generation control.
6. The method according to claim 1, wherein: The method further comprises: In the case of detecting a selection operation on a first content generation control among at least one of the content generation controls, determining whether a note processing area is set in the video playback window; When the note processing area is not set in the video playback window, the note processing area is displayed in the video playback window.
7. The method according to claim 1 or 2, wherein: The first content generation control is a screenshot control; the combining the first content generation control and the current video frame to perform content generation processing to obtain the content to be processed includes: Performing screenshot processing on the current video frame in combination with the screenshot control to obtain a screenshot image; The content to be processed is generated by combining the screenshot image and the timestamp of the current video frame.
8. The method according to claim 1 or 2, wherein: The first content generation control is a subtitle extraction control; the combining the first content generation control and the current video frame to perform content generation processing to obtain the content to be processed includes: Obtaining a current timestamp of the current video frame and voice information of the video in a first time period before the current timestamp; Performing speech recognition processing on the speech information to obtain subtitle text; The content to be processed is generated by combining the subtitle text and the timestamp of the current video frame.
9. The method according to claim 1 or 2, wherein: The first content generation control is a text extraction control; the combining the first content generation control and the current video frame to perform content generation processing to obtain the content to be processed includes: Performing screenshot processing on the current video frame to obtain a screenshot image; Performing text recognition processing on the screenshot image to obtain recognized text; The content to be processed is generated by combining the recognized text and the timestamp of the current video frame.
10. The method according to claim 9, wherein: The step of performing screenshot processing on the current video frame to obtain a screenshot image includes: Pause the video played in the video play window, and stop displaying the video play related controls in the video play window; The current video frame is processed by taking a screenshot according to the detected screenshot frame to obtain the screenshot image.
11. The method according to claim 1, wherein: A pause control is displayed in the video playback area; and the method further includes: In the case where a selection operation on the pause control is detected, determining whether a note assist mode in a note processing area of the video playback window is in an on state; When the note assisting mode is turned on, obtaining a paused video frame of the video; Performing at least two of screenshot, subtitle extraction and text extraction on the paused video frame to obtain extracted content; The content to be processed is generated by combining the extracted content.
12. The method according to claim 11, wherein: The extracted content includes a screenshot image obtained by taking a screenshot and a subtitle text obtained by extracting subtitles; and combining the extracted content to generate the content to be processed includes: Obtaining the timestamp of the paused video frame; The content to be processed is generated by combining the screenshot image, the timestamp of the paused video frame and the subtitle text.
13. The method according to claim 1 or 11, wherein: The method further comprises: In the case where no operation on the content to be processed in the editing component is detected and new content to be processed is generated, abandoning the content to be processed; The new content to be processed is displayed in the editing component in the note processing area.
14. The method according to claim 1 or 11, wherein: The method further comprises: If no operation on the content to be processed in the editing component is detected and the content of the video note changes, the content to be processed is abandoned.
15. The method according to claim 1 or 11, wherein: The step of determining the video note corresponding to the video according to the operation on the content to be processed in the editing component and the content to be processed includes: In a case where the operation on the to-be-processed content is an insert operation, determining the to-be-processed content as the content in the video note; In a case where the operation on the to-be-processed content is a abandonment operation, abandoning determination of the to-be-processed content as the content in the video note; In the case where the operation on the content to be processed is a regeneration operation, the determination of the content to be processed as the content in the video note is abandoned, and the first content generation control and the current video frame are recombined to perform the processing of generating the content to be processed.
16. The method according to claim 15, wherein: The method further comprises: Determine the number of consecutive abandonment operations performed on the content to be processed in the editing component in the video playback window; When the number of times is greater than or equal to the first number threshold, the note auxiliary mode in the note processing area is switched to a closed state, and a closing prompt text is displayed in a pop-up window or floating layer.
17. The method according to claim 1 or 11, wherein: The method further comprises: When the content to be processed is not displayed in the note processing area, during the process of generating the content to be processed, a generating prompt text is displayed in the editing component; and / or, In the event that content generation fails, generation failure text is displayed within the editing component.
18. A device for generating video notes, comprising: Window display module, used to display the video playback window; The video play window is provided with a video play area; At least one content generation control is displayed in the video playback area; A video frame acquisition module, configured to acquire a current video frame of the video played in the video playback window when a selection operation is detected for a first content generation control in at least one of the content generation controls; A content generation module, configured to perform content generation processing in combination with the first content generation control and the current video frame to obtain content to be processed, and to display the content to be processed in an editing component in a note processing area of the video playback window; The note determination module is used to determine the video notes corresponding to the video based on the operation on the content to be processed in the editing component and the content to be processed.
19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 17.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 17.
21. A computer program product, comprising a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 17.