Video note generation method, electronic equipment and computer readable storage medium
By detecting user operations on an electronic device and entering the note mode, video notes are generated based on user operations on video and screen text content, the problem of video notes in the prior art is solved, and the effect of users quickly viewing the content of attention is achieved.
Patent Information
- Application Number
- CN202311787266.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, when a user generates video notes, it is not convenient to view subsequently through screenshots or copying, and cannot meet the actual needs of users, and the user experience is poor.
By detecting the user's operation on the electronic device, entering the note mode, and determining the target text content based on the user's second operation on the video and the text content in the screen, video notes are generated. The video notes may include the target text content and related content, and highlight the focus content.
It enables users to quickly and accurately view the content they are concerned about when viewing video notes, which facilitates users to view video notes, meets users' actual needs and improves user experience.
Smart Images

Figure CN120201250A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of terminals, and particularly relates to a method for generating video notes, an electronic device, and a computer-readable storage medium. Background Art
[0002] Currently, when a user views a video picture played by an electronic device, if the user is interested in the played video picture, the user can open a note application or a memo application, take a screenshot of the video picture to the note application or the memo application, or copy the content in the video picture to the note application or the memo application to form a video note. The video note formed in this way by screenshot or copy is not convenient for the user to view later, cannot meet the actual needs of the user, and the user experience is poor. Summary of the Invention
[0003] Embodiments of this application provide a method for generating video notes, an electronic device, and a computer-readable storage medium, which can generate video notes by combining the target text content determined by the operations of the user when viewing the first video, so that when the user views the video note, the user can quickly and accurately view the content he / she is concerned about, facilitating the user to view the video note, meeting the actual needs of the user, and improving the user experience.
[0004] In a first aspect, embodiments of this application provide a method for generating video notes, including:
[0005] The electronic device displays a first picture of the first video;
[0006] When detecting a first operation, the electronic device enters a note mode;
[0007] When, in the note mode, the electronic device detects a second operation on the first video, it determines a second picture, and the second picture is a picture in the first video;
[0008] The electronic device determines target text content according to the second operation and the text content in the second picture, and generates a video note according to the target text content.
[0009] In the above-provided video note generation method, when the electronic device displays the first frame of the first video, if the user wants to generate a video note corresponding to the first video, the user can perform a first operation. When the electronic device detects the first operation, it can enter the note mode. In the note mode, when the user performs a second operation on the first video, the electronic device can determine the second frame and can determine the target text content based on the second operation and the text content in the second frame, so as to generate a video note according to the target text content, that is, generate a video note based on the target text content determined by the relevant operations when the user watches the first video. The generated video note can include the target text content and / or content related to the target text content, and the target text content and / or content related to the target text content in the generated video note can be highlighted, so that when the user views the video note, the key content that the user is concerned about can be quickly and accurately viewed, which is convenient for the user to view the video note, meets the actual needs of the user, and improves the user experience.
[0010] It should be understood that the first video can be any video. For example, the first video can be a live video without subtitles, such as a live online course or a live sports event without subtitles. For example, the first video can be a non-live video with subtitles, such as a TV or a movie with subtitles.
[0011] Optionally, the first frame and the second frame can be any frame in the first video. Among them, the first frame and the second frame can be the same frame in the first video, that is, the first frame and the second frame can be the same frame in the first video. Or, the first frame and the second frame can be different frames in the first video, that is, the first frame and the second frame can be different frames in the first video.
[0012] Optionally, the first operation can be an operation to turn on the AI subtitles. That is, when the electronic device displays the first frame of the first video, if the user wants to generate a video note corresponding to the first video, the user can turn on the AI subtitles. When the electronic device detects that the AI subtitles are turned on, it can determine that the user wants to generate a video note corresponding to the first video. At this time, the electronic device can enter the note mode. Or, the first operation can be an operation to turn on the extraction function. That is, when the electronic device displays the first frame of the first video, if the user wants to generate a video note corresponding to the first video, the user can turn on the extraction function. When the electronic device detects that the extraction function is turned on, it can determine that the user wants to generate a video note corresponding to the first video. At this time, the electronic device can enter the note mode.
[0013] Exemplarily, the second frame is the frame displayed by the electronic device when the second operation is detected.
[0014] It should be understood that the second operation can be an operation performed on the screen of the first video. The second screen can be the screen that the electronic device is displaying when the second operation is detected. The second operation can be a screenshot operation, a content marking operation, a content extraction operation, or a content input operation.
[0015] In a possible implementation, the second operation is a screenshot operation, and the target text content is all or part of the text content in the second screen.
[0016] In the video note generation method provided by the implementation, after the electronic device enters the note mode, if a screenshot operation is detected, the electronic device can determine the second screen according to the time when the screenshot operation is detected, and can determine the target text content according to the text content in the second screen. For example, all the text content in the second screen can be determined as the target text content, or part of the text content in the second screen can be determined as the target text content.
[0017] Optionally, the text content in the second picture includes subtitles corresponding to the second picture. Alternatively, the text content in the second picture may include text content carried by the picture itself (eg, picture text) and subtitles corresponding to the second picture.
[0018] Among them, the subtitles corresponding to the second picture can be subtitles that come with the second picture itself, or subtitles added to the second picture through AI subtitles.
[0019] For example, after the electronic device enters the note-taking mode based on the first operation of turning on the AI subtitles, the electronic device can obtain the audio corresponding to the first video, convert the obtained audio into text content to obtain subtitles corresponding to the first video, and can display the AI subtitles corresponding to the picture when the picture of the first video is displayed in the display interface, that is, the subtitles corresponding to the picture obtained by converting the audio can be displayed in the picture.
[0020] In another possible implementation, the second operation is a partial screenshot operation, and the electronic device determines target text content according to the second operation and text content in the second screen, including:
[0021] The electronic device determines a screenshot image corresponding to the second operation;
[0022] The electronic device determines the target text content according to the text content in the screenshot image, where the text content in the screenshot image is all or part of the text content in the second screen.
[0023] In the video note generation method provided by this implementation manner, the screenshot operation may be a full-screen screenshot operation or a partial screenshot operation. The full-screen screenshot operation refers to an operation of taking a screenshot of the entire area of the display interface, that is, the screenshot image obtained by the full-screen screenshot operation may include the image corresponding to the entire display interface. The partial screenshot operation refers to an operation of taking a screenshot of a partial area of the display interface, that is, the screenshot image obtained by the partial screenshot operation may include the image corresponding to a certain partial area of the display interface. When the screenshot operation is a partial screenshot operation, the electronic device may determine the target text content according to the text content in the screenshot image, so as to accurately determine the content that the user focuses on according to the second operation performed by the user.
[0024] In another possible implementation manner, the second operation is a content marking operation, and the electronic device determines the target text content according to the second operation and the text content in the second screen, including:
[0025] The electronic device determines the text content marked by the second operation, and determines the target text content according to the text content marked by the second operation. The text content marked by the second operation is all or part of the text content in the second screen.
[0026] In the video note generation method provided by this implementation manner, after the electronic device enters the note mode, if a content marking operation is detected, the electronic device may determine the target text content according to the text content marked by the content marking operation. For example, the text content marked by the user may be determined as the target text content to improve the accuracy of the target text content.
[0027] It should be understood that the content marking operation may refer to an operation of marking the text content in the screen. For example, an operation of adding a color to the text content in the screen, or an operation of modifying the font color of the text content in the screen, or an operation of making the text content in the screen bold or enlarged, or an operation of drawing a line at the bottom of the text content in the screen, an operation of drawing a border for the text content in the screen, and so on.
[0028] In another possible implementation manner, the second operation is a content extraction operation, and the electronic device determines the target text content according to the second operation and the text content in the second screen, including:
[0029] The electronic device determines the text content extracted by the second operation, and determines the target text content according to the text content extracted by the second operation. The text content extracted by the second operation is all or part of the text content in the second screen.
[0030] In the video note generation method provided by this implementation manner, after the electronic device enters the note mode, if a content extraction operation is detected, the electronic device can determine the target text content according to the text content extracted by the content extraction operation. For example, the text content extracted by the user can be determined as the target text content to improve the accuracy of the target text content.
[0031] In another possible implementation manner, the second operation is a content input operation, and the electronic device determines the target text content according to the second operation and the text content in the second screen, including:
[0032] The electronic device determines the text content input by the second operation, and determines the target text content according to the text content in the second screen and the text content input by the second operation.
[0033] In the video note generation method provided by this implementation manner, after the electronic device enters the note mode, if a content input operation is detected, the electronic device can determine the target text content according to the text content input by the content input operation and the text content in the second screen. For example, the text content input by the user and the text content in the second screen can be determined together as the target text content to improve the accuracy of the target text content.
[0034] It should be understood that the content input operation may refer to an operation of inputting text content in the screen. For example, the content input operation may be an operation of adding a comment to the screen text in the screen, or an operation of adding a comment to the subtitle corresponding to the screen, and so on.
[0035] In one example, the method further includes:
[0036] When the electronic device detects a second operation on the first video, it determines a third screen, and the third screen is a screen in the first video, and the third screen is a screen before the second screen;
[0037] The electronic device obtains the text content in the third screen, and displays the text content in the third screen and the text content in the second screen.
[0038] In the video note generation method provided in this example, when the second operation is detected, the electronic device can determine the time T when the second operation is detected, and based on the time T, determine the text content in the N third frames before the second frame. That is, these N third frames can be the frames in the first video before the second frame. That is to say, during the playback of the first video, these N third frames can be displayed on the display interface before the second frame. Subsequently, the electronic device can simultaneously display the text content in these N third frames and the text content in the second frame on the display interface, so that the user can reselect the text content for the second operation from the text content in these N + 1 frames. The electronic device can determine the target text content based on the text content selected by the user to accurately determine the content that the user focuses on, reduce the possibility of incorrect determination of the target text content due to reasons such as slow hand speed, and improve the user experience.
[0039] Optionally, when the second operation is a screenshot operation, the electronic device determines the target text content according to the second operation and the text content in the second frame, including:
[0040] The electronic device obtains the text content corresponding to the third operation, and determines the target text content according to the text content corresponding to the third operation, where the third operation is an operation of selecting the text content in the second frame and the text content in the third frame.
[0041] Optionally, when the second operation is a content marking operation, a content extraction operation, or a content input operation, the electronic device determines the target text content according to the second operation and the text content in the second frame, including:
[0042] The electronic device detects a fourth operation on the text content in the second frame and / or the text content in the third frame, where the fourth operation is a content marking operation, a content extraction operation, or a content input operation;
[0043] The electronic device determines the target text content according to the fourth operation, the text content in the second frame, and / or the text content in the third frame.
[0044] In the video note generation method provided in this optional manner, the user can reselect, re-mark, re-extract, or re-enter the text content according to the text content in multiple frames displayed on the display interface. The electronic device can use the text content reselected, re-marked, re-extracted, or re-entered by the user to determine the target text content, so as to improve the accuracy of the target text content, ensure that the generated video note meets the actual needs of the user, and improve the user experience.
[0045] In a possible implementation, the electronic device generates a video note according to the target text content, including:
[0046] The electronic device obtains the text content in the target screen, and generates the video note according to the target text content and the text content in the target screen.
[0047] Optionally, the target screen includes the screen displayed after the electronic device enters the note mode.
[0048] Optionally, the target screen includes the screen displayed after the electronic device enters the note mode, and the screen displayed within a preset time before the electronic device enters the note mode.
[0049] In the video note generation method provided by this implementation, the electronic device can generate a video note according to the text content in the screen displayed on the display interface after the electronic device enters the note mode and the target text content, so as to generate a video note only for the content that the user is interested in, which is convenient for the user to view the video note later. Alternatively, the electronic device can generate a video note according to the text content in the screen displayed on the display interface after the electronic device enters the note mode, the text content in the screen displayed within a preset time before entering the note mode, and the target text content, so that when the user wants to generate a video note corresponding to the content before a certain content while watching the video, there is no need to perform a video return operation, improving the user experience.
[0050] It should be understood that the preset time can be default set by the electronic device or can be custom set by the user. Optionally, the electronic device can default set the preset time to any value such as 5 minutes, 10 minutes, 20 minutes, or 30 minutes according to the actual scenario, or can default set the preset time according to the start playing time and the current time of the electronic device playing the first video. For example, the preset time can be the time difference between the current time and the start playing time. For example, when the start playing time of the electronic device playing the first video is 19:30:00 and the current time is 19:44:00, the electronic device can default set the preset time to 14 minutes. Optionally, the user can custom set the preset time to any value such as 3 minutes, 5 minutes, or 10 minutes according to actual needs.
[0051] For example, when the electronic device detects a first operation, the electronic device may output a prompt message. Alternatively, after the electronic device enters the note mode based on the first operation, if an operation to close the note mode is detected, the electronic device may output a prompt message. Among them, the prompt message may be used to prompt the user whether to generate a video note corresponding to the video played before the electronic device enters the note mode. The prompt message may also be used to prompt the user to perform a custom setting or selection for a preset time when it is necessary to generate a video note corresponding to the video played before the electronic device enters the note mode.
[0052] In one example, the electronic device generating the video note according to the target text content and the text content in the target screen includes:
[0053] The electronic device generates an initial note according to the text content in the target screen;
[0054] The electronic device highlights a first content in the initial note to obtain the video note, and the first content is determined according to the target text content.
[0055] Optionally, the first content includes one or more of the target text content, the context corresponding to the target text content, or the text content having the same content view as the target text content.
[0056] In the video note generation method provided in this example, the electronic device may generate an initial note according to the text content in the target screen, and may highlight one or more of the target text content, the context corresponding to the target text content, or the text content having the same content view as the target text content in the initial note to obtain the video note. That is, the key content (such as the target text content, the context corresponding to the target text content, or the text content having the same content view as the target text content, etc.) that the user is concerned about can be highlighted in the video note, so that when the user views the video note, the text content that the user is concerned about can be quickly seen, which is convenient for the user to view the video note and improves the user experience.
[0057] Optionally, the highlighting may be performed by color marking, or may be performed by font enlargement, or may be performed by font bolding, etc.
[0058] Optionally, the video note includes the target text content, or includes a summary corresponding to the target text content.
[0059] Exemplarily, when the video note includes a summary corresponding to the target text content, the method further includes:
[0060] The electronic device detects a fifth operation on the summary corresponding to the target text content;
[0061] The electronic device displays the original text content corresponding to the summary according to the fifth operation, where the original text content is the text content in the target screen.
[0062] In the video note generation method provided by this alternative approach, the electronic device can directly organize the text content in the target screen to obtain video notes, that is, the video notes can include the target text content. Alternatively, the electronic device can determine one or more of the full text summary, chapter summary, keywords, etc. based on the text content in the target screen, and can generate video notes based on one or more of the full text summary, chapter summary, keywords, etc. That is to say, the video notes can include the content obtained by summarizing the text content in the target screen. Among them, when the video notes include the content obtained by summarizing the text content in the target screen, the electronic device can determine the original text content corresponding to each summary content, and can associate and save each summary content with the corresponding original text content. Subsequently, when the user views the video notes, the user can view the original text content through relevant operations to facilitate the user's understanding of the original text content.
[0063] It should be understood that the original text content can be the text content in the screen. For example, it can be the subtitle corresponding to the screen and / or the on-screen text in the screen.
[0064] For example, the video notes can include a summary corresponding to the target text content, and this summary can be obtained by summarizing the target text content. The electronic device can determine the original text content corresponding to this summary, and associate and save this summary with the target text content. When the user views the video notes, the user can view the original text content associated with this summary through relevant operations.
[0065] In another possible implementation, the electronic device generating video notes according to the target text content includes:
[0066] The electronic device obtains a screenshot image corresponding to the target text content, and generates the video notes according to the target text content and the screenshot image.
[0067] In the video note generation method provided by this implementation, the electronic device can also obtain a screenshot image corresponding to the target text content, and can generate video notes based on the screenshot image and the target text content to enrich the content of the video notes, facilitate the user to view the video notes, and improve the user experience.
[0068] Exemplarily, the electronic device generates the video note according to the target text content and the screenshot image, including:
[0069] The electronic device determines the position of the target text content in the initial note;
[0070] The electronic device inserts the screenshot image into the initial note according to the position of the target text content in the initial note to obtain the video note.
[0071] In the video note generation method provided in this example, the electronic device can insert a screenshot image in front of or behind the target text content according to the position of the target text content in the initial note to obtain a video note. Alternatively, the electronic device can insert an image annotation behind the target text content according to the position of the target text content in the initial note to obtain a video note. When the user views the image annotation corresponding to the target text content, for example, when the user clicks on the image annotation, the electronic device can display the screenshot image corresponding to the target text content on the display interface.
[0072] Optionally, the position of the target text content in the initial note may refer to the paragraph of the target text content in the initial note. After determining the position of the target text content in the initial note, the electronic device can insert a screenshot image in front of or behind the paragraph where the target text content is located according to the position of the target text content in the initial note to obtain a video note.
[0073] In a second aspect, an embodiment of the present application provides a video note generation device, including:
[0074] A first screen display module, configured to display a first screen of a first video;
[0075] A first operation detection module, configured to enter the note mode when a first operation is detected;
[0076] A second screen determination module, configured to determine a second screen in the note mode when a second operation on the first video is detected, where the second screen is a screen in the first video;
[0077] A video note generation module, configured to determine target text content according to the second operation and the text content in the second screen, and generate a video note according to the target text content.
[0078] Exemplarily, the second screen is the screen displayed by the electronic device when the second operation is detected.
[0079] In a possible implementation manner, the second operation is a screenshot operation, and the target text content is all or part of the text content in the second screen.
[0080] Optionally, the text content in the second screen includes the subtitles corresponding to the second screen.
[0081] In another possible implementation, the video note generation module is configured to determine a screenshot image corresponding to the second operation; determine the target text content according to the text content in the screenshot image, where the text content in the screenshot image is all or part of the text content in the second screen.
[0082] In another possible implementation, the second operation is a content marking operation;
[0083] The video note generation module is further configured to determine the text content marked by the second operation, and determine the target text content according to the text content marked by the second operation, where the text content marked by the second operation is all or part of the text content in the second screen.
[0084] In another possible implementation, the second operation is a content extraction operation;
[0085] The video note generation module is further configured to determine the text content extracted by the second operation, and determine the target text content according to the text content extracted by the second operation, where the text content extracted by the second operation is all or part of the text content in the second screen.
[0086] In another possible implementation, the second operation is a content input operation;
[0087] The video note generation module is further configured to determine the text content input by the second operation, and determine the target text content according to the text content in the second screen and the text content input by the second operation.
[0088] In one example, the apparatus further includes:
[0089] A third screen determination module, configured to determine a third screen when a second operation on the first video is detected, where the third screen is a screen in the first video, and the third screen is a screen before the second screen;
[0090] A text content display module, configured to obtain the text content in the third screen, and display the text content in the third screen and the text content in the second screen.
[0091] Optionally, when the second operation is a screenshot operation, the video note generation module is further configured to obtain the text content corresponding to the third operation, and determine the target text content according to the text content corresponding to the third operation, where the third operation is an operation of selecting the text content in the second screen and the text content in the third screen.
[0092] Optionally, when the second operation is a content marking operation, a content extraction operation, or a content input operation, the video note generation module is further configured to detect a fourth operation on the text content in the second screen and / or the text content in the third screen, where the fourth operation is a content marking operation, a content extraction operation, or a content input operation; and determine the target text content according to the fourth operation, the text content in the second screen, and / or the text content in the third screen.
[0093] In a possible implementation manner, the video note generation module is further configured to obtain the text content in the target screen, and generate the video note according to the target text content and the text content in the target screen.
[0094] Optionally, the target screen includes the screen displayed after entering the note mode.
[0095] Optionally, the target screen includes the screen displayed after entering the note mode and the screen displayed within a preset time before entering the note mode.
[0096] In an example, the video note generation module is further configured to generate an initial note according to the text content in the target screen; and highlight a first content in the initial note to obtain the video note, where the first content is determined according to the target text content.
[0097] Optionally, the first content includes one or more of the target text content, the context corresponding to the target text content, or the text content having the same content view as the target text content.
[0098] Optionally, the video note includes the target text content, or includes a summary corresponding to the target text content.
[0099] Exemplarily, when the video note includes a summary corresponding to the target text content, the apparatus further includes:
[0100] A fifth operation detection module, configured to detect a fifth operation on the summary corresponding to the target text content;
[0101] An original text content display module, configured to display the original text content corresponding to the summary according to the fifth operation, where the original text content is the text content in the target screen.
[0102] In another possible implementation, the video note generation module is further configured to obtain a screenshot image corresponding to the target text content, and generate the video note according to the target text content and the screenshot image.
[0103] Exemplarily, the video note generation module is further configured to determine the position of the target text content in the initial note, and insert the screenshot image into the initial note according to the position of the target text content in the initial note to obtain the video note.
[0104] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the video note generation method according to any one of the above first aspects.
[0105] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a computer, causes the computer to implement the video note generation method according to any one of the above first aspects.
[0106] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on an electronic device, causes the electronic device to execute the video note generation method according to any one of the above first aspects.
[0107] It can be understood that the beneficial effects of the above second to fifth aspects can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0109] Figure 2 is a schematic software architecture diagram applicable to the video generation method provided by an embodiment of the present application;
[0110] Figure 3 is a schematic flowchart of a video generation method provided by an embodiment of the present application;
[0111] Figure 4 is a schematic diagram of an application scenario provided by an embodiment of the present application Figure 1 ;
[0112] Figure 5 It is a schematic diagram of the application scenario provided by the embodiments of the present application Figure 2 ;
[0113] Figure 6 It is a schematic diagram of the application scenario provided by the embodiments of the present application Figure 3 ;
[0114] Figure 7 It is a schematic diagram of the application scenario provided by the embodiments of the present application Figure 4 ;
[0115] Figure 8 It is a schematic diagram of the application scenario provided by the embodiments of the present application Figure 5 ;
[0116] Figure 9 It is a schematic diagram of the application scenario provided by the embodiments of the present application Figure 6 。 Detailed implementation manners
[0117] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0118] It should also be understood that the term "and / or" as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0119] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when...", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0120] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0121] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "include but not limited to", unless otherwise specifically emphasized in other ways.
[0122] In addition, "a plurality of" mentioned in the embodiments of this application should be construed as two or more.
[0123] The steps involved in the video note generation method provided in the embodiments of this application are merely examples, not all steps are necessarily to be executed, or not all the content in each piece of information or message is mandatory. During use, it can be increased or decreased as needed. The same step or steps with the same function or messages in different embodiments of this application can be referred to and borrowed from each other.
[0124] The business scenarios described in the embodiments of this application are for more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation to the technical solutions provided by the embodiments of this application. As known to those of ordinary skill in the art, with the evolution of the network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0125] When a user is watching a video screen played on an electronic device, if the user is interested in the played video screen, the user can open a note application or a memo application, take a screenshot of the video screen to the note application or the memo application, or copy the content in the video screen to the note application or the memo application to form a video note. The video note formed in this way by taking a screenshot or copying is not convenient for the user to view later, cannot meet the actual needs of the user, and the user experience is poor.
[0126] To solve the above problems, an embodiment of the present application provides a method for generating a video note, an electronic device, and a computer-readable storage medium. In this method, when the electronic device displays the first frame of the first video, if the user wants to generate a video note corresponding to the first video, the user can perform a first operation. When the electronic device detects the first operation, it can enter the note mode. In the note mode, when the user performs a second operation on the first video, the electronic device can determine the second frame and can determine the target text content according to the second operation and the text content in the second frame, so as to generate a video note according to the target text content. Among them, the generated video note may include the target text content and / or content related to the target text content, and the target text content and / or content related to the target text content in the generated video note may be highlighted, so that when the user views the video note, the key content that the user is concerned about can be quickly and accurately viewed, which is convenient for the user to view the video note, meets the actual needs of the user, improves the user experience, and has strong usability and practicality.
[0127] In an embodiment of the present application, the electronic device may be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a notebook computer, an ultra-mobile personal computer (UMPC), a smart screen, a netbook, a personal digital assistant (PDA), a desktop computer, etc., which can play videos. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device.
[0128] First, the electronic device involved in the embodiment of the present application will be introduced below. Please refer to Figure 1 , Figure 1 which shows a schematic structural diagram of the electronic device provided by the embodiment of the present application.
[0129] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, an antenna 1, an antenna 2, a mobile communication module 140, a wireless communication module 150, a sensor module 160, and a display screen 170, etc. Among them, the sensor module 160 may include a pressure sensor 160A, a gyroscope sensor 160B, and a touch sensor 160C, etc.
[0130] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0131] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0132] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0133] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may store the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0134] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0135] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 may be coupled to the touch sensor 160C through the I2C interface, enabling the processor 110 and the touch sensor 160C to communicate through the I2C bus interface to implement the touch function of the electronic device 100.
[0136] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 150.
[0137] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 170. The MIPI interface includes a display serial interface (DSI), etc. In some embodiments, the processor 110 and the display screen 170 communicate through the DSI interface to implement the display function of the electronic device 100.
[0138] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the display screen 170, the wireless communication module 150, the sensor module 160, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0139] The USB interface 130 is an interface compliant with the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to transfer data between the electronic device 100 and peripheral devices. This interface can also be used to connect to other electronic devices, such as AR devices, etc.
[0140] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0141] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 140, the wireless communication module 150, the modulation and demodulation processor, and the baseband processor, etc.
[0142] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0143] The mobile communication module 140 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 140 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 140 can receive electromagnetic waves by the antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 140 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves through the antenna 1 and radiate it out. In some embodiments, at least some functional modules of the mobile communication module 140 can be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 140 and at least some modules of the processor 110 can be provided in the same device.
[0144] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor displays images or videos through the display screen 170. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be disposed in the same device as the mobile communication module 140 or other functional modules.
[0145] The wireless communication module 150 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 150 may be one or more devices integrating at least one communication processing module. The wireless communication module 150 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 150 may also receive the signal to be transmitted from the processor 110, perform frequency modulation on it, amplify it, and convert it into electromagnetic waves through the antenna 2 and radiate it out.
[0146] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 140, and antenna 2 is coupled to wireless communication module 150, such that electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).
[0147] Electronic device 100 implements a display function through a GPU, display screen 170, and an application processor, etc. The GPU is a microprocessor for image processing, and is connected to display screen 170 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0148] The display screen 170 is used to display images, videos, etc. The display screen 170 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 170, where N is a positive integer greater than 1.
[0149] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0150] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0151] The NPU is a neural-network (NN) computing processor. By learning from the biological neural network structure, such as learning from the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.
[0152] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, save files such as videos in the external memory card.
[0153] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121, and / or the instructions stored in the memory provided in the processor.
[0154] The pressure sensor 160A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 160A can be disposed on the display screen 170. There are many types of pressure sensors 160A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates having conductive materials. When a force acts on the pressure sensor 160A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 170, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 160A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 160A. In some embodiments, touch operations with the same touch position but different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0155] The gyroscope sensor 160B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 160B. Touch sensor 160C, also known as a "touch device". The touch sensor 160C can be arranged on the display screen 170, and the touch sensor 160C and the display screen 170 form a touch screen, also known as a "touch screen". The touch sensor 160C is used to detect touch operations acting on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 170. In other embodiments, the touch sensor 160C can also be arranged on the surface of the electronic device 100, which is different from the position of the display screen 170.
[0156] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. For example, the software system of the electronic device 100 may adopt an Android operating system (OS), a Harmony OS, or an IOS with a layered architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to exemplify the software structure of the electronic device 100.
[0157] Figure 2 It is a software structure block diagram of the electronic device 100 according to an embodiment of the present application.
[0158] The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system library, and the kernel layer.
[0159] The application layer can include a series of application packages.
[0160] like Figure 2 As shown, the application package may include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message and other applications.
[0161] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0162] like Figure 2As shown, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.
[0163] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0164] The content provider is used to store and obtain data, and make this data accessible to application programs. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone books, etc.
[0165] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build application programs. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures.
[0166] The telephone manager is used to provide the communication function of the electronic device 100. For example, the management of call states (including connection, disconnection, etc.).
[0167] The resource manager provides various resources for application programs, such as localized strings, icons, pictures, layout files, video files, etc.
[0168] The notification manager enables application programs to display notification information in the status bar, can be used to convey notification-type messages, can disappear automatically after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or a scroll bar text, such as the notification of a background-running application program, and can also be a notification that appears on the screen in the form of a dialogue window. For example, prompt text information in the status bar, emit a prompt sound, the electronic device vibrates, the indicator light flashes, etc.
[0169] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system.
[0170] The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.
[0171] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0172] The system library may include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc.
[0173] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.
[0174] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0175] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0176] The 2D graphics engine is a drawing engine for 2D drawing.
[0177] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0178] Next, the video note generation method provided by the embodiments of the present application will be described in detail in combination with the accompanying drawings and specific application scenarios.
[0179] Please refer to Figure 3 , Figure 3 , which shows a schematic flowchart of a video note generation method provided by the embodiments of the present application. This method can be applied to an electronic device. As Figure 3 shown, this method may include:
[0180] S301. The electronic device displays the first frame of the first video.
[0181] S302. When detecting the first operation, the electronic device enters the note mode.
[0182] S303. When detecting the second operation on the first video in the note mode, determine the second frame, and the second frame is a frame in the first video.
[0183] S304. The electronic device determines the target text content according to the second operation and the text content in the second frame, and generates a video note according to the target text content.
[0184] In an embodiment of the present application, when the electronic device displays the first frame of the first video, if the user wants to generate a video note, the user can perform a first operation. When the electronic device detects the first operation, it can enter the note mode. In the note mode, when the user performs a second operation on the first video, the electronic device can determine the second frame, and can determine the target text content according to the second operation and the text content in the second frame, so as to generate a video note according to the target text content. Wherein, the video note may include the target text content and / or content related to the target text content, and the target text content and / or content related to the target text content in the video note may be highlighted, so that when the user views the video note, the key content that the user is concerned about can be quickly and accurately viewed, which is convenient for the user to view the video note, meets the actual needs of the user, and improves the user experience.
[0185] It can be understood that the first video can be any video. For example, the first video can be a live video without subtitles, such as a live online class or a live sports event without subtitles. For example, the first video can be a non-live video with subtitles, such as a TV or a movie with subtitles.
[0186] Optionally, the first frame and the second frame can be any frame in the first video. Wherein, the first frame and the second frame can be the same frame in the first video, that is, the first frame and the second frame can be the same frame in the first video. Or, the first frame and the second frame can be different frames in the first video, that is, the first frame and the second frame can be different frames in the first video. The embodiments of the present application do not impose any restrictions on this.
[0187] Optionally, the electronic device can play the first video in full screen mode, that is, the first frame and the second frame can be full screen frames. Or the electronic device can play the first video in non-full screen mode, for example, it can play the first video in split screen mode, or it can play the first video in a small window (such as a pop-up window) mode, etc., that is, the first frame and the second frame can be non-full screen frames. It should be understood that the embodiments of the present application do not impose any restrictions on the playback form of the first video.
[0188] Exemplarily, the first operation can be an operation to trigger the electronic device to enter the note mode (which can also be called the video co-viewing mode). That is, when the user is interested in the content in the first video and wants to generate a video note corresponding to the first video for convenient subsequent viewing, the user can perform the first operation when the electronic device displays any frame (i.e., the first frame) of the first video, so that the electronic device enters the note mode to generate a video note corresponding to the first video.
[0189] It should be understood that the first operation can be specifically determined according to the actual scenario, and the embodiments of the present application do not impose any restrictions on this.
[0190] Optionally, the first operation may be an operation of turning on the artificial intelligence (AI) subtitles. That is, when the electronic device displays the first frame of the first video, if the user wants to generate a video note corresponding to the first video, the user can turn on the AI subtitles. When the electronic device detects that the AI subtitles are turned on, it can determine that the user wants to generate a video note corresponding to the first video. At this time, the electronic device can enter the note mode.
[0191] It should be noted that the embodiments of the present application do not impose any restrictions on the specific manner in which the user turns on the AI subtitles, and it can be specifically determined according to the actual scenario.
[0192] Optionally, the first operation may be an operation of turning on the excerpt function. That is, when the electronic device displays the first frame of the first video, if the user wants to generate a video note corresponding to the first video, the user can turn on the excerpt function. When the electronic device detects that the excerpt function is turned on, it can determine that the user wants to generate a video note corresponding to the first video. At this time, the electronic device can enter the note mode.
[0193] It should be understood that the embodiments of the present application do not impose any restrictions on the specific manner in which the user turns on the excerpt function, and it can be determined according to the actual scenario. For example, the user can swipe inwards or double-tap the stylus in the upper right corner of the electronic device to open a menu including the excerpt function, and can click on the menu item corresponding to the excerpt function in the menu to turn on the excerpt function. For example, the user can directly swipe inwards or double-tap the stylus in the upper right corner of the electronic device to turn on the excerpt function.
[0194] The following will detail the process of determining the second screen when the electronic device detects a second operation on the first video, and determining the target text content based on the second operation and the second screen. The process of determining the target text content based on the second operation and the second screen will be described in detail.
[0195] In the embodiments of the present application, the second operation may be an operation performed on the frame of the first video. Among them, the second operation can be specifically determined according to the actual scenario.
[0196] Exemplarily, the second operation may be a screenshot operation, a content marking operation, a content extraction operation, or a content input operation, etc.
[0197] Optionally, the screenshot operation may be a full-screen screenshot operation or a partial screenshot operation. The full-screen screenshot operation refers to an operation of taking a screenshot of the entire area of the display interface, that is, the screenshot image obtained by the full-screen screenshot operation may include the image corresponding to the entire display interface. The partial screenshot operation refers to an operation of taking a screenshot of a partial area of the display interface, that is, the screenshot image obtained by the partial screenshot operation may include the image corresponding to a certain partial area of the display interface.
[0198] For example, when a split-screen display of an electronic device shows the screen of a first video and an application interface corresponding to a certain application (such as Application A), the partial screenshot operation may be an operation of taking a screenshot of the screen of the first video (such as Screen A). At this time, the screenshot image obtained by the partial screenshot operation may include the image corresponding to Screen A. For example, when the electronic device full-screen displays the screen A of the first video, the partial screenshot operation may be an operation of taking a screenshot of a certain partial area in Screen A. At this time, the screenshot image obtained by the partial screenshot operation may include the image corresponding to this partial area in Screen A.
[0199] It can be understood that the full-screen screenshot operation or the partial screenshot operation can be specifically determined according to the actual scenario, and the embodiments of the present application do not impose any restrictions on this. For example, the full-screen screenshot operation may be an operation of double-tapping the display interface with a knuckle, that is, the user can double-tap the display interface with a knuckle to take a full-screen screenshot. For example, the partial screenshot operation may be an operation of tapping the display interface with a knuckle and drawing a closed pattern in the display interface, that is, the user can tap the display interface with a knuckle and draw a closed pattern in the display interface to take a partial screenshot of the area included in the drawn closed pattern. For example, when the extraction function is started, the screenshot operation (such as the full-screen screenshot operation or the partial screenshot operation) may be an operation of drawing a closed pattern in the display interface with a finger or a stylus.
[0200] Optionally, the content marking operation may be an operation of marking the text content in the screen.
[0201] For example, the content marking operation may be an operation of adding a color to the text content in the screen of the first video (such as Screen A), or an operation of modifying the font color of the text content in Screen A. For example, the content marking operation may be an operation of bolding or enlarging the text content in Screen A. For example, the content marking operation may be an operation of drawing a line at the bottom of the text content in Screen A. For example, the content marking operation may be an operation of drawing a border for the text content in Screen A, and so on.
[0202] In some embodiments, the text content in the screen may only include the subtitles corresponding to the screen.
[0203] In other embodiments, the text content in the screen may include the subtitles corresponding to the screen and the text content carried by the screen itself (hereinafter may be referred to as screen text), etc.
[0204] It is understandable that the electronic device can perform text recognition on the screen to determine the screen text in the screen. Among them, the embodiment of the present application does not impose any restrictions on the specific manner in which the electronic device performs text recognition on the screen, and it can be specifically determined according to the actual scenario. For example, the electronic device can perform text recognition on the screen through optical character recognition (OCR) technology to determine the screen text in the screen.
[0205] Optionally, the subtitles corresponding to the picture can be subtitles that come with the picture itself, or subtitles added to the picture by AI subtitles.
[0206] For example, after the electronic device enters the note-taking mode based on the first operation of turning on AI subtitles, the electronic device can obtain the audio corresponding to the first video, convert the obtained audio into text content to obtain subtitles corresponding to the first video, and display the AI subtitles corresponding to the picture when displaying the picture of the first video in the display interface.
[0207] It should be noted that the embodiment of the present application does not impose any restrictions on the manner in which the electronic device obtains the audio corresponding to the first video, and can be specifically determined according to the actual scenario. For example, when the electronic device plays the audio corresponding to the first video, the electronic device can record the played audio to obtain the audio corresponding to the first video. For example, the electronic device can obtain the source file corresponding to the first video, and can obtain the audio corresponding to the first video from the source file corresponding to the first video.
[0208] Optionally, the content extraction operation may refer to an operation of extracting text content in a picture, wherein extracting text content in a picture may refer to copying text content in a picture.
[0209] For example, the content extraction operation may be an operation of copying all or part of the subtitles corresponding to the screen (e.g., screen A) of the first video to a certain application. For example, the content extraction operation may be an operation of copying all or part of the screen text in screen A to a certain application, and so on.
[0210] Optionally, the content input operation may refer to an operation of inputting text content in the screen.
[0211] For example, the content input operation may be an operation of adding annotations to the screen text in the screen of the first video (eg, screen A). For example, the content input operation may be an operation of adding annotations to the subtitles corresponding to screen A, and so on.
[0212] Exemplarily, the second screen may be a screen displayed in a display interface of the electronic device when the second operation is detected.
[0213] Optionally, when the second operation is detected, the electronic device may determine the time T when the second operation is detected, and may determine the second screen according to the time T. That is, it is determined that at the time T, the screen displayed on the display interface in the screen of the first video, and this screen may be the second screen.
[0214] It should be noted that the embodiments of the present application do not specifically limit the manner in which the electronic device determines the screen displayed on the display interface in the screen of the first video at the time T, and may be specifically determined according to the actual scenario. For example, the electronic device may determine the time when each screen of the first video is displayed on the display interface, and may find the time (such as time T1) that matches the time T from the times when each screen of the first video is displayed on the display interface, so as to determine the screen corresponding to the time T1 as the second screen.
[0215] In one example, when the second operation is a screenshot operation, the electronic device may directly determine the target text content according to the text content in the second screen.
[0216] Optionally, the target text content may include all or part of the text content in the second screen.
[0217] For example, the target text content may include all the subtitles corresponding to the second screen and all the on-screen text in the second screen. For example, the target text content may include all or part of the subtitles corresponding to the second screen. For example, the target text content may include all or part of the on-screen text in the second screen.
[0218] In another example, when the second operation is a screenshot operation, the electronic device may obtain a screenshot image according to the screenshot operation, and determine the target text content according to the text content included in the screenshot image.
[0219] It should be understood that the screenshot image may be an image including the second screen, or may be an image including a partial area in the second screen. Among them, the text content included in the screenshot image may be all the text content in the second screen, or may be part of the text content in the second screen.
[0220] For example, the text content included in the screenshot image may include all or part of the subtitles corresponding to the second screen.
[0221] For example, the text content included in the screenshot image may include all or part of the on-screen text in the second screen.
[0222] Optionally, the target text content may include all or part of the text content included in the screenshot image.
[0223] For example, when the text content included in the screenshot image includes all the subtitles corresponding to the second frame, the target text content may include all or part of all the subtitles included in the screenshot image, that is, the target text content may include all the subtitles corresponding to the second frame or some of the subtitles.
[0224] For example, when the text content included in the screenshot image includes some of the subtitles corresponding to the second frame, the target text content may include all or part of some of the subtitles included in the screenshot image, that is, the target text content may include some of the subtitles corresponding to the second frame.
[0225] For example, when the text content included in the screenshot image includes all the frame text in the second frame, the target text content may include all or part of all the frame text included in the screenshot image, that is, the target text content may include all the frame text in the second frame or some of the frame text.
[0226] For example, when the text content included in the screenshot image includes some of the frame text in the second frame, the target text content may include all or part of some of the frame text included in the screenshot image, that is, the target text content may include some of the frame text in the second frame.
[0227] In another example, when the second operation is a content marking operation, the electronic device may determine the target text content according to the text content marked in the second frame by the content marking operation. Optionally, the target text content may include the text content marked in the second frame by the content marking operation.
[0228] In another example, when the second operation is a content extraction operation, the electronic device may determine the target text content according to the text content extracted in the second frame by the content extraction operation. Optionally, the target text content may include the text content extracted in the second frame by the content extraction operation.
[0229] In another example, when the second operation is a content input operation, the electronic device may determine the target text content according to the text content in the second frame (which may be referred to as text content A for example) and the text content input in the second frame by the content input operation (which may be referred to as text content B for example). Optionally, the target text content may include all or part of text content A and may include all or part of text content B.
[0230] That is to say, when the electronic device displays the second screen of the first video, the user can perform operations such as taking a screenshot, marking, extracting, or inputting on the second screen. When the electronic device detects operations such as taking a screenshot, marking, extracting, or inputting on the second screen, it can determine the target text content, that is, determine the content that the user focuses on, so that the generation of video notes can be performed according to the target text content. In the generated video notes, the content that the user focuses on (such as the target text content) can be highlighted, so that when the user views the video notes, the content that is focused on can be quickly viewed, improving the user experience.
[0231] Please refer to Figure 4 , Figure 4 which shows the application scenario provided by the embodiment of the present application Figure 1 . This application scenario is exemplarily described by taking the text content in the screen as only including the subtitles corresponding to the screen as an example.
[0232] As Figure 4 shown in (a) of Figure 4 , when the electronic device displays screen A of the first video in full screen, assuming that the subtitle corresponding to screen A is "Marcel Breuer who entered school in Weimar in 1920", if a local screenshot operation on screen A is detected, for example, as
[0233] shown in (a) of
[0234] , if it is detected that after the knuckle taps the display interface, an operation of drawing a closed pattern 400 in the display interface, the electronic device can determine the time when the local screenshot operation is detected, and can determine the second screen according to the time when the local screenshot operation is detected. For example, it can be determined that the second screen is screen A. Subsequently, the electronic device can obtain the subtitle corresponding to the second screen, that is, it can obtain "Marcel Breuer who entered school in Weimar in 1920", and can determine the target text content according to the subtitle corresponding to the second screen. Figure 4 For example, the electronic device can determine the entire subtitle corresponding to the second screen, that is, "Marcel Breuer who entered school in Weimar in 1920" as the target text content. For example, the electronic device can determine a part of the subtitle corresponding to the second screen, such as "Marcel Breuer", as the target text content.
[0235] As Figure 4 shown in (b) below, when the electronic device displays the picture A of the first video in full screen, if a marking operation on the text content in the picture A is detected, for example, detecting an operation of drawing a straight line at the bottom of "Marcel Breuer" in the subtitle "He is Marcel Breuer who entered school in Weimar in 1920" corresponding to the picture A, the electronic device can determine that the second picture is the picture A, and can determine that the target text content includes the content "Marcel Breuer" marked by the content marking operation.
[0236] As Figure 4 shown in (c) below, when the electronic device displays the picture A of the first video and the application interface corresponding to a certain application (such as a browser application) in split screen, if an extraction operation on the text content in the picture A is detected, for example, detecting an operation of copying "Marcel Breuer" in the subtitle "He is Marcel Breuer who entered school in Weimar in 1920" corresponding to the picture A to the search box 410 of the browser application, the electronic device can determine that the second picture is the picture A, and can determine that the target text content includes the content "Marcel Breuer" extracted by the content extraction operation.
[0237] As Figure 4 shown in (d) below, when the electronic device displays the picture A of the first video in full screen, if a content input operation on the picture A is detected, for example, detecting an operation of inputting "designer" next to "Marcel Breuer" in the subtitle "He is Marcel Breuer who entered school in Weimar in 1920" corresponding to the picture A, the electronic device can determine that the second picture is the picture A, and can determine that the target text content includes the subtitle "He is Marcel Breuer who entered school in Weimar in 1920" corresponding to the second picture and the content "designer" input by the content input operation in the second picture.
[0238] In a possible implementation, during the playback of the first video, the speed of screen switching is generally relatively fast. When the user's hand speed is slow, it may cause the screen displayed on the display interface (hereinafter referred to as screen B) when the second operation is detected to not be the screen that the user actually wants to perform the second operation on. Therefore, when the second operation is detected, the electronic device can determine the time T when the second operation is detected, and based on the time T, determine the text content in the N frames of screens (hereinafter referred to as screen C) before screen B. That is, these N frames of screen C can be the screens in the first video before screen B. That is to say, during the playback of the first video, these N frames of screen C can be displayed on the display interface before screen B. Subsequently, the electronic device can simultaneously display the text content in these N frames of screen C and the text content in screen B on the display interface, so that the user can select the text content for the second operation from the text content in these N + 1 frames of screens. The electronic device can determine the target text content based on the text content selected by the user to accurately determine the content that the user focuses on, reduce the possibility of incorrect determination of the target text content due to slow hand speed, and improve the user experience.
[0239] In some embodiments, when the second operation is a screenshot operation, the electronic device can determine the time T when the screenshot operation is detected, and can, based on the time T, determine the text content in the N frames of screen C before screen B. Subsequently, the electronic device can simultaneously display the text content in these N frames of screen C and the text content in screen B on the display interface, so that the user can select the text content from the text content in these N + 1 frames of screens.
[0240] In other embodiments, when the second operation is a content marking operation, a content extraction operation, or a content input operation, the electronic device can determine the time T when the second operation is detected, and can, based on the time T, determine the text content in the N frames of screen C before screen B. Subsequently, the electronic device can simultaneously display the text content in these N frames of screen C and the text content in screen B on the display interface, so that the user can re-perform content marking, content extraction, or content input from the text content in these N + 1 frames of screens.
[0241] Optionally, the text content in these N frames of screen C is different from each other, and the text content in these N frames of screen C is respectively different from the text content in screen B. For example, when the text content in the screen only includes the subtitles corresponding to the screen, the subtitles corresponding to these N frames of screen C are different from each other, and the subtitles corresponding to these N frames of screen C are respectively different from the subtitles corresponding to screen B.
[0242] Optionally, N can be greater than or equal to 1. The specific value of N can be determined according to the actual scenario, and the embodiments of the present application do not impose any restrictions on this.
[0243] It should be understood that the manner in which the electronic device simultaneously displays the text content in these N frames of picture C and the text content in picture B in the display interface can be specifically determined according to the actual scenario, and the embodiments of the present application do not impose any restrictions thereon.
[0244] In one example, when the electronic device activates the AI subtitle, the electronic device can simultaneously display the text content in these N frames of picture C and the text content in picture B at the position where the subtitle corresponding to picture B is located in the display interface.
[0245] For example, when the text content in the picture only includes the subtitle corresponding to the picture and the electronic device activates the AI subtitle, when detecting a screenshot operation on picture B, the electronic device can determine the time T corresponding to the screenshot operation and determine N frames of picture C before time T. Subsequently, the electronic device can obtain the subtitles corresponding to these N frames of picture C and simultaneously display the subtitles corresponding to these N frames of picture C and the subtitle corresponding to picture B at the position where the subtitle corresponding to picture B is located.
[0246] In another example, when the electronic device does not activate the AI subtitle, the electronic device can simultaneously display the text content in these N frames of picture C and the text content in picture B in a pop-up window manner.
[0247] Exemplarily, when simultaneously displaying the text content in these N frames of picture C and the text content in picture B in the display interface, the electronic device can perform the simultaneous display of the text content according to the time of each picture in the first video.
[0248] Optionally, the electronic device can simultaneously display the text content in these N frames of picture C and the text content in picture B in the display interface in the order from earliest to latest time.
[0249] For example, when N is 3, assuming that the time of the first frame of picture C in the first video is 18 minutes and 35 seconds, the time of the second frame of picture C in the first video is 18 minutes and 50 seconds, the time of the third frame of picture C in the first video is 19 minutes, and the time of picture B in the first video is 19 minutes and 5 seconds, the electronic device can perform the simultaneous display of the subtitles in the display interface in the order of the subtitle corresponding to the first frame of picture C - the subtitle corresponding to the second frame of picture C - the subtitle corresponding to the third frame of picture C - the subtitle corresponding to picture B.
[0250] Please refer to Figure 5 , Figure 5 which shows the schematic application scenario provided by the embodiments of the present application Figure 2 . This application scenario is exemplarily described by taking the text content in the picture only including the subtitle corresponding to the picture and N being 3 as an example.
[0251] Such as Figure 5As shown in (a), when the electronic device displays the screen B of the first video in full screen, assuming that the subtitle corresponding to the screen B is "Marcel Breuer entered school in Weimar in 1920", if a full-screen screenshot operation on the screen B is detected, for example, a double-tap operation of the knuckle on the display interface is detected, the electronic device can determine the time when the full-screen screenshot operation is detected (for example, time T0), and can obtain the subtitle corresponding to the screen B, that is, "Marcel Breuer entered school in Weimar in 1920". Subsequently, the electronic device can obtain the subtitles corresponding to the three frames of the screen C before the screen B according to the time T0.
[0252] Assume that the three frames of the screen C before the screen B are the screen C1, the screen C2, and the screen C3 respectively. The subtitle corresponding to the screen C1 is "This uses steel pipes", the subtitle corresponding to the screen C2 is "Look, these steel pipes are a kind of metal material", and the subtitle corresponding to the screen C3 is "The designer is the first graduate student of Bauhaus", and the time of the screen C1 in the first video is earlier than the time of the screen C2 in the first video, and the time of the screen C2 in the first video is earlier than the time of the screen C3 in the first video.
[0253] As Figure 5 As shown in (b), when the electronic device activates the AI subtitle, after obtaining the subtitles corresponding to the screen C1, the screen C2, and the screen C3, the electronic device can simultaneously display the subtitles corresponding to these four frames at the display position of the subtitle corresponding to the screen B in the display interface, that is, it can display "This uses steel pipes", "Look, these steel pipes are a kind of metal material", "The designer is the first graduate student of Bauhaus", and "Marcel Breuer entered school in Weimar in 1920" at the position of the subtitle corresponding to the screen B in the display interface.
[0254] As Figure 5 As shown in (c), when the electronic device does not activate the AI subtitle, after obtaining the subtitles corresponding to the screen C1, the screen C2, and the screen C3, the electronic device can display a pop-up window 500 in the display interface. Among them, the pop-up window 500 can display "This uses steel pipes", "Look, these steel pipes are a kind of metal material", "The designer is the first graduate student of Bauhaus", and "Marcel Breuer entered school in Weimar in 1920".
[0255] As Figure 5 As shown in (b) and Figure 5As shown in (c) thereof, when a full-screen screenshot operation is detected, the electronic device may default the text content selected by the user to the caption corresponding to Screen B, that is, the caption corresponding to Screen B, "Marcel Breuer entered school in Weimar in 1920", may be default marked as the text content selected by the user. Among them, after the captions corresponding to these four frames are simultaneously displayed on the display interface, the user can select the text content he is concerned about from the captions corresponding to these four frames according to actual needs, that is, the user can adjust the text content default selected by the electronic device. The electronic device may determine the target text content according to the text content selected by the user. For example, the text content selected by the user may be determined as the target text content.
[0256] For example, when the captions corresponding to the four frames, "This uses steel pipes", "Look, these steel pipes are a kind of metal material", "The designer is the first graduate of the Bauhaus", and "Marcel Breuer entered school in Weimar in 1920", are simultaneously displayed at the position of the caption corresponding to Screen B on the display interface, if the user adjusts the selected text content to Figure 5 As shown in (d) thereof, the electronic device may determine that the text content selected by the user is "The designer is the first graduate of the Bauhaus". At this time, the electronic device may determine "The designer is the first graduate of the Bauhaus" as the target text content.
[0257] The following will detail the process of the electronic device generating a video note based on the target text content.
[0258] In the embodiment of the present application, after the electronic device enters the note mode, the electronic device may obtain the text content in the target screen and detect a second operation on the first video. When the second operation on the first video is detected, the electronic device may determine the second screen and may determine the target text content according to the second operation and the text content in the second screen. After determining the target text content, the electronic device may generate a video note according to the target text content and the text content in the target screen. It should be understood that the target screen may include one or more frames in the first video.
[0259] Optionally, the target screen may include one or more frames displayed on the display interface after the electronic device enters the note mode. That is to say, the electronic device may generate a video note according to the text content in one or more frames displayed after the electronic device enters the note mode.
[0260] For example, the target screen may include one or more frames displayed on the display interface during the period from when the electronic device enters the note mode to when it exits the note mode.
[0261] Optionally, the target screen may include one or more frames of the display interface after the electronic device enters the note mode, and one or more frames of the display interface within a preset time before the electronic device enters the note mode. That is, the electronic device may generate a video note based on the text content in one or more frames of the display interface after the electronic device enters the note mode and the text content in one or more frames of the display interface within a preset time before the electronic device enters the note mode.
[0262] It should be understood that when the user is watching the first video, the user may find that a certain content (such as content B) before content A is also very helpful when watching a certain content (such as content A) of the first video. At this time, if the user wants to generate a video note including content B, the video played by the electronic device needs to be returned to content B or the content before content B, so that when the electronic device plays content B, the electronic device enters the note mode through the first operation to generate a video note including content B, and the operation is relatively cumbersome and the user experience is poor. Therefore, to avoid the user having to perform a content return operation to generate the desired video note to improve the user experience, when the electronic device plays the first video and enters the note mode based on the first operation, the electronic device can obtain the text content (such as text content A) in the screen displayed within the preset time before the electronic device enters the note mode and the text content (such as text content A) in the screen displayed after entering the note mode, and can generate a video note based on text content A and text content B, so that the user does not need to perform a content return operation and can also generate a video note including the previously played screen.
[0263] It should be noted that the preset time can be specifically determined according to the actual scenario, and the embodiments of the present application do not impose any restrictions on this. Among them, the preset time can be set by default by the electronic device or can be set by the user customarily. Optionally, the electronic device can default to set the preset time to any value such as 5 minutes, 10 minutes, 20 minutes, or 30 minutes according to the actual scenario, or can default to set the preset time according to the start playing time and the current time of the electronic device playing the first video. For example, the preset time can be the time difference between the current time and the start playing time. For example, when the start playing time of the electronic device playing the first video is 19:30:00 and the current time is 19:44:00, the electronic device can default to set the preset time to 14 minutes. Optionally, the user can customarily set the preset time to any value such as 3 minutes, 5 minutes, or 10 minutes according to actual needs.
[0264] For example, when the electronic device detects a first operation, the electronic device may output a prompt message. Alternatively, after the electronic device enters the note mode based on the first operation, if an operation to close the note mode is detected, the electronic device may output a prompt message. The prompt message may be used to prompt the user whether to generate a video note corresponding to the video played before the electronic device enters the note mode. The prompt message may also be used to prompt the user to perform a custom setting or selection for a preset time when it is necessary to generate a video note corresponding to the video played before the electronic device enters the note mode.
[0265] In an embodiment of the present application, after the electronic device enters the note mode, the electronic device may obtain the text content in the target screen, and may generate an initial note according to the text content in the target screen. When a second operation on the second screen is detected, the electronic device may determine the target text content, and may generate a final video note according to the target text content and the initial note.
[0266] The following will be described separately: First, generating an initial note according to the text content in the target screen; Second, generating a video note according to the target text content and the initial note.
[0267] First, generating an initial note according to the text content in the target screen.
[0268] In one example, after obtaining the text content in the target screen, the electronic device may divide the text content in the target screen into paragraphs according to content relevance to obtain an initial note. That is to say, the initial note may be a note obtained by organizing the text content in the target screen.
[0269] It can be understood that when dividing the text content in the target screen into paragraphs according to content relevance, punctuation marks may be added to make the generated note convenient for the user to view.
[0270] For example, when the text content in the screen only includes the subtitles corresponding to the screen, after obtaining the subtitles corresponding to the target screen, the electronic device may divide the subtitles corresponding to the target screen into paragraphs, and when dividing the paragraphs, punctuation marks may be added to generate an initial note.
[0271] In another example, after obtaining the text content in the target screen, the electronic device may determine one or more of the full text summary, chapter summary, keywords, etc. according to the text content in the target screen, and may generate an initial note according to one or more of the full text summary, chapter summary, keywords, etc. That is to say, the initial note may include the content obtained by summarizing the text content in the target screen.
[0272] Exemplarily, after obtaining the text content in the target screen, the electronic device may determine the full-text summary based on the text content in the target screen, and may generate an initial note based on the full-text summary and the text content in the target screen. That is, the initial note may include the full-text summary and the text content in the target screen.
[0273] For example, the electronic device may determine the full-text summary based on the text content in the target screen, and may divide the text content in the target screen into paragraphs according to content relevance to obtain text content including one or more paragraphs (hereinafter referred to as text content C). Subsequently, the electronic device may add the full-text summary to the front of the text content C to obtain the initial note. That is to say, in the initial note, the full-text summary may be located at the beginning, and the text content in the target screen may be located behind the full-text summary to facilitate the user to quickly understand the full content of the note.
[0274] Exemplarily, after obtaining the text content in the target screen, the electronic device may determine the chapter summary based on the text content in the target screen, and may generate an initial note based on the chapter summary and the text content in the target screen. That is, the initial note may include the chapter summary and the text content in the target screen.
[0275] Optionally, the chapter summary may include a chapter title and / or a chapter abstract. It should be understood that the chapter abstract may be a summary of the text content corresponding to the chapter.
[0276] For example, the electronic device may determine one or more chapter summaries based on the text content in the target screen. Assume that the chapter summary includes a chapter title and a chapter abstract. In addition, the electronic device may also divide the text content in the target screen into paragraphs according to content relevance to obtain text content C including one or more paragraphs. Subsequently, the electronic device may add each chapter summary to the front of the text content C to obtain the initial note. That is to say, in the initial note, the chapter summary may be located at the beginning, and the text content in the target screen may be located behind the chapter summary to facilitate the user to quickly understand the content of each chapter in the note.
[0277] For example, the electronic device may determine one or more chapter summaries based on the text content in the target screen. Assume that the chapter summary includes a chapter title. In addition, the electronic device may also divide the text content in the target screen into paragraphs according to content relevance to obtain text content C including one or more paragraphs. For each chapter summary (i.e., the chapter title), the electronic device may determine the text content in the text content C corresponding to the chapter summary (hereinafter referred to as text content D), and may add the chapter summary (i.e., the chapter title) to the front of the text content D to obtain the initial note to facilitate the user to consult the initial note.
[0278] Optionally, when generating an initial note based on the chapter summary and the text content in the target screen, after the electronic device enters the note mode, if a second operation on the first video is detected, the electronic device can determine the chapter summary according to the text content in the target screen and the target text content. The chapter summary corresponding to the target text content may include the target text content or an abstract corresponding to the target text content, so that when the user views the note, the key content of interest can be quickly viewed, improving the user experience. That is to say, when generating an initial note based on the chapter summary and the text content in the target screen, after the electronic device enters the note mode, if a second operation on the first video is not detected, the electronic device can directly determine the chapter summary according to the text content in the target screen. If a second operation on the first video is detected, the electronic device can determine the chapter summary according to the target text content and the text content in the target screen, so as to display the key content of interest to the user in the chapter summary for the user to view conveniently.
[0279] Exemplarily, after obtaining the text content in the target screen, the electronic device can determine the chapter summary (such as including a chapter title and a chapter abstract) according to the text content in the target screen, and can directly generate an initial note based on the chapter summary. That is, the initial note may include the chapter summary.
[0280] Similarly, when directly generating an initial note based on the chapter summary, after the electronic device enters the note mode, if a second operation on the first video is detected, the electronic device can determine the chapter summary according to the text content in the target screen and the target text content. The chapter summary corresponding to the target text content may include the target text content or an abstract corresponding to the target text content, so that when the user views the note, the key content of interest can be quickly viewed, improving the user experience. It should be understood that the abstract corresponding to the target text content refers to the content obtained by summarizing the target text content.
[0281] That is to say, in the scenario of directly generating an initial note based on the chapter summary, after the electronic device enters the note mode, if a second operation on the first video is not detected, the electronic device can directly determine the chapter summary according to the text content in the target screen. If a second operation on the first video is detected, the electronic device can determine the chapter summary according to the target text content and the text content in the target screen, so as to display the key content of interest to the user in the chapter summary for the user to view conveniently.
[0282] Optionally, when generating the initial note according to the chapter summary, for each chapter summary, the electronic device may determine the text content D in the target screen that corresponds to the chapter summary, and may associate and save the text content D with the corresponding chapter summary. Among them, when the user views the note, the user may view the text content D corresponding to a certain chapter summary through a specified operation, that is, may view the original text content associated with the chapter summary. It should be understood that the original text content may refer to the text content in the target screen.
[0283] Optionally, when associating and saving the text content D with the corresponding chapter summary, the electronic device may determine the text content (hereinafter referred to as text content E) in the text content D that corresponds to each content in the chapter summary, and may associate and save the text content E with the corresponding content in the chapter summary. Among them, when the user views the note, the user may view the text content E corresponding to a certain content in a certain chapter summary through a specified operation, that is, may view the original text content associated with the content in the chapter summary.
[0284] It can be understood that the specified operation may be specifically determined according to the actual scenario, and the embodiments of the present application do not make any restrictions on this. For example, the specified operation may be determined to be a long-press operation according to the actual scenario. That is to say, when the user views the note, if the user long-presses a certain content in a certain chapter summary, the electronic device may display the original text content corresponding to the content in the display interface, or may display the original text content corresponding to the chapter summary.
[0285] Exemplarily, after obtaining the text content in the target screen, the electronic device may determine one or more keywords according to the text content in the target screen, and may generate an initial note according to the keywords and the text content in the target screen. That is, the initial note may include one or more keywords and the text content in the target screen.
[0286] Optionally, when the initial note includes keywords, the user may search for relevant content through the keywords.
[0287] For example, the electronic device may determine keyword A, keyword B, and keyword C according to the text content in the target screen, and perform paragraph division on the text content in the target screen according to content relevance to obtain text content C including one or more paragraphs. Subsequently, the electronic device may add keyword A, keyword B, and keyword C to the front of the text content C to obtain the initial note. That is to say, in the initial note, keyword A, keyword B, and keyword C may be located at the beginning, and the text content in the target screen may be located behind keyword A, keyword B, and keyword C to facilitate content search according to the keywords.
[0288] Exemplarily, after obtaining the text content in the target screen, the electronic device can determine multiple ones of the full-text abstract, chapter summary, and keywords based on the text content in the target screen, and can generate an initial note based on multiple ones of the full-text abstract, chapter summary, and keywords. That is, the initial note can include multiple ones of the full-text abstract, chapter summary, and keywords.
[0289] Optionally, after obtaining the text content in the target screen, the electronic device can determine the full-text abstract and keywords based on the text content in the target screen, and can generate an initial note based on the full-text abstract, keywords, and the text content in the target screen. That is, the initial note can include the full-text abstract, keywords, and the text content in the target screen.
[0290] For example, the electronic device can determine the full-text abstract, and keywords A, B, and C based on the text content in the target screen, and can divide the text content in the target screen into paragraphs according to content relevance to obtain text content C including one or more paragraphs. Subsequently, the electronic device can add the full-text abstract and keywords A, B, and C to the front of text content C to obtain the initial note. Among them, the full-text abstract can be in front of keywords A, B, and C, or can be behind keywords A, B, and C.
[0291] Optionally, after obtaining the text content in the target screen, the electronic device can determine the full-text abstract and chapter summary based on the text content in the target screen, and can generate an initial note based on the full-text abstract, chapter summary, and the text content in the target screen. That is, the initial note can include the full-text abstract, chapter summary, and the text content in the target screen.
[0292] For example, the electronic device can determine the full-text abstract and chapter summary based on the text content in the target screen. Assume that the chapter summary includes chapter titles. In addition, the electronic device can also divide the text content in the target screen into paragraphs according to content relevance to obtain text content C including one or more paragraphs. Subsequently, the electronic device can add the full-text abstract to the front of text content C. For each chapter summary (i.e., chapter title), the electronic device can determine the text content D in text content C corresponding to the chapter title, and can add the chapter title to the front of text content D to obtain the initial note.
[0293] Optionally, after obtaining the text content in the target screen, the electronic device may determine the full text summary and chapter summaries (such as including chapter titles and chapter summaries) based on the text content in the target screen, and may directly generate an initial note based on the full text summary and chapter summaries. That is, the initial note may include the full text summary and chapter summaries. Among them, the full text summary may be located in front of the chapter summaries, facilitating the user to quickly understand the full text content of the note and the content of each chapter.
[0294] Optionally, when directly generating an initial note based on the full text summary and chapter summaries, when the second operation is not detected, the electronic device may directly determine the chapter summaries based on the text content in the target screen. When the second operation is detected, the electronic device may determine the chapter summaries based on the target text content and the text content in the target screen, so as to display the key content concerned by the user in the chapter summaries, facilitating the user to view.
[0295] It should be understood that the embodiments of the present application do not specifically limit the determination methods of the full text summary, chapter summaries, and keywords, and may be determined according to the actual scenario.
[0296] Please refer to Figure 6 , Figure 6 which shows the schematic diagram of the application scenario provided by the embodiments of the present application. Figure 3 This application scenario is exemplarily described by taking the text content in the screen including the subtitles corresponding to the screen and the second operation not being detected as an example.
[0297] Suppose that after the electronic device enters the note mode, the subtitles corresponding to the target screen obtained by the electronic device may include "We can't talk about modern design history without this chair", "But that's a bit overstated", "How can modern design history be separated from a chair", "If we don't talk about this chair, we can't talk about Bauhaus", "Can't talk about Bauhaus", "Can't start the design history", "So according to everyone's understanding, this chair", "Is the most important chair in the world of modern design", "If we take 10 pictures of this one, it should be the first", "This chair actually looks a bit strange", "It may not be very comfortable to sit on", "But it was the first one to be explored", "It changed the material of the furniture", "This uses steel pipes", "Look, these steel pipes are a kind of metal material", "The designer is the first graduate student of Bauhaus", "He entered the school in Weimar in 1920, Marcel Breuer", "Marcel Breuer graduated", "6 graduate students stayed at the school", "One of the most important", "He followed the founder of Bauhaus", "Walter Gropius to Harvard", "He taught at Harvard University", "His direct student is I.M. Pei", "So this person is a very important designer", "He rented a place next to Bauhaus", and "Rode a bicycle to school every day", and so on.
[0298] like Figure 6 As shown in (a), the electronic device can divide the text content in the target screen into paragraphs according to the content relevance, and add punctuation marks when dividing the paragraphs to obtain initial notes.
[0299] like Figure 6 As shown in (b) of FIG. 1 , the electronic device can determine the full text summary based on the text content in the target screen, and can add the full text summary to the beginning of the initial note. For example, the full text summary can be added to Figure 6 The full text summary can be "The video introduces three of the most classic chairs in the history of modern design - Wassily Chair, Barcelona Chair and Red and Blue Chair, and analyzes their design background and design concept."
[0300] like Figure 6 As shown in (c) in the figure, the electronic device can determine the full text summary and chapter summary based on the text content in the target screen, and can generate initial notes based on the full text summary and chapter summary. The chapter summary can include chapter titles and chapter summaries. For example, chapter titles can include Wassily Chair, Barcelona Chair, and Red and Blue Chair. The chapter summary corresponding to the Wassily Chair can include "The Wassily Chair was designed by Marcel Breuer, the first graduate student of the Bauhaus who entered the school in 1920. It is made of steel pipes. Inspired by the abstract painter Kandinsky, Marcel Breuer used the steel pipes of bicycle handlebars to design this chair. This chair can be assembled and has a simple structure. The Wassily Chair not only takes into account the support of the back, but also pays attention to the support of the waist, which is more ergonomic", and the chapter summary corresponding to the Barcelona Chair can include "The Barcelona Chair was designed by Mies van der Rohe for the 1929 Barcelona World Expo." The design of the Barcelona chair is a milestone in the history of architecture. This chair is large, luxurious, and made of metal and leather. There is a steel cylinder underneath it, which is different from Marcel Breuer's steel tube chair, showing the difference in design concepts". The chapter summary corresponding to the red and blue chair can include "The red and blue chair was designed by Rietveld. It is one of the most creative classics in the history of Western modern art design in the 20th century. It has distinctive Dutch style characteristics. It uses square and rectangular wooden strips and boards, which are combined in modular combinations. The red and blue colors are very bright and eye-catching, and have highly Cubist symbolic characteristics."
[0301] When the user wants to view a chapter summary or the original subtitles corresponding to a content in a chapter summary, the user can long press a content in the chapter summary, for example, Figure 6 As shown in (c) in the figure, when the user wants to view the original subtitles corresponding to the Wassily Chair, the user can long press Marcel Breuer in the Wassily Chair. Figure 6As shown in (d) in [the reference], when the electronic device detects a long - press operation on the Marcel Breuer in the Wassily chair, it can display the original caption corresponding to the Wassily chair on the display interface. For example, it can display "The designer is the first graduate student of Bauhaus. He entered the school in Weimar in 1920. Marcel Breuer is one of the 6 graduate students who stayed at the school after graduation and is the most important one. He followed the founder of Bauhaus, Walter Gropius, to Harvard. He taught at Harvard University, and his direct student is I.M. Pei. So this person is a very important designer. He rented a place next to Bauhaus and rode a bicycle to school every day."
[0302] Second, generate video notes based on the target text content and the initial notes.
[0303] In one example, the electronic device can highlight the target text content in the initial notes to obtain video notes. That is, it can highlight the key content that the user is concerned about, so that when the user views the video notes, they can quickly see the target text content and improve the user experience.
[0304] It can be understood that the embodiments of this application do not limit the way of highlighting, which can be specifically determined according to the actual scenario. For example, highlighting can be performed by color marking. For example, highlighting can be performed by enlarging the font. For example, highlighting can be performed by bolding the font, and so on.
[0305] In another example, after determining the target text content, the electronic device can obtain the context corresponding to the target text content according to the text content in the target screen, and highlight the target text content in the initial notes and the context corresponding to the target text content to obtain video notes. That is, it can determine the context corresponding to the target text content as the key content that the user is concerned about and perform highlighting to facilitate the user's viewing.
[0306] Optionally, when highlighting the target text content in the initial notes and the context corresponding to the target text content, the electronic device can use different methods to highlight the target text content and the context corresponding to the target text content to facilitate the user to accurately distinguish the key content they are concerned about.
[0307] For example, the electronic device can highlight the target text content with color A and highlight the context corresponding to the target text content with color B. Among them, color A and color B are different. For example, color A can be dark yellow and color B can be light yellow.
[0308] In another example, after obtaining the context corresponding to the target text content, the electronic device can determine the content view corresponding to the target text content according to the context corresponding to the target text content. Subsequently, the electronic device can determine the target content A according to the content view corresponding to the target text content, and can highlight the target content A in the initial note to obtain a video note. Among them, the content view corresponding to the target content A is the same as the content view corresponding to the target text content.
[0309] That is to say, the electronic device can, based on the second operation of the user on the first video, highlight all the content in the text content of the target screen that has the same content view as the content view corresponding to the target text content, so as to automatically mark the key content that the user focuses on and facilitate the user to view.
[0310] For example, when it is determined that the context corresponding to the target text content is the content introducing the life of Marcel Breuer, the electronic device can determine that the content view corresponding to the target text content is the life of Marcel Breuer. At this time, the electronic device can determine all the content in the text content of the target screen that introduces the life of Marcel Breuer, and automatically highlight all the content that introduces the life of Marcel Breuer to facilitate the user to view and improve the user experience.
[0311] Optionally, when highlighting the target text content and the target content A in the initial note, the electronic device can highlight the target text content and the target content A in different ways to facilitate the user to accurately distinguish the key content they focus on.
[0312] Optionally, the electronic device can highlight the context corresponding to the target text content and the target content A in the same way.
[0313] For example, the electronic device can highlight the target text content with color A, and can highlight the context corresponding to the target text content and the target content A with color B. Among them, color A and color B are different. For example, color A can be dark yellow and color B can be light yellow.
[0314] In another example, when the target text content includes the content input by the user (hereinafter referred to as target content B), the electronic device can also search for the content related to the target content B (hereinafter referred to as target content C) in the initial note, and can highlight the target content C in the initial note to obtain a video note.
[0315] For example, when the target text content includes "Marcel Breuer" and the user input is "What are the works", the electronic device can search for the works corresponding to Marcel Breuer in the initial note and can highlight the works corresponding to Marcel Breuer in the initial note to facilitate the user's viewing.
[0316] Exemplarily, when the initial note is a note generated according to the chapter summary, the electronic device can determine the chapter summary according to the target text content and the text content in the target screen. Among them, the chapter summary can include the target text content or the abstract corresponding to the target text content. At this time, the electronic device can highlight the target text content or the abstract corresponding to the target text content in the chapter summary to obtain a video note.
[0317] Please refer to Figure 7 , Figure 7 which shows the application scenario provided by the embodiment of the present application. Figure 4 . This application scenario is exemplarily described by taking the text content in the screen including the subtitles corresponding to the screen as an example. Among them, in this application scenario, the subtitles corresponding to the screen include Figure 6 the subtitles shown in
[0318] When the electronic device full-screen displays the second screen of the first video, assuming that the subtitles corresponding to the second screen are "He entered the school in Weimar in 1920, Marcel Breuer", if it detects a full-screen screenshot operation on the second screen, the electronic device can obtain the subtitles corresponding to the second screen, that is, it can obtain "He entered the school in Weimar in 1920, Marcel Breuer", and can determine the target text content according to the subtitles corresponding to the second screen. Assuming that the target text content is "He entered the school in Weimar in 1920, Marcel Breuer". In addition, the electronic device can also obtain the context corresponding to the target text content. For example, it can obtain "The designer is the first graduate student of Bauhaus", "He entered the school in Weimar in 1920, Marcel Breuer", "Marcel Breuer graduated", "6 graduate students stayed at school", "One of them is the most important", "He was followed by the founder of Bauhaus", "Walter Gropius went to Harvard", "He taught at Harvard University", "His direct student is I.M. Pei", "So this person is a very important designer", "He rented a place next to Bauhaus", and "Rode a bicycle to school every day", and so on.
[0319] Such as Figure 7As shown in (a), after the electronic device enters the note mode, the electronic device can obtain the text content in the target screen, can divide the text content in the target screen into paragraphs according to content relevance, and add punctuation marks during paragraph division to obtain an initial note, and can highlight the target text content in the initial note, and can also highlight the context corresponding to the target text content in the initial note to obtain a video note. For example, the target text content and the context corresponding to the target text content can be highlighted in different colors.
[0320] As Figure 7 shown in (b), the electronic device can determine the full-text summary and chapter summary according to the target text content and the text content in the target screen, and can generate an initial note according to the full-text summary and chapter summary. For example, the full-text summary can be "The video introduces three of the most classic chairs in the history of modern design - the Wassily Chair, the Barcelona Chair, and the Red and Blue Chair, and analyzes their design backgrounds and design concepts". Among them, the chapter summary can include a chapter title and a chapter summary.
[0321] For example, the chapter titles can include the Wassily Chair, the Barcelona Chair, and the Red and Blue Chair. Among them, since the target text content is the content in the chapter - Wassily Chair. Therefore, the electronic device can determine that the chapter summary corresponding to the Wassily Chair can include the target text content or the summary corresponding to the target text content, and can highlight the target text content or the summary corresponding to the target text content in the Wassily Chair to obtain a video note. For example, the electronic device can determine that the chapter summary corresponding to the Wassily Chair can include "The Wassily Chair was designed by Marcel Breuer, the first graduate of the Bauhaus and who entered the school in Weimar in 1920. This is made of steel pipes. Inspired by the abstract painter Wassily Kandinsky, Marcel Breuer used the steel pipes of a bicycle handlebar to design this chair. This chair is assembled and has a simple structure. The Wassily Chair not only considers the support of the back but also pays attention to the support of the waist, which is more ergonomic", and can highlight "Marcel Breuer who entered the school in Weimar in 1920" in the chapter summary corresponding to the Wassily Chair to obtain a video note.
[0322] It should be understood that the chapter summary corresponding to the Barcelona Chair can refer to Figure 6 the chapter summary corresponding to the Barcelona Chair shown in Application Scenario 3 of Figure 6 . Similarly, the chapter summary corresponding to the Red and Blue Chair can refer to the chapter summary corresponding to the Red and Blue Chair shown in Application Scenario 3 of , which will not be elaborated here.
[0323] As Figure 7 shown in (b), when the user wants to view the original text content and / or context corresponding to the highlighted content, the user can long-press the highlighted content. AsFigure 7 As shown in (c) in , when the electronic device detects the long - press operation, it can display the corresponding original text content and / or context in the display interface.
[0324] In a possible implementation manner, when the electronic device detects a second operation on the first video, it can obtain a screenshot image corresponding to the target text content and can generate a video note according to the screenshot image and the initial note.
[0325] Exemplarily, the electronic device can determine the position of the target text content in the initial note, and according to the position of the target text content in the initial note, insert the screenshot image into the initial note to obtain a video note.
[0326] Optionally, the electronic device can insert the screenshot image in front of or behind the target text content according to the position of the target text content in the initial note to obtain a video note. Or, the electronic device can insert an image annotation in front of or behind the target text content according to the position of the target text content in the initial note to obtain a video note. When the user views the image annotation corresponding to the target text content, the electronic device can display the screenshot image corresponding to the target text content in the display interface.
[0327] Optionally, the position of the target text content in the initial note can refer to the paragraph of the target text content in the initial note. After determining the position of the target text content in the initial note, the electronic device can insert the screenshot image in front of or behind the paragraph where the target text content is located according to the position of the target text content in the initial note to obtain a video note.
[0328] In one example, when the second operation is a screenshot operation, the electronic device can directly obtain the screenshot image obtained by taking a screenshot of the second screen. That is, the screenshot image corresponding to the target text content can be the image obtained by the screenshot operation.
[0329] In another example, when the second operation is a content marking operation, a content extraction operation, or a content input operation, the electronic device can automatically take a screenshot of the second screen when detecting the second operation to obtain a screenshot image. That is, the screenshot image can include the image corresponding to the second screen.
[0330] For example, when the electronic device displays the second screen in full screen, when a content marking operation, a content extraction operation, or a content input operation on the second screen is detected, the electronic device can automatically take a full-screen screenshot of the second screen to obtain a screenshot image corresponding to the second screen. When the electronic device displays the second screen and the application interface corresponding to a certain application (such as Application A) in a split-screen manner, when a content marking operation, a content extraction operation, or a content input operation on the second screen is detected, the electronic device can obtain the screenshot image corresponding to the second screen by means of partial screenshot.
[0331] Optionally, when the electronic device does not have the function of partial screenshot, the electronic device can only perform a screenshot operation when displaying the second screen in full screen. That is to say, when the electronic device does not have the function of partial screenshot, when a content marking operation, a content extraction operation, or a content input operation on the second screen is detected, the electronic device can determine whether the electronic device is currently displaying the second screen in full screen. If the electronic device is currently displaying the second screen in full screen, the electronic device can automatically take a screenshot of the second screen to obtain a screenshot image corresponding to the second screen. If the electronic device is not currently displaying the second screen in full screen, for example, the electronic device displays the second screen and the application interface corresponding to Application A in a split-screen manner, the electronic device can not take a screenshot of the second screen.
[0332] Please refer to Figure 8 , Figure 8 which shows the application scenario provided by the embodiment of the present application. Figure 5 This application scenario is exemplarily described by taking the text content in the screen including the subtitles corresponding to the screen as an example. Among them, in this application scenario, the subtitles corresponding to the screen include Figure 6 the subtitles shown in
[0333] After the electronic device enters the note mode, if a second operation on the first video is detected by the electronic device, the electronic device can determine the target text content and obtain the screenshot image corresponding to the second screen. As shown in (a) of Figure 8 , when the electronic device divides the text content in the target screen according to content relevance and adds punctuation marks during the paragraph division to obtain the initial note, the electronic device can determine the paragraph of the target text content in the initial note and can insert the screenshot image in front of or behind the paragraph. Figure 8 In (a) of Figure 8 , taking inserting the screenshot image in front of the paragraph as an example. In addition, as shown in (a) of
[0334] As shown in Figure 8As shown in (b) therein, after obtaining the screenshot image corresponding to the second screen, when the electronic device divides the text content in the target screen into paragraphs according to content relevance to obtain the initial note, the electronic device may insert the image annotation corresponding to the screenshot image before or after the target text content. For example, the image annotation may be [1] to obtain the video note. Among them, Figure 8 in (b) therein takes inserting the image annotation after the target text content as an example.
[0335] Such as Figure 8 shown in (b) therein, when the user wants to view the screenshot image, the user can click on the image annotation corresponding to the screenshot image. Such as Figure 8 shown in (c) therein, when detecting a click operation on the image annotation corresponding to the screenshot image, the electronic device may display the screenshot image on the display interface.
[0336] In a possible implementation manner, during the generation process of the video note, the electronic device may determine whether the note application or the memo application is opened, that is, it may determine whether the application interface corresponding to the note application or the memo application is displayed on the display interface. When the note application or the memo application is opened, the electronic device may save the generated video note to the note application or the memo application, or may save the text content in the second screen and / or the screenshot image corresponding to the second screen to the note application or the memo application for the user to view conveniently. When the note application is not opened and the memo application is not opened either, the electronic device may save the generated video note to the note application or the memo application after exiting the note mode. When the user subsequently opens the note application or the memo application, the user may view the video note in the note application or the memo application.
[0337] That is to say, when the user is watching the first video and taking notes through the note application or the memo application, after the electronic device enters the note mode based on the first operation, if it detects the second operation of the user on the second screen of the first video, the electronic device may determine the target text content and may obtain the screenshot image corresponding to the second screen. Among them, after obtaining the target text content and / or the screenshot image corresponding to the second screen, the electronic device may save the target text content and / or the screenshot image corresponding to the second screen to the note application or the memo application. Or, after the electronic device generates a video note according to the target text content and / or the screenshot image corresponding to the target text content, it may save the video note to the note application or the memo application.
[0338] Please refer to Figure 9 , Figure 9 which shows the application scenario schematic provided by the embodiment of the present application Figure 6This application scenario will be exemplarily described by taking the example of saving the target text content and the screenshot image corresponding to the target text content to a note application.
[0339] As Figure 9 shown in (a) of [reference], when the electronic device displays the screen of the first video and the application interface corresponding to the note application in split screen, that is, when the user is watching the video and taking notes through the note application, after the electronic device enters the note mode based on the first operation, if it detects a second operation on the second screen of the first video. For example, as Figure 9 shown in (a) of [reference], if it detects a local screenshot operation of drawing a closed pattern 900 on the display screen after detecting a knuckle tap on the display screen, the electronic device can determine the target text content, such as "Marcel Breuer who entered school in Weimar in 1920", and can obtain the screenshot image corresponding to the second screen. As Figure 9 shown in (b) of [reference], after obtaining the target text content and the screenshot image corresponding to the second screen, the electronic device can save the target text content and the screenshot image corresponding to the second screen to the note application.
[0340] It should be noted that after the electronic device enters the note mode, if it does not detect a second operation on the first video, the electronic device can directly determine the initial note generated according to the text content in the target screen as the video note corresponding to the first video. That is, the initial note without highlighted content can be determined as the video note corresponding to the first video.
[0341] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0342] Corresponding to the video note generation method described in the above embodiments, an embodiment of the present application further provides a video note generation device. Each module of the device can correspondingly implement each step of the video note generation method.
[0343] It should be noted that the information interaction, execution process, etc. between the above devices / units, because they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.
[0344] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0345] An embodiment of this application also provides an electronic device. The electronic device includes at least one memory, at least one processor, and a computer program stored in the at least one memory and executable on the at least one processor. When the processor executes the computer program, the electronic device implements the steps in any of the foregoing method embodiments. Exemplarily, the structure of the electronic device can be as Figure 1 shown.
[0346] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a computer, the computer implements the steps in any of the foregoing method embodiments.
[0347] An embodiment of this application provides a computer program product. When the computer program product runs on an electronic device, the electronic device implements the steps in any of the foregoing method embodiments.
[0348] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can at least include: any entity or device that can carry the computer program code to the device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a portable hard drive, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable storage medium cannot be an electrical carrier signal and a telecommunication signal.
[0349] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0350] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0351] In the embodiments provided in this application, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.
[0352] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0353] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A method for generating video notes, characterized in that, including: The electronic device displays a first screen of a first video; When a first operation is detected, the electronic device enters a note mode; In the note mode, when a second operation on the first video is detected by the electronic device, a second screen is determined, and the second screen is a screen in the first video; The electronic device determines target text content according to the second operation and the text content in the second screen, and generates a video note according to the target text content.
2. The method according to claim 1, characterized in that, The second screen is the screen displayed by the electronic device when the second operation is detected.
3. The method according to claim 1 or 2, characterized in that, The second operation is a screenshot operation, and the target text content is all or part of the text content in the second screen.
4. The method according to claim 3, characterized in that The text content in the second screen includes the subtitle corresponding to the second screen.
5. The method according to claim 1 or 2, characterized in that, The second operation is a partial screenshot operation. When the electronic device determines the target text content according to the second operation and the text content in the second screen, it includes: The electronic device determines a screenshot image corresponding to the second operation; The electronic device determines the target text content according to the text content in the screenshot image, and the text content in the screenshot image is all or part of the text content in the second screen.
6. The method according to claim 1 or 2, characterized in that, The second operation is a content marking operation. When the electronic device determines the target text content according to the second operation and the text content in the second screen, it includes: The electronic device determines the text content marked by the second operation, and determines the target text content according to the text content marked by the second operation. The text content marked by the second operation is all or part of the text content in the second screen.
7. The method according to claim 1 or 2, characterized in that, The second operation is a content extraction operation. When the electronic device determines the target text content according to the second operation and the text content in the second screen, it includes: The electronic device determines the text content extracted by the second operation, and determines the target text content according to the text content extracted by the second operation. The text content extracted by the second operation is all or part of the text content in the second screen.
8. The method according to claim 1 or 2, characterized in that, The second operation is a content input operation. When the electronic device determines the target text content according to the second operation and the text content in the second screen, it includes: The electronic device determines the text content input by the second operation, and determines the target text content according to the text content in the second screen and the text content input by the second operation.
9. The method according to any one of claims 1 to 8, characterized in that The method further includes: When the electronic device detects a second operation on the first video, a third screen is determined. The third screen is a screen in the first video, and the third screen is a screen before the second screen; The electronic device obtains the text content in the third screen, and displays the text content in the third screen and the text content in the second screen.
10. The method according to claim 9, characterized in that, When the second operation is a screenshot operation, when the electronic device determines the target text content according to the second operation and the text content in the second screen, it includes: The electronic device obtains the text content corresponding to the third operation, and determines the target text content according to the text content corresponding to the third operation, where the third operation is an operation of selecting text content from the text content in the second screen and the text content in the third screen.
11. The method according to claim 9, wherein When the second operation is a content marking operation, a content extraction operation, or a content input operation, the electronic device determines the target text content according to the second operation and the text content in the second screen, including: The electronic device detects a fourth operation on the text content in the second screen and / or the text content in the third screen, where the fourth operation is a content marking operation, a content extraction operation, or a content input operation; The electronic device determines the target text content according to the fourth operation, the text content in the second screen, and / or the text content in the third screen.
12. The method according to any one of claims 1 to 11, characterized in that, The electronic device generates a video note according to the target text content, including: The electronic device obtains the text content in the target screen, and generates the video note according to the target text content and the text content in the target screen.
13. The method according to claim 12, wherein The target screen includes the screen displayed after the electronic device enters the note mode.
14. The method according to claim 12, wherein The target screen includes the screen displayed after the electronic device enters the note mode, and the screen displayed within a preset time before the electronic device enters the note mode.
15. The method according to any one of claims 12 to 14, characterized in that, The electronic device generates the video note according to the target text content and the text content in the target screen, including: The electronic device generates an initial note according to the text content in the target screen; The electronic device highlights a first content in the initial note to obtain the video note, where the first content is determined according to the target text content.
16. The method according to claim 15, characterized in that, The first content includes one or more of the target text content, the context corresponding to the target text content, or the text content with the same content view as the target text content.
17. The method according to any one of claims 1 to 16, characterized in that, The video note includes the target text content, or includes a summary corresponding to the target text content.
18. The method according to claim 17, wherein When the video note includes a summary corresponding to the target text content, the method further includes: The electronic device detects a fifth operation on the summary corresponding to the target text content; The electronic device displays the original text content corresponding to the summary according to the fifth operation, where the original text content is the text content in the target screen.
19. The method according to any one of claims 1 to 18, characterized in that, The electronic device generates a video note according to the target text content, including: The electronic device obtains a screenshot image corresponding to the target text content, and generates the video note according to the target text content and the screenshot image.
20. The method according to claim 19, wherein The electronic device generates the video note according to the target text content and the screenshot image, including: The electronic device determines the position of the target text content in the initial note; The electronic device inserts the screenshot image into the initial note according to the position of the target text content in the initial note to obtain the video note.
21. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the electronic device implements the video note generation method described in any one of claims 1 to 20.
22. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a computer, the computer implements the video note generation method described in any one of claims 1 to 20.
23. A computer program product, characterized in that, When the computer program product runs on an electronic device, the electronic device is caused to execute the video note generation method described in any one of claims 1 to 20.