Video processing method and device, equipment and storage medium

By displaying the canvas mask on the video playback area, users can take notes on the canvas mask, and combined with the graphic and text recognition to generate graphic and text articles, the problem of single content in the existing technology is solved, and the quality and user experience of graphic and text articles are improved.

CN120543693APending Publication Date: 2025-08-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410210853.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing video to graphics and text functions only support graphics and text recognition, resulting in the converted graphics and text articles with a single content and poor quality of the articles.

Method used

The canvas mask is displayed on the video playback area, allowing users to take notes on the canvas mask, and generate graphic and text articles based on graphic and text recognition, including electronic notes for user notes.

Benefits of technology

It enriches the content of graphic and text articles, improves the quality of articles and user customization, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543693A_ABST
    Figure CN120543693A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method and device, equipment and a storage medium, and the method comprises the steps: playing a video which needs to be subjected to image-text conversion, and displaying a canvas covering layer on a video playing region of the video; drawing and displaying an electronic note of the current video frame in the canvas mask layer according to a first note recording operation detected on the canvas mask layer; in response to an image-text conversion operation for the video, displaying an image-text article corresponding to the video; the image-text article comprises a target image-text obtained by carrying out image-text recognition on the video, a video frame with an electronic note and the corresponding electronic note. According to the method and the device, the article content of the image-text article obtained by performing image-text conversion on the video can be enriched, so that the article quality of the image-text article is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, specifically to the field of multimedia processing technology, and in particular to a video processing method, apparatus, device and storage medium. Background Art

[0002] With the development of multimedia technology, video-to-text conversion has become widely used. This refers to converting a video into text and image to create a corresponding text and image article. Currently, existing video-to-text conversion functions only support image-to-text recognition within the video. This results in the converted image-to-text article consisting solely of the image and image identified through image-to-text recognition, resulting in a relatively monotonous article content and poor quality. Summary of the Invention

[0003] The embodiments of the present application provide a video processing method, apparatus, device, and storage medium, which can realize the image-text conversion of videos, enrich the content of the converted image-text articles, and thus improve the article quality of the image-text articles.

[0004] In one aspect, an embodiment of the present application provides a video processing method, the method comprising:

[0005] Play the video that needs to be converted from image to text, and display a canvas mask on the video playback area of ​​the video;

[0006] Drawing an electronic note displaying a current video frame in the canvas mask according to a first note recording operation detected on the canvas mask; the current video frame refers to the video frame displayed in the video playback area when the first note recording operation is detected;

[0007] In response to the image-text conversion operation for the video, the image-text article corresponding to the video is displayed; the image-text article includes: the target image-text obtained by image-text recognition of the video, the video frame with electronic notes and the corresponding electronic notes.

[0008] On the other hand, an embodiment of the present application provides a video processing device, comprising:

[0009] A playback unit, used to play videos that require image-to-text conversion;

[0010] a processing unit, configured to display a canvas mask on a video playback area of ​​the video;

[0011] The processing unit is further configured to draw an electronic note displaying a current video frame in the canvas mask based on a first note recording operation detected on the canvas mask; the current video frame refers to the video frame displayed in the video playback area when the first note recording operation is detected;

[0012] The processing unit is also used to display the graphic article corresponding to the video in response to the graphic-text conversion operation for the video; the graphic article includes: the target graphic obtained by graphic-text recognition of the video, the video frame with electronic notes and the corresponding electronic notes.

[0013] In another aspect, an embodiment of the present application provides a computer device, the computer device including an input interface and an output interface, and the computer device further including:

[0014] processors and computer storage media;

[0015] The processor is suitable for implementing one or more instructions, the computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded by the processor and executing the above-mentioned video processing method.

[0016] On the other hand, an embodiment of the present application provides a computer storage medium, which stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the above-mentioned video processing method.

[0017] On the other hand, an embodiment of the present application provides a computer program product, which includes one or more instructions; when the one or more instructions in the computer program product are executed by a processor, the above-mentioned video processing method is implemented.

[0018] In the embodiment of the present application, for a video that needs to be converted from image to text, the video can be played and a canvas mask can be displayed on the video playback area, so that the user (object) can enter a first note recording operation on the canvas mask to obtain an electronic note of the current video frame. When a text-to-image conversion operation for the video is detected, the image-to-text article corresponding to the video can be generated based on the target image-to-text obtained by image-to-text recognition of the video and the video frame with the electronic note and the corresponding electronic note. The image-to-text article finally displayed includes not only the target image-to-text, but also some video frames and corresponding electronic notes. This can enrich the content of the converted image-to-text article and improve the quality of the image-to-text article. In addition, by introducing the electronic notes recorded by the user into the image-to-text article, the image-to-text article can be made more valuable in information, and the user customization of the image-to-text article can be enhanced, thereby improving the user experience of converting video to image-to-text. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1a This is a schematic diagram of an operation flow for a user to convert a video into a graphic article provided by an embodiment of the present application;

[0021] Figure 1b This is a flow chart of a video processing solution jointly executed by a terminal and a server provided in an embodiment of the present application;

[0022] Figure 2 This is a flow chart of a video processing method provided in an embodiment of the present application;

[0023] Figure 3a This is a schematic diagram of setting a video link address provided by an embodiment of the present application;

[0024] Figure 3b This is a schematic diagram of uploading a video provided in an embodiment of the present application;

[0025] Figure 3c This is a schematic diagram of the positional relationship between a canvas mask and a video playback area provided in an embodiment of the present application;

[0026] Figure 3d This is a schematic diagram of triggering a terminal to pause video playback and display a canvas mask, provided by an embodiment of the present application;

[0027] Figure 3e is a schematic diagram of drawing an electronic note of the current video frame provided by an embodiment of the present application;

[0028] Figure 3f This is a schematic diagram of adding note recording information provided by an embodiment of the present application;

[0029] Figure 3g is a schematic diagram of reproducing an electronic note on a first video frame provided by an embodiment of the present application;

[0030] Figure 4a This is a schematic diagram of a process for generating a graphic article provided by an embodiment of the present application;

[0031] Figure 4b This is a schematic diagram of cutting N key video frames provided by an embodiment of the present application;

[0032] Figure 5 is a flowchart of a video processing method provided by another embodiment of the present application;

[0033] Figure 6a This is a schematic diagram of updating and displaying video summary information provided by an embodiment of the present application;

[0034] Figure 6b is a schematic diagram of displaying a note list provided in an embodiment of the present application;

[0035] Figure 6c is a schematic diagram of drawing an electronic note of a second video frame provided by an embodiment of the present application;

[0036] Figure 6d is a schematic diagram of a document editing interface provided in an embodiment of the present application;

[0037] Figure 6e This is a schematic diagram showing a platform logo provided in an embodiment of the present application;

[0038] Figure 7a This is a schematic diagram of the implementation logic of a video-to-article conversion method provided in an embodiment of the present application;

[0039] Figure 7b This is a flowchart of the processing logic of a video player provided in an embodiment of the present application;

[0040] Figure 7c This is a flowchart of a video pre-processing provided by an embodiment of the present application;

[0041] Figure 7d This is a flow chart of material understanding provided by an embodiment of the present application;

[0042] Figure 8 is a structural diagram of a video processing device provided in an embodiment of the present application;

[0043] Figure 9 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0045] The embodiment of the present application proposes a video processing solution, which designs an operable canvas mask on the upper layer of the video playback area, so that the user (object) can take notes on part or all of the video frames in the video through the canvas mask according to actual needs while watching the video, so that the graphic article obtained by converting the video into text and image not only includes the text and image obtained by the video text and image recognition, but also includes the video frames and corresponding electronic notes recorded by the user. This can enrich the content of the converted graphic article and thus improve the quality of the graphic article. In addition, by introducing the electronic notes recorded by the user into the graphic article, the graphic article can be made more valuable in information, and the user customization of the graphic article can be enhanced, thereby improving the user experience of converting video to text and image.

[0046] Specifically, the video processing solution works as follows: A video to be converted from image to text is obtained. This video can be a local video uploaded by a user (object) or an online video indicated by a video address entered by the user, without limitation. After obtaining the video, the video can be played and a canvas overlay can be displayed on the video playback area, allowing the user (object) to take notes on the canvas overlay. Accordingly, if a first note-taking operation is detected on the canvas overlay, an electronic note for the current video frame (i.e., the video frame displayed in the video playback area when the first note-taking operation was detected) can be drawn and displayed on the canvas overlay based on the detected first note-taking operation. Furthermore, after the user completes taking notes on the current video frame, electronic notes for other video frames can be drawn based on the same processing logic based on other note-taking operations entered by the user on the canvas overlay. When the user wishes to export the image to text, the target image obtained through image-text recognition on the video, the video frame with the electronic note, and the corresponding electronic note can be used to generate the image-text article corresponding to the video, thereby exporting and displaying the image-text article.

[0047] Optionally, the exported and displayed graphic articles can also support editing. That is, after exporting the graphic article, users can manually edit the graphic article according to actual needs; or, when the user has the need to rewrite the article, the graphic article can be rewritten in the title, the full text can be polished (i.e., full text optimization and adjustment), and text error correction can be performed based on AI (Artificial Intelligence) technology. Among them, AI technology refers to the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science; it mainly produces a new type of intelligent machine that can respond in a similar way to human intelligence by understanding the essence of intelligence, so that the intelligent machine has multiple functions such as perception, reasoning and decision-making. Accordingly, AI technology is an interdisciplinary subject, which mainly includes several major directions such as computer vision technology (Computer Vision, CV), speech processing technology, natural language processing technology and machine learning (Machine Learning, ML) / deep learning. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question-answering, and knowledge graphs. Machine learning is a multidisciplinary discipline that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of AI and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Deep learning, on the other hand, is a machine learning technique that utilizes deep neural network systems.

[0048] Based on the above, it can be seen that in the specific implementation process of the video processing solution proposed in the embodiment of the present application, the flowchart of the user converting the video into a graphic article can be roughly referred to. Figure 1aAs shown, it may include the following steps: 1. The user uploads a local video or video address; 2. The user records electronic notes of the video frames while watching the video; 3. The user exports pictures and texts; 4. The user edits the picture and text article or performs AI title / article rewriting (i.e., rewriting the title or polishing the full text of the picture and text article based on AI technology). Optionally, after the user uploads the local video or video address, he or she may directly export the picture and text article without taking notes; and, when the user edits the picture and text article, he or she may also perform AI title / article rewriting. It is worth emphasizing that in the embodiments of the present application, related data such as user information (such as local videos or video addresses uploaded by users, any operations performed by users, etc.) are involved. When any method embodiment proposed in the embodiments of the present application is applied to a specific product or technology, these relevant data are collected with the permission or consent of the user, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant regions.

[0049] In a specific implementation, the video processing solution mentioned above can be executed by a computer device, which can be a terminal or a server. Alternatively, the video processing solution mentioned above can be executed by the terminal and the server together; for example, see Figure 1b As shown: the terminal can be responsible for obtaining a video that needs to be converted into text and image, playing the video and displaying a canvas overlay on the video playback area of ​​the video, and drawing an electronic note of the current video frame in the canvas overlay based on the first note recording operation detected, and uploading the current video frame and the corresponding electronic note to the server, so that when the user wants to export text and image, the terminal can send a text and image conversion request to the server, so that the server responds to the text and image conversion request, uses the target text and image obtained by performing text and image recognition on the video, the video frame with the electronic note, and the corresponding electronic note to generate a text and image article corresponding to the video, and returns the generated text and image article to the terminal, which displays the text and image article. Optionally, in other embodiments, for any video frame with an electronic note, the terminal can also copy the corresponding electronic note on the corresponding video frame, thereby taking a screenshot of the video frame with the copied electronic note, obtaining a video frame screenshot, and then uploading the video frame screenshot to the server, so that the server uses the target text and image and the video frame screenshots corresponding to each video frame with an electronic note to generate a text and image article.

[0050] The aforementioned terminals may include smartphones, computers (such as tablets, laptops, and desktop computers), smart wearable devices (such as smart watches and smart glasses), intelligent voice interaction devices, smart home appliances (such as smart TVs), in-vehicle terminals, or aircraft. The servers may be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDNs (Content Delivery Networks), and big data and artificial intelligence platforms. Furthermore, the terminals and servers may be located within or outside the blockchain network, without limitation. Furthermore, the terminals and servers may upload any internally stored data to the blockchain network for storage, thereby preventing tampering with the internally stored data and improving data security.

[0051] Based on the above description of the video processing solution, the embodiment of the present application proposes a video processing method; the video processing method can be executed by the terminal, or by the terminal and the server together, without limitation. Figure 2 As shown, the video processing method may include the following steps S201-S204:

[0052] S201, playing a video that needs to be converted from image to text.

[0053] Among them, the video that needs to be converted from image to text may include but is not limited to any of the following: film and television drama videos, game videos (videos obtained by recording the game process), game advertising videos (advertising videos used to promote a certain game), vlogs (a video that records and shares personal life or topics), tutorial videos of any learning content (such as certain mathematical knowledge, certain camera techniques, etc.), videos obtained by recording any live content (such as the content of a concert or a PPT (slideshow) of a speech), etc. It is understandable that no matter what kind of video is required to be converted from image to text in the embodiments of this application, it can include audio data and a video frame set; the so-called video frame set refers to a collection of multiple video frames, and a video frame is a frame of image.

[0054] In a specific implementation, the user can trigger the terminal to display a workbench interface for the video-to-text function, which may include a video settings area; after the terminal displays the workbench interface, the user can set the video that needs to be converted into text in the video settings area. Specifically, the video settings area 31 in the workbench interface may include multiple tabs such as a video address tab 311 and a local video tab 312. Different tabs correspond to different video settings entries. The video settings entry corresponding to the video address tab 311 is used to enter the video link address of any online video (i.e., a video not stored in the local space of the terminal), and the video settings entry corresponding to the local video tab 312 is used to upload any local video (i.e., a video stored in the local space of the terminal). In this case, the user can select a tab in the video settings area 31 according to their own needs to trigger the terminal to display the video settings entry corresponding to the selected tab in the video settings area, and then perform video settings through the video settings entry displayed by the terminal.

[0055] For example, if a user wants to convert an online video into text, the user can select the video address tab 311, so that the terminal displays a video setting entry 321 for inputting a video link address in the video setting area; then, the user can input a video link address of an online video (such as http: / / xxx.com) in the video setting entry 321, and click or press the confirmation component 322 (such as Figure 3a As shown), the terminal is triggered to obtain the video link address currently displayed in the video setting entry 321, and download the corresponding online video from the video link address as the video to be converted into text. For another example: if the user wants to convert a local video into text, the user can select the local video tab 312, so that the terminal displays the video setting entry 323 for uploading the local video in the video setting area, as shown. Figure 3b Then, the user can select a local video in the local space of the terminal through the video setting entrance 323 and perform an upload operation to trigger the terminal to use the local video uploaded by the user as a video that needs to be converted into text.

[0056] Based on the above description, it can be seen that the embodiment of the present application can obtain local videos as videos that need to be converted from images to text, and can also download online videos as videos that need to be converted from images to text, which has high flexibility. It should be noted that the above is only an exemplary description of a product form of the video setting area, and does not limit it. For example, in other embodiments, the video setting area may not display a tab, and directly display a video setting entrance for entering a video link address and a video setting entrance for uploading a local video at the same time, so that the user can directly perform video settings through the corresponding video setting entrance according to actual needs. Alternatively, the video setting area may also include an address input port, so that the user can enter the storage address of the local video or the video link address of the online video in the address input port according to actual needs, thereby triggering the terminal to obtain the corresponding video from the local space or the video link address based on the address in the address input port as the video that needs to be converted from images to text.

[0057] After the terminal obtains the video that needs to be converted from image to text, it can determine the video playback area on the display screen. The video playback area can be located at any position on the display screen, such as the center, the upper left corner, the lower right corner, and so on. After determining the video playback area, the terminal can play the video in the video playback area through the target application. Among them, the target application mentioned here refers to an application that provides a video-to-text function. For example, if a web application can provide a video-to-text function, the target application can be a web application; for example, if a social application can provide a video-to-text function, the target application can be a social application; for example, if a video application can provide a video-to-text function, the target application can be a video application, and so on.

[0058] S202: Display a canvas mask on the video playback area of ​​the video.

[0059] Among them, the canvas mask refers to: a view window with transparency and supporting the drawing of any element (such as text, graphics, etc.). It should be noted that the embodiment of the present application does not limit the transparency of the canvas mask, and its transparency can be set according to actual needs. For example, the transparency of the canvas mask can be set to 100%, in which case the canvas mask is a completely transparent view window; for another example, the transparency of the canvas mask can be set to 50%, in which case the canvas mask is a semi-transparent view window. It is understandable that after the canvas mask is displayed on the video playback area, the canvas mask is covered on the video playback area (such as Figure 3c As shown), this will prevent users from performing any operations on the video playback area (such as pausing the video or dragging the progress bar). Therefore, even if the canvas overlay is a completely transparent view window, users can still perceive that the canvas overlay is displayed on the video playback area.

[0060] In one specific implementation, the terminal may, by default, display a canvas overlay on the video playback area of ​​the video during video playback. In this specific implementation, when the canvas overlay is displayed on the video playback area, the terminal may maintain video playback in the video playback area. In another specific implementation, considering that the user may need to perform one or more operations on the video playback area during video playback, in order not to affect the user's video playback experience, the terminal may detect the user's note triggering operation (used to indicate the user's intention to take notes) in real time during video playback, and then display the canvas overlay on the video playback area after detecting the note triggering operation. In this specific implementation, when the canvas overlay is displayed on the video playback area, the terminal may maintain video playback in the video playback area to ensure that video playback is not interrupted; or, considering that the user enters the note triggering operation because they want to take notes on the currently displayed video frame through the canvas overlay, in order to facilitate the user to take notes while watching the corresponding video frame, the terminal may pause the video in the video playback area when the canvas overlay is displayed on the video playback area, so that the video remains on the corresponding video frame, thereby improving the user's note-taking experience.

[0061] It should be noted that the embodiments of the present application do not limit the specific implementation of the note trigger operation mentioned above. For example, the note trigger operation may be an operation of inputting a preset gesture or voice instruction for indicating the display of a canvas mask. For another example, the note trigger operation may be an operation of clicking or pressing a blank area in the display screen. For another example, the terminal may display the note recording component 33 during the playback of the video, so that the user can input the note trigger operation by performing a trigger operation (such as a click operation, a press operation, etc.) on the note recording component 33, that is, the note trigger operation may be a trigger operation for the note recording component 33; it can be seen that in this case, the terminal responds to the trigger operation for the note recording component 33, pauses the playback of the video, and triggers the execution of the step of displaying the canvas mask on the video playback area of ​​the video. Among them, the note recording component 33 mentioned here can be located in the video playback area or outside the video playback area, and there is no limitation on this. For example, the note recording component 33 is located in the video playback area 34, and the canvas mask is a completely transparent view window. The user triggers the terminal to pause the video playback and display the canvas mask by performing a trigger operation (such as a click operation) on the note recording component 33. Figure 3d shown; based on Figure 3d It can be seen that when the terminal pauses playing the video, the display form of the play button 35 switches from the first state (the form used to indicate that the video is in the playing state) to the second state (the form used to indicate that the video is in the paused state). The user can also know that the terminal has displayed the canvas mask through the change in the display form of the play button 35.

[0062] based on Figure 3d As can be seen, to facilitate the user's subsequent note-taking on the canvas overlay, the terminal may also display a toolbar 36 on the display screen while the canvas overlay is displayed on the video playback area. This toolbar includes one or more note-taking tools. The so-called note-taking tools are tools that can be used to draw notes on the canvas overlay. Exemplarily, the note-taking tools displayed on the display screen include, but are not limited to, graphic marking tools (such as a rectangle marking tool 361, an ellipse marking tool 362, an arrow marking tool 363, etc.) for marking at least one video content in a video frame, a text recording tool 364 for recording text, a text setting tool 365 for setting text attributes (such as text color, text font, text size, etc.), an undo tool 366 for undoing a note, and so on. It is understood that in other embodiments, if the terminal's display screen supports touch input, the user may also use a touch tool (such as a finger, a stylus, etc.) to draw notes directly on the canvas overlay. In this case, the terminal may not display the note-taking tools while the canvas overlay is displayed.

[0063] Furthermore, after displaying the canvas overlay, the terminal may detect a first note-taking operation on the canvas overlay. Specifically, if the terminal displays one or more note-taking tools on the display screen, the terminal may detect a tool selection operation by the subject (user). After detecting the tool selection operation, the terminal may select a corresponding note-taking tool based on the tool selection operation, and treat the operation entered by the subject (user) on the canvas overlay using the selected note-taking tool as the first note-taking operation detected on the canvas overlay. Specifically, when the selected note-taking tool is a graphic marking tool, the operation entered by the subject (user) using the selected note-taking tool (i.e., the first note-taking operation) may be a graphic marking operation (i.e., an operation of drawing a graphic (e.g., a rectangle, ellipse, or line) on at least one video content) on at least one video content. When the selected note-taking tool is a text recording tool, the operation entered by the subject (user) using the selected note-taking tool (i.e., the first note-taking operation) may be a text input operation. If the display screen of the terminal supports touch input, the terminal can detect the note drawing operation (graphic marking operation, text input operation for at least one video content) input by the object (user) using the touch tool, and use the detected note drawing operation as the first note recording operation.

[0064] S203: Drawing and displaying an electronic note of the current video frame in the canvas mask according to the first note recording operation detected on the canvas mask.

[0065] As can be seen from the foregoing, the first note-recording operation can be a graphic marking operation for at least one video content, or a text input operation. When the first note-recording operation is a graphic marking operation for at least one video content, the terminal may draw and display an electronic note for the current video frame in the canvas overlay according to the first note-recording operation by drawing a corresponding graphic in a target display area (an area covering the at least one video content to be marked) in the canvas overlay according to the graphic marking operation for the at least one video content, thereby obtaining a graphic mark (i.e., an electronic note) for the at least one video content in the current video frame. When the first note-recording operation is a text input operation, the terminal may draw and display an electronic note for the current video frame in the canvas overlay according to the first note-recording operation by drawing corresponding characters at a target location (a preset default display location or an input location for a text input operation, etc.) in the canvas overlay according to the text input operation, thereby obtaining at least one text (i.e., an electronic note) for the current video frame.

[0066] Among them, the current video frame mentioned above refers to: the video frame displayed in the video playback area when the first note recording operation is detected; for example, when the first note recording operation is detected, the video frame displayed in the video playback area is the 10th video frame in the video, then the current video frame is the 10th video frame in the video. It can be understood that: ① The number of first note recording operations is at least one, and each time the terminal detects a first note recording operation, it can draw an electronic note of the current video frame in the canvas mask according to the currently detected first note recording operation; for example, assuming that the user inputs two first note recording operations, which are a graphic mark operation and a text input operation for at least one video content, then the terminal can draw and display two electronic notes such as a graphic mark 371 (assuming a rectangle) and text 372 (assuming "key content") in the canvas mask, such as Figure 3e ② If the terminal does not pause the video when displaying the canvas overlay, the terminal may pause the video when first detecting the first note recording operation, allowing the user to watch the current video frame while taking notes; if the terminal has already paused the video when displaying the canvas overlay, the terminal may continue to pause the video when first detecting the first note recording operation.

[0067] Furthermore, after drawing and displaying the electronic note of the current video frame, the terminal may detect a note completion event for the current video frame. The note completion event may include, but is not limited to, any of the following: an event of clicking a blank display position in an area of ​​the display screen other than the canvas mask, an event of not detecting the first note recording operation again within a preset time period, an event of inputting a gesture or voice command for indicating the end of note recording, etc. Alternatively, the terminal may also display a note completion component on the display screen. In this case, the note completion event may include an event of performing a trigger operation (such as a click operation, a press operation, etc.) on the note completion component displayed on the display screen (in other embodiments, on the display screen).

[0068] If a note completion event is detected for the current video frame, the terminal may associate and record the various electronic notes currently displayed in the canvas mask with the current video frame. Furthermore, since the video is in a paused playback state while the terminal is drawing the electronic note for the current video frame, if the terminal detects a note completion event for the current video frame, the terminal may continue to play the video based on the current video frame in the video playback area (i.e., starting from the current video frame, playing the various video frames following the current video frame in sequence), and the display form of the play button may be switched from the second state to the first state.

[0069] Optionally, if a note completion event for the current video frame is detected, the terminal may also generate note recording information based on the current video frame and the various electronic notes currently displayed in the canvas mask; add the generated note recording information to the note list, and clear the canvas mask, so that the user can subsequently record notes on other video frames on the canvas mask. It can be seen that the embodiment of the present application can support the user to display the corresponding note recording information on the display screen after completing the note recording of the current video frame, so that the user can intuitively see the electronic notes he has recorded, thereby improving the efficiency of electronic note tracing. It is understandable that in other embodiments, the terminal may also clear the canvas mask instead of displaying the note list and note recording information.

[0070] Among them, ① the generated note recording information may include: video frame indication information of the current video frame, and note indication information of each electronic note currently displayed in the canvas overlay. Specifically, the video frame indication information of the current video frame may include at least one of the following: the playback time point of the current video frame, and a thumbnail of the current video frame. As can be seen from the above, any electronic note may include: a graphic mark for at least one video content in the video frame, or at least one text; then, accordingly, the note indication information of any electronic note includes: an indicator symbol of the graphic mark in the corresponding electronic note (i.e., a symbol used to indicate the graphic mark), or indication information of the text in the corresponding electronic note. The text indication information may include the entire content of the corresponding text, or a portion of the content of the corresponding text (such as the first H characters in the corresponding text, where H is a positive integer). ② The note list is located in an area of ​​the display screen other than the video playback area. The note list may be displayed on the display screen synchronously when the canvas overlay is displayed; or, the note list may be displayed on the display screen after the terminal generates the note recording information for the first time, without limitation.

[0071] Exemplary: Following the above Figure 3e In the example shown, the various electronic notes currently displayed in the canvas mask include a graphic mark 371 (assuming a rectangle) and text 372 (assuming "key content") as an example, and assuming that the generated note record information is the first note record information in the note list. After generating the note record information 38, the terminal displays the note list on the display screen, and the video frame indication information 380 in the generated note record information 38 includes the playback time point of the current video frame (assuming it is 00:04), and the note indication information in the generated note record information 38 includes the indication symbol 381 of the graphic mark 371 and the indication information 382 of the text 372; the user clicks on a blank display position in the area outside the canvas mask on the display screen to trigger the terminal to add the generated note record information to the note list and clear the canvas mask. See the schematic diagram for Figure 3f shown.

[0072] Based on the above description, it should be noted that: as can be seen from the above, when the terminal displays the canvas overlay on the video playback area, it may keep the video playing in the video playback area, or it may pause the video in the video playback area. If the video keeps playing when the canvas overlay is displayed on the video playback area, it can be said that the display of the canvas overlay does not affect the playback status of the video (that is, the display of the canvas overlay does not trigger the pause of the video); in this case, after detecting the note completion event for the current video frame, the terminal can keep the canvas overlay displayed on the video playback area, so that the user can subsequently take notes on other video frames directly on the canvas overlay. If the video is paused while the canvas overlay is displayed on the video playback area, it can be understood that the display of the canvas overlay can affect the playback state of the video (i.e., the display of the canvas overlay can trigger the pause of the video); in this case, after detecting the note completion event for the current video frame, the terminal can also respond to the note completion event for the current video frame by canceling the display of the canvas overlay, thereby executing the step of continuing to play the video based on the current video frame in the video playback area. In this case, the user can subsequently trigger the terminal to display the canvas overlay again to take notes on other video frames. It is understandable that if the terminal displays a toolbar on the display screen during the display of the canvas overlay, the terminal can also simultaneously cancel the display of the toolbar when canceling the display of the canvas overlay.

[0073] It's worth emphasizing that, in practice, after completing a note-taking event for the current video frame, the user can, following the same workflow, trigger the terminal to draw and display electronic notes for other video frames on the canvas overlay by inputting a note-taking action on the canvas overlay. Similarly to the processing flow for the current video frame, upon detecting a note-taking completion event for another video frame, the terminal can also generate note-taking information for the other video frame and add the note-taking information for the other video frame to the note list.

[0074] Optionally, when the note list is displayed, if the user wants to view a certain video frame and the corresponding electronic note that have been noted again, the user can perform a selection operation (such as a click operation, a press operation, etc.) on the note record information of the corresponding video frame in the note list. Accordingly, when note record information is selected in the note list, the terminal can jump from the currently displayed video frame to the first video frame in the video playback area; and reproduce (i.e., re-display) the corresponding electronic note on the first video frame according to the triggered note record information. Among them, the currently displayed video frame refers to: the video frame displayed in the video playback area when it is detected that the note record information is selected; the first video frame refers to: the video frame corresponding to the triggered note record information, specifically, it can be the video frame corresponding to the playback time point in the triggered note record information. It can be understood that if the triggered note record information includes an indicator symbol of a graphic mark, the reproduced electronic note includes the graphic mark corresponding to the indicator symbol; if the triggered note record information includes text indicator information, the reproduced electronic note includes the text corresponding to the indicator information. Furthermore, the reproduction position of any electronic note is the same as the drawing display position of the corresponding electronic note; the drawing display position of any electronic note refers to the display position of the corresponding electronic note when the note record information is generated.

[0075] Based on the above description, see for example Figure 3g As shown, after the terminal adds note recording information 38 corresponding to the current video frame to the note list in response to a note completion event for the current video frame, it can also generate note recording information 39 and add note recording information 39 to the note list in response to note completion events for other video frames during subsequent video playback. After detecting that note recording information 38 corresponding to the current video frame has been selected in the note list, the terminal can determine the current video frame as the first video frame, jump from the currently displayed video frame to the first video frame in the video playback area, and reproduce the corresponding electronic note on the first video frame.

[0076] Among them, one implementation method of the terminal reproducing the corresponding electronic notes on the first video frame may be: directly reproducing the corresponding electronic notes on the first video frame; that is, in this case, the reproduced electronic notes are located on the surface of the first video frame. Alternatively, another implementation method of the terminal reproducing the corresponding electronic notes on the first video frame may be: reproducing the corresponding electronic notes on a canvas mask covering the first video frame; that is, in this case, the reproduced electronic notes are located on the surface of the canvas mask. Since the canvas mask covers the first video frame, it can create the phenomenon that the electronic notes are reproduced on the first video frame. It is understandable that in this case, if the canvas mask is not displayed on the video playback area before the electronic notes are reproduced, then when reproducing the electronic notes, the terminal can first display the canvas mask on the video playback area so that the canvas mask covers the first video frame, and then trigger the execution of the step of reproducing the corresponding electronic notes on the canvas mask covering the first video frame. Optionally, if the terminal reproduces the corresponding electronic notes on the canvas mask covering the first video frame, the terminal may also display a toolbar on the display screen when reproducing the electronic notes, so that the user can use the note-taking tools in the toolbar to re-edit the electronic notes of the first video frame, thereby improving the convenience and efficiency of note editing.

[0077] Based on the above description, it can be seen that when displaying a note list, the embodiment of the present application can also support users to select any note recording information according to their own needs, so that the terminal jumps to the corresponding video frame in the video playback area and reproduces the corresponding electronic notes, thereby facilitating users to better trace back their historical electronic notes and re-edit the reproduced electronic notes. This not only improves the efficiency of tracing back electronic notes, but also improves the efficiency and convenience of editing electronic notes.

[0078] S204: In response to the image-text conversion operation for the video, the image-text article corresponding to the video is displayed.

[0079] Among them, the image-text conversion operation refers to the operation used to trigger the image-text conversion of the video; for example, the image-text conversion operation can be a trigger operation (such as a click operation, a press operation, etc.) for the image-text export component (or image-text conversion component) in the display screen, or it can be an operation of inputting a voice command or gesture to instruct the image-text conversion of the video, and so on.

[0080] After the terminal detects the image-text conversion operation for the video, it can respond to the image-text conversion operation, determine the image-text article corresponding to the video, and display the image-text article corresponding to the video. In a specific implementation, the image-text article can be generated by the terminal, or it can be generated by the terminal requesting the server, and there is no limitation on this. The image-text article may include: target images and texts obtained by image-text recognition of the video, video frames with electronic notes, and corresponding electronic notes. Among them, image-text recognition may include at least one of the following processes: recognition processing of key video frames for the video, text recognition for key video frames (i.e., the process of identifying text information in key video frames), and text recognition for the audio data of the video (i.e., the process of converting audio data into text information), etc.

[0081] Accordingly, the target graphic may include at least one of the following: a first text message obtained by performing text recognition on the audio data of the video, and key frame data; the key frame data includes: at least one key video frame, and a second text message obtained by performing text recognition on the corresponding key video frame. The key video frame mentioned here may be referred to as a key frame for short, which refers to a video frame that plays a decisive role in the multiple video frames that constitute a video (an animation); for example, for a video containing an action object, the key frame can determine the starting state and the ending state of the action performed by the action object, and the time between two key frames can also determine the rhythm of the action. For conventional actions such as speaking, walking, running, and fighting, 2-3 key frames can determine the basic appearance of the entire action that the action object needs to perform.

[0082] It should be noted that in the above-mentioned graphic articles, the video frames with electronic notes are presented in the form of a single video frame, or in the form of a video animation; any video animation refers to a dynamic image composed of at least two video frames, and there is at least one video frame with electronic notes in any video animation. It is understandable that when the video frames with electronic notes are presented in the form of video animations, the graphic article includes the target graphic, at least one video animation, and the electronic notes of the corresponding video frames in each video animation. By presenting the video frames with electronic notes in the form of video animations, the playback status of the video when the user records the electronic notes can be better reproduced, thereby further enriching the information presented in the graphic article and improving the article quality of the graphic article.

[0083] In addition, the embodiments of the present application do not limit the method for generating video animations. For example, taking the current video frame mentioned above as an example: if the video is in a paused playback state during the process of drawing the electronic note of the current video frame, a video clip containing the current video frame can be cut out from the video, and the cut out video clip can be used as a video animation for presenting the current video frame; in this case, the position of the current video frame in the corresponding video animation can be any position, such as the first frame position, the last frame position, or the middle frame position, etc. If the video is not paused during the process of drawing the electronic notes of the current video frame, the video frames played in the video during the process of drawing the electronic notes of the current video frame can be used to generate a video animation for presenting the current video frame; in this case, the current video frame is the first frame in the corresponding video animation, and the last frame in the corresponding video animation is the video frame played in the video when the note completion event for the current video frame is detected. For example, when the terminal plays the 11th video frame, the first note recording operation is detected, then the current video frame is the 11th video frame in the video, and when the note completion event for the current video frame is detected, the terminal is playing the 20th video frame in the video, then it can be determined that the video frames played in the video during the process of drawing the electronic notes of the current video frame are the 11th-20th video frames, and therefore the 11th-20th video frames can be used to generate a video animation for presenting the current video frame.

[0084] In the embodiment of the present application, for a video that needs to be converted from image to text, the video can be played and a canvas mask can be displayed on the video playback area, so that the user (object) can enter a first note recording operation on the canvas mask to obtain an electronic note of the current video frame. When a text-to-image conversion operation for the video is detected, the image-to-text article corresponding to the video can be generated based on the target image-to-text obtained by image-to-text recognition of the video and the video frame with the electronic note and the corresponding electronic note. The image-to-text article finally displayed includes not only the target image-to-text, but also some video frames and corresponding electronic notes. This can enrich the content of the converted image-to-text article and improve the quality of the image-to-text article. In addition, by introducing the electronic notes recorded by the user into the image-to-text article, the image-to-text article can be made more valuable in information, and the user customization of the image-to-text article can be enhanced, thereby improving the user experience of converting video to image-to-text.

[0085] Based on the above Figure 2 The following is a description of the video processing method embodiment shown in FIG. Figure 4a The flowchart shown illustrates how to generate the graphic article corresponding to the video; see Figure 4a As shown, the method for generating a graphic article corresponding to a video includes the following steps S401-S403:

[0086] S401, obtaining target images and texts obtained by performing image and text recognition on a video.

[0087] From the relevant description of the aforementioned step S204, it can be seen that the target image and text include: the first text information obtained by performing text recognition on the audio data of the video, and key frame data; the key frame data includes: at least one key video frame, and the second text information obtained by performing text recognition on the corresponding key video frame.

[0088] in:

[0089] (1) The process of generating the first text information in the target image and text may include:

[0090] s11. Acquire audio data from the video. This audio data may be obtained by performing audio and video separation on the video using a multimedia processing framework. The multimedia processing framework here may be, for example, FFmpeg. FFmpeg is an open-source multimedia processing framework that provides a set of tools and libraries for processing audio, video, and other multimedia content. FFmpeg's key features are high customizability, cross-platform compatibility, and rich functionality. It can be used in various multimedia applications, such as video encoding, decoding, transcoding, streaming media servers, and real-time audio and video processing. s12. Perform text recognition on the audio data using an artificial intelligence-based speech recognition library to obtain audio text information. The speech recognition library here may be, for example, whisper, a large pre-trained speech model that integrates multilingual ASR (speech recognition technology), speech translation, and language recognition capabilities. s13. Use the audio text information as the first text information. Alternatively, invoke a first multimodal language model to perform content summarization processing on the audio text information to obtain the first text information. The first multimodal language model here may be a model obtained by fine-tuning (i.e., model training) a multimodal language model (MLLM).

[0091] The so-called multimodal language model refers to a language model that can simultaneously process multiple types of information such as text, pictures, audio and video, and a language model is a neural network model that can focus on processing language tasks (such as natural language generation, machine translation, text summarization, etc.). It is understandable that the multimodal language model can perform different operations based on different task description instructions, and any task description instruction can be a text based on question and answer format; based on this, the specific method of calling the first multimodal language model to summarize the content of the audio text information and obtain the first text information can be: generating a first task description instruction based on the audio text information, and the first task description instruction is used to instruct: performing content summary processing on the audio text information (i.e., key content extraction processing); calling the first multimodal language model to perform task processing based on the first task description instruction to obtain the first text information.

[0092] Regarding the above-mentioned process of generating the first text information, it should be noted that: ① The process of generating the first text information can be executed in the terminal or in the server, and there is no limitation on this. ② If the process of generating the first text information is executed in the server, the terminal may determine the video that needs to be converted from image to text, upload the video to the server, and the server may call the multimedia processing framework to separate the audio and video of the video to obtain audio data, thereby executing the process of generating the first text information (i.e., the above-mentioned steps s11-s13). Alternatively, the terminal may determine the video that needs to be converted from image to text, first call the multimedia processing framework to separate the audio and video of the video to obtain audio data, and then upload the first audio data to the server, so that the server executes the process of generating the first text information. ③ If the terminal is responsible for separating the audio and video of the video, the processing logic adopted by the terminal may be related to the specific type of the target application (the application used to play the video).

[0093] For example, the target application is a web application, that is, the video is played through a web application. Considering that ffmpeg contains an excellent C / C++ audio and video processing library, it can be packaged into the form of wasm+js through emscripten (a compiler) to support running in a web application. The wasm mentioned here is the abbreviation of WebAssemly, which is a binary code format that aims to compile high-level language (such as JavaScript (js), Python, etc.) code into a format that can be efficiently executed by web applications. It uses bytecode and stack machines to execute at a speed close to that of native code, which can improve the performance of web applications, reduce latency and improve scalability. It can be seen that wasm (binary code format) essentially refers to: a code format that supports execution by web applications and supports enhancing the performance of web applications. Based on this, the embodiment of the present application can use wasm+webworker (multi-threaded js) to run the wasm version of ffmpeg in a web application, so that the terminal can achieve audio and video separation by calling the ffmpeg. In this case, the terminal's processing logic may be: obtaining a first compilation result obtained by compiling a multimedia processing framework (ffmpeg) in binary code format, the first compilation result including a programming interface (API) for invoking the multimedia processing framework; the first compilation result may be a compilation result obtained by the terminal by compiling the multimedia processing framework in binary code format, or a pre-compiled version obtained by the terminal from an open source library, without limitation. The terminal may then, through the API in the first compilation result, invoke the multimedia processing framework to perform audio and video separation on the video, obtaining audio data and a set of video frames; and upload the audio data to a server, causing the server to execute the process of generating the first text information.

[0094] (2) The process of generating key frame data in the target graph and text may include:

[0095] s21, obtain N key video frames in the video, and the video text information of each key video frame in the N key video frames; wherein the N key video frames are arranged in the order of the playback time points, N is a positive integer, and the video text information of any key video frame is obtained by performing text recognition on the corresponding key video frame. The text recognition for the key video frames can be achieved by a character recognition tool, that is, by calling the character recognition tool to perform text recognition on the key video frames, the character recognition tool can be, for example, tesseract (an OCR (Optical Character Recognition) recognition tool based on LSTM (Long Short-Term Memory Recurrent Neural Network)). s22, the video text information of each key video frame in the N key video frames is used as the second text information of the corresponding key video frame, thereby constructing the key frame data using the second text information of the N key video frames. Alternatively, based on the video content of the N key video frames, the N key video frames are cut and processed to obtain multiple sub-screens; a sub-screen includes at least one key video frame, and the video content of each key video frame in any sub-screen is continuous. Then, the first multimodal language model is called to perform a text description of the corresponding sub-caption based on the video text information of the key video frame in each sub-caption, thereby obtaining the text description information of the corresponding sub-caption; a key video frame is selected from each sub-caption (which can be randomly selected or the kth key video frame is selected by default, where k is a positive integer (e.g., the value 1)), and the text description information of each sub-caption is used as the second text information of the key video frame selected from the corresponding sub-caption to obtain key frame data; wherein the key frame data includes: the key video frame selected from each sub-caption and the second text information of the corresponding key video frame. Through this embodiment, the second text information of the key video frame can be generated by comprehensively considering the video text information of multiple key video frames with continuous content, thereby improving the accuracy of the second text information.

[0096] Among them, based on the video content of N key video frames, a specific implementation method for cutting N key video frames to obtain multiple sub-captions can be: calling a video cutting model (such as a TransNetV2 model) to identify the cutting frames of the N key video frames based on the video content of the N key video frames, and obtain cutting frame indication information, where the cutting frame indication information is used to indicate: among the N key video frames, the key video frames to be used as cutting frames; based on the cutting frame indication information, cutting the N key video frames to obtain M sub-captions, and the cutting process includes: performing a cutting operation every time a cutting frame is detected. For example, see Figure 4bAs shown: suppose there are 10 key video frames in total (i.e., N=10), and the cutting frame indication information obtained by the video cutting model indicates that the key video frames to be used as cutting frames include: the 3rd key video frame and the 8th key video frame; then, the cutting operation can be performed when the 3rd key video frame and the 8th key video frame are detected, respectively, to obtain 3 sub-screens. The embodiment of the present application implements the cutting process by calling the video cutting model, which can improve the efficiency of the cutting process. Optionally, based on the video content of N key video frames, another specific implementation method of cutting the N key video frames to obtain multiple sub-screens can be: according to the video content of the N key video frames, respectively calculate the content similarity between any two adjacent key video frames; according to the size relationship between the content similarity between any two adjacent key video frames and the similarity threshold, detect the continuity between the video content of the corresponding two key video frames, and then cut the N key video frames based on the continuity detection result to obtain multiple sub-screens. When the content similarity between the two key video frames is greater than a similarity threshold, it is determined that continuity is detected between the video contents of the corresponding two key video frames; otherwise, it is determined that no continuity is detected.

[0097] In addition, similar to the aforementioned implementation method of obtaining the first text information through the first multimodal language model, when calling the first multimodal language model and performing a text description of the corresponding sub-caption based on the video text information of the key video frames in each sub-caption to obtain the text description information of the corresponding sub-caption, a second task description instruction can be generated based on the video text information of the key video frames in each sub-caption. The second task description instruction is used to instruct: by understanding the video text information of the key video frames in the sub-caption, perform a text description of the corresponding sub-caption and output the corresponding text description information; then, the first multimodal language model can be called to perform task processing based on the second task description instruction to obtain the text description information of each sub-caption.

[0098] Regarding the above-mentioned key frame data generation process, it should be noted that: ① The key frame data generation process can be executed in the terminal or in the server, and there is no limitation on this. ② If the key frame data generation process is executed in the server, the server can extract N key video frames and corresponding video text information based on the video frame set of the video, and then execute the key frame data generation process (i.e., the above steps s21-s22). Alternatively, the terminal can extract N key video frames and corresponding video text information based on the video frame set of the video, and then upload the N key video frames and corresponding video text information to the server, so that the server executes the key frame data generation process; by realizing the interception and text recognition of key video frames on the terminal side, it is possible to avoid relying on the server-side interface, reduce the cost of network transmission and video processing, and improve the data upload speed and parsing speed, thereby improving the user experience.

[0099] Among them, the processing logic of the terminal uploading N key video frames and corresponding video text information to the server based on the video frame set can be as follows: after obtaining the video frame set of the video, the key frame set is extracted from the video frame set, and the key frame set includes multiple key video frames, which can be extracted from the video frame set by calling the multimedia processing framework. Then, text recognition can be performed on each key video frame in the key frame set to obtain the video text information of the corresponding key video frame, and the multiple key video frames in the key frame set are used as N key video frames, so that the N key video frames and the corresponding video text information are uploaded to the server, so that the server executes the key frame data generation process. Alternatively, considering that the key frame set extracted by the multimedia processing framework may contain some repeated key video frames, then in order to save the network resources required for the subsequent transmission of key frame data and corresponding video text information, as well as the processing resources required for the subsequent server to generate key frame data, the terminal can upload the non-repeated key video frames and corresponding video text information in the key frame set to the server. Specifically:

[0100] The terminal can call the classification model to perform category labeling processing on each key video frame in the key frame set to obtain the label information of each key video frame. The label information of any key video frame may include one or more category labels, and a category label is used to indicate the category of an object element in the key video frame (such as a box, a person, etc.). According to the repetition rate between the label information of each key video frame in the key frame set, the repeated key video frames are filtered out in the key frame set to obtain the remaining N key video frames. Among them, for the i-th key video frame in the key frame set (i is a positive integer), if the repetition rate between the label information of the i-th key video frame and the label information of the i-1-th key video frame is greater than the repetition rate threshold, then the i-th key video frame can be determined to be a repeated key video frame. The repetition rate between any two pieces of label information can be the ratio between the number of identical category labels contained in the corresponding two pieces of label information and the number of category labels in one of the label information. For example, label information 1 contains four category labels, such as "blackboard," "table," "chair," and "student," and label information 2 contains four category labels, such as "blackboard," "table," "chair," and "teacher." Label information 1 and label information 2 contain three identical category labels, and the number of category labels in one of the label information is four. The repetition rate between label information 1 and label information 2 is 3 / 4 = 0.75. After obtaining N key video frames, each of the N key video frames can be subjected to image-text recognition to obtain video text information of the corresponding key video frame. The N key video frames and the corresponding video text information are then uploaded to a server, causing the server to execute a key frame data generation process.

[0101] S402, obtaining note content, where the note content is generated based on the video frame with the electronic note and the corresponding electronic note.

[0102] In a specific implementation, the terminal or server may integrate the video frame with the electronic note and the corresponding electronic note to obtain the note content; in this case, the note content includes: the video frame with the electronic note and the corresponding electronic note. Alternatively, the terminal or server may optimize the corresponding electronic note based on the video frame with the electronic note to improve the corresponding electronic note, thereby obtaining the note content; in this case, the note content includes: the video frame with the electronic note and the corresponding optimized electronic note.

[0103] Among them, for the terminal, the terminal can call the first multimodal language model to optimize the corresponding electronic notes based on the video frame with the electronic notes to obtain the note content. For the server, the server receives the video content and corresponding electronic notes of the video frame with the electronic notes uploaded by the terminal, and then calls the first multimodal language model to optimize the corresponding electronic notes based on the video frame with the electronic notes to obtain the note content. Alternatively, the server can call the first multimodal language model to optimize the note material to obtain the note content; the note material mentioned here can include at least one video frame screenshot uploaded by the terminal, and a video frame screenshot includes: the video content of a video frame with an electronic note and the corresponding electronic note. By uploading the video frame screenshot to the server to optimize the electronic note, the amount of data transmitted over the network can be reduced and network transmission resources can be saved.

[0104] It is understandable that if the note content is obtained by the server calling the first multimodal language model to optimize the note material, then after the terminal draws the electronic note displaying the current video frame in the canvas mask according to the first note recording operation, if a note completion event for the current video frame is detected, a video frame screenshot of the current video frame can be generated and uploaded to the server to facilitate the server to subsequently optimize the note material. Specifically, if the terminal detects a note completion event for the current video frame, it can call the canvas video loading interface (canvas.getContext('2d').drawImage(video, 0, 0, width, height) to load the video into the canvas mask; in the canvas mask, the video playback screen is adjusted to the current video frame, and the display level of the video in the canvas mask is adjusted. The adjusted display level is lower than the display level of the electronic notes of the current video frame in the canvas mask to avoid the current video frame obscuring the electronic notes, thereby ensuring the integrity of the electronic notes; then, the canvas screenshot interface (canvas.toDataURL) can be called to take a screenshot of the canvas mask to obtain a video frame screenshot of the current video frame, and the video frame screenshot of the current video frame is uploaded to the server.

[0105] Exemplarily, taking the note material as an example, the method of obtaining the note content through the first multimodal language model can be: based on each video frame screenshot in the note material, a third task description instruction is generated, and the third task description instruction is used to instruct: to understand the content of each video frame screenshot to generate content description information of the corresponding video frame screenshot; to call the first multimodal language model to perform task processing based on the third task description instruction to obtain the content description information of each video frame screenshot. Based on the content description information of each video frame screenshot and the corresponding electronic note, a fourth task description instruction is generated, and the fourth task description instruction is used to instruct: based on the content description information of each video frame screenshot and the corresponding electronic note, to improve and summarize the note content; then, the first multimodal language model can be called to perform task processing based on the fourth task description instruction to obtain the note content.

[0106] Based on the description of steps S401-S402 above, it can be seen that the embodiment of the present application pre-trains a first multimodal language model, and then uses this first multimodal language model to implement the following functions: summarizing the audio text information to obtain the first text information, providing text descriptions of the sub-captions to obtain text description information of the corresponding sub-captions, and optimizing the note materials to obtain the note content. In this regard, the following points need to be explained:

[0107] First point: During the model pre-training phase, corresponding question-and-answer samples can be pre-constructed for different video types. The question-and-answer samples corresponding to any video type include questions and answers. Questions may include sample task description instructions, which are consistent in form with any of the aforementioned task description instructions. The difference is that the sample task description instructions are generated based on sample data (such as sample audio and text information, sample video frame screenshots, etc.), while any of the aforementioned task description instructions are generated based on relevant data of the target video (such as audio and text information, note materials, etc.). Then, the corresponding question-and-answer samples can be used to train the multimodal language model under different video types to fine-tune the multimodal language model, thereby obtaining a fine-tuned multimodal language model for each video type. Furthermore, the number of multimodal language models can be one or more; when there are multiple multimodal language models, for any video type, the question and answer samples corresponding to the video type can be used to fine-tune each multimodal language model to obtain multiple candidate multimodal language models, and questions and corresponding answers of different dimensions can be used to measure the fine-tuning effect of each candidate multimodal language model, so as to select the candidate multimodal language model with the best fine-tuning effect as the fine-tuned multimodal language model for the corresponding video type.

[0108] In this case, the first multimodal language model mentioned above specifically refers to: a fine-tuned multimodal language model under the video type of the target video; in this way, the compatibility between the first multimodal language model and the target video can be enhanced, so that when the first multimodal language model performs corresponding processing based on the relevant data of the target video (such as audio text information, note materials, etc.), it can more accurately learn the characteristics of the relevant data, thereby improving the accuracy of the processing results. Optionally, in other embodiments, the video type can be ignored, and the multimodal language model can be trained by constructing question and answer samples, so that the trained multimodal language model can be directly used as the first multimodal language model.

[0109] Second point: In other embodiments, corresponding language models can be trained based on different functional requirements, so that corresponding language models can be called to obtain corresponding information; that is, a first language model for content summarization, a second language model for textual description of sub-captions, and a third language model for optimizing note materials can be trained respectively, so that the first language model is called to perform content summary processing on the audio text information to obtain the first text information, the second language model is called to perform textual description of the corresponding sub-caption according to the video text information of the key video frames in each sub-caption to obtain the text description information of the corresponding sub-caption, and the third language model is called to optimize the note materials to obtain the note content.

[0110] S403: Integrate the target image and text with the note content to obtain the image and text article corresponding to the video.

[0111] In a specific implementation, the target image and text and the note content can be spliced ​​to obtain the image and text article corresponding to the video. Alternatively, the target image and text and the note content can be integrated through a multimodal language model to obtain the image and text article corresponding to the video, so as to improve the readability and logic of the content of the image and text article. Specifically, a second multimodal language model with the ability to generate articles can be determined. Using the target image and text and the note content, a target task description instruction of the second multimodal language model is constructed, and the target task description instruction is used to indicate: perform article generation processing based on the target image and text and the note content; then, call the second multimodal language model to perform task processing based on the target task description instruction to obtain the image and text article corresponding to the video. Among them, the second multimodal language model mentioned here and the first multimodal language model mentioned above can be the same model or different models, and there is no limitation on this.

[0112] The embodiment of the present application can generate target graphics and text by combining audio-to-text conversion and text recognition of key video frames, which can effectively improve the accuracy of the target graphics and text. In addition, by calling a multimodal language model to integrate the target graphics and text with the note content, it can quickly and automatically generate the corresponding graphics and text articles for the video, improving the efficiency and accuracy of the graphics and text article generation.

[0113] Based on the relevant descriptions of the above-mentioned various method embodiments, the present application embodiment proposes another video processing method; the video processing method can be executed by the terminal, or by the terminal and the server together, without limitation. Figure 5 As shown, the video processing method may include the following steps S501-S505:

[0114] S501, playing a video that needs to be converted from image to text, and displaying video summary information corresponding to the video during the video playback process.

[0115] In a specific implementation, after determining the video that needs to be converted from image to text, the video summary information corresponding to the video can be obtained, so that during the playback of the video, the video summary information corresponding to the video can be displayed in a preset area on the display screen. It should be noted that the terminal can start displaying the video summary information corresponding to the video in the preset area when the video starts playing; or the terminal can also detect the operation for triggering the display of the video summary information after starting to play the video, and after detecting the corresponding operation, start displaying the video summary information corresponding to the video in the preset area, and there is no limitation on this. Among them, the preset area can be any area of ​​the display screen other than the video playback area, such as the area on the display screen to the left of the video playback area, or the area on the display screen to the right of the video playback area, and so on.

[0116] The video summary information corresponding to the video may include: frame description information of at least one key video frame in the video; the frame description information of any key video frame may include at least one of the following: the playback time point of the corresponding key video frame, text information related to the video content of the corresponding key video frame, etc. Optionally, the video summary information may also include video introduction information, which refers to text information used to indicate the video content and theme of the entire video; for example, for a tutorial video on language A algorithm, its corresponding video introduction information may be, for example, "This is a class on language A algorithm. The video explains in detail the process of the specified operation, including defining pointers, comparing elements, moving pointers, etc.; at the same time, the video also provides examples of code implementation."

[0117] Exemplarily, the video summary information may be generated in any of the following ways:

[0118] Method 1: Identify key video frames in a video to determine the playback time point of at least one key video frame; and call a pre-trained content understanding model to perform content understanding on the video content of each key video frame to obtain text information related to the video content of the corresponding key video frame, thereby using the playback time point of each key video frame and the text information related to the video content of the corresponding key video frame to generate frame description information for each key video frame. In addition, the content understanding model can also be called to perform content understanding on the video content of other video frames in the video except the key video frames to obtain text information related to the video content of other video frames, thereby summarizing the content of each text information obtained by the content understanding model to obtain video introduction information. After obtaining the video introduction information and the frame description information of each key video frame, the frame description information and video introduction information of each key video frame can be used to generate video summary information corresponding to the video.

[0119] Method 2: Target graphics obtained by performing graphics and text recognition on a video can be obtained, and then video summary information corresponding to the video can be generated based on the target graphics. The target graphics include: first text information obtained by performing text recognition on the audio data of the video, and key frame data; and the key frame data includes: at least one key video frame, and second text information obtained by performing text recognition on the corresponding key video frame. Accordingly: ① The text information related to the video content of any key video frame in the video summary information can be generated based on the second text information corresponding to the corresponding key video frame contained in the target graphics; for example, it can be generated by copying the second text information corresponding to the corresponding key video frame, or by summarizing or simplifying the second text information corresponding to the corresponding key video frame. ② The video introduction information in the video summary information can be generated based on at least one text information in the target graphics (such as at least one of the first text information and the second text information corresponding to each key video frame); for example, it can be generated by combining at least one text information in the target graphics, or by calling a summary generator to perform summary generation based on at least one text information in the target graphics, etc.

[0120] Optionally, after displaying the video summary information, the terminal may also support the user to modify the video summary information; specifically, the terminal may detect modification operations on the video summary information, and the modification operations may include but are not limited to at least one of the following: deleting at least one frame description information, adding at least one frame description information, modifying at least one frame description information, and modifying the video introduction information, etc. If a modification operation on the video summary information is detected, the terminal may update the displayed video summary information in response to the modification operation on the video summary information. Taking the example that the video summary information includes video introduction information 60 and frame description information of multiple key video frames, and the modification operation includes deleting frame description information 61, a schematic diagram of the terminal updating the displayed video summary information can be seen in Figure 6a shown.

[0121] Furthermore, if the video summary information is generated based on the target image and text, the terminal can also update the target image and text based on the updated video summary information. For example, the key frame data in the target image and text can be updated based on the text information related to the video content of the key video frame contained in the updated video summary information, so that the key video frames involved in the updated key frame data are consistent with the key video frames involved in the updated video summary information; that is, relative to the original video summary information, if at least one frame description information is deleted / added / modified in the updated video summary information, then in the key frame data contained in the target image and text, the key video frames and related data corresponding to the corresponding frame description information are deleted / added / modified. For another example, the corresponding text information in the target image and text can be updated based on the video introduction information in the updated video summary information.

[0122] S502: Display a canvas mask on the video playback area of ​​the video, and draw an electronic note displaying the current video frame in the canvas mask according to a first note recording operation detected on the canvas mask.

[0123] The current video frame refers to the video frame displayed in the video playback area when the first note recording operation is detected. In a specific implementation, the specific implementation of step S502 can refer to the relevant description of steps S202-S203 in the above method embodiment, which will not be repeated here.

[0124] Furthermore, considering that there is at least one first note recording operation, a first note recording operation is used to draw an electronic note of the current video frame, and after the user performs a first note recording operation, there may be a situation where the electronic note generated by the first note recording operation is canceled. Therefore, in order to avoid accidentally canceling the electronic notes generated by other first note recording operations, thereby ensuring the accuracy of the electronic notes finally obtained, each time the terminal detects a first note recording operation, the operation indication information of the currently detected first note recording operation (i.e., the information used to indicate the first note recording operation) can be added to the history record queue, so as to record the various first note recording operations performed by the user in sequence through the history record queue; if a note cancellation operation is detected, the last operation indication information in the history record queue is deleted, and the electronic note drawn based on the last first note recording operation is deleted in the canvas mask. The last note recording operation refers to: the first note recording operation corresponding to the last operation indication information.

[0125] If a note completion event is detected for the current video frame, note record information is generated based on the current video frame and the various electronic notes currently displayed in the canvas mask, and the generated note record information is added to the note list. For example, when the terminal displays video summary information, a schematic diagram of further displaying the note list can be seen in Figure 6b In addition, you can clear the canvas overlay; when clearing the canvas overlay, the terminal also clears the history queue.

[0126] S503: When frame description information is selected in the video summary information, jump from the currently displayed video frame to the second video frame in the video play area.

[0127] In a specific implementation, after the terminal displays the video summary information, it can detect whether frame description information is selected in the video summary information; when frame description information is selected in the video summary information, the terminal can determine the second video frame based on the selected frame description information and jump from the currently displayed video frame to the second video frame in the video playback area. The second video frame refers to the video frame corresponding to the selected frame description information. Therefore, this embodiment of the present application can support users to select frame description information in the video summary information according to their own needs, thereby triggering the terminal to quickly jump to the corresponding video frame, thereby improving the viewing efficiency of video frames.

[0128] S504: If a second note recording operation is detected on the canvas mask when the second video frame is displayed, an electronic note displaying the second video frame is drawn on the canvas mask according to the second note recording operation.

[0129] Among them, the number of second note recording operations can be one or more; the specific implementation of drawing and displaying the electronic note of the second video frame in the canvas mask according to the second note recording operation is similar to the specific implementation of drawing and displaying the electronic note of the current video frame in the canvas mask according to the first note recording operation mentioned above, and will not be described in detail here. It is understandable that in actual application, before the terminal jumps to display the second video frame in the video playback area, the canvas mask may be displayed on the video playback area, or the canvas mask may not be displayed. In the case where the canvas mask is not displayed, the terminal can synchronously display the canvas mask on the video playback area when jumping to display the second video frame; or it can detect the note trigger operation after jumping to display the second video frame, so that the canvas mask is displayed on the video playback area after detecting the note trigger operation.

[0130] For example, see Figure 6c As shown, a user can trigger the terminal to jump from the currently displayed video frame to a second video frame in the video playback area by selecting frame description information 62. At this time, the canvas overlay is not displayed on the video playback area. Then, the user can trigger an operation by inputting a note, causing the terminal to display a canvas overlay on the video playback area. Assuming that the user inputs two second note recording operations on the canvas overlay, namely, a graphic mark operation and a text input operation for at least one video content, the terminal can then draw and display two electronic notes, namely, a graphic mark 63 (assuming an underline) and text 64 (assuming "Key Point!"), on the canvas overlay.

[0131] S505 , in response to the image-text conversion operation on the video, displaying the image-text article corresponding to the video.

[0132] Based on the aforementioned Figure 4a As can be seen from the description of the embodiment of the method shown, the graphic article corresponding to the video can be obtained by integrating the target graphic and note content. The specific generation method can be found in the aforementioned Figure 4a However, it should be understood that if the target image and text are updated before executing step S505, then the above Figure 4a The target image and text used in the illustrated method embodiment refers to the updated target image and text.

[0133] In a specific implementation, the graphic article corresponding to the video can be displayed in a document viewing interface or in a document editing interface, and there is no limitation on this. If the graphic article is displayed in the article editing interface, the terminal can also obtain the user's article editing operation and perform corresponding editing processing on the graphic article according to the article editing operation. Among them, the article editing operation may include at least one of the following: an operation of manually editing the article and an operation of automatically editing the article. The so-called manual editing operation of the article refers to the operation of the user manually modifying the article, which may include but is not limited to: the operation of modifying any text information (such as the title, text information in the text), and the operation of replacing, deleting or adding pictures in the article. Correspondingly, the operation of automatically editing the article refers to the operation used to trigger the terminal to modify the graphic article based on AI technology, which may include but is not limited to: a first operation for triggering title rewriting based on AI technology, a second operation for triggering full-text polishing based on AI technology, and a third operation for triggering text error correction based on AI technology, and so on.

[0134] Among them, if the article editing operation includes a first operation, the specific way in which the terminal performs corresponding editing processing on the graphic article according to the article editing operation may include: calling a title rewriting model generated based on AI technology to rewrite the title of the graphic article, obtaining a rewritten title, and updating the title displayed in the document editing interface to the rewritten title. If the article editing operation includes a second operation, the specific way in which the terminal performs corresponding editing processing on the graphic article according to the article editing operation may include: calling an article polishing model generated based on AI technology to polish the entire text of the graphic article, obtaining a polished graphic article, and using the polished graphic article to update the graphic article displayed in the document editing interface. If the article editing operation includes a third operation, the specific way in which the terminal performs corresponding editing processing on the graphic article according to the article editing operation may include: calling a text error correction model generated based on AI technology to perform text error correction processing on the graphic article, obtaining a corrected graphic article, and using the corrected graphic article to update the graphic article displayed in the document editing interface.

[0135] Optionally, to facilitate the user to input the operation of automatically editing the article, the terminal can provide the following components in the document editing interface: title rewriting component 651, full text polishing component 652 and text error correction component 653, such as Figure 6dIn this case, the first operation mentioned above can be a triggering operation (such as a click operation, a press operation, etc.) for the title rewriting component, the second operation mentioned above can be a triggering operation (such as a click operation, a press operation, etc.) for the full text polishing component, and the first operation mentioned above can be a triggering operation (such as a click operation, a press operation, etc.) for the text error correction component. Of course, in other embodiments, the first, second, and third operations mentioned above can also be gesture operations or operations of inputting voice commands, etc.

[0136] Optionally, the document editing interface may also include: article publishing components (such as Figure 6d If the terminal detects a trigger operation for the article publishing component, it can display the platform logo of at least one content sharing platform (or new media platform). After the user selects at least one platform logo, the text and image article currently displayed in the document editing interface is published (i.e., forwarded) to the content sharing platform indicated by the selected platform logo, so that the user can forward the text and image article to various content sharing platforms with one click.

[0137] Optionally, the document editing interface may also include: article export components (such as Figure 6d If the terminal detects a trigger operation for the article export component, it can export the graphic article currently displayed in the document editing interface, and save the article identifier 66 of the graphic article to be exported to the conversion record, so that the user can view the corresponding graphic article by selecting the article identifier in the conversion record later, and support the user to display the platform identifier 67 of at least one content sharing platform in response to the user's article publishing operation after viewing the graphic article, so as to support the user to publish the graphic article at any time, such as Figure 6e shown.

[0138] It should be noted that when the terminal exports a graphic article, it can export the entire content of the graphic article by default; or, in order to improve the flexibility of article export and to export the high-quality graphic content required by the user with maximum flexibility, the terminal can also support the user to customize the export configuration of the graphic article, so as to set it to export only the video frames with electronic marks and recorded electronic notes in the graphic article, or export the entire content of the graphic article, etc. In this case, when the terminal exports the graphic article, it can determine the target content to be exported from the graphic article according to the custom configuration, and thus export the target content.

[0139] In the embodiment of the present application, for a video that needs to be converted from image to text, the video can be played and a canvas mask can be displayed on the video playback area, so that the user (object) can draw an electronic note of at least one video frame on the canvas mask. When a text-to-image conversion operation for the video is detected, the target image and text obtained by image-to-text recognition of the video and the video frame with the electronic note and the corresponding electronic note can be used to generate a graphic article corresponding to the video. The graphic article finally displayed includes not only the target image and text, but also some video frames and corresponding electronic notes. This can enrich the content of the converted graphic article and improve the quality of the graphic article. In addition, by introducing the electronic notes recorded by the user into the graphic article, the graphic article can be made more valuable in information, and the user customization of the graphic article can be enhanced, thereby improving the user experience of converting video to image.

[0140] Based on the relevant descriptions of the above-mentioned method embodiments, the embodiment of the present application also proposes a video-to-article method based on wasm, ffmpeg and AI; this method can realize the extraction and processing of video frames in web applications, and by combining audio-to-text, OCR image recognition text and multimodal language models and other technologies, it can generate graphic articles with high-quality graphic content, and can support one-click local editing and sharing (publishing) of the graphic article.

[0141] See also Figure 7aAs shown, the implementation logic of the video-to-article method is roughly as follows: after the user enters the workbench of the video-to-text assistant webpage and chooses to upload a local video or paste a video address, the terminal can determine whether the video to be converted from text to image is a local video. If so, the local video uploaded by the user can be obtained as the video to be converted from text to image. If not, the online video can be parsed and downloaded based on the video address pasted by the user to obtain the video to be converted from text to image. Then, the terminal can display a transparent operation mask of the webpage (i.e., a transparent canvas mask) on the playback area of ​​the video to support the user to input at least one note recording operation for at least one video frame on the canvas mask, thereby drawing an electronic note of the corresponding video frame according to the note recording operation input by the user. After playing the video, the terminal can determine whether the text-to-image article of the video needs to be exported by detecting whether there is a text-to-image conversion operation for the video. If a text-to-image conversion operation is detected, the text-to-image article of the video to be exported can be determined. At this time, the user's operation points (i.e., each video frame where the user has performed a note-taking operation) can be traversed to obtain the note content; and the key video frames can be subjected to text recognition through OCR to obtain key frame data, and the audio-to-text processing of the video (i.e., text recognition of the audio data of the video) can be performed to obtain the first text information. Based on the obtained note content, key frame data, and first text information, the text-to-image article corresponding to the video can be generated. Optionally, the generated text-to-image article can also be rewritten in title or polished in full.

[0142] The specific implementation of the above-mentioned video-to-article method mainly involves implementing a real-time tagging video player in a web application, using wasm technology to pre-process the video using an open source multimedia processing framework (including an open source video processing library), an OCR library, and tensorflow.js (a library for machine learning development using JavaScript), and using various models derived from AI technology to understand, organize, and polish the video content, thereby generating high-quality graphic articles that allow users to quickly and easily understand the video content and absorb the essence of the video. The following is a detailed explanation of the relevant technologies involved in this video-to-article method:

[0143] (1) Video player with real-time tagging:

[0144] The video-to-article method proposed in the embodiment of the present application adds a transparent canvas mask to the video player to display the canvas mask on the video playback area, thereby supporting users to quickly and conveniently perform note-taking operations such as graphic marking and commenting (i.e., entering text) on at least one video frame through the canvas mask while watching the video. It also supports the real-time generation of corresponding note-taking information after detecting the note completion event for the video frame, and adds the generated note-taking information to the note list for easy tracing. Figure 7b As shown, the specific implementation process of the processing logic is as follows:

[0145] 1. Use fabricjs (a canvas tool library) to add a transparent canvas mask to the video playback area of ​​the video player for drawing electronic notes on the video frames. At the same time, it provides common marking tools such as brushes, lines, and arrows for users to draw directly on the canvas mask. A note list can be displayed on the right side of the video playback area.

[0146] 2. After loading the video, read the attributes of the video tag (video tag) to obtain the current video's width, height, play time and other attribute information for video playback;

[0147] 3. Register the monitoring event of the canvas mask. This event is used to monitor the user's note-taking operations on the canvas mask, such as text input and graphic markup. When a note-taking operation is detected, the player interface video.pause() can be called to automatically pause the video.

[0148] 4. After the user completes tagging and commenting (that is, completing the note taking for a certain video frame):

[0149] a. Export the user's electronic notes and the playback time of the corresponding video frames through the toJSON interface (data export interface) of fabricjs. Generate note record information based on the exported data and add it to the note list, so that users can intuitively view their operation history;

[0150] b. Load the video into the canvas mask through canvas.getContext('2d').drawImage(video,0,0,width,height) (video loading interface), automatically adjust the video playback screen to the current video frame, and adjust the video display level to the lowest display level of the canvas mask. Use canvas.toDataURL (screenshot interface) to obtain a video frame screenshot of the current video frame marked by the user, and upload the video frame screenshot to the server to be used as the basic material for subsequent AI-generated graphic articles;

[0151] c. Clear the canvas mask and call the player interface video.play() to continue playing the video;

[0152] 5. When the user clicks on a note record in the note list, the video is automatically positioned (i.e., jumped) to the video frame corresponding to the clicked note record by setting video.currentTime (the value of the video playback time), and the corresponding electronic note is loaded based on the corresponding note record information. The loaded electronic note is reproduced on the video player, thereby reproducing the corresponding electronic note on the corresponding video frame.

[0153] (2) Video pre-processing:

[0154] A video is composed of many continuous video frames. The video-to-article method proposed in the embodiment of the present application focuses on understanding the video content and generating corresponding graphic articles. Therefore, it has only the minimum requirements for the quality of the video itself transmitted to the server. Before understanding the image content, the video can be pre-processed to process the original video provided by the user into the basic material required for understanding the image content, thereby reducing the cost of network transmission and video processing and improving the user experience. The video-to-article method proposed in the embodiment of the present application can use wasm+webworker to run ffmpeg in a web application, thereby realizing the pre-processing of the video based on ffmpeg, so that the pre-processing process does not need to rely on the server capabilities and does not block the main thread of the web application. For details, see Figure 7c As shown, the specific implementation process of the processing logic is as follows:

[0155] 1. Compile the wasm version of ffmpeg or use the compiled version of the open source library to obtain the wasm version of ffmpeg program interface (API) that can run in the web application for audio and video separation and video frame extraction;

[0156] 2. Compile the wasm version of Tesseract or use the compiled version of the open source library to obtain a text interface that can run on the web application for extracting text information from the video frame;

[0157] 3. Use the API to call the wasm version of ffmpeg to perform audio and video separation and key frame extraction on the video to obtain audio data and key frame sets. The audio data is uploaded to the server as the basic material for subsequent speech recognition.

[0158] 4. Call the classification model (such as a trained TensorFlow model or an open source mature deep learning algorithm model) to perform category labeling on each key video frame in the key frame set to obtain the label information of each key video frame. Compare the label information of each key video frame in the key frame set to obtain the repetition rate between each label information. Then, based on the repetition rate between each label information, filter out the repeated key video frames in the key frame set to obtain the remaining N key video frames;

[0159] 5. Call tesseract through the text interface to perform text recognition on N key video frames, obtain the text information appearing in each key video, such as subtitles, text in PPT, etc., and upload it to the server as video materials for subsequent image content understanding; specifically, N video frames can be merged into video clips with lower picture quality and frame number using ffmpeg, but containing key information (such as text information appearing in the identified key video frames), and uploaded to the server as video materials for subsequent image content understanding.

[0160] (3) Material understanding (including audio data processing and image content understanding):

[0161] The video-to-article method proposed in the embodiment of the present application can identify the continuous content of N key video frames of the video to determine the cutting frames, and then quickly cut the N key video frames into multiple sub-scenes (i.e., cut video sub-segments) based on the cutting frames, and then combine speech recognition, multimodal language model and other technologies to interpret and analyze the various materials mentioned in the above steps to generate graphic articles. For details, see Figure 7d As shown, the specific implementation process of the processing logic is as follows:

[0162] 1. Use the OpenAI open-source speech recognition library whisper to build a local speech recognition service and perform text recognition on the audio data to convert the audio data in the original video into text and obtain audio and text information for subsequent processing and use in generating graphic articles;

[0163] 2. Use the TransNetV2 model (video segmentation model) to build a local segmentation key information recognition service. This service identifies the segmentation frames of N key video frames in the video material and returns segmentation frame indication information. Use ffmpeg to divide the N key video frames in the video material into multiple sub-segments with continuous content based on the segmentation frame indication information, so as to facilitate the subsequent overall understanding of the video content.

[0164] 3. Fine-tune the established multimodal language model (MLLM) to varying degrees for various pre-defined video types, such as popular science tutorials, ghost videos, dance, and documentaries. This involves pre-constructing question-and-answer samples and using them for alignment training across different video types. The effectiveness of the fine-tuned model is measured by the quality of answers to questions across different dimensions, ultimately selecting and training a fine-tuned multimodal language model suitable for each video type.

[0165] 4. According to the video type of the video selected by the user, the fine-tuned multimodal language model under the corresponding video type is applied as the first multimodal language model. The audio text information, each sub-caption and the video frame screenshot corresponding to each key video frame are respectively input into the first multimodal language model, so as to call the first multimodal language model to perform the following processing: ① Summarize the content of the audio text information to obtain the first text information; ② According to the video text information of the key video frame in each sub-caption, the corresponding sub-caption is described in text to obtain the text description information of the corresponding sub-caption, and a key video frame is selected from each sub-caption respectively, and the text description information of each sub-caption is used as the second text information of the key video frame selected from the corresponding sub-caption to obtain the key frame data (including the key video frame selected from each sub-caption and the second text information of the corresponding key video frame); ③ Optimize the note material composed of the screenshots of each video frame uploaded by the terminal to obtain the note content;

[0166] 5. Infer the first text information, keyframe data, and note content to obtain a graphic article. Specifically, natural language processing technology can be used to call the second multimodal language model to deeply understand the content in the relevant data, analyze the theme, scene, characters, and other information of the picture, and generate high-quality graphic notes for users to download and read.

[0167] Based on the above description, the video-to-article method based on wasm, ffmpeg, and AI proposed in the embodiment of the present application can have the following beneficial effects:

[0168] ① Based on wasm and ffmpeg technology, video frame stream interception and processing can be implemented on the web application side, avoiding dependence on the server-side interface, improving parsing speed and user experience; and the interception and generation of notes of user-marked video frames can be processed through worker multi-threading. By supporting background parsing of videos while users are watching videos, it can greatly save time and lower the threshold of use.

[0169] ② Design a transparent operation mask for the web page (i.e., a transparent canvas mask) so that users can enter notes and record operations while watching a video without having to manually pause the video. This simplifies the operation process and effectively improves the user experience. The generated graphic articles contain the user's recorded electronic notes and the corresponding video frames, which can improve the article quality and enhance the user customization of the graphic articles.

[0170] ③ Combining technologies such as audio-to-text and OCR image recognition to quickly extract high-quality graphic content from videos, reducing the user's editing workload; and automatically editing graphic articles based on AI technology, further improving the quality of graphic articles and the convenience of article editing.

[0171] ④ Provide custom configuration function to support users to personalize and customize the generated graphic articles, including online editing of text information, pictures, etc., to meet the needs of different users, so that users can adjust the operation of intercepting graphic articles according to their needs and achieve more accurate graphic generation; and built-in AI technology to realize AI-assisted optimization of graphic articles, automatic title generation and full-text rewriting of graphic articles, etc., to improve the quality and readability of graphic articles, and further reduce the editing burden of users.

[0172] ⑤ Make full use of the web application interface to realize quick forwarding, uploading, document linkage import and other operations of graphic articles, thereby improving operational convenience; and support simultaneous publishing on multiple platforms to achieve one-stop management, improve the publishing efficiency of graphic articles, and save users' operation time between different content sharing platforms.

[0173] ⑥ It has a wide range of application scenarios, that is, it can be applied to various scenarios, such as game advertisements, learning tutorials, speeches, etc., and has strong practicality and universality. For example:

[0174] a. In the context of game advertising, corresponding graphic articles can be quickly generated through local web applications. For example, while watching a game advertisement video, users can quickly select at least one video frame from the video and selectively enter auxiliary information (such as electronic notes), thereby quickly generating corresponding graphic articles based on AI technology.

[0175] b. When watching a PPT or other tutorial video, a corresponding graphic article can be quickly generated through a local web application. For example, the local web application can load and play the video, and allow users to quickly navigate to the explanation PPT with a single click while watching the video. By comparing the changes in the video frames and the content before and after, the appropriate video frames can be captured and exported as an outline. The corresponding audio section can be converted into text to form a note segment.

[0176] c. When rehearsing a PPT speech, record the scene to obtain a video, and then quickly generate a graphic article corresponding to the video as a speech manuscript through a local web application. If the local web application loads the video, the speech manuscript can be quickly generated through voice-to-text conversion and video frame export. The speech manuscript can also be quickly imported into an online PPT document.

[0177] In summary, the video-to-article method proposed in the embodiment of the present application can provide users with efficient and convenient video-to-article services, optimize user experience, and play an important role in various application scenarios.

[0178] Based on the description of the above video processing method embodiment, the present application also discloses a video processing device; the video processing device can be a computer program (including one or more instructions) running on a computer device, and the video processing device can execute each step in the above method flow. Figure 8 , the video processing device can run the following units:

[0179] Playback unit 801, used for playing the video that needs to be converted from image to text;

[0180] The processing unit 802 is configured to display a canvas mask on the video playback area of ​​the video;

[0181] The processing unit 802 is further configured to draw an electronic note displaying a current video frame in the canvas mask according to a first note recording operation detected on the canvas mask; the current video frame refers to the video frame displayed in the video playback area when the first note recording operation is detected;

[0182] The processing unit 802 is also used to display the graphic article corresponding to the video in response to the graphic-text conversion operation for the video; the graphic article includes: the target graphic obtained by performing graphic-text recognition on the video, the video frame with electronic notes and the corresponding electronic notes.

[0183] In one embodiment, the processing unit 802 may further be configured to:

[0184] During playback of the video, displaying a note-taking component;

[0185] In response to a triggering operation on the note-recording component, the video is paused, and the step of displaying a canvas mask on the video playing area of ​​the video is triggered.

[0186] In another implementation, the processing unit 802 may further be configured to:

[0187] During the process of displaying the canvas overlay on the video playback area, displaying a toolbar on the display screen, wherein the toolbar includes one or more note-taking tools;

[0188] A corresponding note-recording tool is selected according to a tool selection operation, and an operation input by an object on the canvas mask through the selected note-recording tool is regarded as a first note-recording operation detected on the canvas mask.

[0189] In another implementation, the processing unit 802 may further be configured to:

[0190] If a note completion event for the current video frame is detected, generating note record information according to the current video frame and each electronic note currently displayed in the canvas mask;

[0191] The generated note record information is added to the note list, and the canvas mask is cleared; wherein the note list is located in an area of ​​the display screen except the video playback area.

[0192] In another embodiment, when the canvas mask is displayed on the video playback area, the video is paused; accordingly, the processing unit 802 may further be configured to:

[0193] In response to a note completion event for the current video frame, canceling display of the canvas mask;

[0194] In the video playback area, the video continues to be played based on the current video frame.

[0195] In another implementation, the processing unit 802 may further be configured to:

[0196] When a note record information is selected in the note list, in the video playback area, the video frame currently displayed is jumped to the first video frame; the first video frame refers to the video frame corresponding to the triggered note record information;

[0197] According to the triggered note recording information, the corresponding electronic note is reproduced on the first video frame.

[0198] In another embodiment, the number of the first note recording operation is at least one, and one first note recording operation is used to draw an electronic note of the current video frame; accordingly, the processing unit 802 may further be configured to:

[0199] Whenever a first note recording operation is detected, operation instruction information of the currently detected first note recording operation is added to the history record queue;

[0200] If a note undo operation is detected, deleting the last operation instruction information in the history record queue, and deleting the electronic note drawn based on the last first note recording operation in the canvas mask; wherein the last note recording operation refers to: the first note recording operation corresponding to the last operation instruction information;

[0201] When the canvas mask is cleared, the history record queue is cleared.

[0202] In another implementation, the processing unit 802 may further be configured to:

[0203] During playback of the video, displaying video summary information corresponding to the video; the video summary information includes: frame description information of at least one key video frame in the video;

[0204] When frame description information is selected in the video summary information, in the video playback area, the video frame currently displayed is jumped to a second video frame; the second video frame is the video frame corresponding to the selected frame description information;

[0205] If a second note recording operation is detected on the canvas mask when the second video frame is displayed, an electronic note displaying the second video frame is drawn in the canvas mask according to the second note recording operation.

[0206] In another embodiment, the video summary information is generated based on the target image and text; accordingly, the processing unit 802 may further be configured to:

[0207] In response to a modification operation on the video summary information, updating and displaying the video summary information;

[0208] Based on the updated video summary information, the target image and text are updated.

[0209] In another implementation, the processing unit 802 may further be configured to:

[0210] Obtaining target text and image obtained by performing text and image recognition on the video, wherein the target text and image include: first text information obtained by performing text recognition on audio data of the video, and key frame data; the key frame data includes: at least one key video frame, and second text information obtained by performing text recognition on the corresponding key video frame;

[0211] Obtaining note content, where the note content is generated based on the video frame with the electronic note and the corresponding electronic note;

[0212] The target image and text and the note content are integrated to obtain the image and text article corresponding to the video.

[0213] In another implementation, the processing unit 802 may further be configured to:

[0214] Acquire audio data of the video, where the audio data is obtained by calling a multimedia processing framework to perform audio and video separation on the video;

[0215] Using a speech recognition library based on artificial intelligence technology to perform text recognition on the audio data to obtain audio text information;

[0216] The first multimodal language model is called to perform content summarization processing on the audio text information to obtain first text information.

[0217] In another embodiment, the generation process of the first text information is executed in the server, and the video is played through a web application; accordingly, the processing unit 802 may also be configured to:

[0218] Obtaining a first compilation result obtained by compiling a binary code format version of the multimedia processing framework, wherein the first compilation result includes a program interface for invoking the multimedia processing framework; the binary code format is a code format that supports execution by the web application and enhances the performance of the web application;

[0219] Calling the multimedia processing framework to perform audio and video separation on the video through the program interface in the first compilation result to obtain audio data and a video frame set of the video;

[0220] The audio data is uploaded to a server, so that the server executes a process of generating the first text information.

[0221] In another implementation, the processing unit 802 may further be configured to:

[0222] Obtaining N key video frames in the video and video text information of each of the N key video frames; the N key video frames are arranged in chronological order according to playback time points, where N is a positive integer, and the video text information of any key video frame is obtained by performing text recognition on the corresponding key video frame;

[0223] Based on the video contents of the N key video frames, the N key video frames are segmented to obtain a plurality of sub-captions; each sub-caption includes at least one key video frame, and the video contents of the key video frames in any sub-caption are continuous;

[0224] Calling the first multimodal language model to perform text description on the corresponding sub-caption according to the video text information of the key video frame in each sub-caption, thereby obtaining text description information of the corresponding sub-caption;

[0225] A key video frame is selected from each sub-caption, and textual description information of each sub-caption is used as second textual information of the key video frame selected from the corresponding sub-caption, thereby obtaining key frame data; wherein the key frame data includes: the key video frame selected from each sub-caption and the second textual information of the corresponding key video frame.

[0226] In another embodiment, when the processing unit 802 is configured to segment the N key video frames based on the video contents of the N key video frames to obtain a plurality of sub-scenes, it may be specifically configured to:

[0227] The video segmentation model is called to perform segmentation frame identification on the N key video frames based on the video contents of the N key video frames to obtain segmentation frame indication information; the segmentation frame indication information is used to indicate: a key video frame among the N key video frames to be used as a segmentation frame;

[0228] The N key video frames are cut based on the cut frame indication information to obtain M sub-scenes; wherein the cut includes: performing a cut operation once each time a cut frame is detected.

[0229] In another embodiment, the key frame data generation process is executed in the server; accordingly, the processing unit 802 may also be configured to:

[0230] After obtaining a video frame set of the video, extracting a key frame set from the video frame set, wherein the key frame set includes a plurality of key video frames;

[0231] Calling the classification model to perform category labeling processing on each key video frame in the key frame set to obtain label information of each key video frame;

[0232] filtering out repeated key video frames in the key frame set according to a repetition rate between label information of each key video frame in the key frame set to obtain the remaining N key video frames;

[0233] Performing text recognition on each of the N key video frames to obtain video text information of the corresponding key video frame;

[0234] The N key video frames and corresponding video text information are uploaded to a server, so that the server executes a process of generating the key frame data.

[0235] In another embodiment, the note content is obtained by the server invoking a first multimodal language model to optimize the note material, and the note material includes at least one video frame screenshot uploaded by the terminal; a video frame screenshot includes: video content of a video frame with an electronic note and a corresponding electronic note; accordingly, the processing unit 802 can also be used to:

[0236] If a note completion event for the current video frame is detected, a canvas video loading interface is called to load the video into the canvas mask;

[0237] Adjusting the playback screen of the video to the current video frame in the canvas mask, and adjusting the display level of the video in the canvas mask, wherein the adjusted display level is lower than the display level of the electronic note of the current video frame in the canvas mask;

[0238] Calling a canvas screenshot interface to perform a screenshot process on the canvas mask to obtain a video frame screenshot of the current video frame;

[0239] Upload the video frame screenshot of the current video frame to the server.

[0240] In another embodiment, when the processing unit 802 is used to integrate the target image and text with the note content to obtain the image and text article corresponding to the video, it can be specifically used to:

[0241] Determine a second multimodal language model with article generation capabilities;

[0242] Using the target image and text and the note content, construct a target task description instruction of the second multimodal language model; wherein the target task description instruction is used to instruct: to perform article generation processing based on the target image and text and the note content;

[0243] The second multimodal language model is called to perform task processing based on the target task description instruction to obtain a graphic article corresponding to the video.

[0244] According to another embodiment of the present application, Figure 8 The various units in the video processing device shown can be separately or all merged into one or several other units to constitute, or a certain unit (or units) therein can also be split into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the function of a unit can also be realized by multiple units, or the function of multiple units can be realized by one unit. In other embodiments of the present application, other units can also be included based on the video processing device. In actual applications, these functions can also be assisted by other units to be realized, and can be realized by the collaboration of multiple units.

[0245] According to another embodiment of the present application, a computer program (including one or more instructions) capable of executing each step involved in the above method embodiment can be constructed by running the computer program (including one or more instructions) on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). Figure 8The video processing device shown in the embodiment of the present application is used to implement the video processing method of the embodiment of the present application. The computer program can be recorded on a computer-readable storage medium, for example, and loaded into the above-mentioned computing device through the computer-readable storage medium and run therein.

[0246] It is worth noting that, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can include a part of the overall module or unit of the module or unit function.

[0247] In the embodiment of the present application, for a video that needs to be converted from image to text, the video can be played and a canvas mask can be displayed on the video playback area, so that the user (object) can enter a first note recording operation on the canvas mask to obtain an electronic note of the current video frame. When a text-to-image conversion operation for the video is detected, the image-to-text article corresponding to the video can be generated based on the target image-to-text obtained by image-to-text recognition of the video and the video frame with the electronic note and the corresponding electronic note. The image-to-text article finally displayed includes not only the target image-to-text, but also some video frames and corresponding electronic notes. This can enrich the content of the converted image-to-text article and improve the quality of the image-to-text article. In addition, by introducing the electronic notes recorded by the user into the image-to-text article, the image-to-text article can be made more valuable in information, and the user customization of the image-to-text article can be enhanced, thereby improving the user experience of converting video to image-to-text.

[0248] Based on the description of the above method embodiment and apparatus embodiment, the present application embodiment also provides a computer device. Figure 9, the computer device at least includes a processor 901, an input interface 902, an output interface 903 and a computer storage medium 904. Among them, the processor 901, input interface 902, output interface 903 and computer storage medium 904 in the computer device can be connected via a bus or other means. The computer storage medium 904 can be stored in the memory of the computer device, and the computer storage medium 904 is used to store a computer program, and the computer program includes one or more instructions. The processor 901 is used to execute one or more instructions in the computer program stored in the computer storage medium 904. The processor 901 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to realize the corresponding method flow or corresponding function.

[0249] In one embodiment, the processor 901 described in the embodiment of the present application can be used to perform a series of video processing on the video that needs to be converted from image to text, specifically including: playing the video that needs to be converted from image to text, and displaying a canvas mask on the video playback area of ​​the video; drawing an electronic note displaying the current video frame in the canvas mask according to a first note recording operation detected on the canvas mask; the current video frame refers to: the video frame displayed in the video playback area when the first note recording operation is detected; in response to the image-text conversion operation for the video, displaying the image-text article corresponding to the video; the image-text article includes: the target image and text obtained by image-text recognition of the video, the video frame with the electronic note and the corresponding electronic note, etc.

[0250] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in a computer device for storing computer programs and data. It is understandable that the computer storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer storage medium provides a storage space, which stores the operating system of the computer device. In addition, a computer program is also stored in the storage space, which includes one or more instructions suitable for being loaded and executed by the processor 901, and these instructions can be one or more program codes. It should be noted that the computer storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor.

[0251] In one embodiment, a processor may load and execute one or more instructions stored in a computer storage medium to implement the corresponding steps in the above-mentioned video processing method embodiment. In a specific implementation, the processor may load and execute the following steps:

[0252] Play videos that require image-to-text conversion;

[0253] Displaying a canvas overlay on a video playback area of ​​the video;

[0254] Drawing an electronic note displaying a current video frame in the canvas mask according to a first note recording operation detected on the canvas mask; the current video frame refers to the video frame displayed in the video playback area when the first note recording operation is detected;

[0255] In response to the image-text conversion operation for the video, the image-text article corresponding to the video is displayed; the image-text article includes: the target image-text obtained by image-text recognition of the video, the video frame with electronic notes and the corresponding electronic notes.

[0256] In one embodiment, the one or more instructions may be loaded and specifically executed by a processor:

[0257] During playback of the video, displaying a note-taking component;

[0258] In response to a triggering operation on the note-recording component, the video is paused, and the step of displaying a canvas mask on the video playing area of ​​the video is triggered.

[0259] In another embodiment, the one or more instructions may be loaded and specifically executed by a processor:

[0260] During the process of displaying the canvas overlay on the video playback area, displaying a toolbar on the display screen, wherein the toolbar includes one or more note-taking tools;

[0261] A corresponding note-recording tool is selected according to a tool selection operation, and an operation input by an object on the canvas mask through the selected note-recording tool is regarded as a first note-recording operation detected on the canvas mask.

[0262] In another embodiment, the one or more instructions may be loaded and specifically executed by a processor:

[0263] If a note completion event for the current video frame is detected, generating note record information according to the current video frame and each electronic note currently displayed in the canvas mask;

[0264] The generated note record information is added to the note list, and the canvas mask is cleared; wherein the note list is located in an area of ​​the display screen except the video playback area.

[0265] In another embodiment, when the canvas overlay is displayed on the video playback area, the video is paused; accordingly, the one or more instructions may be loaded and specifically executed by the processor:

[0266] In response to a note completion event for the current video frame, canceling display of the canvas mask;

[0267] In the video playback area, the video continues to be played based on the current video frame.

[0268] In another embodiment, the one or more instructions may be loaded and specifically executed by a processor:

[0269] When a note record information is selected in the note list, in the video playback area, the video frame currently displayed is jumped to the first video frame; the first video frame refers to the video frame corresponding to the triggered note record information;

[0270] According to the triggered note recording information, the corresponding electronic note is reproduced on the first video frame.

[0271] In another embodiment, the number of the first note recording operation is at least one, and one first note recording operation is used to draw an electronic note of the current video frame; accordingly, the one or more instructions may be loaded and specifically executed by the processor:

[0272] Whenever a first note recording operation is detected, operation instruction information of the currently detected first note recording operation is added to the history record queue;

[0273] If a note undo operation is detected, deleting the last operation instruction information in the history record queue, and deleting the electronic note drawn based on the last first note recording operation in the canvas mask; wherein the last note recording operation refers to: the first note recording operation corresponding to the last operation instruction information;

[0274] When the canvas mask is cleared, the history record queue is cleared.

[0275] In another embodiment, the one or more instructions may be loaded and specifically executed by a processor:

[0276] During playback of the video, displaying video summary information corresponding to the video; the video summary information includes: frame description information of at least one key video frame in the video;

[0277] When frame description information is selected in the video summary information, in the video playback area, the video frame currently displayed is jumped to a second video frame; the second video frame is the video frame corresponding to the selected frame description information;

[0278] If a second note recording operation is detected on the canvas mask when the second video frame is displayed, an electronic note displaying the second video frame is drawn in the canvas mask according to the second note recording operation.

[0279] In another embodiment, the video summary information is generated based on the target image and text; accordingly, the one or more instructions may be loaded and specifically executed by the processor:

[0280] In response to a modification operation on the video summary information, updating and displaying the video summary information;

[0281] Based on the updated video summary information, the target image and text are updated.

[0282] In another embodiment, the one or more instructions may be loaded and specifically executed by a processor:

[0283] Obtaining target text and image obtained by performing text and image recognition on the video, wherein the target text and image include: first text information obtained by performing text recognition on audio data of the video, and key frame data; the key frame data includes: at least one key video frame, and second text information obtained by performing text recognition on the corresponding key video frame;

[0284] Obtaining note content, where the note content is generated based on the video frame with the electronic note and the corresponding electronic note;

[0285] The target image and text and the note content are integrated to obtain the image and text article corresponding to the video.

[0286] In another embodiment, the one or more instructions may be loaded and specifically executed by a processor:

[0287] Acquire audio data of the video, where the audio data is obtained by calling a multimedia processing framework to perform audio and video separation on the video;

[0288] Using a speech recognition library based on artificial intelligence technology to perform text recognition on the audio data to obtain audio text information;

[0289] The first multimodal language model is called to perform content summarization processing on the audio text information to obtain first text information.

[0290] In another embodiment, the generation process of the first text information is executed in a server, and the video is played through a web application; accordingly, the one or more instructions can be loaded and specifically executed by a processor:

[0291] Obtaining a first compilation result obtained by compiling a binary code format version of the multimedia processing framework, wherein the first compilation result includes a program interface for invoking the multimedia processing framework; the binary code format is a code format that supports execution by the web application and enhances the performance of the web application;

[0292] Calling the multimedia processing framework to perform audio and video separation on the video through the program interface in the first compilation result to obtain audio data and a video frame set of the video;

[0293] The audio data is uploaded to a server, so that the server executes a process of generating the first text information.

[0294] In another embodiment, the one or more instructions may be loaded and specifically executed by a processor:

[0295] Obtaining N key video frames in the video and video text information of each of the N key video frames; the N key video frames are arranged in chronological order according to playback time points, where N is a positive integer, and the video text information of any key video frame is obtained by performing text recognition on the corresponding key video frame;

[0296] Based on the video contents of the N key video frames, the N key video frames are segmented to obtain a plurality of sub-captions; each sub-caption includes at least one key video frame, and the video contents of the key video frames in any sub-caption are continuous;

[0297] Calling the first multimodal language model to perform text description on the corresponding sub-caption according to the video text information of the key video frame in each sub-caption, thereby obtaining text description information of the corresponding sub-caption;

[0298] A key video frame is selected from each sub-caption, and textual description information of each sub-caption is used as second textual information of the key video frame selected from the corresponding sub-caption, thereby obtaining key frame data; wherein the key frame data includes: the key video frame selected from each sub-caption and the second textual information of the corresponding key video frame.

[0299] In another embodiment, when the N key video frames are segmented based on the video contents of the N key video frames to obtain a plurality of sub-scenes, the one or more instructions may be loaded and specifically executed by the processor:

[0300] The video segmentation model is called to perform segmentation frame identification on the N key video frames based on the video contents of the N key video frames to obtain segmentation frame indication information; the segmentation frame indication information is used to indicate: a key video frame among the N key video frames to be used as a segmentation frame;

[0301] The N key video frames are cut based on the cut frame indication information to obtain M sub-scenes; wherein the cut includes: performing a cut operation once each time a cut frame is detected.

[0302] In another embodiment, the key frame data generation process is executed in the server; accordingly, the one or more instructions may be loaded and specifically executed by the processor:

[0303] After obtaining a video frame set of the video, extracting a key frame set from the video frame set, wherein the key frame set includes a plurality of key video frames;

[0304] Calling the classification model to perform category labeling processing on each key video frame in the key frame set to obtain label information of each key video frame;

[0305] filtering out repeated key video frames in the key frame set according to a repetition rate between label information of each key video frame in the key frame set to obtain the remaining N key video frames;

[0306] Performing text recognition on each of the N key video frames to obtain video text information of the corresponding key video frame;

[0307] The N key video frames and corresponding video text information are uploaded to a server, so that the server executes a process of generating the key frame data.

[0308] In another embodiment, the note content is obtained by the server invoking a first multimodal language model to optimize the note material, and the note material includes at least one video frame screenshot uploaded by the terminal; a video frame screenshot includes: video content of a video frame with an electronic note and a corresponding electronic note; accordingly, the one or more instructions can be loaded and specifically executed by the processor:

[0309] If a note completion event for the current video frame is detected, a canvas video loading interface is called to load the video into the canvas mask;

[0310] Adjusting the playback screen of the video to the current video frame in the canvas mask, and adjusting the display level of the video in the canvas mask, wherein the adjusted display level is lower than the display level of the electronic note of the current video frame in the canvas mask;

[0311] Calling a canvas screenshot interface to perform a screenshot process on the canvas mask to obtain a video frame screenshot of the current video frame;

[0312] Upload the video frame screenshot of the current video frame to the server.

[0313] In another embodiment, when the target image and text and the note content are integrated to obtain the image and text article corresponding to the video, the one or more instructions may be loaded and specifically executed by the processor:

[0314] Determine a second multimodal language model with article generation capabilities;

[0315] Using the target image and text and the note content, construct a target task description instruction of the second multimodal language model; wherein the target task description instruction is used to instruct: to perform article generation processing based on the target image and text and the note content;

[0316] The second multimodal language model is called to perform task processing based on the target task description instruction to obtain a graphic article corresponding to the video.

[0317] In the embodiment of the present application, for a video that needs to be converted from image to text, the video can be played and a canvas mask can be displayed on the video playback area, so that the user (object) can enter a first note recording operation on the canvas mask to obtain an electronic note of the current video frame. When a text-to-image conversion operation for the video is detected, the image-to-text article corresponding to the video can be generated based on the target image-to-text obtained by image-to-text recognition of the video and the video frame with the electronic note and the corresponding electronic note. The image-to-text article finally displayed includes not only the target image-to-text, but also some video frames and corresponding electronic notes. This can enrich the content of the converted image-to-text article and improve the quality of the image-to-text article. In addition, by introducing the electronic notes recorded by the user into the image-to-text article, the image-to-text article can be made more valuable in information, and the user customization of the image-to-text article can be enhanced, thereby improving the user experience of converting video to image-to-text.

[0318] It should be noted that, according to one aspect of the present application, a computer program product or computer program is also provided, which includes one or more instructions, and the one or more instructions are stored in a computer storage medium. The processor of the computer device reads the one or more instructions from the computer storage medium, and the processor executes the one or more instructions, so that the computer device performs the methods provided in the various optional embodiments of the video processing method mentioned above. It should be understood that what is disclosed above is only a preferred embodiment of the present application, and it is certainly not used to limit the scope of rights of the present application. Therefore, equivalent changes made in accordance with the claims of the present application are still within the scope covered by the present application.

Claims

1. A video processing method, characterized in that: include: Play the video that needs to be converted from image to text, and display a canvas mask on the video playback area of ​​the video; Drawing an electronic note displaying a current video frame in the canvas mask according to a first note recording operation detected on the canvas mask; the current video frame refers to the video frame displayed in the video playback area when the first note recording operation is detected; In response to a picture-text conversion operation on the video, displaying the picture-text article corresponding to the video; The graphic article includes: target graphics obtained by performing graphic recognition on the video, video frames with electronic notes, and corresponding electronic notes.

2. The method according to claim 1, wherein The method further comprises: During playback of the video, displaying a note-taking component; In response to a triggering operation on the note-recording component, the video is paused, and the step of displaying a canvas mask on the video playing area of ​​the video is triggered.

3. The method according to claim 1, wherein The method further comprises: During the process of displaying the canvas overlay on the video playback area, displaying a toolbar on the display screen, wherein the toolbar includes one or more note-taking tools; A corresponding note-recording tool is selected according to a tool selection operation, and an operation input by an object on the canvas mask through the selected note-recording tool is regarded as a first note-recording operation detected on the canvas mask.

4. The method according to any one of claims 1 to 3, wherein The method further comprises: If a note completion event for the current video frame is detected, generating note record information according to the current video frame and each electronic note currently displayed in the canvas mask; The generated note record information is added to the note list, and the canvas mask is cleared; wherein the note list is located in an area of ​​the display screen except the video playback area.

5. The method according to claim 4, wherein The generated note record information includes: video frame indication information of the current video frame, and note indication information of each electronic note currently displayed in the canvas mask; The video frame indication information of the current video frame includes at least one of the following: a playback time point of the current video frame, and a thumbnail of the current video frame; Any electronic note includes: a graphic mark for at least one video content in a video frame, or at least one text; the note indication information of any electronic note includes: an indication symbol of the graphic mark in the corresponding electronic note, or indication information of the text in the corresponding electronic note.

6. The method according to claim 4, wherein When the canvas mask is displayed on the video playback area, the video is paused; the method further includes: In response to a note completion event for the current video frame, canceling display of the canvas mask; In the video playback area, the video continues to be played based on the current video frame.

7. The method according to claim 4, wherein The method further comprises: When a note record information is selected in the note list, in the video playback area, the video frame currently displayed is jumped to the first video frame; the first video frame refers to the video frame corresponding to the triggered note record information; According to the triggered note recording information, the corresponding electronic note is reproduced on the first video frame.

8. The method according to claim 4, wherein The number of the first note recording operations is at least one, and one first note recording operation is used to draw an electronic note of the current video frame; the method further includes: Whenever a first note recording operation is detected, operation instruction information of the currently detected first note recording operation is added to the history record queue; If a note undo operation is detected, deleting the last operation instruction information in the history record queue, and deleting the electronic note drawn based on the last first note recording operation in the canvas mask; wherein the last note recording operation refers to: the first note recording operation corresponding to the last operation instruction information; When the canvas mask is cleared, the history record queue is cleared.

9. The method according to claim 1, wherein The method further comprises: During playback of the video, displaying video summary information corresponding to the video; the video summary information includes: frame description information of at least one key video frame in the video; When frame description information is selected in the video summary information, in the video playback area, the video frame currently displayed is jumped to a second video frame; the second video frame is the video frame corresponding to the selected frame description information; If a second note recording operation is detected on the canvas mask when the second video frame is displayed, an electronic note displaying the second video frame is drawn in the canvas mask according to the second note recording operation.

10. The method according to claim 9, wherein The video summary information is generated based on the target image and text; the method further includes: In response to a modification operation on the video summary information, updating and displaying the video summary information; Based on the updated video summary information, the target image and text are updated.

11. The method according to claim 1, wherein Methods for generating graphic articles corresponding to the video include: Obtaining target text and image obtained by performing text and image recognition on the video, wherein the target text and image include: first text information obtained by performing text recognition on audio data of the video, and key frame data; the key frame data includes: at least one key video frame, and second text information obtained by performing text recognition on the corresponding key video frame; Obtaining note content, where the note content is generated based on the video frame with the electronic note and the corresponding electronic note; The target image and text and the note content are integrated to obtain the image and text article corresponding to the video.

12. The method according to claim 11, wherein The process of generating the first text information includes: Acquire audio data of the video, where the audio data is obtained by calling a multimedia processing framework to perform audio and video separation on the video; Using a speech recognition library based on artificial intelligence technology to perform text recognition on the audio data to obtain audio text information; The first multimodal language model is called to perform content summarization processing on the audio text information to obtain first text information.

13. The method according to claim 12, wherein: The generation process of the first text information is executed in the server, and the video is played through a web application; the method further includes: Obtaining a first compilation result obtained by compiling a binary code format version of the multimedia processing framework, wherein the first compilation result includes a program interface for invoking the multimedia processing framework; the binary code format is a code format that supports execution by the web application and enhances the performance of the web application; Calling the multimedia processing framework to perform audio and video separation on the video through the program interface in the first compilation result to obtain audio data and a video frame set of the video; The audio data is uploaded to a server, so that the server executes a process of generating the first text information.

14. The method according to claim 11, wherein The key frame data generation process includes: Obtaining N key video frames in the video and video text information of each of the N key video frames; the N key video frames are arranged in chronological order according to playback time points, where N is a positive integer, and the video text information of any key video frame is obtained by performing text recognition on the corresponding key video frame; Based on the video contents of the N key video frames, the N key video frames are segmented to obtain a plurality of sub-captions; each sub-caption includes at least one key video frame, and the video contents of the key video frames in any sub-caption are continuous; Calling the first multimodal language model to perform text description on the corresponding sub-caption according to the video text information of the key video frame in each sub-caption, thereby obtaining text description information of the corresponding sub-caption; A key video frame is selected from each sub-caption, and textual description information of each sub-caption is used as second textual information of the key video frame selected from the corresponding sub-caption, thereby obtaining key frame data; wherein the key frame data includes: the key video frame selected from each sub-caption and the second textual information of the corresponding key video frame.

15. The method according to claim 14, wherein The step of cutting the N key video frames based on the video contents of the N key video frames to obtain a plurality of sub-scenes includes: The video segmentation model is called to perform segmentation frame identification on the N key video frames based on the video contents of the N key video frames to obtain segmentation frame indication information; the segmentation frame indication information is used to indicate: a key video frame among the N key video frames to be used as a segmentation frame; The N key video frames are cut based on the cut frame indication information to obtain M sub-scenes; wherein the cut includes: performing a cut operation once each time a cut frame is detected.

16. The method according to claim 14, wherein The key frame data generation process is executed in the server; the method further includes: After obtaining a video frame set of the video, extracting a key frame set from the video frame set, wherein the key frame set includes a plurality of key video frames; Calling the classification model to perform category labeling processing on each key video frame in the key frame set to obtain label information of each key video frame; filtering out repeated key video frames in the key frame set according to a repetition rate between label information of each key video frame in the key frame set to obtain the remaining N key video frames; Performing text recognition on each of the N key video frames to obtain video text information of the corresponding key video frame; The N key video frames and corresponding video text information are uploaded to a server, so that the server executes a process of generating the key frame data.

17. The method according to claim 11, wherein The note content is obtained by the server invoking the first multimodal language model to optimize the note material, and the note material includes at least one video frame screenshot uploaded by the terminal; A video frame screenshot includes: video content of a video frame with an electronic note and a corresponding electronic note; the method further includes: If a note completion event for the current video frame is detected, a canvas video loading interface is called to load the video into the canvas mask; Adjusting the playback screen of the video to the current video frame in the canvas mask, and adjusting the display level of the video in the canvas mask, wherein the adjusted display level is lower than the display level of the electronic note of the current video frame in the canvas mask; Calling a canvas screenshot interface to perform a screenshot process on the canvas mask to obtain a video frame screenshot of the current video frame; Upload a video frame screenshot of the current video frame to the server.

18. The method according to any one of claims 11 to 17, wherein: The target image and text are integrated with the note content to obtain the image and text article corresponding to the video, including: Determine a second multimodal language model with article generation capabilities; Using the target image and text and the note content, construct a target task description instruction of the second multimodal language model; wherein the target task description instruction is used to instruct: to perform article generation processing based on the target image and text and the note content; The second multimodal language model is called to perform task processing based on the target task description instruction to obtain a graphic article corresponding to the video.

19. The method according to claim 1, wherein In the graphic article, the video frame with the electronic note is presented in the form of a single video frame or in the form of a video animation; Any video animation refers to a dynamic image composed of at least two video frames, and any video animation contains at least one video frame with electronic notes.

20. A video processing device, characterized in that: include: A playback unit, used to play videos that require image-to-text conversion; a processing unit, configured to display a canvas mask on a video playback area of ​​the video; The processing unit is further configured to draw an electronic note displaying a current video frame in the canvas mask based on a first note recording operation detected on the canvas mask; the current video frame refers to the video frame displayed in the video playback area when the first note recording operation is detected; The processing unit is also used to display the graphic article corresponding to the video in response to the graphic-text conversion operation for the video; the graphic article includes: the target graphic obtained by graphic-text recognition of the video, the video frame with electronic notes and the corresponding electronic notes.