Video processing method and device, electronic equipment and storage medium
By extracting document data from video and organizing it into target documents, the problem of inefficient user manual operation of video conversion is solved, and an efficient video conversion process is realized, and the user's work efficiency is improved.
Patent Information
- Application Number
- CN202510402311.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, it is time-consuming and labor-intensive to manually operate the video content transfer process, which is inefficient, and the content integration and format adjustment are complex.
Provide a video processing method, which extracts document data from the target video in response to the video transfer instruction, organizes it according to the target document style, and generates the target document.
It greatly reduces the time consumption of the entire process from watching videos to content refining to document sorting and typesetting, and improves users' work efficiency and information processing capabilities.
Smart Images

Figure CN120472485A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a video processing method, device, electronic device, and storage medium. Background Art
[0002] With the popularity of online teaching and training, users often need to convert video content into learning or summary documents. However, the current process faces multiple challenges: (1) High time cost: Long videos require a long time to watch, and manually recording key information is time-consuming and labor-intensive, reducing learning efficiency. (2) Difficult content integration: Users need to manually take screenshots and match them with the text, which is a tedious and error-prone process and affects document quality. (3) Complex formatting: After completing the text and screenshots, the document format needs to be manually adjusted, which adds additional workload and may affect the document's aesthetics due to improper formatting.
[0003] It can be seen that the current video content conversion process that relies on manual user operations is time-consuming, labor-intensive and inefficient. Summary of the Invention
[0004] The present application provides a video processing method, device, electronic device and storage medium to solve the technical problem that the current video content conversion process that relies on manual operation by the user is time-consuming, labor-intensive and inefficient.
[0005] In a first aspect, the present application provides a video processing method, the method comprising:
[0006] In response to a video-to-text instruction for a target video, extracting document data from the target video;
[0007] The document data is sorted according to the style of the target document to form the target document.
[0008] In a possible implementation, extracting document data from the target video includes:
[0009] In response to an instruction to use a target content module carried in the video-to-text instruction, document data containing the target content module is extracted from the target video.
[0010] In a possible implementation manner, extracting the document data containing the target content module from the target video includes:
[0011] Determining a target video element according to the target content module;
[0012] Determining a target video frame associated with the target video element from the target video;
[0013] Document data is extracted from the target video frame.
[0014] In a possible implementation manner, extracting the document data containing the target content module from the target video includes:
[0015] Extracting text data from the target video;
[0016] Based on the target content module, the text data is formatted and processed to generate a paragraph containing the target content module;
[0017] Determining supplementary content for the paragraph based on the target video;
[0018] The paragraph and its supplementary content are combined to form document data.
[0019] In a possible implementation, extracting text data from the target video includes:
[0020] Extracting audio data from the target video, and converting the audio data into first text data;
[0021] and, extracting image data from the target video, and generating second text data based on the image data;
[0022] The first text data and the second text data are combined to form text data extracted from the target video.
[0023] In a possible implementation, the paragraph supplementary content includes a paragraph illustration, and the method further includes:
[0024] Overlay subtitles on the paragraph image.
[0025] In a possible implementation, the style of the target document is determined in the following manner:
[0026] Identifying the style of documents appearing in the target video;
[0027] The style of the document appearing in the target video is determined as the style of the target document.
[0028] In a possible implementation, the style of the target document is determined in the following manner:
[0029] Determining the content type of the target video;
[0030] The style of the target document is determined according to the content type of the target video.
[0031] In a possible implementation, the style of the target document is determined in the following manner:
[0032] In response to an instruction to use a target document type carried in the video-to-text instruction, a style of the target document is determined according to the target document type.
[0033] In one possible implementation, the method further includes:
[0034] In response to a translation instruction for the target document, text data in the first language in the target document is converted into text data in a second language, and the style of the target document is retained to obtain a translated target document.
[0035] In a second aspect, the present application provides a video processing device, comprising:
[0036] A document data extraction module, configured to extract document data from a target video in response to a video-to-text instruction for the target video;
[0037] The document data arranging module is used to arrange the document data according to the style of the target document to form the target document.
[0038] In a possible implementation, the document data extraction module is specifically configured to:
[0039] In response to an instruction to use a target content module carried in the video-to-text instruction, document data containing the target content module is extracted from the target video.
[0040] In one possible implementation, the document data extraction module includes:
[0041] a video element determination unit, configured to determine a target video element according to the target content module;
[0042] a video frame determining unit, configured to determine a target video frame associated with the target video element from the target video;
[0043] A data extraction unit is used to extract document data from the target video frame.
[0044] In one possible implementation, the document data extraction module includes:
[0045] A text data extraction unit, configured to extract text data from the target video;
[0046] A text layout unit, configured to layout the text data based on the target content module to generate a paragraph containing the target content module;
[0047] a paragraph supplementing unit, configured to determine paragraph supplementing content for the paragraph based on the target video;
[0048] The document data forming unit is used to form document data from the paragraph and its supplementary content.
[0049] In one possible implementation, the text data extraction unit includes:
[0050] an audio-to-text subunit, configured to extract audio data from the target video and convert the audio data into first text data;
[0051] an image-to-text subunit, configured to extract image data from the target video and generate second text data based on the image data;
[0052] The text merging subunit is configured to merge the first text data and the second text data into text data extracted from the target video.
[0053] In a possible implementation, the paragraph supplementary content includes a paragraph illustration, and the device further includes:
[0054] The subtitle overlay module is used to overlay subtitles on the paragraph pictures.
[0055] In a possible implementation, the device further includes:
[0056] The target style determination module is used to determine the style of the target document by:
[0057] Identifying the style of documents appearing in the target video;
[0058] The style of the document appearing in the target video is determined as the style of the target document.
[0059] In a possible implementation, the device further includes:
[0060] The target style determination module is used to determine the style of the target document by:
[0061] Determining the content type of the target video;
[0062] The style of the target document is determined according to the content type of the target video.
[0063] In a possible implementation, the device further includes:
[0064] The target style determination module is used to determine the style of the target document by:
[0065] In response to an instruction to use a target document type carried in the video-to-text instruction, a style of the target document is determined according to the target document type.
[0066] In a possible implementation, the device further includes:
[0067] The translation module is configured to convert the text data in the first language in the target document into text data in the second language in response to a translation instruction for the target document, and retain the style of the target document to obtain a translated target document.
[0068] In a third aspect, the present application provides an electronic device comprising: a processor and a memory, wherein the processor is configured to execute a video processing program stored in the memory to implement the video processing method described in any one of the first aspects.
[0069] In a fourth aspect, the present application provides a storage medium storing one or more programs, which can be executed by one or more processors to implement the video processing method described in any one of the first aspects.
[0070] The technical solution provided by the embodiments of the present application has the following advantages over the prior art: The method provided by the embodiments of the present application, by responding to a video-to-text instruction for a target video, extracts document data from the target video, organizes the document data in the format of the target document, and forms a target document. This allows a user to easily trigger a video-to-text instruction to convert the target video into a target document in the desired format. This solution can significantly reduce the time spent on the entire process from watching the video to extracting the content, and then to organizing and formatting the document. This eliminates the need for users to repeatedly watch videos to find key information, nor does it require them to manually summarize the content and carefully format the document, thereby significantly improving their work efficiency and information processing capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0072] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0073] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0074] Figure 1A flowchart of an embodiment of a video processing method provided in an embodiment of the present application;
[0075] Figure 2 An example of setting up an interface for video-to-text conversion;
[0076] Figure 3 A flowchart of another video processing method provided in an embodiment of the present application;
[0077] Figure 4 A schematic diagram of a video processing case provided in an embodiment of the present application;
[0078] Figure 5 A block diagram of an embodiment of a video processing device provided in an embodiment of the present application;
[0079] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0080] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0081] The disclosure below provides many different embodiments or examples for implementing different structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, these are merely examples and are not intended to limit the present application. In addition, the present application may repeat reference numbers and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.
[0082] In order to solve the technical problem in the prior art that the video content conversion process that relies on manual user operation is time-consuming, labor-intensive and inefficient, the present application provides a video processing method, device, electronic device and storage medium, which can realize a solution in which a user can easily trigger a video-to-text instruction to convert a target video into a target document in the format required by the user, which can greatly improve the user's work efficiency and information processing capabilities.
[0083] Figure 1 A flowchart of an embodiment of a video processing method provided in an embodiment of the present application.
[0084] like Figure 1 As shown, the method includes the following steps:
[0085] Step 101: In response to a video-to-text instruction for a target video, extract document data from the target video.
[0086] Step 102: Arrange the document data according to the style of the target document to form the target document.
[0087] To facilitate understanding of the technical solutions provided in the embodiments of the present application, steps 101 and 102 are explained in a unified manner below:
[0088] First of all, it can be seen from the description of step 101 and step 102 that the technical solution provided by the embodiment of the present application is intended to provide users with a convenient way, allowing users to easily trigger the video-to-text instruction to trigger the execution subject of the embodiment of the present application to convert the target video into a document in the format required by the user. For example, various document formats such as study notes, summary reports, social media articles, and graphics and texts on knowledge sharing platforms. This solution can greatly reduce the time consumed by users from watching videos to content extraction, and then to document organization and typesetting, so that users no longer need to watch videos repeatedly to find key information, nor do they have to manually summarize content and carefully typeset documents, thereby greatly improving users' work efficiency and information processing capabilities.
[0089] Specifically, in one embodiment, the above-mentioned target video can be a finished product video that is pre-recorded and stored, or it can be a live video played in real time. In the scenario where the target video is a finished product video, the user can upload the video to the executive subject of the embodiment of the present application by operating the visual interface, or provide the URL address of the video (that is, the technical solution provided by the embodiment of the present application supports the transfer of cloud video). In the scenario where the target video is a live video, the technical solution provided by the embodiment of the present application can be used to achieve the synchronous generation of documents following the live content, while also allowing the generated document content to be automatically adjusted and optimized as the live broadcast continues, until the final target document is formed. It can be added that the technical solution provided by the embodiment of the present application can achieve the generation of corresponding documents efficiently and flexibly, whether for pre-prepared video materials or instant broadcast video content.
[0090] In one embodiment, the number of the above-mentioned target videos can be one or more. For the case of multiple target videos, as an optional implementation method, the video-to-document operation can be performed separately for each target video, and the user can be allowed to start the document-to-document function of multiple target videos with one click. For example, the user can select multiple target videos through an intuitive video selection interface, then select the style of the target document for these target videos, and finally trigger the document-to-document icon. Once the document-to-document icon is triggered, a video-to-document instruction for the target video will be generated. Thus, the execution entity of the embodiment of the present application responds to the video-to-document instruction for the target video, applies the technical solution provided by the embodiment of the present application, automatically processes these target videos in parallel, and converts these target videos one by one into target documents. Here, the multiple target documents have the same document style. As another optional implementation method, in order to be more flexible, the user can also independently configure the document style, content details and other attributes for each target video. After all target videos have been configured, the document-to-document icon can be triggered. This can also realize the function of starting the document-to-document function of multiple target videos with one click, and also allow target documents of different styles to be generated for different target videos, or target documents of the same style but different content details, thereby combining personalization and flexibility.
[0091] As another optional implementation, multiple target videos are regarded as a whole and transferred to document processing, that is to say, a target document will eventually be generated for multiple target videos. In this way, in the process of transferring to document processing, not only can the document content required by the user be extracted separately from each target video, but also the continuity between videos can be considered at the same time, and the key content in different target videos can be automatically identified and integrated. Under this processing mode, the document data extracted from multiple target videos not only captures the information in each target video, but also includes the document data for maintaining the logic and fluency of the overall narrative of multiple target videos. This shows that under this implementation, the user only needs to simply select the required video and specify unified document output requirements (such as attributes such as the style and content details of the document) to trigger the execution subject of the embodiment of the present application to automatically transfer multiple videos to document processing.
[0092] The above embodiments can improve the user's operational convenience when facing the task of converting multiple videos to documents, and greatly improve the efficiency of document generation, greatly saving the user's time and energy.
[0093] In one embodiment, the composition of the document data includes but is not limited to: voice transcription content, subtitle extraction content, intelligently recognized video content (such as specific features, task behaviors, etc.), role analysis content, scene description content, etc., thereby providing users with more intuitive, detailed and insightful document content.
[0094] Furthermore, in one embodiment, the technical solution provided by this application allows users to customize the content and type of documents. This approach can break the relatively fixed content limitations of traditional AI video products, allowing users to independently splice various content elements to create unique documents.
[0095] Among them, as an optional implementation method, a method is provided to allow users to easily click on the content modules required for the document through a visual interface to realize the content and type of customized documents. This method can greatly simplify the user's operation process and save the user the trouble of conceiving the content of the document. Specifically, the execution subject of the embodiment of the present application displays a preset content module set through a visual interface, and the user can directly click on the required content module in the content module set, and then trigger the text conversion icon. At this time, a video-to-text instruction for the target video will be generated, and the video-to-text instruction carries an instruction to use the target content module (that is, the content module selected by the user). On this basis, the specific implementation of extracting document data from the target video includes: in response to the instruction to use the target content module carried in the video-to-text instruction, extracting document data containing the target content module from the target video.
[0096] For example, see Figure 2 , which is an example of the video-to-text setting interface. Figure 2 The example content modules include the following content modules: adding paragraph titles and outline directories, adding key video screenshots, adding key video GIF animations, adding AI full-text summaries, and adding AI paragraph knowledge point extraction.
[0097] The addition of paragraph headings and an outline table of contents aims to organize document content into logically clear paragraphs, assigning each paragraph a corresponding heading, and generating a comprehensive outline table of contents. This not only makes it easier for users to quickly browse and understand the document structure, but also improves the overall readability and professionalism of the document.
[0098] The Add Key Video Screenshots and Add Key Video GIFs modules allow you to insert images captured from videos into appropriate locations in your document as paragraph illustrations. These paragraph illustrations can intuitively showcase key scenes or details from the video, enhancing the visual appeal of the document. These paragraph illustrations can be in various formats, including static images, animated graphics, and video clips.
[0099] The AI Full-Text Summary module is designed to automatically generate a concise and clear summary of a document using artificial intelligence technology. This summary will extract the core information and key points of the document, helping users quickly grasp the main idea of the document.
[0100] The AI Paragraph Knowledge Point Extraction module leverages AI technology to identify and extract key knowledge points from documents. These points are presented in a structured format, facilitating in-depth learning and understanding. This module not only allows users to quickly access the core information within a document but also allows for targeted review and consolidation of knowledge as needed.
[0101] It will be understood by those skilled in the art that Figure 2 The content module set shown in the is only an example. In actual application, the content module set has extremely high flexibility and scalability, and can be customized and adjusted according to the actual needs of users. For example, after identifying the target video, more detailed content module options can be provided, such as character selection, behavior selection, etc. For example, users can choose to summarize only the speech content of a specific person according to their own focus, or extract only relevant clips of specific behaviors (such as interview dialogues) in the video. In addition, for live videos, specific recording rules can also be preset, such as only recording the scene content of the live broadcast, or filtering content according to specific events or topics in the live broadcast. In this way, even if the amount of information is large and the content is complex during the live broadcast, users can easily obtain the key information they need, which greatly improves the efficiency of watching live broadcasts and organizing documents.
[0102] Based on this, in one embodiment, the specific implementation of extracting document data containing a target content module from a target video includes: determining a target video element based on the target content module, determining a target video frame related to the target video element in the target video; and extracting document data from the target video frame. The target video element may be a specific video scene, character behavior, key dialogue, etc., depending on the target content module selected by the user. For example, if the target content module is to extract interview dialogue behavior in a video, then based on the target content module, the target video element is determined to be an interview dialogue behavior, and then the video frame containing the interview dialogue behavior in the target video is used as the target video frame, and the document data is extracted from the target video frame.
[0103] The above embodiment explains document data from the perspective of content details. In one embodiment, from the perspective of content type, the above document data is not limited to traditional text format, but can include a variety of elements that can be embedded in documents, such as images, animated images, video clips, and audio, thereby providing users with a rich and diverse means of content expression.
[0104] Accordingly, in one embodiment, the specific implementation of extracting document data containing a target content module from a target video includes: extracting text data from the target video, and based on the target content module, formatting and processing the text data to generate paragraphs containing the target content module; based on the target video, determining paragraph supplementary content for the paragraph; and forming document data with the paragraph and its paragraph supplementary content.
[0105] Among them, paragraph supplementary content includes a variety of elements that can be embedded in documents, such as images, animated images, video clips, and voice. It should be noted here that paragraph supplementary content is not limited to appearing only below or to the left of the paragraph text as additional instructions, but can also be interspersed within the paragraph text. For example, inserting small and appropriate emoticon images into the paragraph text at the right time not only adds color to the paragraph, but also can intuitively convey emotions, thereby enhancing the user's reading experience of the target document.
[0106] The process of arranging and processing text data based on the target content module to generate paragraphs containing the target content module is a comprehensive process that not only involves organizing the text data extracted from the target video, but also includes in-depth processing and expansion of this text data. Specifically, based on the content requirements of the target content module, such as the required information points, style and tone, and structural layout, the processing direction and expansion scope of the text data are determined, thereby deeply processing and expanding the original text data.
[0107] For example, if the target content module indicates that a summary is required, the original text data is abstracted and the summary content is generated. Then, according to the logical structure of the target content module, the summary content is arranged and organized to generate a paragraph containing the summary. For another example, if the target content module indicates that knowledge point analysis is required, the original text data is identified by knowledge point recognition to identify the knowledge points therein (which may be a concept, principle, method or fact, etc.), and then the content of the knowledge points is searched out, and the searched knowledge point content is arranged and organized to generate a paragraph containing the knowledge point analysis.
[0108] As for how to extract text data from the target video and how to determine the supplementary content for a paragraph, it will be explained through specific embodiments below and will not be described in detail here.
[0109] Finally, the document data extracted from the target video is organized according to the target document's style to form the target document. This process reorganizes and formats the document data to ensure that the resulting target document meets both the user's content requirements and specific style and formatting standards. The formatting layer primarily focuses on the structure and organization of the document content, including but not limited to organizing the extracted document data into a specific logical order (such as chronological order or order of importance) according to the target document's style requirements; using specific header and footer information (such as page numbers, author names, and dates); and other aspects. The style layer primarily focuses on the document's visual presentation and layout, typically relating to the document's typesetting, fonts, colors, heading structure, and paragraph formatting. For example, the target document's style may specify specific page margins, line spacing, and paragraph spacing; specify the font type, size, and color scheme used in the document; require the use of specific heading levels (such as H1, H2, and H3) to organize the content; and specify paragraph formatting (such as left-aligned, centered, right-aligned, or justified).
[0110] Among them, as an optional implementation method, the document data extracted from the target video is sorted according to the style of the target document and filled into a predefined template, so that the document data extracted from the target video is sorted according to the style of the target document to form a target document. Among them, the predefined template can be a Word document, PDF file or other editable document format, which contains the basic structure and style requirements of the document, such as title, paragraph format, font size, color, margin, header and footer, etc. In the template, specific placeholders (such as {{field_name}}) can be reserved to identify the location where data needs to be filled. Then, based on the placeholders reserved in the template, a mapping relationship between the extracted document data and the template placeholders is established. Use automated scripts or document processing software to fill the extracted data into the corresponding positions in the template according to the mapping relationship to form the target document.
[0111] In one embodiment, the style of the target document is determined in the following manner: in response to an instruction to use the target document type carried in the video-to-text instruction, the style of the target document is determined according to the target document type.
[0112] See also Figure 2 For example, the user can select the target document type from the preset document type set. Figure 2 As shown, the preset document types include the following document types: pure video text content, study notes, and summary reports. Each document type corresponds to specific style requirements. For example, pure video text content requires direct presentation and readability of text, study notes require key points and note areas, and summary reports require formal layout and formatting. In addition, Figure 2 As shown, custom options are also provided, which means that users can flexibly select the target document type by clicking on the custom option to meet more personalized document needs. Figure 2 After selecting the target document type in the interface shown, an instruction to use the target document type will be generated and included in the video-to-text instruction. Then, the execution entity of the embodiment of the present application responds to the instruction to use the target document type included in the video-to-text instruction and determines the target document style based on the target document type.
[0113] In this way, a mechanism can be implemented that allows users to flexibly select the target document style as needed, thereby ensuring that the generated target document meets the type requirements specified by the user, achieving flexibility and personalization of document generation.
[0114] In another embodiment, the target document style is determined by determining the target video's content type and, based on the target video's content type, determining the target document style. The core of this embodiment is analyzing the target video's content and then recommending the most appropriate document style based on the content type. For example, if the target video is an instructional video, a "knowledge point analysis template" might be recommended as the target document style. This template typically includes a clear title, an overview of the knowledge points, detailed analysis, examples, or case studies, making it easier for viewers to understand and absorb the instructional content.
[0115] Among them, content analysis of the target video involves technologies such as speech recognition, keyword extraction, and scene recognition of the video content, in order to accurately determine the theme and type of the video.
[0116] In this way, the most suitable document style can be intelligently selected according to the actual content of the target video, thereby improving the accuracy and practicality of document generation; at the same time, the automation and efficiency of document generation can also be improved.
[0117] In another embodiment, the style of the target document is determined in the following manner: the style of the document that appears in the target video is identified, and the style of the document that appears in the target video is determined as the style of the target document. The core of this embodiment is: identifying and extracting the style of the document that has appeared in the target video, and then directly applying this style to the target document that is finally generated. For example, the target video is an online lecture or meeting minutes, which contains PPT slides (or whiteboard notes, electronic spreadsheets, etc.) used by the speaker, then the style features of these slides, such as font, font size, color, layout, border, etc., can be identified and extracted. Subsequently, when the video content is converted into a document, these style features can be applied to the generated target document, thereby ensuring that the style of the target document is consistent with the document that appears in the video.
[0118] This reduces the need for users to manually set document styles, improves the automation and efficiency of document generation, and ensures that the document style is consistent with the document in the video, improving the accuracy and automation of document generation. In practical applications, the above embodiment is particularly important for application scenarios that require accurate reproduction of video content (such as education, training, meeting records, etc.).
[0119] In addition, in the above embodiment, it is not necessary for the entire target document to strictly follow the styles in the video. Instead, these styles can be applied to specific parts of the target document rather than the entire content. For example, if the target video contains a PPT slide presentation and other video content, the user may want to use the original style of the slide for the text content, while using a user-defined style for the rest of the video content.
[0120] The technical solution provided by the embodiments of this application extracts document data from a target video in response to a video-to-text instruction, organizes the document data in the format of the target document, and forms a target document. This allows users to easily trigger a video-to-text instruction to convert a target video into a target document in the desired format. This solution significantly reduces the time spent on the entire process from watching a video to extracting content, and then to document organization and layout. Users no longer need to repeatedly watch videos to find key information, nor do they need to manually summarize content and carefully layout documents, thereby significantly improving their work efficiency and information processing capabilities.
[0121] Figure 3 This is a flow chart of another video processing method provided in an embodiment of the present application. Figure 3 The process shown in Figure 1 Based on the process shown in the figure, the specific implementation process of video to text conversion is described. Figure 3 As shown, the following steps are included:
[0122] Step 301: extract text data from a target video and determine the timestamp information corresponding to the text data in the target video.
[0123] In one embodiment, the specific implementation of extracting text data from the target video includes: extracting audio data from the target video and converting the audio data into first text data; and extracting image data from the target video and generating second text data based on the image data; merging the first text data and the second text data to form the text data extracted from the target video.
[0124] The embodiment of the present application is different from the traditional method of simply extracting audio data (or subtitle data) from the video and converting it into text data. In this method, image recognition and analysis technology is introduced to extract image data from the target video, and second text data is generated by identifying the content in the image, such as text information, characters, behaviors, scenes, etc. in the image. This expands the source of text data and incorporates the visual elements in the video into the scope of textualization, thereby enriching the dimension and depth of the text data finally extracted, and providing richer and more accurate materials for subsequent text processing and utilization. In addition, in the extreme case where the target video does not have audio information or the audio quality is poor, this embodiment can still effectively extract text data by means of image-to-text, which makes the embodiment of the present application have good flexibility and adaptability when dealing with diverse video content.
[0125] Among them, as an optional implementation method, extracting image data from the target video and generating second text data based on the image data specifically include: extracting image data of video key frames from the target video, deduplicating the extracted image data, and generating second text data based on the deduplicated image data.
[0126] As a highly summarized and representative picture of the video content, the video key frame carries important information and visual focus in the video. Therefore, the image data of the video key frame can be extracted from the target video. The image data of these key frames usually contains key events, turning points or important visual elements in the target video, which is an indispensable part of understanding the video content. Subsequently, since there may be continuous similar or repeated pictures in the video, the extracted image data is deduplicated to remove redundant information. After the deduplication process is completed, the image data is deeply interpreted and converted using image recognition and analysis technology to convert visual elements such as text, symbols, objects, scenes, etc. in the image into text form, thereby generating second text data.
[0127] As another optional implementation method, the specific implementation of extracting image data from the target video and generating the second text data based on the image data includes: extracting image data from the target video at a set time interval, deduplicating the extracted image data, and generating the second text data based on the deduplicated image data.
[0128] With the rapid development of digital media, the amount of video data is exploding. Directly processing and analyzing an entire video is not only time-consuming and labor-intensive, but can also lead to low processing efficiency due to the sheer volume of data. Therefore, video sampling—extracting image data from a target video at set intervals and subsequently processing this image data—can effectively reduce data volume and improve processing efficiency while ensuring information integrity.
[0129] Furthermore, in one embodiment, if the audio quality of the target video is poor (e.g., due to noise interference, unclear speech, or too fast or too slow speech), the first text data extracted directly from the audio data may contain errors or omissions. To address this, an artificial intelligence model, such as a speech recognition model or a natural language processing model, may be used to further process and optimize the extracted first text data. For example, the extracted first text data may be intelligently polished and improved based on contextual information and grammatical rules, thereby improving the accuracy and completeness of the first text data.
[0130] In one embodiment, if the audio content of the target video does not meet the user's language requirements (such as using a different language, dialect, or accent), the extracted first text data can be translated during the text extraction phase to meet the user's language requirements. For example, a machine translation model can be used to translate the extracted first text data into a language specified by the user. This embodiment can achieve a more flexible and intelligent processing of the text extraction and generation process of the target video to cope with different video situations and user needs, and can provide users with a more convenient and personalized service experience.
[0131] For the first text data, the timestamp of its corresponding audio data in the target video is the timestamp of the first text data; for the second text data, the timestamp of its corresponding image data in the target video is the timestamp of the second text data.
[0132] Step 302: Generate document data based on the target video, the above text data and its corresponding timestamp information. The document data includes paragraphs and paragraph supplementary content.
[0133] In one embodiment, based on the target content module, the text data is formatted to generate a paragraph containing the target content module. Based on the timestamp information corresponding to the text data, supplementary content is determined for the paragraph from the target video, and the paragraph and its supplementary content are formed into document data.
[0134] Among them, taking the example of paragraph supplementary content including images or video slices, in the specific implementation, through the mapping relationship between the paragraph and text data and the timestamp information corresponding to the text data, the video part related to the paragraph can be located from the target video, then the video key frames in this video part can be selected for screenshots to form a paragraph illustration, or this video part or the key video clips therein can be selected to form a paragraph illustration in the form of video slices.
[0135] In addition, in one embodiment, if the paragraph image does not have original subtitles, subtitles can be superimposed on the paragraph image. This method can improve the readability and information density of the target document and help users better understand the paragraph image. The following two methods can be used to generate and superimpose subtitles on the paragraph image:
[0136] Method 1: Generate subtitles based on audio data. If the video frame corresponding to the paragraph image has audio data, speech recognition technology can be used to convert the audio into text to generate the corresponding subtitles.
[0137] Method 2: Generate subtitles based on video understanding content. If the video frame corresponding to the paragraph image does not have audio data, video understanding technology can be used to analyze the content of the video frame and extract key information from it to generate subtitles. In practice, this process involves identifying objects, people, actions, or scenes in the video frame and using natural language processing technology to convert this information into coherent text descriptions.
[0138] Furthermore, in actual applications, users can flexibly choose the strategy for generating subtitles for paragraph images based on their needs and preferences. For example, if the video frame corresponding to the paragraph image has audio data, but the user prefers to generate subtitles based on video understanding, the user can manually select the second method described above to generate subtitles for the paragraph image.
[0139] Overlaying subtitles within paragraph images can help readers better understand the video content, improving the readability and information density of the document. By combining audio data and video understanding technology, subtitles can be flexibly generated and overlaid to meet the needs and preferences of different users.
[0140] Step 303: Arrange the document framework data according to the style of the target document to form the target document.
[0141] The specific implementation of step 302 and step 303 is described below:
[0142] As an optional implementation method, the trained Vincent big model can be used to implement step 302. Specifically, the text data extracted from the target video and its corresponding timestamp information are filled into the preset prompt word template to obtain the target prompt word. Among them, the preset prompt word template may contain some placeholders for indicating where the big model inserts text, timestamps or other relevant information. The target prompt word is then input into the trained Vincent big model (such as the GPT model) to obtain the paragraphs and paragraph image index output by the big model, wherein the paragraph image index contains the timestamp information corresponding to the text data in the paragraph. Subsequently, based on the paragraph image index, video frames can be extracted from the target video as paragraph images.
[0143] In practice, the Vincent Wen model typically outputs data in Markdown format. After locating the paragraph image from the target video based on the paragraph image index, the paragraph image can be uploaded to a public storage service, where it will have a unique URL. The URL of the storage service is then used to replace the paragraph image index in the Markdown data output by the Vincent Wen model, and the Markdown data is converted to HTML format for subsequent format conversion.
[0144] Subsequently, based on the characteristics of each tag in the HTML format data obtained above, the nodes are divided into different types, such as paragraphs, text, pictures, tables, lists, etc.; then, these types are abstracted to construct a JSON structure containing key information such as node type, style and content. Then, the data in the JSON structure is used to generate documents in various formats such as Docx, PPTX, mind maps, WPS smart documents, etc., completing the conversion of HTML format data to target documents in various formats.
[0145] See also Figure 4 , is a schematic diagram of a video processing case provided in an embodiment of the present application. Figure 4 In the example, the left picture shows the target video, and the right picture shows the target document generated for the target video. Figure 4 It can be seen that the generated target document includes: outline, paragraphs, paragraph pictures, original text, knowledge point analysis generated by the large model, etc.
[0146] Combine Figure 4 From the example of , the technical solution provided by the embodiment of the present application realizes a solution in which a user can easily trigger a video-to-text command to convert a target video into a target document in the format required by the user. This solution can greatly reduce the time consumed by users in the entire process from watching videos to content extraction, and then to document organization and typesetting, so that users no longer need to watch videos repeatedly to find key information, nor do they need to manually summarize content and carefully typeset documents, thereby greatly improving users' work efficiency and information processing capabilities. At the same time, the technical solution provided by the embodiment of the present application can realize image-text matching, which further enriches the presentation form and content depth of the document, and provides users with a more comprehensive and high-quality video-to-text service.
[0147] Finally, this application also provides the following embodiments:
[0148] In one embodiment, in response to a translation instruction for a target document, text data in a first language in the target document is converted into text data in a second language, while retaining the style of the target document, thereby obtaining a translated target document.
[0149] In the above embodiment, when the user needs to translate the target document, the translation instruction to the target document can be triggered, and the second language specified by the user is carried in the translation instruction. Then, the execution subject of the embodiment of the present application will detect text data from the target document and convert it from the first language to the second language. Moreover, during the conversion process, it is ensured that the style (such as font, font size, paragraph format, position of picture and table, etc.) of the target document remains unchanged, thereby generating a target document after translation that not only includes the translated text but also retains the original style.
[0150] Figure 5 This is a block diagram of an embodiment of a video processing device provided in an embodiment of the present application. Figure 5 As shown, the device includes:
[0151] A document data extraction module 51 is configured to extract document data from a target video in response to a video-to-text instruction for the target video;
[0152] The document data arrangement module 52 is used to arrange the document data according to the style of the target document to form the target document.
[0153] In a possible implementation, the document data extraction module 51 is specifically configured to:
[0154] In response to an instruction to use a target content module carried in the video-to-text instruction, document data containing the target content module is extracted from the target video.
[0155] In one possible implementation, the document data extraction module 52 includes:
[0156] a video element determination unit, configured to determine a target video element according to the target content module;
[0157] a video frame determining unit, configured to determine a target video frame associated with the target video element from the target video;
[0158] A data extraction unit is used to extract document data from the target video frame.
[0159] In a possible implementation, the document data extraction module 51 includes:
[0160] A text data extraction unit, configured to extract text data from the target video;
[0161] A text layout unit, configured to layout the text data based on the target content module to generate a paragraph containing the target content module;
[0162] a paragraph supplementing unit, configured to determine paragraph supplementing content for the paragraph based on the target video;
[0163] The document data forming unit is used to form document data from the paragraph and its supplementary content.
[0164] In one possible implementation, the text data extraction unit includes:
[0165] an audio-to-text subunit, configured to extract audio data from the target video and convert the audio data into first text data;
[0166] an image-to-text subunit, configured to extract image data from the target video and generate second text data based on the image data;
[0167] The text merging subunit is configured to merge the first text data and the second text data into text data extracted from the target video.
[0168] In a possible implementation, the paragraph supplementary content includes a paragraph illustration, and the device further includes:
[0169] The subtitle overlay module is used to overlay subtitles on the paragraph pictures.
[0170] In a possible implementation, the device further includes:
[0171] The target style determination module is used to determine the style of the target document by:
[0172] Identifying the style of documents appearing in the target video;
[0173] The style of the document appearing in the target video is determined as the style of the target document.
[0174] In a possible implementation, the device further includes:
[0175] The target style determination module is used to determine the style of the target document by:
[0176] Determining the content type of the target video;
[0177] The style of the target document is determined according to the content type of the target video.
[0178] In a possible implementation, the device further includes:
[0179] The target style determination module is used to determine the style of the target document by:
[0180] In response to an instruction to use a target document type carried in the video-to-text instruction, a style of the target document is determined according to the target document type.
[0181] In a possible implementation, the device further includes:
[0182] The translation module is configured to convert the text data in the first language in the target document into text data in the second language in response to a translation instruction for the target document, and retain the style of the target document to obtain a translated target document.
[0183] like Figure 6 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.
[0184] Memory 113, for storing computer programs;
[0185] In one embodiment of the present application, the processor 111 is configured to execute a program stored in the memory 113 to implement the video processing method provided by any of the aforementioned method embodiments, including:
[0186] In response to a video-to-text instruction for a target video, extracting document data from the target video;
[0187] The document data is sorted according to the style of the target document to form the target document.
[0188] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the video processing method provided in any of the aforementioned method embodiments are implemented.
[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0191] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0192] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.
Claims
1. A video processing method, characterized in that: The method comprises: In response to a video-to-text instruction for a target video, extracting document data from the target video; The document data is sorted according to the style of the target document to form the target document.
2. The method according to claim 1, characterized in that The extracting document data from the target video includes: In response to an instruction to use a target content module carried in the video-to-text instruction, document data containing the target content module is extracted from the target video.
3. The method according to claim 2, characterized in that The step of extracting document data containing the target content module from the target video includes: Determining a target video element according to the target content module; Determining a target video frame associated with the target video element from the target video; Document data is extracted from the target video frame.
4. The method according to claim 2, characterized in that The step of extracting document data containing the target content module from the target video includes: Extracting text data from the target video; Based on the target content module, the text data is formatted and processed to generate a paragraph containing the target content module; Determining supplementary content for the paragraph based on the target video; The paragraph and its supplementary content are combined to form document data.
5. The method according to claim 4, characterized in that The extracting text data from the target video includes: Extracting audio data from the target video, and converting the audio data into first text data; and, extracting image data from the target video, and generating second text data based on the image data; The first text data and the second text data are combined to form text data extracted from the target video.
6. The method according to claim 4, characterized in that The paragraph supplementary content includes a paragraph illustration, and the method further includes: Overlay subtitles on the paragraph image.
7. The method according to claim 1, characterized in that The style of the target document is determined by: Identifying the style of documents appearing in the target video; The style of the document appearing in the target video is determined as the style of the target document.
8. The method according to claim 1, characterized in that The style of the target document is determined by: Determining the content type of the target video; The style of the target document is determined according to the content type of the target video.
9. The method according to claim 1, characterized in that The style of the target document is determined by: In response to an instruction to use a target document type carried in the video-to-text instruction, a style of the target document is determined according to the target document type.
10. The method according to claim 1, characterized in that The method further comprises: In response to a translation instruction for the target document, text data in the first language in the target document is converted into text data in a second language, and the style of the target document is retained to obtain a translated target document.
11. A video processing device, characterized in that: The device comprises: A document data extraction module, configured to extract document data from a target video in response to a video-to-text instruction for the target video; The document data arranging module is used to arrange the document data according to the style of the target document to form the target document.
12. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the video processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
System and method for converting video into PPT based on video processing
CN110493640A
Video abstract generation method and device, equipment and storage medium
CN114143479A
Video processing method and device, electronic equipment and storage medium
CN117851639A
Video processing method and device, electronic equipment, storage medium and program product
CN119135953A