Video summarization method and apparatus

By extracting text and image recognition results from video files and combining optical character extraction, speech recognition, and natural language processing technologies, high-quality video summaries are generated, solving the problem of existing technologies failing to take into account comprehensive information and improving the quality of video summaries.

CN114078221BActive Publication Date: 2025-10-31ALIBABA GROUP HOLDING LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010808917.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-12
Publication Date
2025-10-31
Estimated Expiration
2040-08-12

AI Technical Summary

Technical Problem

Existing methods for generating video content summaries fail to effectively integrate information from images, audio, and text, making it difficult to generate high-quality video summaries.

Method used

By acquiring video files, extracting text recognition and image recognition results, and generating video summaries based on these results, high-quality video summaries are generated using optical character extraction, speech recognition, and natural language processing technologies combined with computer vision.

Benefits of technology

It achieves the comprehensive collection of information such as images, audio, and text from video files, generating high-quality video summaries and improving the overall quality of video summaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114078221B_ABST
    Figure CN114078221B_ABST
Patent Text Reader

Abstract

This application discloses a video summarization method and apparatus. The method includes: acquiring a video file; extracting text recognition results and image recognition results from the video file; and generating a video summary based on the text recognition results and image recognition results. This application solves the technical problem that existing methods for generating video content summaries based on news text do not take into account comprehensive information such as images, audio, and text, making it difficult to generate high-quality video summaries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and more specifically, to a video summarization method and apparatus. Background Technology

[0002] In today's society, video news remains an important form of news information for people. However, with the accelerated pace of work and life, people have less and less time to watch complete news. More people tend to directly check some key information rather than watch the whole news or summarize the key information themselves from the video news. This creates a demand for the generation of high-quality video news content summaries.

[0003] However, existing methods for generating video content summaries all have some shortcomings, as described below: Method 1: Manual generation of news summaries. Advantages: Reliable quality assurance; Disadvantages: High labor costs and relatively low timeliness. Method 2: AI-based analysis of general video, film / television video, or surveillance video content to generate summaries. Advantages: High efficiency and low cost; Disadvantages: The analysis of general video, film / television video, or surveillance video content does not take into account the important features of news videos and the target result requirements, so the generated videos do not meet the requirements of news summaries. Method 3: Generating summaries based on news text. Advantages: Incorporating news characteristics, resulting in higher quality text summaries; Disadvantages: Existing methods mainly generate summaries based on the text of the news, without considering video, audio, and other information, making it difficult to generate high-quality summaries.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a video summary generation method and apparatus to at least solve the technical problem that existing methods for generating video content summaries based on news text do not take into account comprehensive information such as images, audio, and text, making it difficult to generate high-quality video summaries.

[0006] According to one aspect of the embodiments of this application, a video summary generation method is provided, comprising: acquiring a video file; extracting text recognition results and image recognition results from the video file; and generating a video summary based on the text recognition results and the image recognition results.

[0007] According to another aspect of the embodiments of this application, a video summary generation method is provided, including: receiving a video file sent by a server; extracting text recognition results and image recognition results from the video file; and generating and displaying a video summary on a client based on the text recognition results and the image recognition results.

[0008] According to another aspect of the embodiments of this application, a video summary generation method is provided, including: acquiring a video file; extracting text recognition results and image recognition results from the video file; generating a video summary based on the text recognition results and the image recognition results; and sending the video summary to a client to trigger the client to display the video summary.

[0009] According to another aspect of the embodiments of this application, a video summarization generation apparatus is also provided, comprising: an acquisition module for acquiring a video file; an extraction module for extracting text recognition results and image recognition results from the video file; and a generation module for generating a video summary based on the text recognition results and the image recognition results.

[0010] According to another aspect of the embodiments of this application, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is running, it controls the device where the storage medium is located to execute any of the above-described video summary generation methods.

[0011] According to another aspect of the embodiments of this application, a video summarization generation device is also provided, including: a processor; and a memory connected to the processor, for providing the processor with instructions to perform the following processing steps: acquiring a video file; extracting text recognition results and image recognition results from the video file; and generating a video summary based on the text recognition results and the image recognition results.

[0012] In this embodiment, a video file is acquired; text recognition results and image recognition results are extracted from the video file; and a video summary is generated based on the text recognition results and image recognition results. It is noteworthy that this embodiment uses artificial intelligence to extract text recognition results and image recognition results from the acquired video file, taking into account the comprehensive information of the video file, including images, audio, and text. Based on the text recognition results and image recognition results, a high-quality video summary can be generated.

[0013] Therefore, the embodiments of this application achieve the goal of generating video summaries by taking into account the comprehensive information of video files such as images, voice, and text, thereby improving the quality of video summaries generated based on video files. This solves the technical problem that existing methods for generating video content summaries based on news text do not take into account the comprehensive information of images, voice, and text, making it difficult to generate high-quality video summaries. Attached Figure Description

[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0015] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a video summarization method according to an embodiment of this application;

[0016] Figure 2 This is a flowchart of a video summary generation method according to an embodiment of this application;

[0017] Figure 3 This is a flowchart of an optional video summary generation method according to an embodiment of this application;

[0018] Figure 4 This is a flowchart of another video summary generation method according to an embodiment of this application;

[0019] Figure 5 This is a flowchart of another video summary generation method according to an embodiment of this application;

[0020] Figure 6 This is a schematic diagram of a video summarization device according to an embodiment of this application;

[0021] Figure 7 This is a schematic diagram of the structure of a video summary generation device according to an embodiment of this application;

[0022] Figure 8 This is a structural block diagram of another computer terminal according to an embodiment of this application. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0025] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0026] Video news refers to news broadcasts that use video as the medium, such as CCTV's evening news and morning news programs.

[0027] Optical Character Recognition (OCR): refers to a computer vision technique for extracting text from images.

[0028] Speech recognition (S2T) refers to a technology that extracts text content from speech signals.

[0029] Face recognition: refers to a computer vision technology based on the task of recognizing faces.

[0030] Human body recognition: refers to a computer vision technology based on the task of recognizing human bodies.

[0031] Landmark recognition refers to a computer vision technology that identifies key locations (e.g., famous tourist attractions, amusement parks, hospitals, etc.) and scenes (e.g., outdoor, indoor) in videos.

[0032] Event behavior recognition: refers to a technology that identifies events and behaviors occurring in a video.

[0033] Example 1

[0034] According to an embodiment of this application, an embodiment of a video summary generation method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0035] The method embodiment provided in Embodiment 1 of this application can be executed in a mobile terminal, computer terminal or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a video summarization method is shown, such as... Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0036] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0037] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video summarization generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned video summarization generation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0038] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0039] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0040] Under the aforementioned operating environment, this application provides the following: Figure 2 The above describes a video summarization method. Figure 2 This is a flowchart of a video summary generation method according to an embodiment of this application, such as... Figure 2 As shown, the video summary generation method described above includes the following steps:

[0041] Step S202: Obtain the video file;

[0042] Step S204: Extract text recognition results and image recognition results from the above video file;

[0043] Step S206: Generate a video summary based on the above text recognition results and the above image recognition results.

[0044] In this embodiment, a video file is acquired; text recognition results and image recognition results are extracted from the video file; and a video summary is generated based on the text recognition results and image recognition results. It is noteworthy that this embodiment uses artificial intelligence to extract text recognition results and image recognition results from the acquired video file, taking into account the comprehensive information of the video file, including images, audio, and text. Based on the text recognition results and image recognition results, a high-quality video summary can be generated.

[0045] Therefore, the embodiments of this application achieve the goal of generating video summaries by taking into account the comprehensive information of video files such as images, voice, and text, thereby improving the quality of video summaries generated based on video files. This solves the technical problem that existing methods for generating video content summaries based on news text do not take into account the comprehensive information of images, voice, and text, making it difficult to generate high-quality video summaries.

[0046] In one optional embodiment, the video files include one of the following: news videos; scientific exploration videos; historical documentary videos; arts and entertainment videos; and sports event videos.

[0047] It should be noted that the video summary generation method provided in this application embodiment can be applied to, but is not limited to, the following application scenarios: current affairs news videos; scientific exploration videos; historical documentary videos; arts and entertainment videos; and sports event videos.

[0048] The video summarization method provided in this application embodiment can be applied, but is not limited to, to an AI-based image and text summarization system. This system, targeting video files, utilizes the unique characteristics of video files and, by implementing the video summarization method provided in this application embodiment, comprehensively considers the images, audio, and text of the video file to generate a high-quality video summary.

[0049] In an optional embodiment, extracting the text recognition result from the video file includes:

[0050] Step S302: Extract the first recognition result from the subtitle data of the above video file using optical character extraction method;

[0051] Step S304: Extract the second recognition result from the audio data of the above video file using a speech recognition extraction method;

[0052] Step S306: Perform natural language processing on the first recognition result and the second recognition result to obtain the text recognition result.

[0053] The video files in this embodiment of the application have significant differences and characteristics compared to other ordinary videos (e.g., short videos of the Vlog type). For example, taking the above-mentioned video file as a news video file, the news video file has at least the following characteristics compared to other ordinary videos: 1) The theme is clear. Each part of the news video file has a clear theme (e.g., the flood control and flood fighting situation is severe, a leader visits a certain region, etc.); 2) There is a clear distinction between different themes in the same news video file. This distinction includes text content, audio content, image content, etc., and even the appearance of the host / anchor can be used as one of the features to distinguish different themes in the same news video file; 3) The natural language content (e.g., text subtitles, audio, etc.) has a clear connection and guiding significance to the entire video image content; 4) The expression of the natural language content is relatively formal. For example, the subject-verb-object structure is clear, and written language is used. For example, someone made a certain statement, something happened in a certain place, etc.

[0054] Based on the characteristics of the video files in this application embodiment compared to other ordinary videos, this application uses optical character recognition (OCR), speech recognition extraction (S2T), and natural language processing technology to identify the text recognition results in the video files.

[0055] Figure 3 This is a flowchart of an optional video summary generation method according to an embodiment of this application, such as... Figure 3 As shown in this embodiment, an optical character recognition (OCR) method can be used to extract a first recognition result from the subtitle data of the video file, such as subtitles in the video file; a speech recognition extraction method can be used to extract a second recognition result from the speech data of the video file, such as text in the video file; and natural language processing can be performed on the first and second recognition results, for example... Figure 3 As shown, the text summarization method is used to extract text summaries from the first recognition result and the second recognition result to obtain the text recognition result.

[0056] In an optional embodiment, the above method further includes:

[0057] Step S402: Record the time information corresponding to the above text recognition results.

[0058] In an optional embodiment, extracting the image recognition result from the video file includes:

[0059] Step S502: Based on the above text recognition results and the above time information, extract the above image recognition results from the above video file.

[0060] Because the video files in this embodiment have clear themes—each part of the news video file has a clear theme, different themes within the same news video file are clearly distinguishable, and natural language content (e.g., text subtitles, audio, etc.) has a clear connection and guiding significance to the overall video image content—the extraction of image recognition results from the video files can be guided by utilizing the recognized text recognition results and the corresponding time information. That is, based on the aforementioned text recognition results and time information, the image recognition results can be extracted from the video files, thereby improving the quality of the generated video summary.

[0061] In an optional embodiment, extracting the image recognition result from the video file based on the text recognition result and the time information includes:

[0062] Step S602a: Obtain the text information to be adapted based on the above text recognition results;

[0063] Step S604a: Determine the time point corresponding to the above-mentioned text information to be adapted based on the above-mentioned time information.

[0064] Step S606a: Based on the aforementioned time points, extract the image content associated with the aforementioned text information to be adapted from the aforementioned video file using computer vision methods, and obtain the aforementioned image recognition results.

[0065] In the above optional embodiments, the text recognition results include the text and subtitles of the video file. Adapted text information is obtained from the recognized text recognition results, and the time point corresponding to the text information to be adapted is determined according to the time information. Optionally, the time point corresponding to the text information to be adapted can be understood as corresponding each text to a time point in the video file according to the time progress of the video file. For example, the text "flood prevention" in the sentence "The flood prevention and control situation is severe..." in a video file corresponds to a time point. Then, based on the time point, the image content associated with the text information to be adapted is extracted from the video file using computer vision to obtain the image recognition results.

[0066] In an optional embodiment, extracting the image recognition result from the video file based on the text recognition result and the time information includes:

[0067] Step S602b: Obtain the text information to be adapted based on the above text recognition results;

[0068] Step S604b: Determine the time period corresponding to the above-mentioned text information to be adapted based on the above-mentioned time information;

[0069] Step S606b: Based on the aforementioned time period, extract the image content associated with the aforementioned text information to be adapted from the aforementioned video file to obtain the aforementioned image recognition result.

[0070] In the above optional embodiments, the text recognition results include the text and subtitles of the video file. Adapted text information is obtained from the recognized text recognition results, and the time period corresponding to the text information to be adapted is determined according to the time information. Optionally, the time period corresponding to the text information to be adapted can be understood as determining a time period based on the start and end times of a given text. For example, a sentence in a video file, "The flood control and flood fighting situation is severe...", corresponds to a time period. According to the time progress of the video file, computer vision is used to extract the image content that corresponds one-to-one with the text in the time period from the video file to obtain the above image recognition results.

[0071] Based on the characteristics of news video files, this application proposes a video summarization generation scheme that takes into account images, audio, and text. This scheme effectively overcomes the shortcomings of existing solutions. For example, compared with the existing method one, this application can significantly save manpower and improve efficiency. Compared with the existing solution two, by utilizing the characteristics of video news and using text recognition results and corresponding time information to guide the processing of image information, more accurate image recognition results can be obtained. Compared with the existing solution three, by adding news information such as audio and video, higher quality video summaries can be obtained.

[0072] Optionally, the aforementioned computer vision methods include at least one of the following: face recognition extraction method, human body recognition extraction method, scene location recognition extraction method, and event behavior recognition extraction method.

[0073] In one optional embodiment, extracting the image content associated with the text information to be adapted from the video file includes at least one of the following:

[0074] Step S702: Use face recognition extraction to extract the face recognition content associated with the above-mentioned text information to be adapted from the above-mentioned video file;

[0075] Step S704: Use human body recognition extraction method to extract human body recognition content associated with the above-mentioned text information to be adapted from the above video file;

[0076] Step S706: Use scene location recognition extraction method to extract scene location recognition content associated with the above-mentioned text information to be adapted from the above video file;

[0077] Step S708: Use event behavior recognition extraction method to extract event behavior recognition content associated with the above-mentioned text information to be adapted from the above video file.

[0078] In the above optional embodiments, it is still as follows Figure 3 As shown, the computer vision method provided by the computer vision module can be used to extract the image content associated with the text information to be adapted from the video file to obtain the image recognition result.

[0079] As an optional embodiment, computer vision methods such as face recognition, human body recognition, scene location recognition, event behavior recognition, and pure video extraction provided by the computer vision module are used. For example, face recognition extraction is used to extract face recognition content associated with the text information to be adapted from the video file; human body recognition extraction is used to extract human body recognition content associated with the text information to be adapted from the video file; scene location recognition extraction is used to extract scene location recognition content associated with the text information to be adapted from the video file; and event behavior recognition extraction is used to extract event behavior recognition content associated with the text information to be adapted from the video file.

[0080] In an alternative embodiment, it remains as follows Figure 3 As shown, after extracting the image content associated with the text information to be adapted from the video file using the above computer vision method and obtaining the above image recognition results, the image recognition results and text recognition results can be fused to generate a graphic video summary.

[0081] Under the aforementioned operating environment, this application also provides, for example: Figure 4 Another video summarization method is shown. Figure 4 This is a flowchart of another video digest generation method according to an embodiment of this application, such as... Figure 4 As shown, the video summary generation method described above includes the following steps:

[0082] Step S802: Receive the video file sent by the server;

[0083] Step S804: Extract text recognition results and image recognition results from the aforementioned video file;

[0084] Step S806: Based on the above text recognition results and the above image recognition results, generate and display a video summary on the client.

[0085] It should be noted that the video summary generation method provided in steps S802 to S806 of the embodiments of this application can be applied to the client side, but is not limited to. The client receives a video file from the server and, based on artificial intelligence, utilizes the unique characteristics of the video file to extract text recognition results and image recognition results from the video file, taking into account the comprehensive information of the video file such as images, voice, and text, and directly generates a high-quality video summary on the client based on the text recognition results and image recognition results.

[0086] Therefore, the embodiments of this application achieve the goal of generating video summaries by taking into account the comprehensive information of video files such as images, voice, and text, thereby improving the quality of video summaries generated based on video files. This solves the technical problem that existing methods for generating video content summaries based on news text do not take into account the comprehensive information of images, voice, and text, making it difficult to generate high-quality video summaries.

[0087] In one optional embodiment, the video files include one of the following: news videos; scientific exploration videos; historical documentary videos; arts and entertainment videos; and sports event videos.

[0088] It should be noted that the video summary generation method provided in this application embodiment can be applied to, but is not limited to, the following application scenarios: current affairs news videos; scientific exploration videos; historical documentary videos; arts and entertainment videos; and sports event videos.

[0089] Based on the characteristics of the video files in this application embodiment compared to other ordinary videos, this application uses optical character recognition (OCR), speech recognition extraction (S2T), and natural language processing technology to identify the text recognition results in the video files.

[0090] In this embodiment, the client can use optical character recognition (OCR) extraction to extract a first recognition result from the subtitle data of the video file, such as subtitles in the video file; and use speech recognition extraction to extract a second recognition result from the speech data of the video file, such as text in the video file; and perform natural language processing on the first recognition result and the second recognition result. For example, the client can, but is not limited to, use a text summarization extraction method to perform text summarization processing on the first recognition result and the second recognition result to obtain the text recognition result.

[0091] In the above optional embodiments, the text recognition results include the text and subtitles of the video file. The client can obtain the matching text information from the recognized text recognition results, determine the time point corresponding to the text information to be matched based on the time information, and then extract the image content associated with the text information to be matched from the video file using computer vision based on the time point to obtain the image recognition results.

[0092] Because the video files in this embodiment have clear themes—each part of the news video file has a clear theme, different themes within the same news video file are clearly distinguishable, and natural language content (e.g., text subtitles, audio, etc.) has a clear connection and guiding significance to the overall video image content—the client can use the recognized text recognition results and the corresponding time information to guide the extraction of image recognition results when extracting image recognition results from the video file. That is, based on the aforementioned text recognition results and time information, the image recognition results can be extracted from the video file, thereby improving the quality of the generated video summary.

[0093] Under the aforementioned operating environment, this application also provides, for example: Figure 5 Another video summarization method is shown. Figure 5 This is a flowchart of another video digest generation method according to an embodiment of this application, such as... Figure 5 As shown, the video summary generation method described above includes the following steps:

[0094] Step S902: Obtain the video file;

[0095] Step S904: Extract text recognition results and image recognition results from the aforementioned video file;

[0096] Step S906: Generate a video summary based on the above text recognition results and the above image recognition results, and send the video summary to the client to trigger the client to display the video summary.

[0097] It should be noted that the video summary generation method provided in steps S902 to S906 of the embodiments of this application can be applied to the server side, but is not limited to. After the server obtains the video file, it uses artificial intelligence to extract text recognition results and image recognition results from the video file by utilizing the unique characteristics of the video file. It takes into account the comprehensive information of the video file, such as images, voice, and text, and generates a high-quality video summary based on the text recognition results and image recognition results. Then, the video summary is sent to the client to trigger the client to display the video summary.

[0098] Therefore, the embodiments of this application achieve the goal of generating video summaries by taking into account the comprehensive information of video files such as images, voice, and text, thereby improving the quality of video summaries generated based on video files. This solves the technical problem that existing methods for generating video content summaries based on news text do not take into account the comprehensive information of images, voice, and text, making it difficult to generate high-quality video summaries.

[0099] In one optional embodiment, the video files include one of the following: news videos; scientific exploration videos; historical documentary videos; arts and entertainment videos; and sports event videos.

[0100] It should be noted that the video summary generation method provided in this application embodiment can be applied to, but is not limited to, the following application scenarios: current affairs news videos; scientific exploration videos; historical documentary videos; arts and entertainment videos; and sports event videos.

[0101] Based on the characteristics of the video files in this application embodiment compared to other ordinary videos, this application uses optical character recognition (OCR), speech recognition extraction (S2T), and natural language processing technology to identify the text recognition results in the video files.

[0102] In this embodiment, the server can use optical character recognition (OCR) to extract a first recognition result from the subtitle data of the video file, such as subtitles in the video file; and use speech recognition extraction to extract a second recognition result from the speech data of the video file, such as text in the video file; and perform natural language processing on the first and second recognition results. For example, the server can, but is not limited to, use a text summarization extraction method to perform text summarization processing on the first and second recognition results to obtain the text recognition result.

[0103] In the above optional embodiments, the text recognition results include the text and subtitles of the video file. The server can obtain the matching text information from the recognized text recognition results, determine the time point corresponding to the text information to be matched based on the time information, and then extract the image content associated with the text information to be matched from the video file using computer vision based on the time point to obtain the image recognition results.

[0104] Because the video files in this embodiment have clear themes—each part of the news video file has a clear theme, different themes within the same news video file are clearly distinguishable, and natural language content (e.g., text subtitles, audio, etc.) has a clear connection and guiding significance to the overall video image content—the server can use the recognized text recognition results and the corresponding time information to guide the extraction of image recognition results when extracting image recognition results from the video file. That is, based on the aforementioned text recognition results and time information, the image recognition results are extracted from the video file, thereby improving the quality of the generated video summary.

[0105] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0106] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0107] Example 2

[0108] According to an embodiment of this application, an apparatus embodiment for implementing the above-described video summarization generation method is also provided. Figure 6 This is a schematic diagram of a video summarization device according to an embodiment of this application, such as... Figure 6 As shown, the device includes: an acquisition module 40, an extraction module 42, and a generation module 44, wherein:

[0109] The acquisition module 40 is used to acquire video files; the extraction module 42 is used to extract text recognition results and image recognition results from the video files; and the generation module 44 is used to generate video summaries based on the text recognition results and image recognition results.

[0110] It should be noted that the acquisition module 40, extraction module 42, and generation module 44 mentioned above correspond to steps S202 to S206 in Embodiment 1. The three modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run on the computer terminal 10 provided in Embodiment 1.

[0111] This application embodiment is based on artificial intelligence. By extracting text recognition results and image recognition results from the acquired video file, it takes into account the comprehensive information of the video file, such as images, voice, and text. Based on the text recognition results and image recognition results, a high-quality video summary can be generated.

[0112] Therefore, the embodiments of this application achieve the goal of generating video summaries by taking into account the comprehensive information of video files such as images, voice, and text, thereby improving the quality of video summaries generated based on video files. This solves the technical problem that existing methods for generating video content summaries based on news text do not take into account the comprehensive information of images, voice, and text, making it difficult to generate high-quality video summaries.

[0113] It should also be noted that the preferred implementation of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0114] Example 3

[0115] According to an embodiment of this application, an embodiment of a video summarization generation device is also provided, which can be any one of the computing devices in a group of computing devices. Figure 7 This is a schematic diagram of the structure of a video summarization generation device according to an embodiment of this application, such as... Figure 7 As shown, the video summary generation device includes: a processor 500 and a memory 502, wherein:

[0116] A processor 500; and a memory 502, connected to the processor 500, for providing the processor with instructions to perform the following processing steps: acquiring a video file; extracting text recognition results and image recognition results from the video file; and generating a video summary based on the text recognition results and the image recognition results.

[0117] In this embodiment of the application, a video file is obtained; text recognition results and image recognition results are extracted from the video file; and a video summary is generated based on the text recognition results and the image recognition results.

[0118] It is worth noting that the embodiments of this application are based on artificial intelligence. By extracting text recognition results and image recognition results from the acquired video file, and taking into account the comprehensive information of the video file such as images, voice, and text, high-quality video summaries can be generated based on the text recognition results and image recognition results.

[0119] Therefore, the embodiments of this application achieve the goal of generating video summaries by taking into account the comprehensive information of video files such as images, voice, and text, thereby improving the quality of video summaries generated based on video files. This solves the technical problem that existing methods for generating video content summaries based on news text do not take into account the comprehensive information of images, voice, and text, making it difficult to generate high-quality video summaries.

[0120] It should also be noted that the preferred implementation of this embodiment can be found in the relevant description in Embodiment 1, and will not be repeated here.

[0121] Example 4

[0122] According to an embodiment of this application, an embodiment of a computer terminal is also provided. This computer terminal can be any one of a group of computer terminal devices. Optionally, in this embodiment, the aforementioned computer terminal can also be replaced with a mobile terminal or other terminal device.

[0123] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0124] In this embodiment, the computer terminal described above can execute the program code for the following steps in the video summary generation method: acquiring a video file; extracting text recognition results and image recognition results from the video file; and generating a video summary based on the text recognition results and image recognition results.

[0125] Optionally, Figure 8 This is a structural block diagram of another computer terminal according to an embodiment of this application, such as... Figure 8 As shown, the computer terminal may include: one or more (only one is shown in the figure) processors 602, memory 604, and peripheral interfaces 606.

[0126] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the video summarization generation method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned video summarization generation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0127] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: acquire a video file; extract text recognition results and image recognition results from the video file; and generate a video summary based on the text recognition results and image recognition results.

[0128] Optionally, the processor may also execute program code that performs the following steps: extracting a first recognition result from the subtitle data of the video file using optical character extraction; extracting a second recognition result from the audio data of the video file using speech recognition; and performing natural language processing on the first and second recognition results to obtain the character recognition result.

[0129] Optionally, the processor may also execute program code that records the time information corresponding to the text recognition results.

[0130] Optionally, the processor may also execute program code that performs the following steps: extracting the image recognition result from the video file based on the text recognition result and the time information.

[0131] Optionally, the processor may also execute program code that performs the following steps: obtaining text information to be adapted based on the text recognition result; determining the time point corresponding to the text information to be adapted based on the time information; and extracting image content associated with the text information to be adapted from the video file using computer vision based on the time point to obtain the image recognition result.

[0132] Optionally, the processor may also execute program code that performs the following steps: obtaining text information to be adapted based on the text recognition result; determining the time period corresponding to the text information to be adapted based on the time information; and extracting image content associated with the text information to be adapted from the video file using computer vision based on the time period to obtain the image recognition result.

[0133] Optionally, the processor may also execute program code for the following steps: extracting face recognition content associated with the text information to be adapted from the video file using face recognition extraction; extracting human body recognition content associated with the text information to be adapted from the video file using human body recognition extraction; extracting scene location recognition content associated with the text information to be adapted from the video file using scene location recognition extraction; and extracting event behavior recognition content associated with the text information to be adapted from the video file using event behavior recognition extraction.

[0134] This application provides a video summary generation scheme. It involves acquiring a video file; extracting text recognition and image recognition results from the video file; and generating a video summary based on the text recognition and image recognition results. This application uses artificial intelligence to extract text recognition and image recognition results from the acquired video file, taking into account the comprehensive information of the video file, including images, audio, and text. Based on the text recognition and image recognition results, a high-quality video summary can be generated.

[0135] Therefore, the embodiments of this application achieve the goal of generating video summaries by taking into account the comprehensive information of video files such as images, voice, and text, thereby improving the quality of video summaries generated based on video files. This solves the technical problem that existing methods for generating video content summaries based on news text do not take into account the comprehensive information of images, voice, and text, making it difficult to generate high-quality video summaries.

[0136] The processor can also access information and applications stored in the memory via the transmission device to perform the following steps: receiving a video file sent by the server; extracting text recognition results and image recognition results from the video file; and generating and displaying a video summary on the client based on the text recognition results and image recognition results.

[0137] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: obtain a video file; extract text recognition results and image recognition results from the video file; generate a video summary based on the text recognition results and image recognition results, and send the video summary to the client to trigger the client to display the video summary.

[0138] Those skilled in the art will understand that Figure 8 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a mobile internet device (MID), a PAD, and other terminal devices. Figure 8 This does not limit the structure of the aforementioned electronic devices. For example, a computer terminal may also include components that are more... Figure 8 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 8 The different configurations shown.

[0139] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0140] Example 5

[0141] According to an embodiment of this application, an embodiment of a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the video summarization generation method provided in Embodiment 1.

[0142] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0143] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: acquiring a video file; extracting text recognition results and image recognition results from the video file; and generating a video summary based on the text recognition results and image recognition results.

[0144] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: extracting a first recognition result from the subtitle data of the video file using optical character extraction; extracting a second recognition result from the audio data of the video file using speech recognition extraction; and performing natural language processing on the first recognition result and the second recognition result to obtain the character recognition result.

[0145] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: recording time information corresponding to the above-mentioned text recognition results.

[0146] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: extracting the image recognition results from the video file based on the text recognition results and the time information.

[0147] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining text information to be adapted based on the above text recognition results; determining the time point corresponding to the above text information to be adapted based on the above time information; and extracting image content associated with the above text information to be adapted from the above video file using computer vision based on the above time point to obtain the above image recognition results.

[0148] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining text information to be adapted based on the above text recognition results; determining the time period corresponding to the above text information to be adapted based on the above time information; and extracting image content associated with the above text information to be adapted from the above video file using computer vision according to the above time period, thereby obtaining the above image recognition results.

[0149] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: extracting face recognition content associated with the text information to be adapted from the video file using a face recognition extraction method; extracting human body recognition content associated with the text information to be adapted from the video file using a human body recognition extraction method; extracting scene location recognition content associated with the text information to be adapted from the video file using a scene location recognition extraction method; and extracting event behavior recognition content associated with the text information to be adapted from the video file using an event behavior recognition extraction method.

[0150] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a video file sent by the server; extracting text recognition results and image recognition results from the video file; and generating and displaying a video summary on the client based on the text recognition results and image recognition results.

[0151] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: acquiring a video file; extracting text recognition results and image recognition results from the video file; generating a video summary based on the text recognition results and image recognition results, and sending the video summary to the client to trigger the client to display the video summary.

[0152] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0153] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0154] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0155] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0156] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0157] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0158] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A video summarization method, characterized in that, include: Get the video file; The first recognition result is extracted from the subtitle data of the video file using optical character extraction (OCI). A second recognition result is extracted from the audio data of the video file using a speech recognition extraction method. Natural language processing is performed on the first recognition result and the second recognition result to obtain the text recognition result; Based on the text recognition result and the corresponding time information, image recognition results are extracted from the video file; A video summary is generated based on the text recognition results and the image recognition results.

2. The method according to claim 1, characterized in that, Extracting the image recognition result from the video file based on the text recognition result and the time information includes: Based on the text recognition results, obtain the text information to be adapted; Based on the time information, determine the time point corresponding to the text information to be adapted; Based on the time point, image content associated with the text information to be adapted is extracted from the video file to obtain the image recognition result.

3. The method according to claim 1, characterized in that, Extracting the image recognition result from the video file based on the text recognition result and the time information includes: Based on the text recognition results, obtain the text information to be adapted; Based on the time information, determine the time period corresponding to the text information to be adapted; Based on the time period, image content associated with the text information to be adapted is extracted from the video file to obtain the image recognition result.

4. The method according to claim 2 or 3, characterized in that, Extracting the image content associated with the text information to be adapted from the video file includes at least one of the following: The facial recognition extraction method is used to extract facial recognition content associated with the text information to be adapted from the video file; Human body recognition extraction method is used to extract human body recognition content associated with the text information to be adapted from the video file; The scene location recognition extraction method is used to extract the scene location recognition content associated with the text information to be adapted from the video file; The event behavior recognition extraction method is used to extract the event behavior recognition content associated with the text information to be adapted from the video file.

5. The method according to claim 1, characterized in that, The video file includes one of the following: Current affairs news video files; Scientific exploration video files; Historical documentary video files; Arts and entertainment video files; Video files related to sports events.

6. A video summarization method, characterized in that, include: Receive video files sent by the server; The first recognition result is extracted from the subtitle data of the video file using optical character extraction (OCI). A second recognition result is extracted from the audio data of the video file using a speech recognition extraction method. Natural language processing is performed on the first recognition result and the second recognition result to obtain the text recognition result; Based on the text recognition result and the corresponding time information, image recognition results are extracted from the video file; Based on the text recognition results and the image recognition results, a video summary is generated and displayed on the client.

7. A video summarization method, characterized in that, include: Get the video file; The first recognition result is extracted from the subtitle data of the video file using optical character extraction (OCI). A second recognition result is extracted from the audio data of the video file using a speech recognition extraction method. Natural language processing is performed on the first recognition result and the second recognition result to obtain the text recognition result; Based on the text recognition result and the corresponding time information, image recognition results are extracted from the video file; A video summary is generated based on the text recognition results and the image recognition results, and the video summary is sent to the client to trigger the client to display the video summary.

8. A video summarization generation device, characterized in that, include: The acquisition module is used to acquire video files; The extraction module is used to: extract a first recognition result from the subtitle data of the video file using optical character recognition (OCR); and extract a second recognition result from the audio data of the video file using speech recognition. Natural language processing is performed on the first recognition result and the second recognition result to obtain the text recognition result; Based on the text recognition result and the corresponding time information, image recognition results are extracted from the video file; The generation module is used to generate a video summary based on the text recognition results and the image recognition results.

9. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the video summarization generation method according to any one of claims 1 to 7.

10. A video summarization generation device, characterized in that, include: processor; as well as A memory, connected to the processor, for providing the processor with instructions to perform the following processing steps: Get the video file; The first recognition result is extracted from the subtitle data of the video file using optical character extraction (OCI). A second recognition result is extracted from the audio data of the video file using a speech recognition extraction method. Natural language processing is performed on the first recognition result and the second recognition result to obtain the text recognition result; Based on the text recognition result and the corresponding time information, image recognition results are extracted from the video file; A video summary is generated based on the text recognition results and the image recognition results.

Citation Information

Patent Citations

  • Storage method for intelligent monitoring of video data

    CN106878676A

  • Keyword-based video abstract generation method

    CN110442747A

  • Video keyword determination method and device, video retrieval method and device, storage medium and terminal

    CN110795597A

  • Video file interception method based on voice recognition

    CN111385645A