Method for detecting key segment in audio and video, system, and computing device
By extracting multimodal features of audio and video and automatically detecting key clips in audio and video, the problem of low efficiency of user manual selection of key clips in the existing technology is solved, and more efficient key clip detection and subtitle style application are achieved.
Patent Information
- Application Number
- PCT/CN2024/125837
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-10-18
- Publication Date
- 2025-06-12
AI Technical Summary
In the prior art, users need to manually add and select key clips in audio and video, which are inefficient and cost-effective, and most users lack the ability to focus on audio and video and frequency control.
By obtaining the multimodal features of audio and video, including visual features, acoustic features and natural language features, candidate key segments are determined, and the list of key words is obtained in combination with automated speech recognition text, and the key segments in audio and video are finally determined.
It realizes automatic detection of key clips in audio and video, reducing the cost and workload of users to manually add subtitles and select key clips, and improving the ability to discover key clips.
Smart Images

Figure CN2024125837_12062025_PF_FP_ABST
Abstract
Description
Method, system and computing device for detecting key segments in audio and video
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed on December 5, 2023, with application number 202311660482.9 and invention name “Method, system and computing device for detecting key segments in audio and video”. The entire contents of the application are incorporated by reference into this application. Technical Field
[0003] The present disclosure relates to the field of audio and video technology, and more particularly, to a method, system, computing device, computer-readable storage medium, and computer program product for detecting key segments in audio and video. Background Art
[0004] With the rapid development of short video technology, users often need to add subtitles to videos. At the same time, they hope to use differentiated subtitle styles in some key subtitle segments to increase the richness of the video.
[0005] The current mainstream process for detecting audio and video highlights involves users manually adding subtitles to a video, then reviewing the video content, manually selecting key segments based on their preferences and the subtitle content, and modifying the subtitle style one by one. This approach is inefficient and costly, and most users lack the ability to identify and control the frequency of audio and video highlights.
[0006] Summary of the Invention
[0007] In view of this, the present disclosure provides a method, system, computing device, computer-readable storage medium, and computer program product for detecting key segments in audio and video.
[0008] According to a first aspect of the present disclosure, a method for detecting key segments in audio and video is provided, comprising: obtaining multimodal features of the audio and video, the multimodal features comprising visual features, acoustic features and natural language features; determining candidate key segments in the audio and video based on the multimodal features; obtaining a list of key words based on automated speech recognition text of the candidate key segments; and determining the key segments in the audio and video based on the list of key words.
[0009] According to a second aspect of the present disclosure, a system for detecting key segments in audio and video is provided, comprising: a feature extraction unit configured to obtain multimodal features of the audio and video, wherein the multimodal features include visual features, acoustic features, and natural language features; a candidate key segment identification unit configured to determine candidate key segments in the audio and video based on the multimodal features; a key word list acquisition unit configured to obtain a key word list based on automated speech recognition text of the candidate key segments; and a key segment acquisition unit configured to determine the key segments in the audio and video based on the key word list.
[0010] According to a third aspect of the present disclosure, a computing device is provided, comprising: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to execute the method as described in the first aspect of the present disclosure.
[0011] According to a fourth aspect of the present disclosure, a non-transitory computer storage medium is provided, comprising machine-executable instructions, which, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.
[0012] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising machine-executable instructions, which, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.
[0013] It should be understood that the summary of the invention is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above and other objects, features and advantages of the embodiments of the present disclosure will become more readily understood through the following detailed description with reference to the accompanying drawings, in which several embodiments of the present disclosure are illustrated by way of example and not limitation, in which:
[0015] FIG1 illustrates a block diagram of a computing device capable of implementing various embodiments of the present disclosure;
[0016] FIG2 shows a schematic block diagram of a framework of a focus detector according to an embodiment of the present disclosure;
[0017] FIG3 shows a schematic flow chart of a method for detecting key segments in audio and video according to an embodiment of the present disclosure;
[0018] FIG4 shows a schematic diagram of a key word list acquisition unit according to an embodiment of the present disclosure;
[0019] FIG5A shows a schematic diagram of an audio and video input page for detecting key segments in audio and video according to an embodiment of the present disclosure;
[0020] FIG5B is a schematic diagram showing an output result of highlighting for detecting key segments in audio and video according to an embodiment of the present disclosure; and
[0021] FIG6 shows a schematic block diagram of an apparatus for detecting key segments in audio and video according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] The concepts of the present disclosure will now be described with reference to the various exemplary embodiments shown in the accompanying drawings. It should be understood that the description of these embodiments is merely to enable those skilled in the art to better understand and further implement the present disclosure, and is not intended to limit the scope of the present disclosure in any way. It should be noted that similar or identical reference numerals may be used in the figures where possible, and similar or identical reference numerals may represent similar or identical elements. It will be understood by those skilled in the art from the description below that alternative embodiments of the structures and / or methods described herein may be adopted without departing from the principles and concepts of the present disclosure described.
[0023] In the context of this disclosure, the term "including" and its various variations can be understood as open-ended terms, meaning "including but not limited to," the term "based on" can be understood as "based, at least in part, on," the term "one embodiment" can be understood as "at least one embodiment," and the term "another embodiment" can be understood as "at least one other embodiment." Other terms that may appear but are not mentioned here should not be interpreted or limited in a manner that is inconsistent with the concepts underlying the embodiments of this disclosure, unless explicitly stated.
[0024] With the development of short video technology, users often need to add subtitles to their videos, and they hope to use different subtitle styles for key subtitle segments to increase the richness and appeal of the video. The current mainstream audio and video key point detection process requires users to manually add subtitles, then review the video content one by one, manually select important segments based on their preferences and subtitle content, and modify the corresponding subtitle styles. This method is inefficient and costly, and most users lack the ability to discover and control the frequency of key audio and video content.
[0025] To solve or alleviate the above-mentioned problems and / or other potential problems, an embodiment of the present disclosure proposes a method for detecting key segments in audio and video. This method extracts multimodal features from audio and video, analyzes the multimodal features separately, obtains candidate key segments, and further screens the candidate key segments in combination with the automated speech recognition text to determine the final key segments. In this way, key segments in audio and video can be automatically detected, thereby reducing the cost and workload of users manually adding subtitles and selecting key segments, and has better key discovery capabilities compared to users selecting key segments based on personal preferences.
[0026] The following describes the basic principles and implementations of the present disclosure with reference to the accompanying drawings. It should be understood that the exemplary embodiments provided are only intended to enable those skilled in the art to better understand and implement the embodiments of the present disclosure, and are not intended to limit the scope of the present disclosure in any way.
[0027] FIG1 illustrates a block diagram of a computing device 100 capable of implementing various embodiments of the present disclosure. It should be understood that the computing device 100 illustrated in FIG1 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. As shown in FIG1 , the components of the computing device 100 may include, but are not limited to, one or more processors or processing units 110, a memory 120, a storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.
[0028] In some implementations, the computing device 100 can be implemented as various user terminals or service terminals with computing capabilities. The service terminal can be a server, a large computing device, etc. provided by various service providers. The user terminal is such as a mobile terminal, a fixed terminal, or a portable terminal of any type, including a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is also foreseeable that the computing device 100 can support any type of interface for the user (such as a "wearable" circuit, etc.).
[0029] Processing unit 110 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 100. Processing unit 110 may also be referred to as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, a controller, or a microcontroller.
[0030] The computing device 100 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 120 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The memory 120 can include a focus detector 122 implemented as a program module, which can be configured as a program module to perform the audio and video focus detection functions described herein. The focus detector 122 can be accessed and run by the processing unit 110 to implement the corresponding functions.
[0031] The storage device 130 may be a removable or non-removable medium and may include machine-readable media that can be used to store information and / or data and can be accessed within the computing device 100. The computing device 100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG1 , a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces.
[0032] The communication unit 140 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 100 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 100 can operate in a networked environment using logical connections to one or more other servers, personal computers (PCs), or another general network node.
[0033] Input device 150 may be one or more of various input devices, such as a mouse, keyboard, trackball, touch screen, voice input device, etc. Output device 160 may be one or more output devices, such as a display, speaker, printer, etc. Computing device 100 may also communicate with one or more external devices (not shown) via communication unit 140 as needed, such as storage devices, display devices, etc., with one or more devices that allow a user to interact with computing device 100, or with any device that allows computing device 100 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0034] In some implementations, in addition to being integrated on a single device, some or all of the various components of computing device 100 may be configured in the form of a cloud computing architecture. In a cloud computing architecture, these components may be remotely located and work together to implement the functionality described herein. In some implementations, cloud computing provides computing, software, data access, and storage services that do not require the end user to be aware of the physical location or configuration of the systems or hardware providing these services. In various implementations, cloud computing provides services over a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides applications over a wide area network, and these applications can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated at remote data center locations or they may be dispersed. Cloud computing infrastructure can provide services through shared data centers, even though they appear to be a single access point for users. Therefore, the components and functionality described herein can be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they can be provided from traditional servers, or they can be installed directly or otherwise on the client device.
[0035] The computing device 100 can detect key segments in audio and video according to various implementations of the present disclosure. As shown in Figure 1, the computing device 100 can receive audio and video 170 through the input device 150. The audio and video 170 can be a video including voice without subtitles provided by the user. Alternatively, the computing device 100 can also read the audio and video 170 from the storage device 130 or receive the audio and video 170 from other devices (for example, mobile phones, tablets, personal computers, etc.) from the communication device 140. The computing device 100 can transmit the audio and video 170 to the key detector 122. The key detector 122 detects the key segments 180 therein based on the audio and video 170. When generating subtitles, differentiated subtitle styles are applied to the detected key segments 180, thereby increasing the richness of the video.
[0036] For example, if audio / video 170 is a user-recorded video explaining how to cook a dish without subtitles, and can be in various languages, such as English or Chinese, then highlight detector 122 detects highlight segments 180 based on audio / video 170, including the highlight information content of the video and having an appropriate highlight frequency. If audio / video 170 is another video without subtitles and includes speech, highlight segments 180 can also contain the highlight information content therein, and are not limited to specific audio / video content.
[0037] The technical solution described above is only for illustration and does not limit the present invention. In order to more clearly explain the principle of the above solution, the process of detecting the key segment 180 based on the audio and video 170 will be described in more detail with reference to FIG.
[0038] FIG2 shows a schematic block diagram of a framework of a focus detector 200 according to an embodiment of the present disclosure. Focus detector 200 is an example implementation of focus detector 122 of FIG1 . It should be noted that focus detector 200 shown in FIG2 is merely illustrative and can be implemented using different systems or frameworks. For example, some modules can be omitted or modified, and the framework is not limited to that shown in FIG2 .
[0039] As shown in FIG2 , a focus detector 200 can receive audio and video 170 input by a user. The input audio and video 170 can be a video without subtitles and including speech, for example, a video recorded by a user on a mobile phone with cooking instructions. In some embodiments, the focus detector 200 can extract visual features 202 from the audio and video 170 using a visual feature extraction unit 201, extract acoustic features 204 from the audio and video 170 using an acoustic feature extraction unit 203, and extract natural language features 206 from the audio and video 170 using a natural language feature extraction unit 205.
[0040] In some embodiments, visual feature extraction unit 201 can employ object detection technology from the field of computer vision (CV) to extract visual features 202 from audio and video 170. The basic process of object detection involves finding the target of interest within the audio and video 170, determining the target category, and outputting the corresponding coordinate location, i.e., identification and positioning. Visual features 202 include visual angle-based image features. Optionally, image features can include color features, shape features, motion features, and the like.
[0041] In some embodiments, the acoustic feature extraction unit 203 can use audio event detection (AED) technology to extract acoustic features 204 in the audio and video 170. The acoustic feature 204 can be a specific sound event. The sound event detection AED can identify and classify specific sound events in the audio and video 170, such as applause, laughter, collision sounds, etc. Optionally, the sound event detection AED can be based on Mel-Frequency Cepstral Coefficients (MFCC), which can simulate the characteristics of the human auditory system to detect and identify sound events. Optionally, the sound event detection AED can also be based on filter banks (Fbanks), using filter banks to analyze and process sound signals to detect and identify sound events.
[0042] In some embodiments, the natural language feature extraction unit 205 may use natural language processing (NLP) technology to extract natural language features 206 from the audio and video 170. Alternatively, a knowledge graph (KG) and a pre-trained text detection model using bidirectional encoder representation based on transformers (BERT) may be combined to detect the highlighted segments in the automatic speech recognition (ASR) text corresponding to the audio in the audio and video 170.
[0043] As shown in the figure, visual features 202, acoustic features 204, and natural language features 206 extracted from the audio and video 170 may be provided to a candidate key segment identification unit 207. In some embodiments, the input multimodal features may be classified and scored by a multimodal key point classifier.
[0044] A multimodal key point classifier is a classifier that can process multiple modal data simultaneously. It can use technologies such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to classify different types of information, such as audio, video, and text. It can also score each sample based on the characteristics of each modality and the interactions between them to evaluate the quality, similarity, or relevance of the sample. In response to the scoring result exceeding a preset threshold, the candidate key segment identification unit 207 can determine multiple segments that may contain key information, and after further filtering, remove audio and video segments without ASR text, thereby obtaining candidate key segments 208.
[0045] As shown in the figure, the candidate key segments 208 can be provided to the key word list acquisition unit 209. In some embodiments, the key word list acquisition unit 209 can use a recall algorithm to extract candidate key words from the ASR text of the candidate key segments 208. The recall algorithm is a method for screening out items related to user needs from a large number of candidate items. Optionally, the candidate key words can be recalled based on a pre-trained deep learning model, or based on a pre-defined vocabulary or dictionary, or by analyzing data patterns.
[0046] In some embodiments, the key word list acquisition unit 209 can obtain the key word list 210 by sorting the candidate key words. For example, if the audio and video 170 is identified as a tourism theme, the candidate key words under the tourism tag of the first priority are first determined as key words. If the number of key words under the tourism tag does not meet the key word frequency within the unit time interval, the candidate key words under the second priority tag (for example, food) are then determined as key words. If that is still not enough, the candidate key words under the next priority tag (for example, photography) are continued to be determined as key words, and so on, until the number of key words within the unit time interval meets the key word frequency condition.
[0047] Optionally, the theme of the audio and video 170 and the tags of the candidate key segments 208 can be obtained via a multimodal key classifier. Optionally, the priority of each tag can be obtained from a knowledge graph. The knowledge graph contains entities and their corresponding tags, and specifies tag priorities. Optionally, tag priorities can be obtained based on data statistics and machine learning. In some implementations, tag priorities can be user-defined.
[0048] As shown in the figure, the key word list 210 can be provided to the key segment acquisition unit 211 to obtain the key segments 180. The key words in the key word list 210 have associated timestamp information, so the key segment acquisition unit 211 can locate the corresponding time interval list based on the key word list 210, thereby obtaining the corresponding key segments 180.
[0049] FIG3 illustrates a flow diagram of a method 300 for detecting key segments in audio and video, according to some embodiments of the present disclosure. In some embodiments, method 300 may be implemented, for example, by the computing device 100 shown in FIG1 . More specifically, method 300 may be implemented by the key segment detector 122 of FIG1 . It should be understood that method 300 may also include additional actions not shown and / or may omit actions shown, and the scope of the present disclosure is not limited in this respect. For ease of illustration, method 300 will be described with reference to the framework shown in FIG2 .
[0050] As shown in Figure 3, at block 310, computing device 100 obtains multimodal features of audio and video, including visual features 202, acoustic features 203, and natural language features 204. In some embodiments, computing device 100 may be a local device, such as a mobile phone, and a user may operate an application (APP) to input audio and video. In some embodiments, computing device 100 may be a server on the Internet, such as a cloud server, that receives audio and video transmitted from the user's mobile phone via the network.
[0051] In some embodiments, the computing device 100 can extract visual features 202 from the audio and video 170 through object detection. The visual features 202 may include picture features of the audio and video 170. Optionally, the picture features may be color features, shape features, motion features, etc. In some embodiments, the computing device 100 can extract acoustic features 203 from the audio and video 170 through an acoustic feature extraction unit 203. The acoustic features 203 include sound events such as applause and laughter. In some embodiments, the computing device 100 can extract natural language features 204 from the audio and video 170 based on a knowledge graph and a pre-trained text detection model. The natural language features 204 include automated speech recognition text.
[0052] As shown in Figure 3, in box 320, the computing device 100 determines the candidate key segments 208 in the audio and video based on the multimodal features. In some embodiments, the computing device 100 can classify and score the multimodal features through a multimodal key classifier, and in response to the scoring result exceeding a preset threshold, determine the timestamp of the segment containing the key information, thereby determining multiple segments in the audio and video 170 that include key information. Subsequently, the computing device 100 can identify the automated speech recognition text of these segments, and obtain the candidate key segments 208 by filtering out the segments without automated speech recognition text in these segments. The segments without automated speech recognition text, that is, the segments do not contain speech information, and therefore do not need to generate subtitles and corresponding key segments for them. In some embodiments, the computing device 100 can also obtain the topic category labels of the audio and video 170 and the labels of the candidate key segments based on the output results of the multimodal key classifier and the knowledge graph.
[0053] As shown in FIG3 , at block 330 , the computing device 100 may obtain a list of key words 210 based on the automated speech recognition text of the candidate key segments 208 . In some embodiments, the computing device 100 may extract candidate key words based on the automated speech recognition text and then sort the candidate key words based on the knowledge graph-related tags of the candidate segments to determine the list of key words 210 . The process of obtaining the list of key words 210 will be described in more detail below with reference to FIG4 .
[0054] FIG4 shows a schematic diagram of a key word list acquisition unit 400 according to an embodiment of the present disclosure. The key word list acquisition unit 400 may be an exemplary implementation of the key word list acquisition unit 209 shown in FIG2 . In some embodiments, as shown in FIG4 , the automated speech recognition text and language corresponding to the candidate key segments 208 (e.g., acquired using the natural language feature extraction unit 205 ) may be provided to a recall unit 401 in the key word list acquisition unit 209 to obtain a candidate key word list 402 , which may then be filtered using a key word filter 403 to obtain a key word list 210 .
[0055] A variety of methods can be used to mine candidate key words. In some embodiments, recall may include model-based recall. For example, the recall unit 401 can mine candidate key words based on the semantic information of the automated speech recognition text through a pre-trained deep learning model. Additionally or alternatively, recall may include recall based on vocabulary matching. The recall unit 401 can mine candidate key words by querying a pre-defined vocabulary or dictionary for matching, and determine the words that appear in the vocabulary or dictionary as candidate key words. The vocabulary and dictionary can be obtained based on a knowledge graph. Additionally or alternatively, recall may also include pattern matching-based recall. The recall unit 401 can recall the candidate key word list 402 by matching data patterns or structures from large-scale data. For example, information such as time and place can be determined as candidate keywords. It should be noted that the above-mentioned recall methods can be combined in any way, and the present disclosure has no restrictions on this.
[0056] The candidate key word list 402 is further provided to the key word filter 403. The key word filter 403 can be configured with label priority and key word frequency conditions. In some embodiments, the key word filter 403 can sort the candidate key word list 402 according to the labels of the candidate key segments 208 where the candidate key words are located and the label priority. As mentioned above, the labels of the candidate key segments 208 can be obtained by a multimodal focus classifier. Label priority can be determined based on a knowledge graph, which can be a vertical knowledge graph for a specific field (e.g., food, agriculture, tourism, etc.). In the label priority, the label of the theme of the current audio and video can be the first priority, and lower priority labels can be determined based on the relationship between the labels in the knowledge graph (e.g., child labels, parent labels, brother labels, etc.) and the distance between the labels.
[0057] The key word filter 403 can also further screen the sorted candidate key word list 402 according to the key word frequency condition, so as to determine the key word list 210. The key word frequency condition specifies the maximum number or proportion of key words allowed within a period of time. The key word filter 403 first sets the subject tag of the audio and video 170 as the first priority tag, and uses the tag as the target tag, and then determines the candidate key words belonging to the target tag as the key words. If the key word frequency condition is not met at this time, the key word filter 403 uses the second priority tag of the next level as the target tag, and determines the candidate key words belonging to the current target tag as the key words. And so on, until the key word frequency condition is met, the key word list 210 is finally obtained.
[0058] Returning to Figure 3, in box 340, the computing device 100 can determine the key segments in the audio and video based on the key word list. Referring to Figure 2, the key word list 210 can be provided to the key segment acquisition unit 211. The key segment acquisition unit 211 can locate the time interval list of the final key segment 180 based on the timestamp of the key words in the key word list 210, thereby obtaining the key segment 180. In some embodiments, a reminder message can be issued to the user when the key segment 180 is played, for example, the corresponding key words are displayed in a specific style. In some implementations, the style can be related to the label of the key word or key segment, for example, different styles are applied according to the label priority. In some implementations, the user can adjust the style according to his or her needs.
[0059] Figures 5A-5B illustrate the user interaction process of automatically detecting key segments in audio and video according to some embodiments of the present disclosure. Figure 5A shows a schematic diagram of an audio and video input page 500A for detecting key segments in audio and video according to an embodiment of the present disclosure. The audio and video input page 500A includes input audio and video 501 and an automatic highlighting control 502. The input audio and video 501 can be a subtitle-free video including voice. As shown in Figure 5A, the audio and video 501 input by the user is a subtitle-free video explaining the cooking method of Mapo Tofu. After uploading the audio and video, the user can click the automatic highlighting control 502 to enter the highlighting output result page 500B to obtain a video with subtitles and subtitles with key keywords marked.
[0060] Figure 5B shows a schematic diagram of a highlight output result page 500B for detecting key segments in audio and video according to an embodiment of the present disclosure. The highlight output result page 500B includes a generated audio and video with subtitles 503 and real-time subtitles 504 with key words marked. As shown in Figure 5B, compared to the input audio and video 501, the generated audio and video 503 has generated subtitles, and in the real-time subtitles 504, "pepper" as a word under the seasoning tag has been marked as a key word.
[0061] The above reference figures 2 to 5B describe exemplary embodiments of the present disclosure. Compared with the existing subtitle adding solutions, the solution of the present disclosure for detecting key segments in audio and video can establish an automated processing flow, automatically generate subtitles and determine key segments based on the multimodal features of the input audio and video, thereby facilitating the user to perform subsequent subtitle style additions, effectively reducing manpower and time costs. In some implementations, the solution of the present disclosure can have better video focus discovery capabilities by simultaneously introducing multiple recall algorithms to obtain candidate key words and further obtaining the final key words through label priority sorting. In some implementations, the solution of the present disclosure also controls the frequency of key words within a unit time interval by introducing key word frequency conditions, thereby guiding users to use differentiated subtitle styles more discriminatively and increase the richness of the video.
[0062] FIG6 shows a schematic block diagram of an apparatus 600 for detecting key segments in audio or video, according to an embodiment of the present disclosure. Apparatus 600 can be implemented, for example, in the key segment detector 122 of the computing device 100 shown in FIG1 . As shown in FIG6 , apparatus 600 includes a feature extraction unit 610 , a candidate key segment identification unit 620 , a key word list acquisition unit 630 , and a key segment acquisition unit 640 .
[0063] In some embodiments, the feature extraction unit 610 is configured to obtain multimodal features of the audio and video, which include visual features, acoustic features, and natural language features; the candidate key segment identification unit 620 is configured to determine the candidate key segments in the audio and video based on the multimodal features; the key word list acquisition unit 630 is configured to obtain a key word list based on the automated speech recognition text of the candidate key segments; and the key segment acquisition unit 640 is configured to determine the key segments in the audio and video based on the key word list.
[0064] It should be noted that more actions or steps shown with reference to FIG. 2 to FIG. 5B can be implemented by the apparatus 600 shown in FIG. 6 . For example, the apparatus 600 can include more modules or units to implement the actions or steps described above, or some units or modules shown in FIG. 6 can be further configured to implement the actions or steps described above. These details will not be repeated here.
[0065] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0066] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0067] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0068] The computer program instructions for performing the disclosed operation can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or the object code written in any combination of one or more programming languages, wherein the programming languages include object-oriented programming languages, and conventional procedural programming languages.Computer-readable program instructions can be performed completely on the user's computer, partially on the user's computer, performed as an independent software package, partly on the user's computer and partly on a remote computer, or performed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network-including local area network (LAN) or wide area network (WAN), or can be connected to an external computer (such as utilizing an Internet service provider to connect by the Internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to carry out personalized customization electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLA), this electronic circuit can perform computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0069] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0070] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0071] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0072] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technical improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for detecting key segments in audio and video, comprising: Acquire multimodal features of audio and video, wherein the multimodal features include visual features, acoustic features, and natural language features; Based on the multimodal features, determining candidate key segments in the audio and video; Based on the automated speech recognition text of the candidate key segments, obtaining a key word list; as well as Based on the key word list, key segments in the audio and video are determined.
2. The method according to claim 1, wherein: Determining the candidate key segments in the audio and video based on the multimodal features includes: Based on the multimodal features, determining a plurality of segments in the audio and video; identifying automated speech recognition text of the plurality of segments; and The candidate key segments are obtained by filtering out segments without automated speech recognition text from the multiple segments.
3. The method according to claim 2, wherein: Determining the multiple segments in the audio and video includes: Classifying and scoring the multimodal features; and In response to the scoring result exceeding a preset threshold, multiple segments in the audio and video are determined.
4. The method according to claim 1, wherein: The automatic speech recognition text acquisition of the key word list based on the candidate key segments includes: Based on the automated speech recognition text of the candidate key segment, obtaining candidate key words; and Based on the labels of the candidate key segments related to the knowledge graph, the candidate key words are sorted to obtain the key word list.
5. The method according to claim 4, wherein: Obtaining candidate key words includes: The candidate key words are recalled from the automated speech recognition text of the candidate key segments, wherein the recall includes at least one of the following: model-based recall; vocabulary matching-based recall; or data pattern matching-based recall.
6. The method according to claim 4, wherein: Sorting the candidate key words to obtain the key word list includes: sorting the candidate key words based on the tags and tag priorities of the candidate key segments where the candidate key words are located; and Based on the key word frequency condition, the key word list is obtained from the sorted candidate key words.
7. The method according to claim 6, wherein: The focus word frequency condition specifies the maximum number or proportion of focus words allowed within a period of time.
8. The method according to claim 6, wherein: The tag priority is determined based on the knowledge graph.
9. The method according to claim 1, wherein: The multimodal features of audio and video include: The visual features of the audio and video are obtained from the audio and video through target detection, and the visual features include picture features of the audio and video.
10. The method according to claim 1, wherein: The multimodal features of audio and video include: The acoustic features are acquired from the audio and video through sound event detection, and the acoustic features include sound events.
11. The method according to claim 1, wherein: The multimodal features of audio and video include: Based on the knowledge graph and the pre-trained text detection model, natural language features are obtained from the audio and video, and the natural language features include automated speech recognition text.
12. A system for detecting key segments in audio and video, comprising: A feature extraction unit is configured to obtain multimodal features of audio and video, wherein the multimodal features include visual features, acoustic features, and natural language features; A candidate key segment identification unit is configured to determine a candidate key segment in the audio and video based on the multimodal features; A key word list acquisition unit is configured to acquire a key word list based on the automated speech recognition text of the candidate key segment; as well as The key segment acquisition unit is configured to determine the key segments in the audio and video based on the key word list.
13. A computing device comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform the method as claimed in any one of claims 1 to 11.
14. A non-transitory computer storage medium comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 11.
15. A computer program product comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 11.
Citation Information
Patent Citations
Video voice recognition method and device, equipment and storage medium
CN113838460A
Video clip search method and system oriented to open domain query
CN115687687A
Generating text snippets using supervised machine learning algorithm
US20170300563A1