Method, system and computing device for detecting key segments in audio and video
By analyzing the multimodal characteristics of audio and video, combining with automated speech recognition text, and automatically detecting key clips in audio and video, the problem of low efficiency of users manually selecting key clips in the existing technology is solved, and more efficient key discovery and video richness are achieved.
Patent Information
- Application Number
- CN202311660482.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, users need to manually add and select key clips in audio and video, which are inefficient and cost-effective, and most users lack the ability to focus on audio and video and frequency control.
By analyzing the multimodal features of audio and video, key fragments are automatically detected, including visual features, acoustic features and natural language features, combined with automated speech recognition text, obtain a list of key words, and then determine the key fragments.
It realizes automatic detection of key clips in audio and video, reducing the cost of users manually adding subtitles and selecting key clips, has better focus discovery capabilities, and improves the richness of the video.
Smart Images

Figure CN120107838A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of audio and video technology, and more specifically, to a method, system, computing device, computer-readable storage medium, and computer program product for detecting key segments in audio and video. Background Art
[0002] With the rapid development of short video technology, users often need to add subtitles to videos. At the same time, they hope to use differentiated subtitle styles in some key subtitle segments to increase the richness of the video.
[0003] The current mainstream process of audio and video focus detection is: after manually adding video subtitles, users review the video content, manually select the corresponding key segments based on personal preferences and subtitle content, and modify the subtitle style one by one. This method is inefficient and costly, and most users lack the ability to discover audio and video highlights and control frequency. Summary of the invention
[0004] In view of this, the present disclosure provides a method, system, computing device, computer-readable storage medium and computer program product for detecting key segments in audio and video, which can automatically detect key segments by analyzing the multimodal features of audio and video, effectively saving user workload and having good key discovery capabilities.
[0005] According to a first aspect of the present disclosure, a method for detecting key segments in audio and video is provided, comprising: acquiring multimodal features of the audio and video, the multimodal features comprising visual features, acoustic features and natural language features; determining candidate key segments in the audio and video based on the multimodal features; acquiring a key word list based on automated speech recognition text of the candidate key segments; and determining the key segments in the audio and video based on the key word list.
[0006] According to a second aspect of the present disclosure, a system for detecting key segments in audio and video is provided, comprising: a feature extraction unit, configured to obtain multimodal features of the audio and video, the multimodal features including visual features, acoustic features and natural language features; a candidate key segment identification unit, configured to determine candidate key segments in the audio and video based on the multimodal features; a key word list acquisition unit, configured to obtain a key word list based on automated speech recognition text of the candidate key segments; and a key segment acquisition unit, configured to determine the key segments in the audio and video based on the key word list.
[0007] According to a third aspect of the present disclosure, a computing device is provided, comprising: at least one processing unit; and at least one memory, wherein the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, wherein the instructions, when executed by the at least one processing unit, enable the computing device to execute the method as described in the first aspect of the present disclosure.
[0008] According to a fourth aspect of the present disclosure, a non-transitory computer storage medium is provided, comprising machine executable instructions, which, when executed by a device, cause the device to perform the method as described in the first aspect of the present disclosure.
[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising machine executable instructions, which, when executed by a device, cause the device to perform the method as described in the first aspect of the present disclosure.
[0010] It should be understood that the invention summary is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of the embodiments of the present disclosure will become more easily understood through the following detailed description with reference to the accompanying drawings. In the accompanying drawings, various embodiments of the present disclosure will be described in an exemplary and non-limiting manner, in which:
[0012] Figure 1 A block diagram showing a computing device capable of implementing various embodiments of the present disclosure;
[0013] Figure 2 A schematic block diagram showing a framework of a focus detector according to an embodiment of the present disclosure is shown;
[0014] Figure 3 A schematic flow chart of a method for detecting key segments in audio and video according to an embodiment of the present disclosure is shown;
[0015] Figure 4 A schematic diagram of a key word list acquisition unit according to an embodiment of the present disclosure is shown;
[0016] Figure 5A A schematic diagram of an audio and video input page for detecting key segments in audio and video according to an embodiment of the present disclosure is shown;
[0017] Figure 5B A schematic diagram showing an output result of highlighting for detecting key segments in audio and video according to an embodiment of the present disclosure; and
[0018] Figure 6 A schematic block diagram of a device for detecting key segments in audio and video according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0019] The concept of the present disclosure will now be described with reference to the various exemplary embodiments shown in the accompanying drawings. It should be understood that the description of these embodiments is only to enable those skilled in the art to better understand and further implement the present disclosure, and is not intended to limit the scope of the present disclosure in any way. It should be noted that similar or identical reference numerals may be used in the figures where feasible, and similar or identical reference numerals may represent similar or identical elements. Those skilled in the art will understand from the following description that alternative embodiments of the structures and / or methods described herein may be adopted without departing from the principles and concepts of the present disclosure described.
[0020] In the context of the present disclosure, the term "including" and its various variations may be understood as open terms, which means "including but not limited to"; the term "based on" may be understood as "based at least in part on"; the term "one embodiment" may be understood as "at least one embodiment"; the term "another embodiment" may be understood as "at least one other embodiment". Other terms that may appear but are not mentioned here should not be interpreted or limited in a manner contrary to the concept on which the embodiments of the present disclosure are based, unless explicitly stated.
[0021] With the development of short video technology, users often need to add subtitles to videos and hope to use different subtitle styles in key subtitle segments to increase the richness and attractiveness of the video. The current mainstream audio and video key detection process is: after users manually add subtitles, they need to review the video content one by one, manually select important segments according to their preferences and subtitle content, and modify the corresponding subtitle style. This method is inefficient and costly, and most users lack the ability to discover and control the frequency of audio and video key content.
[0022] In order to solve or alleviate the above problems and / or other potential problems, an embodiment of the present disclosure proposes a method for detecting key segments in audio and video. The method extracts multimodal features in audio and video, analyzes the multimodal features separately, obtains candidate key segments, and further screens the candidate key segments in combination with the automated speech recognition text to determine the final key segments. In this way, key segments in audio and video can be automatically detected, thereby reducing the cost of users manually adding subtitles and selecting key segments, and has better key discovery capabilities compared to users selecting key segments based on personal preferences.
[0023] The basic principles and implementations of the present disclosure are described below with reference to the accompanying drawings. It should be understood that the exemplary embodiments given are only to enable those skilled in the art to better understand and implement the embodiments of the present disclosure, and are not intended to limit the scope of the present disclosure in any way.
[0024] Figure 1 1 is a block diagram of a computing device 100 capable of implementing multiple embodiments of the present disclosure. It should be understood that Figure 1 The computing device 100 shown is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described in the present disclosure. Figure 1 As shown, components of computing device 100 may include, but are not limited to, one or more processors or processing units 110 , memory 120 , storage device 130 , one or more communication units 140 , one or more input devices 150 , and one or more output devices 160 .
[0025] In some implementations, the computing device 100 can be implemented as various user terminals or service terminals with computing capabilities. The service terminal can be a server, a large computing device, etc. provided by various service providers. The user terminal is such as any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the computing device 100 can support any type of interface for the user (such as a "wearable" circuit, etc.).
[0026] Processing unit 110 may be a real or virtual processor and may be capable of performing various processes according to a program stored in memory 120. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to increase the parallel processing capabilities of computing device 100. Processing unit 110 may also be referred to as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, a controller, or a microcontroller.
[0027] The computing device 100 typically includes a plurality of computer storage media. Such media may be any available media accessible to the computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 120 may be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The memory 120 may include a focus detector 122 implemented as a program module, and the focus detector 122 may be configured as a program module for performing the audio and video focus detection functions described herein. The focus detector 122 may be accessed and run by the processing unit 110 to implement the corresponding functions.
[0028] Storage device 130 may be a removable or non-removable medium and may include machine-readable media that can be used to store information and / or data and can be accessed within computing device 100. Computing device 100 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 1 As shown in , a disk drive for reading or writing from a removable, nonvolatile disk and an optical drive for reading or writing from a removable, nonvolatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces.
[0029] The communication unit 140 enables communication with another computing device via a communication medium. Additionally, the functions of the components of the computing device 100 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Therefore, the computing device 100 can operate in a networked environment using a logical connection with one or more other servers, a personal computer (PC), or another general network node.
[0030] Input device 150 may be one or more various input devices, such as a mouse, keyboard, trackball, touch screen, voice input device, etc. Output device 160 may be one or more output devices, such as a display, speaker, printer, etc. Computing device 100 may also communicate with one or more external devices (not shown) through communication unit 140 as needed, such as storage devices, display devices, etc., communicate with one or more devices that allow a user to interact with computing device 100, or communicate with any device that allows computing device 100 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0031] In some implementations, in addition to being integrated on a single device, some or all of the various components of the computing device 100 can also be set in the form of a cloud computing architecture. In a cloud computing architecture, these components can be remotely arranged and can work together to implement the functions described in the present disclosure. In some implementations, cloud computing provides computing, software, data access and storage services, which do not require end users to know the physical location or configuration of the system or hardware that provides these services. In various implementations, cloud computing uses appropriate protocols to provide services through a wide area network (such as the Internet). For example, a cloud computing provider provides applications through a wide area network, and they can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be merged at a remote data center location or they can be dispersed. Cloud computing infrastructure can provide services through a shared data center, even if they appear as a single access point for users. Therefore, the components and functions described herein can be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they can also be provided from a traditional server, or they can be installed on a client device directly or otherwise.
[0032] The computing device 100 can detect key segments in audio and video according to various implementations of the present disclosure. Figure 1 As shown, the computing device 100 can receive audio and video 170 through the input device 150, and the audio and video 170 can be a video including voice without subtitles provided by the user. Alternatively, the computing device 100 can also read the audio and video 170 from the storage device 130 or receive the audio and video 170 from other devices (e.g., mobile phones, tablets, personal computers, etc.) from the communication device 140. The computing device 100 can transmit the audio and video 170 to the focus detector 122. The focus detector 122 detects the focus segments 180 therein based on the audio and video 170. When generating subtitles, differentiated subtitle styles are applied to the detected focus segments 180, thereby increasing the richness of the video.
[0033] For example, the audio and video 170 is a video without subtitles recorded by a user explaining how to cook, which can be a video in various languages, such as English, Chinese, etc. Accordingly, the key segment 180 detected by the key detector 122 based on the audio and video 170 contains the key information content of the video and has a suitable key frequency. In the case where the audio and video 170 is other videos without subtitles including voice, the key segment 180 can also contain the key information content therein, without being limited to a specific audio and video.
[0034] The technical solution described above is only used as an example and does not limit the present invention. Figure 2 The process of detecting the key segment 180 according to the audio and video 170 will be described in more detail.
[0035] Figure 2 FIG. 2 is a schematic block diagram showing a framework of a focus detector 200 according to an embodiment of the present disclosure. Figure 1 An example implementation of the focus detector 122. It should be noted that, Figure 2 The illustrated focus detector 200 is only illustrative, and the focus detector 200 may also be implemented by a system or framework different from this, for example, some modules may be omitted or changed, and is not limited to Figure 2 The framework shown.
[0036] like Figure 2 As shown, the focus detector 200 can receive the audio and video 170 input by the user. The input audio and video 170 can be a video without subtitles including voice, for example, a video with cooking instructions recorded by the user through a mobile phone. In some embodiments, the focus detector 200 can extract visual features 202 in the audio and video 170 through a visual feature extraction unit 201, extract acoustic features 204 in the audio and video 170 through an acoustic feature extraction unit 203, and extract natural language features 206 in the audio and video 170 through a natural language feature extraction unit 205.
[0037] In some embodiments, the visual feature extraction unit 201 can use the object detection technology in the field of computer vision (CV) to extract the visual features 202 in the audio and video 170. The basic process of object detection is to find the target of interest in the image of the audio and video 170, determine the target category and output the corresponding coordinate position, that is, recognition and positioning. The visual features 202 include the picture features of the visual angle. Optionally, the picture features can be color features, shape features, motion features, etc.
[0038] In some embodiments, the acoustic feature extraction unit 203 may use the audio event detection (Audio Event Detection, AED) technology to extract the acoustic features 204 in the audio and video 170. The acoustic features 204 may be specific sound events. The sound event detection AED can identify and classify specific sound events in the audio and video 170, such as applause, laughter, impact, etc. Optionally, the sound event detection AED may be based on the Mel-Frequency Cepstral Coefficients (MFCC), which can simulate the characteristics of the human auditory system to detect and identify sound events. Optionally, the sound event detection AED may also be based on filter banks (Filter Banks, Fbanks), using filter banks to analyze and process sound signals to detect and identify sound events.
[0039] In some embodiments, the natural language feature extraction unit 205 may use natural language processing (NLP) technology to extract natural language features 206 in the audio and video 170. Optionally, a knowledge graph (KG) and a pre-trained text detection model using a transformer-based bidirectional encoder representation technology (BERT) may be combined to detect the highlighted segments in the automatic speech recognition (ASR) text corresponding to the audio in the audio and video 170.
[0040] As shown, visual features 202, acoustic features 204, and natural language features 206 extracted from the audio and video 170 may be provided to a candidate key segment identification unit 207. In some embodiments, the input multimodal features may be classified and scored by a multimodal key point classifier.
[0041] The multimodal key point classifier is a classifier that can process multiple modal data at the same time. It can use technologies such as convolutional neural network (CNN) and recurrent neural network (RNN) to classify different types of information such as audio, video and text, and score each sample according to the characteristics of each modality and the interaction between them to evaluate the quality, similarity or relevance of the sample. In response to the scoring result exceeding the preset threshold, the candidate key segment identification unit 207 can determine multiple segments that may contain key information, and remove the audio and video segments without ASR text through further filtering, thereby obtaining the candidate key segment 208.
[0042] As shown in the figure, the candidate key segment 208 can be provided to the key word list acquisition unit 209. In some embodiments, the key word list acquisition unit 209 can use a recall algorithm to extract candidate key words from the ASR text of the candidate key segment 208. The recall algorithm is a method of screening out items related to user needs from a large number of candidate items. Optionally, the candidate key words can be recalled based on a pre-trained deep learning model, or based on a pre-defined vocabulary or dictionary, or by analyzing data patterns.
[0043] In some embodiments, the key word list acquisition unit 209 can obtain the key word list 210 by sorting the candidate key words. For example, if the audio and video 170 is identified as a tourism theme, the candidate key words under the tourism tag of the first priority are first determined as key words. If the number of key words under the tourism tag cannot meet the key word frequency in the unit time interval, the candidate key words under the second priority tag (for example, food) are determined as key words. If it is still not enough, the candidate key words under the next priority tag (for example, photography) are determined as key words, and so on, until the number of key words in the unit time interval meets the condition of the key word frequency.
[0044] Optionally, the subject of the audio and video 170 and the tags of the candidate key segments 208 can be obtained via a multimodal key classifier. Optionally, the priority of each tag can be obtained from a knowledge graph. The knowledge graph contains entities and their corresponding tags, and tag priorities are specified. Optionally, tag priorities can be obtained based on data statistics and machine learning. In some implementations, tag priorities can be user-defined.
[0045] As shown in the figure, the key word list 210 can be provided to the key segment acquisition unit 211 to obtain the key segment 180. The key words in the key word list 210 have associated timestamp information, so the key segment acquisition unit 211 can locate the corresponding time interval list based on the key word list 210, thereby obtaining the corresponding key segment 180.
[0046] Figure 3 FIG. 3 is a flow chart of a method 300 for detecting key segments in audio and video according to some embodiments of the present disclosure. In some embodiments, the method 300 may be performed by, for example, Figure 1 More specifically, the method 300 may be implemented by the computing device 100 shown in FIG. Figure 1 It should be understood that the method 300 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect. Figure 2The framework shown is used to illustrate the method 300 .
[0047] like Figure 3 As shown, in box 310, the computing device 100 obtains multimodal features of audio and video, and the multimodal features include visual features 202, acoustic features 203, and natural language features 204. In some embodiments, the computing device 100 can be a local device, such as a mobile phone, and the user can operate in an application (APP) to input audio and video. In some embodiments, the computing device 100 can be a server on the Internet, such as a cloud server, which receives audio and video transmitted from the user's mobile phone via the network.
[0048] In some embodiments, the computing device 100 may extract visual features 202 in the audio and video 170 through target detection, and the visual features 202 may include picture features of the audio and video 170. Optionally, the picture features may be color features, shape features, motion features, etc. In some embodiments, the computing device 100 may extract acoustic features 203 in the audio and video 170 through an acoustic feature extraction unit 203, and the acoustic features 203 include sound events, such as applause, laughter, etc. In some embodiments, the computing device 100 may extract natural language features 204 in the audio and video 170 based on a knowledge graph and a pre-trained text detection model, and the natural language features 204 include automated speech recognition text.
[0049] like Figure 3 As shown, in box 320, the computing device 100 determines the candidate key segments 208 in the audio and video based on the multimodal features. In some embodiments, the computing device 100 can classify and score the multimodal features through a multimodal key classifier, and in response to the scoring result exceeding a preset threshold, determine the timestamp of the segment containing the key information, thereby determining multiple segments including the key information in the audio and video 170. Subsequently, the computing device 100 can identify the automated speech recognition text of these segments, and obtain the candidate key segments 208 by filtering out the segments without automated speech recognition text in these segments. The segments without automated speech recognition text, that is, the segments do not contain speech information, so there is no need to generate subtitles and corresponding key segments for them. In some embodiments, the computing device 100 can also obtain the topic category labels of the audio and video 170 and the labels of the candidate key segments based on the output results of the multimodal key classifier and the knowledge graph.
[0050] like Figure 3As shown, in block 330, the computing device 100 may obtain the key word list 210 based on the automated speech recognition text of the candidate key segment 208. In some embodiments, the computing device 100 may extract the candidate key words based on the automated speech recognition text, and then sort the candidate key words based on the tags of the candidate segments related to the knowledge graph to determine the key word list 210. Figure 4 The process of obtaining the key word list 210 is described in more detail.
[0051] Figure 4 FIG. 4 is a schematic diagram showing a key word list acquisition unit 400 according to an embodiment of the present disclosure. The key word list acquisition unit 400 may be Figure 2 The exemplary implementation of the key word list acquisition unit 209 is shown in FIG. Figure 4 As shown, the automated speech recognition text and language (for example, obtained using the natural language feature extraction unit 205) corresponding to the candidate key segment 208 can be provided to the recall unit 401 in the key word list acquisition unit 209 to obtain the candidate key word list 402, and then the key word filter 403 can be used to filter to obtain the key word list 210.
[0052] Candidate key words can be mined in a variety of ways. In some embodiments, recall may include model-based recall. For example, the recall unit 401 may mine candidate key words according to the semantic information of the automated speech recognition text through a pre-trained deep learning model. Additionally or alternatively, recall may include recall based on vocabulary matching. The recall unit 401 may mine candidate key words by querying a pre-defined vocabulary or dictionary for matching, and determine the words appearing in the vocabulary or dictionary as candidate key words. The vocabulary and dictionary may be obtained based on a knowledge graph. Additionally or alternatively, recall may also include pattern matching-based recall. The recall unit 401 may recall the candidate key word list 402 by matching the data pattern or structure from large-scale data. For example, information such as time and place may be determined as candidate keywords. It should be noted that the above-mentioned recall methods may be combined in any manner, and the present disclosure is not limited thereto.
[0053] The candidate key word list 402 is further provided to the key word filter 403. The key word filter 403 can be configured with label priority and key word frequency conditions. In some embodiments, the key word filter 403 can sort the candidate key word list 402 according to the labels of the candidate key segments 208 where the candidate key words are located and the label priority. As mentioned above, the labels of the candidate key segments 208 can be obtained by a multimodal focus classifier. The label priority can be determined based on a knowledge graph, which can be a vertical knowledge graph of a specific field (e.g., food, agriculture, tourism, etc.). In the label priority, the label of the theme of the current audio and video can be the first priority, and a lower priority label can be determined based on the relationship of the label in the knowledge graph (e.g., child label, parent label, brother label, etc.) and the distance between the labels.
[0054] The key word filter 403 can also further screen the sorted candidate key word list 402 according to the key word frequency condition, so as to determine the key word list 210. The key word frequency condition specifies the maximum number or proportion of key words allowed within a period of time. The key word filter 403 first sets the theme tag of the audio and video 170 as the first priority tag, and uses the tag as the target tag, and then determines the candidate key words belonging to the target tag as the key words. If the key word frequency condition is not met at this time, the key word filter 403 uses the second priority tag of the next level as the target tag, and determines the candidate key words belonging to the current target tag as the key words. And so on, until the key word frequency condition is met, the key word list 210 is finally obtained.
[0055] return Figure 3 In block 340, the computing device 100 may determine a key segment in the audio or video based on the key word list. Figure 2 , the key word list 210 can be provided to the key segment acquisition unit 211. The key segment acquisition unit 211 can locate the time interval list of the final key segment 180 based on the timestamp of the key words in the key word list 210, thereby obtaining the key segment 180. In some embodiments, a reminder message can be issued to the user when the key segment 180 is played, for example, the corresponding key words are displayed in a specific style. In some implementations, the style can be related to the label of the key word or key segment, for example, different styles are applied according to the label priority. In some implementations, the user can adjust the style according to his or her own needs.
[0056] Figure 5A-Figure 5B The user interaction process of automatically detecting key segments in audio and video according to some embodiments of the present disclosure is shown. Figure 5AA schematic diagram of an audio and video input page 500A for detecting key segments in audio and video according to an embodiment of the present disclosure is shown. The audio and video input page 500A includes input audio and video 501 and an automatic focus control 502. The input audio and video 501 may be a video without subtitles and including voice. Figure 5A As shown, the audio and video 501 input by the user is a video without subtitles that explains the cooking method of Mapo Tofu. After uploading the audio and video, the user can click the automatic highlighting control 502 to enter the highlighting output result page 500B to obtain the video with subtitles and the subtitles with key keywords marked.
[0057] Figure 5B FIG. 5 is a schematic diagram of a highlight output result page 500B for detecting key segments in audio and video according to an embodiment of the present disclosure. The highlight output result page 500B includes the generated audio and video with subtitles 503 and the real-time subtitles 504 with the highlighted key words. Figure 5B As shown, compared with the input audio and video 501, the generated audio and video 503 has generated subtitles, and in the real-time subtitles 504, "pepper" as a word under the seasoning tag has been marked as a key word.
[0058] References Figures 2 to 5B An exemplary embodiment of the present disclosure is described. Compared with the existing subtitle adding scheme, the scheme of the present disclosure for detecting key segments in audio and video can establish an automated processing flow, automatically generate subtitles and determine key segments based on the multimodal features of the input audio and video, thereby facilitating users to perform subsequent subtitle style additions, effectively reducing manpower and time costs. In some implementations, the scheme of the present disclosure can have better video key discovery capabilities by simultaneously introducing multiple recall algorithms to obtain candidate key words and further obtaining the final key words through label priority sorting. In some implementations, the scheme of the present disclosure also controls the frequency of key words within a unit time interval by introducing key word frequency conditions, thereby being able to guide users to use differentiated subtitle styles more discriminatively and increase the richness of the video.
[0059] Figure 6 FIG. 6 is a schematic block diagram of an apparatus 600 for detecting a key segment in an audio or video according to an embodiment of the present disclosure. The apparatus 600 may be implemented in, for example, Figure 1 At the focus detector 122 in the computing device 100 shown. Figure 6 As shown, the device 600 includes: a feature extraction unit 610 , a candidate key segment identification unit 620 , a key word list acquisition unit 630 and a key segment acquisition unit 640 .
[0060] In some embodiments, the feature extraction unit 610 is configured to obtain multimodal features of the audio and video, the multimodal features including visual features, acoustic features and natural language features; the candidate key segment identification unit 620 is configured to determine the candidate key segments in the audio and video based on the multimodal features; the key word list acquisition unit 630 is configured to obtain a key word list based on the automated speech recognition text of the candidate key segments; and the key segment acquisition unit 640 is configured to determine the key segments in the audio and video based on the key word list.
[0061] It should be noted that the reference Figures 2 to 5B Further actions or steps shown can be performed by Figure 6 For example, the device 600 may include more modules or units to implement the actions or steps described above, or Figure 6 Some of the units or modules shown may be further configured to implement the actions or steps described above, which will not be repeated here.
[0062] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0063] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0064] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0065] The computer program instructions for performing the disclosed operation may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, and conventional procedural programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, executed as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In certain embodiments, by utilizing the state information of a computer-readable program instruction to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute a computer-readable program instruction, thereby realizing various aspects of the present disclosure.
[0066] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0067] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0068] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the equipment, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0069] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for detecting key segments in audio and video, include: Acquire multimodal features of audio and video, wherein the multimodal features include visual features, acoustic features, and natural language features; Based on the multimodal features, determining candidate key segments in the audio and video; Based on the automated speech recognition text of the candidate key segments, obtaining a key word list; as well as Based on the key word list, key segments in the audio and video are determined.
2. The method according to claim 1, in, Determining the candidate key segments in the audio and video based on the multimodal features includes: Based on the multimodal features, determining a plurality of segments in the audio and video; identifying automated speech recognition text of the plurality of segments; and The candidate key segments are obtained by filtering out segments without automated speech recognition text from the multiple segments.
3. The method according to claim 2, in, Determining the multiple segments in the audio and video includes: Classifying and scoring the multimodal features; and In response to the scoring result exceeding a preset threshold, multiple segments in the audio and video are determined.
4. The method according to claim 1, in, The automatic speech recognition text acquisition of the key word list based on the candidate key segments includes: Based on the automated speech recognition text of the candidate key segment, obtaining candidate key words; and Based on the labels of the candidate key segments related to the knowledge graph, the candidate key words are sorted to obtain the key word list.
5. The method according to claim 4, in, Obtaining candidate key words includes: The candidate key words are recalled from the automated speech recognition text of the candidate key segments, wherein the recall includes at least one of the following: model-based recall; vocabulary matching-based recall; or data pattern matching-based recall.
6. The method according to claim 4, in, Sorting the candidate key words to obtain the key word list includes: sorting the candidate key words based on the tags and tag priorities of the candidate key segments where the candidate key words are located; and Based on the key word frequency condition, the key word list is obtained from the sorted candidate key words.
7. The method according to claim 6, in, The focus word frequency condition specifies the maximum number or proportion of focus words allowed within a period of time.
8. The method according to claim 6, in, The tag priority is determined based on the knowledge graph.
9. The method according to claim 1, in, The multimodal features of audio and video include: The visual features of the audio and video are obtained from the audio and video through target detection, and the visual features include picture features of the audio and video.
10. The method according to claim 1, in, The multimodal features of audio and video include: The acoustic features are acquired from the audio and video through sound event detection, and the acoustic features include sound events.
11. The method according to claim 1, in, The multimodal features of audio and video include: Based on the knowledge graph and the pre-trained text detection model, natural language features are obtained from the audio and video, and the natural language features include automated speech recognition text.
12. A system for detecting key segments in audio and video, include: A feature extraction unit is configured to obtain multimodal features of audio and video, wherein the multimodal features include visual features, acoustic features, and natural language features; A candidate key segment identification unit is configured to determine a candidate key segment in the audio and video based on the multimodal features; A key word list acquisition unit is configured to acquire a key word list based on the automated speech recognition text of the candidate key segment; as well as The key segment acquisition unit is configured to determine the key segments in the audio and video based on the key word list.
13. A computing device, include: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform the method as claimed in any one of claims 1 to 11.
14. A non-transitory computer storage medium comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 11.
15. A computer program product comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 11.