Retrieval method and device for video content
By converting video content into text scripts and building a knowledge base, combined with natural language processing models, the problems of low efficiency and poor accuracy of video information retrieval are solved, and fast and accurate video content retrieval is achieved, reducing computing resource consumption and threshold.
Patent Information
- Application Number
- CN202510525904.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The current processing and utilization of video information faces many challenges, including the need for users to spend a lot of time manually browsing video content. The existing video processing technology has great limitations in understanding video content and providing accurate information.
By converting video content into text scripts and building a knowledge base, combining ordinary natural language processing models, answering questions based on this knowledge base, and achieving rapid retrieval of video content.
It greatly reduces the consumption of computing resources, improves the efficiency and accuracy of video information retrieval, reduces the threshold for the utilization of video information, and improves the utilization rate of video resources.
Smart Images

Figure CN120067396A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of artificial intelligence and video processing. Specifically, it relates to a method and device for retrieving video content. Background Art
[0002] In the current context of the rapid growth of digital information, videos have become an important carrier for information dissemination. Whether it is the course videos on online education platforms, the film and television works on film and television entertainment platforms, or the training videos within enterprises, their numbers are constantly increasing. It is crucial for users to accurately and quickly obtain the required information from a vast amount of videos.
[0003] However, the current processing and utilization of video information still face many challenges. On the one hand, when users query video-related information, they often need to spend a lot of time manually browsing the video content. On the other hand, existing video processing technologies have significant limitations in understanding video content and providing accurate information. Summary of the Invention
[0004] The present invention provides a method and device for retrieving video content. By converting video content into a text script and constructing a knowledge base, combined with a general natural language processing model, question answering can be performed based on this knowledge base, and users can quickly retrieve the required video content through natural language.
[0005] In a first aspect, the present invention provides a method for retrieving video content, the method comprising: Obtaining the retrieval intention of a target user; Determining the retrieval information of the question to be retrieved according to the retrieval intention; Determining the target type and associated information of the question to be retrieved according to the retrieval information; Determining the analysis result of the question to be retrieved according to the target type and the associated information; Based on a preset knowledge base, determining the script information corresponding to the question to be retrieved according to the analysis result, the knowledge base containing various video file information; Determining the target answer according to the script information.
[0006] Preferably, the determining the target type and associated information of the question to be retrieved according to the retrieval information includes: Performing intention recognition on the retrieval information based on an intention recognition model to determine the target type; Performing semantic analysis on the retrieval information based on natural language processing technology to determine the associated information.
[0007] Preferably, the steps for establishing the preset knowledge base include: Extract and process the video source file to determine the target text information; Determine the target speech text according to the audio information in the video source file; Determine the basic video information associated with the video source file; Determine the information set according to the basic video information, the target text information and the target speech text; Extract the target features in the information set; the target features include scene features, character features and plot features; Determine the target script according to the scene features, the character features and the plot features; Build the knowledge base according to the target script.
[0008] Preferably, the extracting and processing the video source file to determine the target text information includes: Preprocess the video source file to determine the video file to be recognized; Identify the text information in the video file to be recognized based on optical character recognition technology to determine the first text information; Verify and sort the first text information based on natural language processing technology to determine the target text information.
[0009] Preferably, the determining the target speech text according to the audio information in the video source file includes: Extract the audio track in the video source file; determine the audio information according to the audio track; Preprocess the audio information to determine the target audio file; Determine the first transcription result based on the speech-to-text model according to the target audio file; Determine the target transcription result based on natural language processing technology according to the first transcription result; Determine the target speech text according to the target transcription result.
[0010] Preferably, the determining the target speech text according to the target transcription result includes: Perform segmentation processing on the target transcription result to determine the segmentation result; Annotate the corresponding timestamps to the segmentation result to determine the target speech text.
[0011] Preferably, the extracting the target features in the information set includes: Use a pre-trained scene recognition model to recognize and divide the scenes in the information set to determine the scene features; Determine the role relationship and relationship model based on the speech content in the information set; determine the role characteristics based on the role relationship and the relationship model; Based on a deep learning model, extract the plot features in the information set to determine the target features.
[0012] In a second aspect, the present invention provides a retrieval device for video content, including: A retrieval intention acquisition module, configured to acquire the retrieval intention of a target user; A retrieval information determination module, configured to determine the retrieval information of a problem to be retrieved according to the retrieval intention; A target type and associated information determination module, configured to determine the target type and associated information of the problem to be retrieved according to the retrieval information; An analysis result determination module, configured to determine the analysis result of the problem to be retrieved according to the target type and the associated information; A script information determination module, configured to determine the script information corresponding to the problem to be retrieved based on a preset knowledge base according to the analysis result; A target answer determination module, configured to determine a target answer according to the script information.
[0013] In a third aspect, the present invention provides a readable medium, including execution instructions. When a processor of an electronic device executes the execution instructions, the electronic device executes the method according to any one of the first aspects.
[0014] In a fourth aspect, the present invention provides an electronic device, including a processor and a memory storing execution instructions. When the processor executes the execution instructions stored in the memory, the processor executes the method according to any one of the first aspects.
[0015] The present invention provides a method and device for retrieving video content. By obtaining various types of video file information, a knowledge base is established in advance according to the video file information; the retrieval intention of the target user is obtained; the retrieval information of the problem to be retrieved is determined according to the retrieval intention; the target type and associated information of the problem to be retrieved are determined according to the retrieval information; the analysis result of the problem to be retrieved is determined according to the target type and associated information; based on the knowledge base, the script information corresponding to the problem to be retrieved is determined according to the analysis result; and the target answer is determined according to the script information. The present invention converts video content into a text script and constructs a knowledge base, and a general natural language processing model can answer questions based on this knowledge base without the need for a dedicated video understanding model, greatly reducing the consumption of computing resources. The establishment of the knowledge base makes the retrieval of video information more efficient. It avoids the answer errors caused by inaccurate processing of visual and audio information in traditional video understanding models, and improves the accuracy and integrity of the answers. The method of the present invention does not depend on a specific video understanding model and has strong versatility. Ordinary users can obtain video-related information by simply asking questions, which reduces the threshold for using video information and improves the utilization rate of video resources.
[0016] The further effects of the above non-conventional preferred methods will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 Schematic diagram of a method for retrieving video content provided by an embodiment of the present invention; Figure 2 Schematic diagram of the generation process of the target answer in an embodiment of the present invention; Figure 3 Schematic diagram of another method for retrieving video content provided by an embodiment of the present invention; Figure 4 Schematic diagram of a device for retrieving video content provided by an embodiment of the present invention; Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0020] In the current context of the rapid growth of digital information, videos have become an important carrier for information dissemination. Whether it is the course videos on online education platforms, the film and television works on film and television entertainment platforms, or the training videos within enterprises, their numbers are constantly increasing. It is crucial for users to accurately and quickly obtain the required information from a vast amount of videos.
[0021] However, the current processing and utilization of video information still face many challenges. On the one hand, when users query video-related information, they often need to spend a lot of time manually browsing the video content; on the other hand, existing video processing technologies have significant limitations in understanding video content and providing accurate information.
[0022] Traditional video understanding and information retrieval methods mainly rely on specialized video understanding models, and these methods have the following obvious defects: High computational resource consumption. Specialized video understanding models usually need to process multi-modal information such as the vision and audio of videos, involving complex algorithms and large-scale computations, which pose extremely high requirements for hardware resources, not only increasing equipment costs but also raising energy consumption; Poor generality. Different video understanding models are often optimized for specific tasks or scenarios and lack generality. When facing different types of videos or user questions, the models need to be retrained or adjusted, increasing the usage cost and time cost; Low information retrieval efficiency. Due to the complexity of video content, traditional methods are less efficient in retrieving specific information in videos. Users may need to spend a lot of time searching for the required information in the videos, and the accuracy of the search results is also difficult to guarantee; Lack of effective utilization of text information. The text information contained in videos, such as subtitles and voiceovers, is an important clue for understanding video content. However, traditional methods often fail to fully utilize this text information, resulting in incomplete information extraction and affecting the understanding and analysis of video content.
[0023] In view of this, the present invention provides a retrieval method for video content. Refer to Figure 1 As shown, it is a specific embodiment of a retrieval method for video content provided by the present invention. The method includes: Step 101, obtain the retrieval intention of the target user; The target users in this embodiment are users who have a retrieval need for relevant video content. For example, users who need a certain course video, film and television work, or training work. The retrieval intention is the need of the target user's query. The manifestation of the retrieval intention is generally the input of the target user's question. By obtaining the target user's question, the retrieval intention can be obtained.
[0024] Step 102: Determine the retrieval information of the question to be retrieved according to the retrieval intention; Through the understanding and analysis of the retrieval intention, the retrieval information of the question to be retrieved can be determined. When the target user asks a question, this question is the question to be retrieved. The input of the question to be retrieved can be in the form of text input or speech text input. The retrieval information can include information such as the relevant type of the question to be retrieved and the associated semantics.
[0025] Step 103: Determine the target type and associated information of the question to be retrieved according to the retrieval information; Specifically, an intention recognition model is used to perform intention recognition on the retrieval information to determine the target type. Among them, the intention recognition model can be pre-trained through a neural network model. By performing intention recognition on the retrieval information, it is possible to know what the target user specifically wants to do, so as to classify the question to be retrieved. For example, factual questions, opinion questions, reasoning questions, etc. The determined type is the target type. Intention recognition refers to analyzing and judging the purpose or intention of the user from the conversation content (such as text, speech, etc.) input by the target user. For example, in a dialogue system, the user inputs "I want to find the introduction of the third episode of a certain TV drama". Through intention recognition, it can be judged that the intention of the target user is to search for one episode of a certain TV drama. Thus, the target type of this question to be retrieved is determined as a factual question. The specific role of intention recognition is to be able to understand what the user wants to do, so as to determine the next step. If it is correctly recognized that the user wants to query a certain episode of a TV drama, that is, a factual question, the next step can be correctly guided to provide the target user with information related to this TV drama in order to complete the retrieval task.
[0026] Semantic analysis is performed on the retrieval information based on natural language processing technology to determine the associated information. Natural language processing technology (NLP) is an important technology in the field of artificial intelligence, aiming to enable computers to understand and process human language. It realizes tasks such as human-computer interaction, information extraction, and semantic analysis by simulating the human language understanding and analysis capabilities. It can convert natural language into a computer-readable form, and then use various algorithms and models for semantic understanding, information extraction, and text generation. By performing semantic analysis on the retrieval information, such as word segmentation, part-of-speech tagging, named entity recognition, syntactic analysis, etc., the question to be retrieved is parsed to extract information such as keywords, entities, and intentions. The composition of the above relevant information is the associated information.
[0027] Step 104: Determine the analysis result of the problem to be retrieved according to the target type and associated information; Specifically, the composition of the target type and associated information constitutes the extraction of important information of the problem to be retrieved, and the final extracted information is the analysis result, which reflects the ultimate requirement of the problem to be retrieved. For example, if the input of the target user is "I want to find English training course videos at the same level as Teacher xx", through the analysis of the above steps, it can be determined that the target type of the problem to be retrieved by this target user is an inference type problem, and the associated information can be: keywords: Teacher xx, at the same level, English, training video; intention: as long as English training video courses at the same level as Teacher xx are retrieved, push them to me. The above text composition of the target type and associated information forms an analysis result that can be understood by the computer.
[0028] Step 105: Based on the preset knowledge base, determine the script information corresponding to the problem to be retrieved according to the analysis result; The knowledge base in this embodiment is a knowledge base established in advance according to video file information, and this knowledge base contains various video file information. Among them, the video file information includes video source files and video basic information. The knowledge base architecture design adopts a hierarchical architecture, and the knowledge base is divided into a data layer, an index layer, and an interface layer. The data layer is responsible for storing the original data of the script, such as video basic information, script text, time stamps, etc.; the index layer is responsible for establishing various indexes, such as keyword indexes, time indexes, role indexes, etc., to improve the retrieval efficiency of data; the interface layer is responsible for providing interfaces for interacting with the knowledge base, such as query interfaces, update interfaces, etc. Among them, the knowledge base also adopts a distributed storage system to improve the storage capacity and reliability of data. At the same time, the knowledge base is backed up regularly to prevent data loss. The backup data can be stored in a local server or cloud storage. After the data storage is completed, various indexes are established to improve the retrieval efficiency of data. According to different query requirements, keyword indexes, time indexes, role indexes, etc. are established. At the same time, the indexes are optimized regularly, such as rebuilding indexes, deleting invalid indexes, etc., to ensure the effectiveness of the indexes.
[0029] Retrieve in the knowledge base according to the analysis result of the problem to be retrieved. Use indexing technology to quickly locate the script information related to the problem to be retrieved. The script information can be the text of video segments related to the problem to be retrieved, video introductions, video classifications, etc. For complex problems to be retrieved, multiple rounds of retrieval and reasoning may be required, and the information of multiple script segments is combined to answer the problem to be retrieved. At the same time, consider the semantics and context information of the problem to be retrieved to improve the accuracy of retrieval.
[0030] Step 106: Determine the target answer according to the script information.
[0031] The generation process of the target answer is as follows Figure 2 As shown, through the understanding and analysis of the retrieval question to be processed, after indexing retrieval, multi-round retrieval, and context association in the knowledge base, the target answer is finally generated. For simple retrieval questions to be processed, the script information can be directly extracted as the target answer. For example, if the retrieval question to be processed is "Search for the second episode of xx TV drama", the target answer can directly push the video clip content of the second episode of xx TV drama to the target user.
[0032] For complex retrieval questions to be processed, it is necessary to integrate and reason about the retrieved script information to generate a complete target answer. For example, if the retrieval question to be processed is "Search for the growth process of xx actor in xx TV drama", the retrieved script information may include answers of multiple video contents. These answers can be arranged in chronological order and pushed to the target user one by one, so that the target user can understand the growth process of the favorite actor in this TV drama.
[0033] In the process of generating the target answer, natural language generation technology is used to organize the answer into smooth and easy-to-understand text.
[0034] From the above technical solutions, it can be seen that the beneficial effects of this embodiment are as follows: The present invention converts video content into text scripts and constructs a knowledge base. Using a general natural language processing model, questions can be answered based on this knowledge base without the need for a dedicated video understanding model, greatly reducing the consumption of computing resources. The method of the present invention does not depend on a specific video understanding model and has strong versatility. Ordinary users can obtain video-related information by simply asking questions, reducing the threshold for using video information and improving the utilization rate of video resources.
[0035] Figure 1 The above shows only the basic embodiment of the method of the present invention. Based on this, with certain optimizations and expansions, other preferred embodiments of the method can also be obtained.
[0036] As Figure 3 shown, it is another specific embodiment of a retrieval method for video content according to the present invention. This embodiment is further described based on the foregoing embodiment. In this embodiment, the method includes the following steps Step 301: Extract and process the video source file to determine the target text information As can be seen from the previous embodiment, the video file information includes the video source file and the video basic information. The video source file refers to an editable video file. The video basic information includes the video name, classification, content summary, etc. These basic information and the script related to the video source file will be stored together in subsequent steps, providing more information for subsequent video processing and management. The text information contained in the video source file, such as subtitles and narration, is an important clue for understanding the video content. In order to make full use of this text information and improve the understanding and analysis of the video content, it is necessary to extract the text information in the video source file.
[0037] In order to store the script related to the video source file in the knowledge base accurately and with high quality, it is necessary to process the video source file in advance.
[0038] Specifically, preprocess the video source file to determine the video file to be recognized. It is necessary to preprocess the video frames of the video source file to improve the accuracy of recognition. The preprocessing steps include image enhancement (contrast adjustment, brightness adjustment), denoising (Gaussian filtering, median filtering), skew correction, etc. Through these preprocessing operations, the text can be made clearer and more distinguishable, reducing the error in the subsequent recognition steps. The video source file after preprocessing is the video file to be recognized.
[0039] Based on the optical character recognition technology, recognize the text information in the video file to be recognized to determine the first text information. The optical character recognition technology (OCR) is a technology that converts the text in an image into computer text. Use an electronic device (such as a scanner or digital camera) to detect the shape, color and other characteristics of the characters on the screen, and then use methods such as image processing, pattern recognition and machine learning to translate the characters into character codes or text data. In order to accurately extract the text information in the video of the video file to be recognized, the optical character recognition technology is adopted, and the position of the text is located through edge detection, contour detection, connected region analysis, etc. At the same time, through the inter-frame matching and tracking algorithm, the position of the text in different video frames is tracked to ensure the continuity and integrity of the text information. The finally recognized complete and coherent text information is the first text information.
[0040] Based on natural language processing technology, verify and sort the first text information to determine the target text information. Integrate the first text information recognized in different video frames, removing duplicate and redundant parts. Then, use natural language processing technology to verify the integrated first text information, check for grammar errors, spelling mistakes, etc., and make corrections. Finally, sort the verified text information in chronological order to form a complete text record, which is the target text information.
[0041] Step 302, determine the target speech text according to the audio information in the video source file; To improve the accuracy of the description of video source files in the knowledge base, it is also necessary to extract the audio information from the video source files.
[0042] Specifically, extract the audio track from the video source file and determine the audio information based on the audio track. The audio track is the part of the video file that carries the audio data, which defines the format, parameters, and storage method of the audio data. The audio information is the extracted audio track, and the audio information is preprocessed to determine the target audio file. Preprocess the audio in the extracted audio track. The preprocessing steps include audio noise reduction, audio enhancement, audio normalization, etc., to improve the speech quality and reduce the impact of background noise on transcription.
[0043] Determine the first transcription result based on the speech-to-text model according to the target audio file. Input the preprocessed target audio file into the speech-to-text model for speech recognition to obtain the preliminary first transcription result. The speech-to-text model is a technology that uses deep learning technology to convert speech signals into text. For example, the speech-to-text model can be SenseVoice, Whisper, etc. Determine the target transcription result based on the first transcription result using natural language processing technology. Use natural language processing technology to correct the first transcription result to improve the accuracy of transcription and form the target transcription result.
[0044] Determine the target speech text according to the target transcription result. Specifically, perform segmentation processing on the target transcription result to determine the segmentation result, and label the corresponding timestamps for the segmentation result to determine the target speech text. Segment the speech content in the target transcription result after transcription, divide it into different sentences or paragraphs to form the segmentation result. At the same time, label timestamps for each segmentation result, record the start and end times of the segmentation result in the video source file, and form the target speech text. In this way, in the subsequent process of generating the target script, the speech content can be corresponded to the video images according to the timestamps.
[0045] Step 303: Determine the information set according to the video basic information, target text information, and target speech text; Fuse the extracted target text information, the target speech text transcribed from the audio information, and the video basic information to form an information set. This information set is the text information description most relevant to the video source file. The information set contains information about the scene, characters, and plot of the video source file, and can accurately and high-quality express the video source file.
[0046] Step 304: Extract the target features in the information set; the target features include scene features, character features, and plot features; Specifically, a pre-trained scene recognition model is used to recognize and divide the scenes in the information set to determine scene features. The scene recognition model can recognize the scene where the video in the relevant video source file is located, such as indoor, outdoor, meeting room, classroom, etc., based on information such as the picture content of the target text information related to the video and the audio features of the target speech text related to the audio. Then, the video is segmented according to the change of the scene, and each segment corresponds to an independent scene, forming scene features.
[0047] Determine the role relationship and relationship model according to the speech content in the information set; determine the role features according to the role relationship and relationship model. The target speech text contains the speech content in the video. By analyzing the dialogue information, person names, etc. in the speech content of the target speech text in the information set, the role relationship in the video is recognized. At the same time, a relationship model between the roles is constructed to describe the relationship between the roles, such as father-son, teacher-student, colleague, etc., forming role features. Role recognition and relationship modeling help to better understand the plot and dialogue content of the video.
[0048] Based on the deep learning model, extract the plot features in the information set. The information combination also contains the text content related to the plot development of the video. Through the deep learning model, extract the plot development information in the information set, such as relevant features like the starting stage, conflict, climax, turning point, ending, etc., forming plot features.
[0049] Finally, the composition of the scene features, role features, and plot features forms the target features. The target features will serve as the basis for constructing the target script.
[0050] Step 305: Determine the target script according to the scene features, role features, and plot features; According to the extracted scene features, role features, and plot features, use the large language model to generate the plot of the script. The large language model can generate reasonable plot development and dialogue content according to the input feature information. Then, organize the generated plot according to the structure of the script, including information such as scene title, scene description, character lines, time and place, etc., to form a complete target script.
[0051] The generated target script can also be reviewed and optimized to ensure its quality and accuracy. Reviewers can conduct manual checks on the target script to check whether the plot is reasonable, whether the dialogue is smooth, and whether there are errors, etc. At the same time, the system can also use natural language processing technology to automatically evaluate the target script, give evaluation indicators and suggestions to help reviewers optimize it.
[0052] Step 306: Establish a knowledge base according to the target script.
[0053] The target script corresponds one-to-one with the video content in the video source file. The relevant text information in the target script can be used as the accurate answer when the target user retrieves the video content. Therefore, it is necessary to import the target script into the knowledge base to form a large number of accurate resources to be retrieved. The data layer of the knowledge base stores the original data of the target script, such as video basic information, script text, time stamps, etc. The index layer is responsible for establishing various indexes, such as keyword index, time index, character index, etc., to improve the retrieval efficiency of data. The interface layer is responsible for providing interfaces for interacting with the knowledge base, such as query interfaces, update interfaces, etc. After the data is stored, various indexes are established to improve the retrieval efficiency of data. According to different query requirements, keyword indexes, time indexes, character indexes, etc. are established. Through the storage of various target scripts, the update and establishment of the knowledge base are completed.
[0054] As can be seen from the above technical solutions, the beneficial effects of this embodiment are as follows: The establishment of the knowledge base makes the retrieval of video information more efficient. It avoids the answer errors caused by inaccurate processing of visual and audio information in traditional video understanding models, and improves the accuracy and integrity of the answers.
[0055] As Figure 4 shown, it is a specific embodiment of a retrieval device for video content according to the present invention. The device in this embodiment is an entity device for executing Figures 1 - 3 the method. Its technical solution is essentially the same as that of the above embodiment, and the corresponding descriptions in the above embodiment also apply to this embodiment. The device in this embodiment includes: A retrieval intention acquisition module 401, configured to acquire the retrieval intention of the target user; A retrieval information determination module 402, configured to determine the retrieval information of the question to be retrieved according to the retrieval intention; A target type and associated information determination module 403, configured to determine the target type and associated information of the question to be retrieved according to the retrieval information; An analysis result determination module 404, configured to determine the analysis result of the question to be retrieved according to the target type and associated information; A script information determination module 405, configured to determine the script information corresponding to the question to be retrieved based on a preset knowledge base according to the analysis result; A target answer determination module 406, configured to determine the target answer according to the script information.
[0056] Figure 5It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0057] The processor, network interface, and memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 5 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0058] The memory is used to store executable instructions. Specifically, the executable instructions are computer programs that can be executed. The memory can include a memory and a non-volatile memory, and provide the executable instructions and data to the processor.
[0059] In a possible implementation manner, the processor reads the corresponding executable instructions from the non-volatile memory into the memory and then runs them, or can also obtain the corresponding executable instructions from other devices to form a retrieval device for video content at the logical level. The processor executes the executable instructions stored in the memory to implement a retrieval method for video content provided in any embodiment of the present invention through the executed executable instructions.
[0060] As described above in the present invention Figure 4The method executed by a retrieval device for video content provided by the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0061] The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0062] The embodiments of the present invention also propose a readable medium. When the execution instructions stored in the readable storage medium are executed by the processor of an electronic device, the electronic device can be enabled to execute a retrieval method for video content provided in any embodiment of the present invention, and is specifically used to execute as Figure 1 、 Figure 3 the method shown.
[0063] The electronic device in each of the foregoing embodiments may be a computer.
[0064] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method or a computer program product. Therefore, the present invention may adopt a completely hardware embodiment, a completely software embodiment, or a form combining software and hardware.
[0065] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the descriptions in the method embodiments.
[0066] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0067] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and changes can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A retrieval method for video content, characterized in that: The method comprises: Obtain the target user's search intent; Determine the search information of the question to be searched according to the search intention; Determine the target type and related information of the question to be searched according to the search information; Determine the analysis result of the to-be-searched question according to the target type and the associated information; Based on a preset knowledge base, determining the script information corresponding to the question to be searched according to the analysis result, the knowledge base containing various video file information; A target answer is determined based on the script information.
2. The method according to claim 1, characterized in that Determining the target type and associated information of the question to be searched according to the search information includes: Performing intent recognition on the search information based on an intent recognition model to determine the target type; The search information is semantically analyzed based on natural language processing technology to determine the associated information.
3. The method according to claim 1, characterized in that The steps of establishing the preset knowledge base include: Extract and process the video source file to determine the target text information; Determine the target voice text according to the audio information in the video source file; Determining basic video information associated with the video source file; Determine an information set according to the basic video information, the target text information and the target voice text; Extracting target features from the information set; the target features include scene features, character features and plot features; Determine the target script according to the scene features, the character features and the plot features; The knowledge base is established according to the target script.
4. The method according to claim 3, characterized in that The extracting and processing the video source file to determine the target text information includes: Preprocessing the video source file to determine a video file to be identified; Recognize the text information in the video file to be recognized based on optical character recognition technology to determine the first text information; The first text information is verified and sorted based on natural language processing technology to determine the target text information.
5. The method according to claim 3, characterized in that: Determining the target speech text according to the audio information in the video source file comprises: Extracting the audio track in the video source file; determining the audio information according to the audio track; Preprocessing the audio information to determine a target audio file; Determine a first transcription result based on the speech-to-text model according to the target audio file; Determine a target transcription result based on the first transcription result based on natural language processing technology; The target speech text is determined according to the target transcription result.
6. The method according to claim 5, characterized in that Determining the target speech text according to the target transcription result includes: Performing segmentation processing on the target transcription result to determine a segmentation result; The segmentation results are annotated with corresponding timestamps to determine the target speech text.
7. The method according to claim 3, characterized in that The extracting target features from the information set comprises: Using a pre-trained scene recognition model to identify and divide the scenes in the information set to determine the scene features; Determine the role relationship and the relationship model according to the speech content in the information set; determine the role characteristics according to the role relationship and the relationship model; Based on a deep learning model, the plot features in the information set are extracted to determine the target features.
8. A retrieval device for video content, characterized in that: include: A search intention acquisition module is used to obtain the search intention of the target user; A search information determination module, used to determine the search information of the question to be searched according to the search intention; A target type and associated information determination module, used to determine the target type and associated information of the question to be searched according to the search information; An analysis result determination module, used to determine the analysis result of the to-be-searched question according to the target type and the associated information; A script information determination module, used to determine the script information corresponding to the question to be retrieved according to the analysis result based on a preset knowledge base; A target answer determination module is used to determine the target answer based on the script information.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 7 when executed.
10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Information acquisition method and device, equipment and storage medium
CN114281951A
Video plot question and answer method and device based on RAG
CN119106098A
Method and apparatus for retrieving teleplay content
US20210211784A1
Virtual role-based multimodal interaction method, apparatus and system, storage medium, and terminal
WO2022048403A1
Search result display method and apparatus, and computer device and storage medium
WO2023236710A1
Cited By
Method and system for automatically generating video based on multi-agent unstructured knowledge
CN121418636A