Video-based response method and device, equipment, storage medium and program product

By dividing and refining video clips from long videos and determining the response using machine learning models, the information omission problem caused by the amount of data exceeding the threshold in long video Q&A is solved, and the efficiency and accuracy of Q&A is improved.

CN120281963APending Publication Date: 2025-07-08BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400159.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When processing long videos, the prior art directly provides video to the Q&A model or compressed video for Q&A, which may lead to information omission or affect the accuracy of Q&A, especially when the amount of video data exceeds the threshold data volume.

Method used

By determining the video clips associated with the problem from the video, the video is first divided into a plurality of second video clips, and the clips are further refined until the degree of matching meets the response requirements, the response is determined using a machine learning model.

Benefits of technology

The amount of data received by the Q&A model is reduced, the data correlation is improved, and the response efficiency and accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281963A_ABST
    Figure CN120281963A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video-based response method and device, equipment, a storage medium and a program product. The method comprises the steps of determining a first group of video clips associated with a question from a video in response to an obtained question for the video; dividing the video into a plurality of second video clips in response to the condition that the matching degree score between each video clip in the first group of video clips and the question is smaller than a threshold value score; for at least one second video segment of the plurality of second video segments, determining a second group of video segments associated with the question from each second video segment of the at least one second video segment; and determining a response corresponding to the question based on the matching degree score between the second group of video clips associated with the at least one second video clip and the question.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to video-based response methods, devices, electronic devices, computer-readable storage media, and computer program products. Background Art

[0002] With the rapid development of computer technology, currently more and more applications, platforms, or systems, etc. provide services associated with multimedia content, bringing a lot of convenience to the majority of users. Multimedia content may include various types of content such as videos, images, image sets, texts, audios, etc. For example, an application, platform, or system may provide a response service for videos. After an application, platform, or system with such a response service receives a video and a question for the video, it can determine a response corresponding to the question based on the video. Summary of the Invention

[0003] In a first aspect of the present disclosure, a video-based response method is provided. The method includes: in response to obtaining a question for a video, determining a first set of video segments associated with the question from the video; in response to the match score between each video segment in the first set of video segments and the question being less than a threshold score, dividing the video into a plurality of second video segments; for at least one second video segment among the plurality of second video segments, determining a second set of video segments associated with the question from each second video segment in the at least one second video segment; and determining a response corresponding to the question based on the match score between the second set of video segments associated with each of the at least one second video segment and the question.

[0004] In a second aspect of the present disclosure, a video-based response device is provided. The device includes: a first determination module configured to, in response to obtaining a question for a video, determine a first set of video segments associated with the question from the video; a video division module configured to, in response to the match score between each video segment in the first set of video segments and the question being less than a threshold score, divide the video into a plurality of second video segments; a second determination module configured to, for at least one second video segment among the plurality of second video segments, determine a second set of video segments associated with the question from each second video segment in the at least one second video segment; and a response determination module configured to determine a response corresponding to the question based on the match score between the second set of video segments associated with each of the at least one second video segment and the question.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the electronic device to execute the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. Computer instructions are stored on the medium, and when the computer instructions are executed by a processor, the method of the first aspect is implemented.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided. The product includes a computer program, where when the computer program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals represent the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A schematic diagram showing an example interaction interface according to some embodiments of the present disclosure;

[0012] Figure 3 An example of video segmentation according to some embodiments of the present disclosure is shown;

[0013] Figure 4 An example architecture of a video-based response according to some embodiments of the present disclosure is shown;

[0014] Figure 5 An example of an encoder of a first machine learning model according to some embodiments of the present disclosure is shown;

[0015] Figure 6 A flowchart showing a video-based response method according to some embodiments of the present disclosure;

[0016] Figure 7 A schematic block diagram showing an exemplary structure of a video-based response device according to some embodiments of the present disclosure; and

[0017] Figure 8 A block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented is shown. Detailed implementation manners

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0019] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0020] In this document, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.

[0021] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to relevant laws and regulations.

[0023] For example, when receiving a user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0024] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It should be understood that the above-mentioned notice and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0026] As used herein, the term "model" can learn the corresponding association relationship between input and output from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this document, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably herein.

[0027] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and it generally includes an input layer and an output layer, as well as one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications usually include many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.

[0028] Generally, machine learning can be roughly divided into three stages, namely, the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test inputs are applied to the trained model to test whether the model can provide the correct outputs, so as to determine the performance of the model. The testing stage can sometimes be incorporated into the training stage. In the application or inference stage, the trained model can be used to process the actual model inputs based on the parameter values obtained from training, and determine the corresponding model outputs.

[0029] As mentioned above, after an application, platform, or system with a response service for videos receives a video and a question about the video, it can determine a response corresponding to the question based on the video. An application, platform, or system with such a response service can use a question-and-answer model to determine a response corresponding to the question based on the video. The maximum amount of data that the question-and-answer model can receive can be referred to as the threshold data amount, and the model input of the question-and-answer model is usually less than or equal to this threshold data amount.

[0030] Traditionally, a video can be directly provided to a question-and-answer model to use the question-and-answer model to determine a response based on the video. However, in the case where the duration of the video is long (e.g., in hours), the amount of data included in the video may be large (such a video can be referred to as a long video), and the amount of data included in the model input determined based on such a video may exceed the threshold data amount, that is, such a video cannot be entirely provided to the question-and-answer model. For this reason, traditionally, in some scenarios, the video content is summarized, and question-and-answer is performed based on the summarized result in text format. This may lead to information omission and affect the accuracy of question-and-answer. In other scenarios, the video can be compressed, and question-and-answer is performed based on the compression result. However, the compression result may be affected by the compression intensity and storage control, which may affect some details of the video and also affect the accuracy of question-and-answer.

[0031] Some scenarios also propose that multiple video frames can be extracted from the video or a part of the video can be extracted (e.g., frame extraction from the video at uniform intervals), and the question-and-answer model is used to determine a response corresponding to the question based on the extracted multiple video frames or the extracted part of the video. However, the extracted video frames may not contain enough content to determine an accurate response, which will also affect the quality of the response.

[0032] In view of this, according to an embodiment of the present disclosure, an improved solution for responses based on videos is provided. According to the solution of the embodiment of the present disclosure, in response to obtaining a question about a video, a first set of video segments associated with the question is determined from the video. In response to the match score between each video segment in the first set of video segments and the question being less than the threshold score, the video is divided into multiple second video segments. For at least one second video segment among the multiple second video segments, a second set of video segments associated with the question is determined from each second video segment in the at least one second video segment. Based on the match score between the second set of video segments associated with each of the at least one second video segment and the question, a response corresponding to the question is determined.

[0033] In this way, by starting to search for video segments associated with the question at a coarse-grained level, according to the matching score, the potential video segments can be further refined into fine-grained video segments, and then continue to search for even finer-grained video segments from them, until one or more video segments whose matching degree or association degree with the question meet the response requirements are determined from the video, and a response is determined based on such video segments. On the one hand, this reduces the amount of data that the response model needs to receive and improves the association degree of the data used to determine the response, thereby improving the response efficiency and accuracy.

[0034] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 112 and a digital assistant 114 are installed in the client device 110. The user 140 can interact with the application 112 via the client device 110 and / or an attached device of the client device 110. In some implementations, the application 112 may be authorized to collect voice via an audio capture device (such as a microphone) of the client device 110, collect images via an image capture device (such as a camera) of the client device 110, and so on.

[0035] In some embodiments, the application 112 and the digital assistant 114 can be downloaded and installed on the client device 110. In some embodiments, the application 112 and the digital assistant 114 can also be accessed in other ways, such as via a web page.

[0036] In embodiments of the present disclosure, the application 112 can be any suitable application with a response function, which may include, but is not limited to, one or more of the following: a chat application component (also referred to as an instant messaging application component), a browser application component, a planning application component, a document application component, an audio and video conferencing application component, an email application component, a task application component, a calendar application component, an objective and key results (OKR) application component, and so on. It can be understood that although Figure 1 a single application business component is shown, in fact, multiple application business components can be installed on the client device 110. In some embodiments, the application 112 may include a multi-functional collaboration platform. For example, an office collaboration platform (also referred to as an office suite) can provide the integration of multiple types of business components to facilitate activities such as office work and communication. In the multi-functional collaboration platform, people can start different business components as needed to complete corresponding information processing, sharing, communication, etc.

[0037] In some embodiments, the digital assistant 114 may be provided by a separate application business component or may be integrated into an application 112 capable of providing content entities. The application business component for providing the client interface of the digital assistant may correspond to a single-function application business component or a multi-functional collaboration platform, such as an office suite or other collaboration platforms capable of integrating multiple components. It can be understood that, similar to the application business component, although Figure 1 a single digital assistant is shown, in fact, there may be multiple digital assistants.

[0038] In some embodiments, the digital assistant 114 supports the use of plugins. Each plugin can provide one or more functions of an application. Such plugins include, but are not limited to, one or more of the following: search plugin, contact plugin, message plugin, document plugin, table plugin, mail plugin, calendar plugin, schedule plugin, task plugin, and so on.

[0039] The digital assistant 114 is an intelligent assistant for users, with intelligent conversation and information processing capabilities. In the embodiments of the present disclosure, the digital assistant 114 is used for interacting with the user 140 to assist the user 140 in using the terminal device or application. In some embodiments, multiple interaction modes between the user 140 and the digital assistant 114 may be provided, and flexible switching between the multiple interaction modes is possible. When a certain interaction mode is triggered, a corresponding interaction area is presented to facilitate the interaction between the user 140 and the digital assistant 114. The interaction methods between the user 140 and the digital assistant 114 are different in different interaction modes, so that the interaction requirements in different application scenarios can be flexibly adapted.

[0040] In the environment 100, in response to the application 112 being launched, the client device 110 may present the interface 150 of the application 112 and / or the digital assistant 114. The interface 150 may, for example, include the interaction interface of the application 112 and the digital assistant 114. In some embodiments, an interaction window between the user 140 and the digital assistant 114 may be presented in the interface 150. In the interaction window, the user 140 can communicate with the digital assistant 114 by inputting natural language, pictures, audio files, video files, web page files, etc., to instruct the digital assistant to assist in completing various tasks.

[0041] In some embodiments, a communication connection is established between the client device 110 and the server device 120. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this regard. In the embodiments of the present disclosure, the client device 110 and the server device 120 can implement signaling interaction through the communication connection therebetween to implement the supply of services for the application 112 and / or the digital assistant 114.

[0042] As Figure 1 shown, the server device 120 can invoke the machine learning model 130 to support the task processing and / or query response functions of the application 112 and / or the digital assistant 114 based on the output of the machine learning model 130. The machine learning model 130 can include one or more machine learning models. For convenience of description, one or more machine learning models can be collectively referred to as the machine learning model 130 in this article. The machine learning model 130 can be deployed on the server device 120 or on other devices.

[0043] The machine learning model 130 can be based on any suitable model structure, including but not limited to a Transformer model, a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Deep Neural Network (DNN), etc. In some embodiments, the machine learning model 130 can be based on a Language Model (LM). By learning from a large amount of corpus, the language model can possess the ability to answer questions. The machine learning model 130 can also be based on other suitable models.

[0044] It should be noted that if the machine learning model 130 includes multiple machine learning models, the functions, structures, uses, etc. of these multiple machine learning models can be the same or different. In some embodiments, when the client device 110 or the application 112 can provide voice processing services to the user 140, the machine learning model 130 can at least include multiple machine learning models related to voice, such as a machine learning model for performing Text-to-Speech (TTS) (which can be simply referred to as a TTS model), a machine learning model for performing Automatic Speech Recognition (ASR) (which can be simply referred to as an ASR model), and a machine learning model for performing question answering (which can be simply referred to as a question answering model). The input of the ASR model is voice, and the output is text. The input of the TTS model is text, and the output is the corresponding voice. The input of the question answering model is the question text, and the output is the corresponding response text.

[0045] The client device 110 can be any suitable type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the client device 110 is also capable of supporting any type of user interface (such as a "wearable" circuit, etc.).

[0046] The server device 120 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server device 120 can, for example, include computing systems / servers, such as mainframes, edge computing nodes, computing devices in a cloud environment, and so on.

[0047] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.

[0048] The client device 110 can present an interactive interface for receiving questions and providing answers. Such an interactive interface can be, for example, a session interface between a user and a digital assistant. Refer to Figure 2 , Figure 2 FIG. 200 shows an example interactive interface 200 (which can also be simply referred to as Example 200) according to some embodiments of the present disclosure. The client device 110 can receive the video 201 provided by the user and the question 202 for the video 201 via the example 200. The client device 110 can receive the video 201 and the question 202 via the input box 210, for example.

[0049] The video 201 can be pre-stored locally in the client device 110 and provided in the example 200 in response to a user selection, or can be captured in real time by the client device 110. The video 201 can be in any suitable format. The video 201 can have any suitable duration. The embodiments of the present disclosure do not limit the specific content and duration of the video 201.

[0050] The question 202 received by the client device 110 can be of text type (such a question can be called question text), or of non-text type (such as voice type, and a question of voice type can be called question voice). The question type of the question 202 can be a multiple-choice question and an open-ended question. It can be understood that if the question 202 is a multiple-choice question, the question 202 will also include corresponding multiple options. Of course, the question type can also include any other appropriate type, and the present disclosure does not limit the specific content of the question 202.

[0051] The client device 110 can also present the response 203 in the example 200 in response to obtaining a response to the question. The response 203 can include any appropriate type of content such as text, image, voice, etc. It should be noted that although in the example 200 one video corresponds to one question and one question corresponds to one response, in fact, one video can correspond to multiple questions and each question can correspond to multiple responses.

[0052] Video-based question answering, or visual question answering (VAQ) can include a video understanding task and a video temporal grounding (TG) task. The video understanding task refers to assisting in determining the response corresponding to the question based on the understanding of the content in the video. The video temporal grounding task involves retrieving a specific moment in the video according to the user's question and giving a matching timestamp.

[0053] The response can be determined locally by the client device 110 based on the video 201 and the question 202, or can be determined by the server device 120 and sent to the client device 110. For example, the client device 110 can, in response to obtaining the video 201 and the question 202, provide the video 201 and the question 202 to the server device 120. The server device 120 can, based on the obtained video 201 and question 202, determine the response to the question and send the response to the client device 110.

[0054] In some embodiments, a machine learning model (such as the machine learning model 130) can be used to determine the response to the question based on the video 201 and the question 202. This machine learning model can be, for example, based on an autoregressive large vision language model (LVLM), which can predict information based on visual information and text context. For a video v and text x, the machine learning model can generate an output sequence y = (y1, y2, ……, y L ), and each element in the sequence can be determined based on the previously determined elements. Assuming that the LVLM is parameterized by θ, the conditional probability distribution of generating the sequence y from the video v and the text x is:

[0055]

[0056] where y <i is non - zero, and y <t =(y1, y2, ……, y t-1 ). Taking visual question answering as an example, the distribution of the predicted answer a by a machine learning model can be expressed as p θ (a|v, q, I q ), where q is the question, I q is an instruction (such as "the following questions related to this video"), and v represents a set of frames of a fixed number of frames T downsampled from the original video, and these frames are converted into a visual feature representation through a separate visual encoder.

[0057] It should be understood that the interfaces shown in the drawings are only examples, and various interface designs may actually exist. Each graphical element in the interface may have different arrangements and different visual representations, one or more of the elements may be omitted or replaced, and there may also be one or more other elements. The embodiments of the present disclosure are not limited in this regard.

[0058] Some example embodiments of the present disclosure will be further described below with reference to the drawings. The video - based answering method related to the present disclosure can be implemented on the client device 110 and / or the server device 120. Here, only the case where the video - based answering method is implemented at the server device 120 is taken as an example for illustration. It should be noted that if the video - based answering method is implemented at the client device 110, some operations described with reference to the client device 110 may require the assistance of the server device 120 to complete. It should be noted that the operations performed by the client device 110 can specifically be performed by relevant applications and / or digital assistants installed on the client device 110.

[0059] The server device 120 can, in response to obtaining a question for a video, determine a first set of video segments associated with the question from the video, and the first set of video segments can include one or more video segments. The server device 120 can determine the matching score between each video segment in the first set of video segments and the question, and determine the comparison result between the matching score corresponding to each video segment and the threshold score.

[0060] The threshold score can be a predetermined and arbitrarily appropriate score. The server device 120 can, in response to the matching score between at least one video segment in the first set of video segments and the question reaching the threshold score, determine the answer corresponding to the question based on the at least one video segment. The server device 120 can determine the answer corresponding to the question based on the at least one video segment, or can also determine the answer corresponding to the question based on some of the video segments in the at least one video segment. Only as an example, the server device 120 can determine the answer corresponding to the question based on the video segment with the highest matching score in the at least one video segment.

[0061] The server device 120 may divide the video into multiple second video segments in response to the match score between each video segment in the first set of video segments and the question being less than the threshold score. The number of second video segments included in the multiple second video segments may be any appropriate number, which is not limited in the present disclosure. Refer to Figure 3 , Figure 3 Fig. 300 shows an example of video division according to some embodiments of the present disclosure. The server device 120 may divide the video 310 into 3 second video segments (e.g., video segment 322, video segment 324, and video segment 326) in response to the match score between each video segment in the first set of video segments corresponding to the video 310 and the question being less than the threshold score. It can be understood that this is only an example here, and actually the video can be divided into any number of second video segments.

[0062] The server device 120 may, for example, divide the video evenly into multiple second video segments. In this case, the duration of each second video segment in the multiple second video segments is the same. The server device 120 may also, for example, divide the video based on a predetermined duration. In this case, the duration of each of the second video segments except the last one in the multiple second video segments is the predetermined duration, and the duration of the last second video segment may be less than or equal to the predetermined duration. The server device 120 may also divide the video based on the video content. For example, the video may include multiple scenes, and each second video segment after division may correspond to one scene. It can be understood that the server device 120 may also divide the video in any other appropriate manner to obtain multiple second video segments, and the present disclosure does not limit the specific division method.

[0063] In some embodiments, to avoid the content included in the divided second video segments being too little, the server device 120 also determines whether the duration of each second video segment in the multiple second video segments obtained by division reaches a threshold duration. The threshold duration may be any appropriate duration, which may indicate the minimum value of the video duration used to determine the response. The server device 120 may determine that the durations of the multiple second video segments after division are too short to generate a response based on such second video segments in response to the durations of the multiple second video segments not reaching the threshold duration.

[0064] In this case, the server device 120 may determine a response based on the video before partitioning, that is, determine a response based on the first set of video segments in the video before partitioning. It can be understood that the server device 120 may, in response to the duration of at least some of the second video segments among the multiple second video segments reaching a threshold duration, determine a response based on at least this at least part of the second video segments (for example, may determine a response only based on at least part of the second video segments whose duration reaches the threshold duration, or may determine a response based on the multiple second video segments).

[0065] In some embodiments, the server device 120 may extract a second set of video segments associated with the question from each of the multiple second video segments. The second set of video segments may also include one or more video segments. Continuing to refer to Figure 3 , the server device 120 may extract a second set of video segments 332 associated with the question from the video segment 322, may extract a second set of video segments 334 associated with the question from the video segment 324, and may extract a second set of video segments 336 associated with the question from the video segment 326.

[0066] The server device 120 may, for example, determine a matching score between the second set of video segments corresponding to each of the multiple second video segments and the question. It can be understood that a set of video segments corresponding to each second video segment (that is, the second set of video segments) may correspond to a set of matching scores, that is, each second video segment may correspond to a set of matching scores. The server device 120 may determine a response corresponding to the question based on the matching scores corresponding to each second video segment.

[0067] In some embodiments, the server device 120 may sequentially determine the matching scores corresponding to the multiple second video segments. In this case, the server device 120 may stop determining the second set of video segments for the remaining second video segments among the multiple second video segments in response to the matching score corresponding to the second video segment for which the matching score is currently being determined reaching a threshold score (for example, at least one of the video segments in the second set of video segments corresponding to the current second video segment has a matching score with the question reaching the threshold score). The server device 120 may directly determine a response based on the current second video segment. Thus, as long as the matching score corresponding to the current second video segment reaches the threshold score, the server device 120 does not need to determine the matching scores corresponding to other second video segments, which can reduce the computational load of the server device 120 and improve the response efficiency.

[0068] In some embodiments, the server device 120 may also determine the matching degree scores corresponding to the respective multiple second video segments, and determine the comparison results between the matching degree scores corresponding to the respective multiple second video segments and the threshold score. The server device 120 may, in response to the matching degree score corresponding to at least one of the multiple second video segments reaching the threshold score, determine the response corresponding to the question based on this at least one second video segment. For example, the server device 120 may determine the response corresponding to the question based on this at least one second video segment. For example, the server device 120 may also determine the response corresponding to the question based on some of the second video segments among the at least one second video segment. The present disclosure does not make any limitation thereto. Thus, the server device 120 can respectively determine the matching degree scores corresponding to the respective multiple second video segments, which can ensure the comprehensiveness of the determination of the multiple second video segments and help improve the accuracy of the response.

[0069] In some embodiments, if the server device 120 determines the matching degree scores corresponding to the respective multiple second video segments, the server device 120 may determine the processing priorities corresponding to the respective multiple second video segments according to the matching degree scores corresponding to the respective multiple second video segments. It can be understood that the higher the matching degree score corresponding to the second video segment, the higher its corresponding processing priority. For example, the server device 120 may determine a second set of video segments associated with the question from the second video segments among the multiple second video segments based on the processing priorities corresponding to the respective multiple second video segments. Only as an example, the server device 120 may determine a second set of video segments associated with the question from the second video segment with the highest corresponding matching degree score, that is, the highest corresponding processing priority.

[0070] In some embodiments, the server device 120 may also, in response to the matching degree scores corresponding to the multiple second video segments all being less than the threshold score, continue to divide the multiple second video segments. For example, the server device 120 may divide the second video segment with the highest corresponding processing priority based on the processing priorities corresponding to the respective multiple second video segments. For example, the server device 120 may divide the second video segment with the highest corresponding processing priority into multiple third video segments. Similar to the multiple second video segments, for at least one of the multiple third video segments, the server device 120 may determine a third set of video segments associated with the question from each of the at least one third video segment, and may determine the response corresponding to the question based on the matching degree scores between the respective third sets of video segments associated with the at least one third video segment and the question.

[0071] Regarding the specific manner of determining the video segments associated with the question and determining the matching score between the video segments and the question, in the embodiments of the present disclosure, the server device 120 may adopt any suitable manner to determine the video segments and the matching score. In some embodiments, the server device 120 may utilize a first machine learning model to determine the video segments and may utilize a second machine learning model to determine the matching score.

[0072] Continue to describe below in conjunction with Figure 4 the specific manner of extracting a set of video segments from the video / video segments, determining the matching score between the video segments and the question, and answering. Figure 4 FIG. 400 shows an exemplary architecture of video-based answering according to some embodiments of the present disclosure. The exemplary architecture 400 may be implemented at the server device 120. The exemplary architecture 400 involves a first machine learning model 410, a second machine learning model 420, a judgment module 430, and a question-answering model 440. The first machine learning model 410, the second machine learning model 420, and the question-answering model 440 may all be based on any suitable model structure.

[0073] The server device 120 may, in response to obtaining a video or a video segment 401 obtained by dividing the video / video segments and a question 404 for the video, extract a set of video frames 403 from the video or the divided video segments. The video segments obtained by dividing the video / video segments may be, for example, multiple second video segments obtained by dividing the video, multiple third video segments obtained by dividing the second video segments, etc. The server device 120 may adopt any suitable manner to extract a set of video frames 403. For example, the server device 120 may extract video frames at a predetermined interval, may randomly extract video frames from the video, may extract video frames according to the video content, etc.

[0074] The server device 120 may determine the timestamp 402 of each video frame in the extracted set of video frames 403 in the video. The timestamp 402 may be represented, for example, as an absolute timestamp in the video. In this case, the timestamp of each video frame may indicate, for example, the time information of the video frame in the video. For example, the timestamp of video frame A may be XX seconds, which may indicate that the video frame is at the XXth second of the video. In some embodiments, the absolute timestamps corresponding to a set of video frames 403 may be represented with the same number of digits. For example, if the absolute timestamp of the last video frame in a set of video frames 403 is a three-digit timestamp (such as 999, 778, or any other three-digit number), then the absolute timestamps of other video frames in the set of video frames 403 are also represented as three digits. The server device 120 may represent the absolute timestamp with less than 3 digits as 3 digits by padding 0 on the left. For example, if the absolute timestamp of the first timestamp is 3, it may be represented as 003.

[0075] In some embodiments, the timestamp 402 of each video frame in a set of video frames 403 in the video can also be referred to as the Temporal-Augmented Frame Representation (TAFR). Introducing TAFR can reduce the difficulty in understanding and generating numerical timestamps during the question-and-answer process. For a set of video frames (f1, f2, ……, f T ), this set of video frames can correspond to a set of timestamps (t1, t2, ……, t T ), for example (0.00, 3.33, 6.67, 10.00). The server device 120 can round these timestamps to the nearest integer based on the following formula and ensure the same number of digits:

[0076]

[0077] The server device 120 can represent this set of timestamps as (00, 03, 07, 10). Thus, the model input can be determined based on the processed timestamps (i.e., the rounded timestamps with the same number of digits), which can ensure the unity of the timestamps of different video frames, contribute to the accurate understanding and use of timestamp information by the machine learning model, and ensure the accuracy of the output.

[0078] The server device 120 can also obtain a first prompt word template 405 for the first machine learning model 410. The first prompt word template 405 can be configured, for example, to guide the first machine learning model 410 to determine a set of video segments associated with the question from the video or the divided video segments based on a set of video frames 403. The server device 120 can generate a model input for the first machine learning model 410 based on a set of video frames 403, timestamps 402, question 404, and the first prompt word template 405, and provide the model input to the first machine learning model 410.

[0079] The first machine learning model 410 can also be referred to as the Temporal Spotlight Grounding (TSG) model. It can identify the most relevant time window according to the question and model the continuous numerical timestamps as discrete digital generations. The first machine learning model 410 can be based on LVLM, for example, which can predict a sequence p q (such as "find the relevant window") for the question q and instruction I θ (w|v, q, I q ). The sequence w can be converted into a set of time ranges W of length K = [(s1, e1), ……, (s K , e K)], where s K , e K respectively represent the start timestamp and end timestamp of the K-th video segment.

[0080] The first machine learning model 410 can generate a corresponding model output 415 based on the received model input. The model output 415 of the first machine learning model 410 can indicate a set of video segments associated with the question 404 in the video or the divided video segments. In some embodiments, the model output 415 can indicate the time range in which each of this set of video segments is located in the video. It can be understood that if the set of video segments associated with the question 404 includes only one video segment, the model output 415 can be as Figure 4 shown, indicating the time range in which the video segment is located in the video. If the set of video segments associated with the question 404 includes multiple video segments, the model output 415 can also indicate the time ranges in which the multiple video segments are respectively located in the video.

[0081] In some embodiments, the first machine learning model 410 can at least include a visual encoder and a text encoder. The visual encoder in the first machine learning model 410 can be configured to determine the visual feature representations respectively corresponding to a set of video frames 403, and the text encoder in the first machine learning model 410 is configured to determine the text feature representations corresponding to the timestamps 402 of a set of video frames 403 respectively. Refer to Figure 5 , Figure 5 illustrates an example 500 of the encoder of the first machine learning model 410 according to some embodiments of the present disclosure. The first machine learning model 410 includes a visual encoder 510 and a text encoder 520.

[0082] As Figure 5 shown, a set of video frames 403 can include video frames f(1)-f(n) shown in example 501, and the timestamps 402 can be represented as the absolute timestamps of each video frame, as shown by the timestamps t(1)-t(n) in example 502. The timestamps t(1)-t(n) in example 502 can be obtained by rounding the time information in the timestamps 402 and padding with 0 on the left to ensure the unified representation of the timestamps of each video frame. The video frames f(1)-f(n) shown in example 501 are provided to the visual encoder 510, and the visual encoder 510 can determine the visual features respectively corresponding to the video frames f(1)-f(n). The timestamps t(1)-t(n) shown in example 502 can be provided to the text encoder 520, and the text encoder can determine the text features respectively corresponding to the timestamps t(1)-t(n).

[0083] Accordingly, the visual encoder 510 and the text encoder 520 can determine the respective features 530 of the video frames f(1)-f(n), such as the feature 530-1 corresponding to the video frame f(1), the feature 530-2 corresponding to the video frame f(2), the feature 530-3 corresponding to the video frame f(3), ……, the feature 530-n corresponding to the video frame f(n). For example, the features corresponding to each video frame can be determined based on the following formula:

[0084]

[0085] where T represents the embedding layer, D represents the embedding dimension, N and P respectively represent the number of video frames and timestamps, represents the feature of the i-th video frame. The feature 530 corresponding to each video frame includes the visual feature and the text feature corresponding to the video frame. The first machine learning model 410 may further include a decoder (not shown in the figure), and the decoder can determine a set of video segments associated with the question 404 from the video / divided video segments based on the features of each video frame in a set of video frames 403.

[0086] The server device 120 can at least determine the model input for the second machine learning model 420 based on a set of video segments associated with the question 404, the question 404, and the second prompt word template 406 for the second machine learning model 420 (for example, the model input can also be determined based on the video / divided video segments 401). After this model input is provided to the second machine learning model 420, the second machine learning model can generate a corresponding model output 425 based on the received model input. The model output 425 of the second machine learning model 420 can, for example, indicate the matching degree score between each video segment in a set of video segments associated with the question 404 and the question 404. The second machine learning model 420 can also be referred to as a Temporal Spotlight Reflection (TSR) model, which can evaluate the accuracy of a set of video segments output by the first machine learning model 410 based on a reflection mechanism (for example, the higher the matching degree score, the higher the accuracy of the corresponding video segment), and the TSR module is used to evaluate the accuracy of the output of the TSG model. It should be noted that although the first machine learning model 410 and the second machine learning model 420 are shown as two models here, in fact, the two can be included in the same machine learning model.

[0087] For the open-ended question TF scenario, for the question q, the instruction I TF (such as "Is the time window accurate") and a time window sequence W output by the TSG model, the TSR model can perform matching degree prediction based on the following formula:

[0088] c = pθ (yes|v,q,W,I TF ) (4)

[0089] And for the selected question scenario, for question q and instruction I MC (such as "direct answer option") and a sequence W output by the TSG, the prediction can be performed based on the following formula:

[0090] c = max{p θ (o|v,q,W,I MC )}, o ∈ ("A", "B", ……) (5)

[0091] where ("A", "B", ……) are multiple given options, and c is the matching degree score.

[0092] The matching degree scores of each video segment can be provided to the judgment module 430, and the judgment module 430 can determine the comparison results between the matching degree scores corresponding to each video segment in a group of video segments and the threshold score. For example, the judgment module 430 can determine that the current group of video segments does not meet the response requirement and cannot generate a response based on such a group of video segments in response to the situation that the matching degree scores corresponding to each video segment in a group of video segments do not reach the threshold score, and then instruct the server device 120 to divide the current video / video segment, and instruct the first machine learning model 410 to process the divided multiple video segments to re-determine a group of video segments with a finer granularity.

[0093] The judgment module 430 can also determine that at least one video segment in a group of video segments meets the response requirement in response to the situation that the matching degree score corresponding to at least one video segment in a group of video segments reaches the threshold score, and then instruct the response model 440 to determine the response corresponding to question 404 based on this at least one video segment. As mentioned above, the response can be determined based on all the video segments in this at least one video segment, or can be determined based on some of the video segments in at least one video segment (for example, one video segment with the highest corresponding matching degree score). Taking the case of determining the response based on one video segment with the highest corresponding matching degree score as an example, the judgment module 430 can provide this video segment and question 404 to the response model 440 together, and the response model 440 can determine the response 445 to the question based on the video segment.

[0094] Thus, for video v and question q, two machine learning models can be used to iteratively perform video segment localization. If the output of the second machine learning model indicates that the matching degree score of the video segment is too low, the video content to which the video segment belongs can be further divided, and the operation can be repeated based on the video segments with a finer granularity until the video segments with corresponding matching degree scores meeting the requirements are determined.

[0095] In summary, according to various embodiments of the present disclosure, it is possible to start searching for video segments associated with a question at a coarse-grained level, further refine potential video segments into fine-grained video segments based on a matching score, and then continue to search for even finer-grained video segments therefrom until one or more video segments whose matching degree or association degree with the question in the video meet the response requirements are determined, and a response is determined based on such video segments. This reduces the amount of data that the response model needs to receive on the one hand and improves the association degree of the data used to determine the response, thereby improving the response efficiency and accuracy.

[0096] Figure 6 FIG. 4 shows a flowchart of a video-based response method according to some embodiments of the present disclosure. Method 300 may be implemented at the server device 120.

[0097] At block 610, in response to obtaining a question for a video, the server device 120 determines a first set of video segments associated with the question from the video.

[0098] At block 620, in response to the matching score between each video segment in the first set of video segments and the question being less than a threshold score, the server device 120 divides the video into a plurality of second video segments.

[0099] At block 630, for at least one second video segment among the plurality of second video segments, the server device 120 determines a second set of video segments associated with the question from each second video segment in the at least one second video segment.

[0100] At block 640, the server device 120 determines a response corresponding to the question based on the matching scores between the second set of video segments associated with each of the at least one second video segment and the question.

[0101] In some embodiments, determining a second set of video segments associated with the question from each second video segment in the at least one second video segment includes: determining the matching scores between the plurality of second video segments and the question; determining the respective processing priorities of the plurality of second video segments according to the matching scores between the plurality of second video segments and the question; and determining a second set of video segments associated with the question from the second video segments in the plurality of second video segments based on the respective processing priorities of the plurality of second video segments.

[0102] In some embodiments, method 600 further includes: in response to the matching score of at least one video segment in the second set of video segments associated with a second video segment among the plurality of second video segments reaching the threshold score, stopping determining the second set of video segments associated with the question for the remaining second video segments among the plurality of second video segments.

[0103] In some embodiments, method 600 further includes: in response to a match score between at least one video segment in the first set of video segments and the question reaching a threshold score, determining a response corresponding to the question based on the at least one video segment.

[0104] In some embodiments, determining a first set of video segments associated with a question from a video includes: extracting a set of video frames from the video; and using a first machine learning model to determine the first set of video segments from the video based on the set of video frames, wherein the model input of the first machine learning model includes the set of video frames, the timestamps of the respective video frames in the video, the question, and prompt word information for the first set of video segments, and wherein the model output of the first machine learning model indicates the time ranges in the video where the respective first set of video segments are located.

[0105] In some embodiments, the timestamps of the respective video frames in the video are represented as absolute timestamps in the video, and the absolute timestamps corresponding to the set of video frames are represented with the same number of digits.

[0106] In some embodiments, the first machine learning model includes at least a visual encoder and a text encoder, the visual encoder being configured to determine a visual feature representation corresponding to each of the set of video frames, and the text encoder being configured to determine a text feature representation corresponding to the timestamps of the respective video frames.

[0107] In some embodiments, the match score between each video segment and the question is determined based on the following: using a second machine learning model to determine the match score between the video segment and the question based at least on the video segment and the question.

[0108] In some embodiments, dividing the video into a plurality of second video segments includes: evenly dividing the video into a plurality of second video segments.

[0109] In some embodiments, method 600 further includes: determining whether the duration of each second video segment in the plurality of second video segments reaches a threshold duration; and in response to none of the durations of the plurality of second video segments reaching the threshold duration, determining a response based on the first set of video segments.

[0110] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 7 An exemplary structural block diagram of a video-based response apparatus 700 according to some embodiments of the present disclosure is shown. Apparatus 700 may be implemented as or included in a server device 120. Each module / component in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0111] As Figure 7As shown, device 700 includes a first determination module 710 configured to determine a first set of video segments associated with a question from a video in response to obtaining the question for the video. Device 700 further includes a video division module 720 configured to divide the video into a plurality of second video segments in response to the match score between each video segment in the first set of video segments and the question being less than a threshold score. Device 700 further includes a second determination module 730 configured to determine, for at least one of the plurality of second video segments, a second set of video segments associated with the question from each of the at least one second video segment. Device 700 further includes an answer determination module 740 configured to determine an answer corresponding to the question based on the match scores between the second set of video segments associated with each of the at least one second video segment and the question.

[0112] In some embodiments, the second determination module 730 is further configured to: determine a match score between the plurality of second video segments and the question; determine the respective processing priorities of the plurality of second video segments according to the match scores between the plurality of second video segments and the question; and determine, based on the respective processing priorities of the plurality of second video segments, a second set of video segments associated with the question from the second video segments among the plurality of second video segments.

[0113] In some embodiments, device 700 further includes: a stop determination module configured to stop determining a second set of video segments associated with the question for the remaining second video segments among the plurality of second video segments in response to the match score of at least one video segment in the second set of video segments associated with one of the plurality of second video segments reaching the threshold score.

[0114] In some embodiments, device 700 further includes: a second answer determination module configured to determine an answer corresponding to the question based on at least one video segment in response to the match score between at least one video segment in the first set of video segments and the question reaching the threshold score.

[0115] In some embodiments, the first determination module 710 is further configured to: extract a set of video frames from the video; and use a first machine learning model to determine a first set of video segments from the video based on the set of video frames, wherein the model input of the first machine learning model includes a set of video frames, the respective timestamps of the set of video frames in the video, the question, and prompt word information for the first set of video segments, and wherein the model output of the first machine learning model indicates the time range in the video where each of the first set of video segments is located.

[0116] In some embodiments, the respective timestamps of the set of video frames in the video are represented as absolute timestamps in the video, and the absolute timestamps corresponding to the set of video frames are represented with the same number of digits.

[0117] In some embodiments, the first machine learning model includes at least a visual encoder and a text encoder. The visual encoder is configured to determine visual feature representations corresponding to respective video frames in a set of video frames, and the text encoder is configured to determine text feature representations corresponding to respective timestamps of the set of video frames.

[0118] In some embodiments, the matching score between each video segment and the question is determined based on the following, including: using a second machine learning model, determining the matching score between the video segment and the question based at least on the video segment and the question.

[0119] In some embodiments, the video partitioning module 720 is further configured to: evenly partition the video into a plurality of second video segments.

[0120] In some embodiments, the apparatus 700 further includes: a duration determination module configured to determine whether the duration of each second video segment in the plurality of second video segments reaches a threshold duration; and a third response determination module configured to, in response to the durations of the plurality of second video segments not reaching the threshold duration, determine a response based on the first set of video segments.

[0121] The modules included in the apparatus 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the modules in the apparatus 700 can be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0122] It should be understood that one or more steps in the above methods can be performed by an appropriate electronic device or a combination of electronic devices. Such an electronic device or a combination of electronic devices can include, for example, Figure 1 the client device 110 and the server device 120 in

[0123] Figure 8 A block diagram of an electronic device 800 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 8 the illustrated electronic device 800 is merely exemplary and should not constitute any limitation to the functions and scopes of the embodiments described herein. Figure 8 The illustrated electronic device 800 can be used to implement Figure 1 the server device 120 orFigure 7 Apparatus 700.

[0124] As Figure 8 shown, the electronic device 800 is in the form of a general-purpose electronic device. The components of the electronic device 800 may include, but are not limited to, one or more processors or processors 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processor 810 may be a physical or virtual processor and is capable of performing various processes according to the programs stored in the memory 820. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 800.

[0125] The electronic device 800 generally includes multiple computer storage media. Such media may be any accessible media that can be obtained by the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 may be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 may be a removable or non-removable medium and may include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.

[0126] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 8 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules that are configured to perform the various methods or actions of the various embodiments of the present disclosure.

[0127] The communication unit 840 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 800 may be implemented by a single computing cluster or multiple computer machines that are capable of communicating through a communication connection. Thus, the electronic device 800 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0128] The input device 850 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 860 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 800 can also communicate with one or more external devices (not shown) as needed through the communication unit 840. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 800, or communicate with any device that enables the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0129] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0130] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0131] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing apparatus, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing apparatus, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0132] The computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operation steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other devices to implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0134] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled in the art to understand the various implementation manners disclosed herein.

Claims

1. A video-based response method, comprising: In response to obtaining a question regarding a video, determining a first set of video segments associated with the question from the video; In response to the match score between each video segment in the first set of video segments and the question being less than a threshold score, dividing the video into a plurality of second video segments; For at least one of the plurality of second video segments, determining a second set of video segments associated with the question from each of the at least one second video segment; And Based on the match scores between the second sets of video segments respectively associated with the at least one second video segment and the question, determining a response corresponding to the question.

2. The method according to claim 1, wherein determining a second set of video segments associated with the question from each of the at least one second video segment comprises: Determining the match scores between the plurality of second video segments and the question; Determining the respective processing priorities of the plurality of second video segments according to the match scores between the plurality of second video segments and the question; And Based on the respective processing priorities of the plurality of second video segments, determining a second set of video segments associated with the question from the second video segments among the plurality of second video segments.

3. The method according to claim 1, further comprising: In response to the match score of at least one video segment in the second set of video segments associated with one of the plurality of second video segments reaching the threshold score, stopping determining second sets of video segments associated with the question for the remaining second video segments among the plurality of second video segments.

4. The method according to claim 1, further comprising: In response to the match score of at least one video segment in the first set of video segments and the question reaching the threshold score, determining the response corresponding to the question based on the at least one video segment.

5. The method according to claim 1, wherein determining a first set of video segments associated with the question from the video comprises: Extracting a set of video frames from the video; And Using a first machine learning model, determining the first set of video segments from the video based on the set of video frames, wherein the model input of the first machine learning model includes the set of video frames, the timestamps of the set of video frames in the video respectively, the question, and prompt word information for the first set of video segments, and wherein the model output of the first machine learning model indicates the time ranges in the video where the first set of video segments are respectively located.

6. The method according to claim 5, wherein the timestamps of the set of video frames in the video are represented as absolute timestamps in the video, and the absolute timestamps corresponding to the set of video frames are represented with the same number of digits.

7. The method according to claim 5, wherein the first machine learning model includes at least a visual encoder and a text encoder, the visual encoder being configured to determine visual feature representations corresponding to the respective video frames of the set of video frames, and the text encoder being configured to determine text feature representations corresponding to the respective timestamps of the set of video frames.

8. The method according to claim 1, wherein the matching degree score between each video segment and the question is determined based on the following, including: Using a second machine learning model, determining the matching degree score between the video segment and the question based at least on the video segment and the question.

9. The method according to claim 1, wherein dividing the video into a plurality of second video segments includes: Dividing the video evenly into the plurality of second video segments.

10. The method according to claim 1, further comprising: Determining whether the duration of each of the plurality of second video segments reaches a threshold duration; And In response to the durations of the plurality of second video segments not reaching the threshold duration, determining the response based on the first set of video segments.

11. A video-based response device, comprising: A first determination module configured to, in response to obtaining a question for a video, determine a first set of video segments associated with the question from the video; A video division module configured to, in response to the matching degree scores between each video segment in the first set of video segments and the question being less than a threshold score, divide the video into a plurality of second video segments; A second determination module configured to, for at least one of the plurality of second video segments, determine a second set of video segments associated with the question from each of the at least one second video segment; And A response determination module configured to determine a response corresponding to the question based on the matching degree scores between the second set of video segments associated with each of the at least one second video segment and the question.

12. An electronic device, comprising: At least one processor; And At least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having stored thereon computer-executable instructions, the computer-executable instructions being executable by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product, comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.