Video processing method and device, storage medium and equipment

By extracting and matching query text from videos, the characteristics and event relationships of video segments are determined, and content descriptions are generated. This solves the problem of insufficient video retrieval response capabilities and achieves efficient retrieval and detailed descriptions.

CN120910306APending Publication Date: 2025-11-07CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511014055.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently retrieve video clips containing specific content or events in video processing, resulting in insufficient responsiveness.

Method used

By extracting and matching query text from the video, the video and text features of the target video segment are determined, and combined with event relationships, a content description of the video segment is generated.

Benefits of technology

It improves the responsiveness of video retrieval, enabling it to not only retrieve matching video clips but also provide detailed content descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910306A_ABST
    Figure CN120910306A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method and apparatus, a storage medium and a device. The method comprises the steps of obtaining a target video clip matched with a query text from a video; determining a first video feature of the target video clip based on the target video clip; based on the query text, determining a first text feature of the query text and an association relationship between events contained in the query text; and determining the content description of the target video clip on the basis of the first video feature, the first text feature and the incidence relation among the events contained in the query text, so that the target video clip matched with the query text can be retrieved and the content description of the target video clip can also be obtained when video retrieval is executed, and the video retrieval efficiency is improved. And the response capability of video retrieval can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of video processing, and in particular, to a video processing method and device, a storage medium, and an apparatus. BACKGROUND

[0002] At present, in the process of video processing, there is often a demand for retrieving a video segment containing specific content or specific events from a video, especially a long video. Based on this demand, the related technology can detect and query a video segment matching a query text from a video based on the query text of a user. In this case, how to improve the response capability of video retrieval is a technical problem to be solved. SUMMARY

[0003] The purpose of the present disclosure is to provide a video processing method and device, a storage medium, and an apparatus to improve the response capability of video retrieval.

[0004] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.

[0005] According to a first aspect of the present disclosure, a video processing method is provided, comprising: obtaining a target video segment matching a query text from a video; determining a first video feature of the target video segment based on the target video segment; determining a first text feature of the query text and an association relationship between events contained in the query text based on the query text; and determining a content description of the target video segment based on the first video feature, the first text feature, and the association relationship between the events.

[0006] In some exemplary embodiments of the present disclosure, the obtaining of the target video segment matching the query text from the video comprises: obtaining a plurality of video segments contained in the video; determining a global feature similarity and a local feature similarity between the video segments and the query text based on the video segments and the query text; and determining the target video segment from the plurality of video segments based on the global feature similarity and the local feature similarity between the video segments and the query text.

[0007] In some example embodiments of the present disclosure, the determining the global feature similarity and the local feature similarity between the video segment and the query text based on the video segment and the query text comprises: determining a second video feature of the video segment and a second text feature of the query text based on the video segment and the query text; determining a first similarity between the second text feature and the second video feature, and taking the first similarity as the global feature similarity between the video segment and the query text; determining a second similarity between entities contained in the video segment and the query text based on the second text feature and the second video feature, and taking the second similarity as the local feature similarity between the video segment and the query text.

[0008] In some example embodiments of the present disclosure, the determining the target video segment from the plurality of video segments based on the global feature similarity and the local feature similarity between the video segment and the query text comprises: for each video segment in the plurality of video segments, performing weighted summation processing on the global feature similarity and the local feature similarity between the video segment and the query text to obtain a weighted summation result; and determining the video segment with the maximum weighted summation result in the plurality of video segments as the target video segment.

[0009] In some example embodiments of the present disclosure, the determining the first video feature of the target video segment based on the target video segment comprises: performing dependency syntax analysis processing on the query text to obtain a dependency syntax tree of the query text; determining a depth of the dependency syntax tree based on the dependency syntax tree; determining a feature extraction manner of the target video segment based on the depth of the dependency syntax tree; and extracting the first video feature from the target video segment based on the feature extraction manner.

[0010] In some example embodiments of the present disclosure, the determining the feature extraction manner of the target video segment based on the depth of the dependency syntax tree comprises: when the depth of the dependency syntax tree is less than a first preset threshold, determining to extract the feature of the target video segment by a sliding window; and when the depth of the dependency syntax tree is greater than or equal to the first preset threshold, determining to extract the feature of the target video segment by a TimeSformer model, wherein a feature extraction density of the sliding window is less than a feature extraction density of the TimeSformer model.

[0011] In some example embodiments of the present disclosure, determining the first text feature of the query text and the association relationship between events included in the query text based on the query text comprises: determining a syntactic structure between words included in the query text based on the query text; matching events included in the query text from the syntactic structure based on a semantic dependency template; and matching the association relationship between events from the events included in the query text based on an event relationship template.

[0012] In some example embodiments of the present disclosure, determining the first text feature of the query text and the association relationship between events included in the query text based on the query text comprises: extracting the first text feature from the query text based on a first language model.

[0013] In some example embodiments of the present disclosure, determining the content description of the target video segment based on the first video feature, the first text feature, and the association relationship between events comprises: determining the content description of the target video segment based on a timestamp of the target video segment, the first video feature, the first text feature, and the association relationship between events.

[0014] In some example embodiments of the present disclosure, determining the content description of the target video segment based on the timestamp of the target video segment, the first video feature, the first text feature, and the association relationship between events comprises: generating a causal graph based on the association relationship between events, wherein a node in the causal graph is used to represent an event in the association relationship, and an edge in the causal graph is used to represent the association relationship between events; encoding the causal graph into an adjacency matrix; mapping the adjacency matrix, the timestamp, the first video feature, and the first text feature into a joint embedding feature; and determining the content description of the target video segment based on the joint embedding feature.

[0015] According to a second aspect of the present disclosure, a video processing apparatus is provided, comprising:

[0016] An acquisition module is configured to acquire a target video segment matched with a query text from a video.

[0017] A first determination module is configured to determine a first video feature of the target video segment based on the target video segment.

[0018] A second determination module is configured to determine a first text feature of the query text and an association relationship between events included in the query text based on the query text.

[0019] The third determining module is configured to determine a content description of the target video segment based on the first video feature, the first text feature, and the association relationship between the event.

[0020] In some example embodiments of the present disclosure, the obtaining module is configured to:

[0021] obtain a plurality of video segments included in the video; determine a global feature similarity and a local feature similarity between the video segment and the query text based on the video segment and the query text; and determine a target video segment from the plurality of video segments based on the global feature similarity and the local feature similarity between the video segment and the query text.

[0022] In some example embodiments of the present disclosure, the obtaining module is configured to:

[0023] determine a second video feature of the video segment and a second text feature of the query text based on the video segment and the query text; determine a first similarity between the second text feature and the second video feature, and take the first similarity as the global feature similarity between the video segment and the query text; and determine a second similarity between entities included in the video segment and the query text based on the second text feature and the second video feature, and take the second similarity as the local feature similarity between the video segment and the query text.

[0024] In some example embodiments of the present disclosure, the obtaining module is configured to:

[0025] for each video segment of the plurality of video segments, perform weighted summation processing on the global feature similarity and the local feature similarity between the video segment and the query text to obtain a weighted summation result; and determine a video segment with the maximum weighted summation result from the plurality of video segments as the target video segment.

[0026] In some example embodiments of the present disclosure, the first determining module is configured to:

[0027] perform dependency syntax analysis processing on the query text to obtain a dependency syntax tree of the query text; determine a depth of the dependency syntax tree based on the dependency syntax tree; determine a feature extraction manner of the target video segment based on the depth of the dependency syntax tree; and extract the first video feature from the target video segment based on the feature extraction manner.

[0028] In some example embodiments of the present disclosure, the first determining module is configured to:

[0029] When the depth of the dependency syntax tree is less than a first preset threshold, it is determined to extract the feature of the target video clip by a sliding window; when the depth of the dependency syntax tree is greater than or equal to the first preset threshold, it is determined to extract the feature of the target video clip by a TimeSformer model, and the feature extraction density of the sliding window is less than the feature extraction density of the TimeSformer model.

[0030] In some example embodiments of the present disclosure, the second determining module is configured to:

[0031] Based on the query text, a syntactic structure between words contained in the query text is determined; based on a semantic dependency template, an event contained in the query text is matched from the syntactic structure; and based on an event relationship template, an association relationship between events contained in the query text is matched.

[0032] In some example embodiments of the present disclosure, the second determining module is configured to extract a first text feature from the query text based on a first language model.

[0033] In some example embodiments of the present disclosure, the third determining module is configured to determine a content description of the target video clip based on a timestamp of the target video clip, the first video feature, the first text feature, and the association relationship between events.

[0034] In some example embodiments of the present disclosure, the third determining module is configured to:

[0035] Based on the association relationship between events, a causal graph is generated, wherein a node in the causal graph is used to represent an event in the association relationship, and an edge in the causal graph is used to represent an association relationship between events; the causal graph is encoded into an adjacency matrix; the adjacency matrix, the timestamp, the first video feature, and the first text feature are mapped into a joint embedding feature; and based on the joint embedding feature, a content description of the target video clip is determined.

[0036] According to a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory, the memory being configured to store executable instructions of the processor; wherein the processor is configured to execute the method of the first aspect by executing the executable instructions.

[0037] According to a fourth aspect of the present disclosure, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.

[0038] The video processing method and device, the storage medium and the equipment provided by the embodiments of the present disclosure can obtain and query a target video segment matching a query text from a video, determine a first video feature of the target video segment based on the target video segment, determine a first text feature of the query text and an association relationship between events contained in the query text based on the query text, and determine a content description of the target video segment based on the first video feature, the first text feature and the association relationship between the events contained in the query text. Therefore, when performing video retrieval, not only the target video segment matching the query text can be retrieved, but also the content description of the target video segment can be obtained, which is beneficial to improving the response capability of video retrieval.

[0039] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0040] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0041] Figure 1 A flowchart of a video processing method in an embodiment of the present disclosure is shown;

[0042] Figure 2 A flowchart of a method for obtaining a target video segment in an embodiment of the present disclosure is shown;

[0043] Figure 3 A flowchart of a method for determining global feature similarity and local feature similarity in an embodiment of the present disclosure is shown;

[0044] Figure 4 A flowchart of a method for extracting a video feature in an embodiment of the present disclosure is shown;

[0045] Figure 5 A flowchart of a method for determining an association relationship in an embodiment of the present disclosure is shown;

[0046] Figure 6 A flowchart of a method for determining a content description in an embodiment of the present disclosure is shown;

[0047] Figure 7 A schematic diagram of a video processing device in an embodiment of the present disclosure is shown;

[0048] Figure 8 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown. Detailed Implementation

[0049] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0050] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0051] To facilitate understanding, some of the terms used in the embodiments of this disclosure will be explained first.

[0052] Multimodal Large Language Model (MLLM) is a language model that combines information from multiple modalities (such as text, images, audio, video, etc.). Through cross-modal learning and reasoning, it can provide rich and comprehensive understanding capabilities.

[0053] Cross-modal alignment (CA) refers to the process of effectively matching information from different modalities (such as text, images, and audio), aiming to achieve information sharing and collaboration among different modalities in multimodal learning. Through shared representation spaces or shared features, it can eliminate differences between different modalities, enabling data from different modalities to be compared and jointly analyzed at the same semantic level.

[0054] CLIP (Contrastive Language-Image Pre-training) models can map images and text to a shared vector space through contrastive learning, enabling cross-modal semantic understanding.

[0055] The BMN (Boundary-Matching Network) model can be used to detect the boundaries of potential events in a video, segment the video into multiple video segments, and output the confidence score of each video segment.

[0056] The TimeSformer model refers to a deep learning model used for processing video-related tasks. It introduces the Transformer architecture into the field of video processing, mainly for modeling the temporal sequence information in videos. The TimeSformer model can effectively capture long-range dependencies between video frames through self-attention mechanisms, better understanding the complex patterns and dynamic information that change over time in videos.

[0057] The scheme provided by the embodiments of the present disclosure will be described below in conjunction with exemplary embodiments.

[0058] Figure 1 A flowchart of a video processing method in the embodiments of the present disclosure is shown. The video processing method can be exemplarily executed by an electronic device. In the embodiments of the present disclosure, the electronic device can be exemplarily understood as any device with video and text processing capabilities. For example, a smart phone, a tablet computer, a portable computer, a desktop computer, a vehicle-mounted wireless terminal device, a medical device, an industrial device, etc., but not limited to the devices listed here.

[0059] Referring to Figure 1 In some exemplary embodiments, the video processing method provided by the embodiments of the present disclosure includes the following steps.

[0060] In step S101, a target video segment matching the query text is obtained from the video.

[0061] In the embodiments of the present disclosure, the video can be a short video with a relatively short video time, or a long video with a relatively long video time.

[0062] In some embodiments, the video referred to in the embodiments of the present disclosure can be any type of video, such as game videos, animation videos, movie and TV series videos, and surveillance videos, etc., but not limited to the videos listed here.

[0063] The query text referred to in the embodiments of the present disclosure is used to query a video segment containing a specified content or event in the video. The query text can exemplarily include a description of the specified content or event.

[0064] In some embodiments, the query text can be obtained through a human-computer interaction interface or interface. For example, in some examples, the query text input by the user in the search box can be obtained through the human-computer interaction interface. For another example, in some other examples, the query sentence input by the user's voice can be obtained through an audio acquisition device, and the query sentence is converted into a query text.

[0065] In some embodiments, the target video segment matching the query text can be understood as a video segment containing the specified content or event described in the query text. For example, in some examples, the video can be segmented into multiple video segments by video segmentation technology, and then the video segments and the query text are matched in features to obtain the target video segment matching the query text. Of course, this is only an example of the method of obtaining the target video segment, and is not the only method.

[0066] In step S103, based on the target video segment, a first video feature of the target video segment is determined.

[0067] The first video feature can include at least one of the following features of the target video segment.

[0068] Visual features, such as static features of video frames in the target video segment, such as color (e.g., color histogram, dominant color), texture, shape (e.g., edge contour), and dynamic features such as motion information (e.g., optical flow vector, object motion direction and speed).

[0069] Audio features, such as attributes in the video soundtrack, such as volume, tone, spectral features (e.g., mel-frequency cepstral coefficients), voice / music / environment sound distinction, etc.

[0070] Spacetime features, including features in time and space dimensions, such as the regularity of the position change of objects in different video frames, time structure features such as lens switching (e.g., cut, fade-in and fade-out), and the evolution of the spatial layout of objects in the scene over time.

[0071] In some examples, the first video feature of the target video segment can be extracted from the target video segment by an artificial intelligence (AI) / machine learning (ML) model. The artificial intelligence (AI) / machine learning (ML) model may, for example, be a CLIP model, but is not limited to CLIP.

[0072] It is worth noting that there are many methods for determining the first video feature, and the method in the above example is only one possible method, not the only method.

[0073] In step S105, based on the query text, a first text feature of the query text and a relationship between events contained in the query text are determined.

[0074] The first text feature can include at least one of the following features:

[0075] Semantic features, such as concepts, entities, relationships, and intentions and viewpoints contained in the query text.

[0076] Structural features, such as sentence structure, etc.

[0077] In some examples, the first text feature can be extracted from the query text based on a first language model, such as a CLIP model, but not limited to the CLIP model. The first language model can be trained based on a model training method of the related art. The training sample can be the query text, and the sample label can be the text feature corresponding to the query text.

[0078] In some embodiments, the query text can include one or more events. The association relationship between events can be understood as a dependency relationship or a causal relationship between events, such as the occurrence of event A depends on the sending of event B, and the like.

[0079] In some examples, the association relationship between events contained in the query text can be extracted from the query text based on a second language model. The second language model and the first language model can be the same type of model or different types of models, and the disclosure does not make specific limitations. The second language model can be trained using a model training method provided by the related art. The training sample can be the query text, and the sample label can be the association relationship between events contained in the query text.

[0080] It should be noted that there can be multiple methods for determining the first text feature and the association relationship between events, and the method in the above example is only one possible method, but not the only method.

[0081] In step S107, the content description of the target video segment is determined based on the first video feature, the first text feature, and the association relationship between events.

[0082] In the embodiments of the disclosure, the content description of the target video segment can be exemplarily understood as a text for describing the content of the target video segment. The content description can include a description of fine-grained information such as the association relationship or the causal relationship between objects and / or events in the target video segment.

[0083] In some examples, the content description of the target video segment can be output by a multi-modal large language model by inputting the first video feature, the first text feature, and the association relationship between events into the pre-set multi-modal large language model. The multi-modal large language model can be trained using a model training method provided by the related art. The training sample can include the video feature of the video and the text feature of the query text, and the sample label can include the content description of the object and / or event association relationship description.

[0084] The above embodiments of the present disclosure can obtain and query a target video segment matching the text from the video, determine a first video feature of the target video segment based on the target video segment, determine a first text feature of the query text and a correlation between events contained in the query text based on the query text, and determine a content description of the target video segment based on the first video feature, the first text feature, and the correlation between events contained in the query text. Therefore, when performing video retrieval, not only the target video segment matching the query text can be retrieved, but also the content description of the target video segment can be obtained, which is beneficial to improve the response capability of video retrieval.

[0085] In some example embodiments, the target video segment can be obtained by the following steps. Figure 2 A flowchart of a method for obtaining a target video segment is shown in the embodiments of the present disclosure. As shown in the figure, Figure 2 In some example embodiments, the target video segment referred to in the embodiments of the present disclosure can be obtained by the following steps.

[0086] In step S201, a plurality of video segments contained in a video are obtained.

[0087] In some examples, the plurality of video segments referred to in step S201 can be understood as video segments directly obtained by video segmentation technology. For example, if a video is segmented into five video segments by video segmentation technology, then all the five video segments obtained by the video segmentation technology can be obtained in step S201.

[0088] In some examples, the plurality of video segments referred to in step S201 can also be understood as part of the video segments obtained by video segmentation technology. For example, in some embodiments, the video can be segmented into a plurality of video segments by a BMN model to obtain a set of video segments wherein N is the number of video segments, c i is the confidence of the i-th video segment, is the start time of the i-th video segment, is the end time of the i-th video segment. Then, the plurality of video segments obtained from the video segments segmented by the BMN model can be represented as wherein θ is a confidence threshold, is a video segment with a confidence greater than the confidence threshold.

[0089] In step S203, based on the video segment and the query text, a global feature similarity and a local feature similarity between the video segment and the query text are determined.

[0090] In the embodiments of the present disclosure, the global feature similarity is used to measure the feature matching degree of the video segment and the query text at the overall level.

[0091] The local feature similarity is used to measure the similarity of the video clip and the query text in local details or specific parts.

[0092] In some examples, the video clip and the query text can be input into a preset recognition model, and the global feature similarity and the local feature similarity of the video clip and the query text input by the recognition model. The recognition model referred to in the embodiments of the present disclosure can be understood as a model trained based on a multi-modal large language model. The model can be trained by using a model training method provided by related technologies. The training samples can include video clips and query texts, and the sample labels can include the global feature similarity and the local feature similarity of the video clips and the query texts.

[0093] It is worth noting that the method in the above examples is only one possible method for determining the global feature similarity and the local feature similarity, but not the only method. In other embodiments, the global feature similarity and the local feature similarity of the video clip and the query text can also be determined by other methods.

[0094] In step S205, the target video clip is determined from the plurality of video clips based on the global feature similarity and the local feature similarity of the video clip and the query text.

[0095] For example, in some examples, the global feature similarity and the local feature similarity of each video clip obtained in step S201 can be weighted and summed to obtain a weighted sum result. The video clip with the maximum weighted sum result among the plurality of video clips obtained in step S201 is determined as the target video clip. The expression for weighting and summing the global feature similarity and the local feature similarity of the video clip and the query text is as follows.

[0096] S total = λS global +(1-λ)S local (1)

[0097] Wherein, S global represents the global feature similarity of the video clip and the query text, S local represents the local feature similarity of the video clip and the query text, and λ is a constant greater than 0 and less than 1.

[0098] The above embodiments of the present disclosure, by acquiring a plurality of video clips contained in a video clip, and calculating a global feature similarity and a local feature similarity of each video clip and a query text, determining a target video clip from the plurality of video clips based on the global feature similarity and the local feature similarity of the video clip and the query text, takes into account the global feature and the local feature of the video clip and the query text, and is conducive to improving the accuracy of video retrieval.

[0099] An example of the determination method of the global feature similarity and the local feature similarity is shown in the flowchart of FIG. 3. As shown in FIG. 3, in some embodiments, the global feature similarity and the local feature similarity of the video clip and the query text can be determined by the following steps. Figure 3 An example of the determination method of the global feature similarity and the local feature similarity is shown in the flowchart of FIG. 3. As shown in FIG. 3, in some embodiments, the global feature similarity and the local feature similarity of the video clip and the query text can be determined by the following steps. Figure 3 As shown in FIG. 3, in some embodiments, the global feature similarity and the local feature similarity of the video clip and the query text can be determined by the following steps.

[0100] In step S301, based on the video clip and the query text, the second video feature of the video clip and the second text feature of the query text are determined.

[0101] Wherein the understanding of the second video feature can refer to the first video feature, and the understanding of the second text feature can refer to the first text feature, which will not be repeated here.

[0102] In some exemplary embodiments of the present disclosure, the video clip and the query text can be input into the CLIP model (but not limited to the CLIP model), and the second video feature of the video clip and the second text feature of the query text are extracted by the CLIP model. Wherein the expression of extracting the second video feature and the second text feature by the CLIP model is as follows.

[0103]

[0104] Wherein v represents the second video feature of the video clip , and t represents the second text feature of the query text Q.

[0105] In step S303, the first similarity of the second text feature and the second video feature is determined, and the first similarity is taken as the global feature similarity of the video clip and the query text.

[0106] In some embodiments, the cosine similarity calculation can be performed on the second text feature and the second video feature, and the calculation result of the cosine similarity is taken as the first similarity.

[0107] Suppose the second video feature and the second text feature are represented as expression (2). Then the first similarity (i.e. the global feature similarity) can be calculated by the following expression.

[0108]

[0109] wherein "||" represents norm operation.

[0110] In step S305, based on the second text feature and the second video feature, a second similarity of the entities contained in the video segment and the query text is determined, and the second similarity is taken as a local feature similarity of the video segment and the query text.

[0111] In some embodiments, the entities contained in the video segment and the query text can be processed by cross-modal alignment through a self-attention mechanism to obtain the second similarity of the entities contained in the video segment and the query text. The weight in the self-attention mechanism can be calculated by the following expression.

[0112]

[0113] wherein exp represents exponential operation. FC represents processing of a full connection layer. K represents the number of video frames contained in the video segment, and l represents the number of entities contained in the query text. i is the feature of the i-th video frame in the video segment, t j is the feature of the j-th entity in the query text. The second similarity (local feature similarity) can be calculated by the following expression.

[0114] S local =∑ k,l α ij ·sim(v i ,t j )(5)

[0115] wherein sim represents similarity calculation.

[0116] The above embodiments of the present disclosure determine the second video feature of the video segment, the second text feature of the query text, determine the global feature similarity of the video segment and the query text based on the second text feature and the second video feature, and determine the second similarity of the entities contained in the video segment and the query text; and match the video segment and the query text by combining the similarity of the entities in the video segment and the query text and the global similarity of the video segment and the query text, which can take into account the global features and entity features of the video segment and the query text, and improve the accuracy of matching.

[0117] An example of Figure 4 A flowchart of a method for extracting a video feature in an embodiment of the present disclosure is shown. As Figure 4 shown, in some embodiments, the first video feature of the target video segment can be extracted by the following steps.

[0118] In step S401, dependency syntax analysis processing is performed on the query text to obtain a dependency syntax tree of the query text.

[0119] For example, in some examples, the dependency syntax analysis processing can be performed on the query text by SpaCy tool to obtain the dependency syntax tree of the query text. SpaCy is an open source natural language processing (NLP) tool library, mainly used for processing and understanding human language text. The functions of SpaCy include word segmentation, part-of-speech tagging, named entity recognition (recognizing names, place names, etc.), dependency syntax analysis (analyzing the grammatical relationship of words in a sentence), identifying grammatical structure, and dependency syntax tree, etc.

[0120] In step S403, the depth of the dependency syntax tree is determined based on the dependency syntax tree.

[0121] The depth of the dependency syntax tree refers to the number of edges from the root node of the tree (usually the core verb or predicate of the sentence) to the deepest leaf node (i.e. the last word in the syntactic relationship), which reflects the hierarchical nesting degree of the sentence syntactic structure.

[0122] For example, in the simple sentence "he eats", the root node is "eat", "he" and "food" are directly dependent on "eat", and the depth of the tree is 1; while in the complex sentence "I know he likes to eat fruit", the root node is "know", "like" depends on "know", "eat" depends on "like", "fruit" depends on "eat", and the depth of the deepest path is 3, which embodies a more complex nested relationship.

[0123] In step S405, the feature extraction method of the target video segment is determined based on the depth of the dependency syntax tree.

[0124] In some embodiments, the feature extraction method of the target video segment can be determined based on the size relationship between the depth of the dependency syntax tree and the first preset threshold. The first preset threshold can be set as needed, and the embodiments of the present disclosure do not make specific limitations.

[0125] For example, in some examples, when the depth of the dependency syntax tree is less than the first preset threshold, the features of the target video segment can be extracted by a sliding window. The sliding window can be understood as a small step sliding window, that is, the sliding window slides at a small step each time, and extracts the features of one or more video frames at each position.

[0126] When the depth of the dependency syntax tree is less than the first preset threshold, it can be understood that the event contained in the query text is a short event (such as falling down, picking up a package, collision), and the short event lasts for a short time. By using a small step sliding window to extract features of the target video segment, the feature extraction density of the short event can be improved, and feature omission can be avoided.

[0127] In some examples, when the depth of the dependency syntax tree is greater than or equal to a first preset threshold, the feature of the target video segment can be extracted by a TimeSformer model. The feature extraction density of the sliding window is less than the feature extraction density of the TimeSformer model.

[0128] When the depth of the dependency syntax tree is greater than or equal to the first preset threshold, it can be understood that the event contained in the query text is a long event (such as riding on a certain road section), and the long event lasts for a long time. By extracting the features of the target video segment by the TimeSformer model, the features of the event in the time dimension and the space dimension, and the dependency relationship between the features can be obtained.

[0129] In some examples, when the depth of the dependency syntax tree is greater than or equal to the first preset threshold, the features of the target video segment can also be extracted by a sliding window with a large step length. By extracting the features of the target video segment by the sliding window with a large step length, the feature extraction density can be reduced, and the processing resources can be saved.

[0130] In step S407, based on the determined feature extraction manner, the first video features are extracted from the target video segment.

[0131] For example, for short events (such as falling down, picking up a package, and colliding), a sliding window with a small step length is used to extract features from the target video segment; for long events (such as entering a hall to leaving), a TimeSformer model is used to extract features from the target video segment. The feature extraction of the target video segment by the TimeSformer model can be represented as follows.

[0132] F long =TimeSfomer(V * )(6)

[0133] Wherein, V * is the target video segment, and F long is the feature extracted by the TimeSformer model. Assuming that the feature extracted by the sliding window is F short , the result of the feature extraction of the target video segment (i.e. the first video features) can be represented as follows:

[0134]

[0135] Wherein, F v represents the first video features.

[0136] The above embodiments of the present disclosure, by performing dependency syntax analysis processing on the query text, obtain a dependency syntax tree of the query text; based on the dependency syntax tree, determine the depth of the dependency syntax tree, based on the depth of the dependency syntax tree, determine the feature extraction manner of the target video segment; based on the feature extraction manner, extract first video features from the target video segment, which is conducive to improving the quality of feature extraction.

[0137] Examples of, Figure 5 A flowchart of a method for determining an association relationship in an embodiment of the present disclosure is shown. As Figure 5 As shown, in some example embodiments, the association relationship between events contained in the query text can be determined by the following steps.

[0138] In step S501, based on the query text, the grammatical structure between the words contained in the query text is determined.

[0139] For example, in some examples, the SpaCy tool can be used to perform dependency syntax analysis processing on the query text to obtain the grammatical structure between the words contained in the query text.

[0140] In step S503, based on the semantic dependency template, the events contained in the query text are matched from the grammatical structure.

[0141] The semantic dependency template is a structured framework for describing the semantic association between language units (such as words, phrases). It presents the semantic connection of each component in the sentence through pre-set semantic roles and relationship types (such as agent, patient, time, place, causality, etc.).

[0142] For example, in "Xiao A eats apples in the park", the template may make it clear that "Xiao A" is the agent of "eating", "apples" is the patient of "eating", and "the park" is the place of "eating".

[0143] In an embodiment of the present disclosure, the dependency relationship of a group of words matched by the semantic dependency template corresponds to an event. For example, in the above example "Xiao A eats apples in the park" can be understood as an event.

[0144] In an embodiment of the present disclosure, there can be one or more events matched from the grammatical structure by the semantic dependency template.

[0145] In step S505, based on the event relationship template, the association relationship between events contained in the query text is matched.

[0146] In an embodiment of the present disclosure, the event relationship template refers to a structured framework for sorting the association relationship between different events, aiming to clearly present the logical connection between events.

[0147] In the embodiments of the present disclosure, the association relationship between events can include at least one of the following relationships.

[0148] A cause-effect relationship, such as "earthquake causes house collapse", and the corresponding event relationship template can be "[event A] triggers [event B]".

[0149] A time sequence relationship, such as "get up before washing", and the corresponding event relationship template can be "[event A] occurs before [event B]".

[0150] A containing relationship, such as "graduation ceremony contains the certificate awarding link", and the corresponding event relationship template can be "[event A] contains sub-event [event B]".

[0151] Of course, the above is only an example for illustration and is not the only limitation.

[0152] Based on the semantic dependency template, the above embodiments of the present disclosure are beneficial to the extraction of events contained in the query text, and based on the event relationship template, the extraction of the association relationship between events is beneficial.

[0153] In some example embodiments of the present disclosure, based on the first video feature, the first text feature, and the association relationship between the events contained in the query text, the content description of the target video segment can include: based on the timestamp of the target video segment, the first video feature, the first text feature, and the association relationship between the events contained in the query text, determining the content description of the target video segment.

[0154] In some example embodiments of the present disclosure, based on the first video feature, the first text feature, and the association relationship between the events contained in the query text, the content description of the target video segment can include: based on the timestamp of the target video segment, the first video feature, the first text feature, and the association relationship between the events contained in the query text, determining the content description of the target video segment. Figure 6 A flowchart of a method for determining a content description in an embodiment of the present disclosure is shown. As shown in some embodiments, the content description of the target video segment can be determined by the following steps. Figure 6 As shown in some embodiments, the content description of the target video segment can be determined by the following steps.

[0155] In step S601, a causal graph is generated based on the association relationship between events, and the nodes in the causal graph are used to represent the events in the association relationship, and the edges in the causal graph are used to represent the association relationship between events.

[0156] In the embodiments of the present disclosure, each event contained in the query text can be regarded as a node. According to the association relationship between events, the nodes corresponding to the event pairs with the association relationship are connected by an edge, and after all edges are connected according to the association relationship between events, a causal graph is obtained.

[0157] In step S603, the causal graph is encoded into an adjacency matrix.

[0158] In the embodiments of the present disclosure, the adjacency matrix is used to represent a matrix form of the causal graph, and mainly describes the connection relationship between the nodes in the causal graph.

[0159] For a causal graph containing n nodes, the adjacency matrix is an n x n square matrix, and the rows and columns correspond to the nodes in the causal graph respectively. The element a ij Generally represents the connection state between node i and node j.

[0160] In step S605, the adjacency matrix, the timestamp of the target video segment, the first video feature, and the first text feature are mapped into joint embedding features.

[0161] Wherein, mapping the adjacency matrix, the timestamp of the target video segment, the first video feature, and the first text feature into joint embedding features can be understood as encoding the adjacency matrix, the timestamp of the target video segment, the first video feature, and the first text feature into the same vector space to obtain the joint embedding features of the adjacency matrix, the timestamp, the first video feature, and the first text feature.

[0162] Wherein, joint embedding features are a feature representation method that maps data of different sources and different types (such as text, image, audio, etc.) into the same low-dimensional vector space.

[0163] In some embodiments, after obtaining the first video feature and the first text feature, the first text feature and the first video feature can also be spliced to obtain a multi-modal feature M, and the expression of M is as follows.

[0164]

[0165] Wherein, t h represents the first text feature.

[0166] In this case, the multi-modal feature M, the adjacency matrix, and the timestamp of the target video segment are encoded into the same vector space to obtain joint embedding features, which can be represented as follows.

[0167]

[0168] Wherein, Prompt_feature is the joint embedding feature, Linear(M) represents the vector representation of the multi-modal feature M in the vector space, GraphEncoder(A) represents the vector representation of the adjacency matrix A in the vector space, represents the vector representation of the timestamp of the target video segment in the vector space.

[0169] In step S607, based on the joint embedding features, the content description of the target video segment is determined.

[0170] In some embodiments, the joint embedding feature can be input into a pre-trained multi-modal large language model, and a content description of the target video segment can be output by the multi-modal large language model. The content description can include a correlation between events, a timestamp, and other fine-grained information.

[0171] The above embodiments of the present disclosure encode the adjacency matrix of the causal graph, the timestamp of the target video segment, the first video feature, and the first text feature into the same vector space to obtain a joint embedding feature of these features, use the joint embedding feature as a prompt feature of a multi-modal large language model, and output a content description of the target video segment by the multi-modal large language model. Thus, the content description of the target video segment can also be used as one of the responses of the video retrieval, which is beneficial to enhancing the response capability of the video retrieval.

[0172] Figure 7 A schematic diagram of a video processing apparatus in an embodiment of the present disclosure is shown. As shown in Figure 7 In some embodiments, the video processing apparatus 700 can include:

[0173] The acquisition module 701 is configured to acquire a target video segment matching the query text from a video.

[0174] The first determination module 702 is configured to determine a first video feature of the target video segment based on the target video segment.

[0175] The second determination module 703 is configured to determine a first text feature of the query text and a correlation between events contained in the query text based on the query text.

[0176] The third determination module 704 is configured to determine a content description of the target video segment based on the first video feature, the first text feature, and the correlation between events.

[0177] In some exemplary embodiments of the present disclosure, the acquisition module 701 is configured to:

[0178] acquire a plurality of video segments contained in the video; determine a global feature similarity and a local feature similarity between the video segments and the query text based on the video segments and the query text; and determine a target video segment from the plurality of video segments based on the global feature similarity and the local feature similarity between the video segments and the query text.

[0179] In some exemplary embodiments of the present disclosure, the acquisition module 701 is configured to:

[0180] determine a second video feature of the video clip based on the video clip and the query text, determine a second text feature of the query text, determine a first similarity of the second text feature and the second video feature, and take the first similarity as a global feature similarity of the video clip and the query text; determine a second similarity of entities contained in the video clip and the query text based on the second text feature and the second video feature, and take the second similarity as a local feature similarity of the video clip and the query text.

[0181] In some example embodiments of the present disclosure, the obtaining module 701 is configured to:

[0182] For each of the plurality of video clips, the global feature similarity and the local feature similarity of the video clip and the query text are processed by weighted summation to obtain a weighted summation result, and the video clip with the maximum weighted summation result in the plurality of video clips is determined as the target video clip.

[0183] In some example embodiments of the present disclosure, the first determining module 702 is configured to:

[0184] perform dependency syntax analysis on the query text to obtain a dependency syntax tree of the query text, determine a depth of the dependency syntax tree based on the dependency syntax tree, determine a feature extraction manner of the target video clip based on the depth of the dependency syntax tree, and extract the first video feature from the target video clip based on the feature extraction manner.

[0185] In some example embodiments of the present disclosure, the first determining module 702 is configured to:

[0186] When the depth of the dependency syntax tree is less than a first preset threshold, it is determined that the features of the target video clip are extracted by a sliding window.

[0187] When the depth of the dependency syntax tree is greater than or equal to the first preset threshold, it is determined that the features of the target video clip are extracted by a TimeSformer model, and the feature extraction density of the sliding window is less than the feature extraction density of the TimeSformer model.

[0188] In some example embodiments of the present disclosure, the second determining module 703 is configured to:

[0189] determine a grammatical structure between words contained in the query text based on the query text, match events contained in the query text from the grammatical structure based on a semantic dependency template, and match an association relationship between events from the events contained in the query text based on an event relationship template.

[0190] In some example embodiments of the present disclosure, the second determining module 703 is configured to: extract first text features from the query text based on a first language model.

[0191] In some example embodiments of the present disclosure, the third determining module 704 is configured to: determine the content description of the target video segment based on the timestamps of the target video segment, the first video features, the first text features, and the association relationships between the events.

[0192] In some example embodiments of the present disclosure, the third determining module 704 is configured to:

[0193] generate a causal graph based on the association relationships between the events, wherein nodes in the causal graph are used to represent events in the association relationships, and edges in the causal graph are used to represent the association relationships between the events; encode the causal graph into an adjacency matrix; map the adjacency matrix, the timestamps, the first video features, and the first text features into joint embedding features; and determine the content description of the target video segment based on the joint embedding features.

[0194] The video processing apparatus 700 provided by the embodiments of the present disclosure can perform the method of any of the above method embodiments, and has similar implementation modes and beneficial effects, which will not be described here in detail.

[0195] In some embodiments, the embodiments of the present disclosure further provide an electronic device, including a processor and a memory, the memory is configured to store executable instructions of the processor; wherein the processor is configured to perform the method of any of the above method embodiments by executing the executable instructions.

[0196] Figure 8 A structural block diagram of an electronic device in the embodiments of the present disclosure is shown. The electronic device 800 according to this embodiment of the present disclosure will be described below with reference to Figure 8 Figure 8 The electronic device 800 shown is merely an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0197] As shown in Figure 8 The electronic device 800 is in the form of a general computing device. The components of the electronic device 800 can include, but are not limited to: at least one processing unit 810 (included in one or more processors), at least one storage unit 820 (included in one or more memories), and a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810).

[0198] ​The storage unit stores program codes which can be executed by the processing unit 810, so that the processing unit 810 performs the steps described in the above "Exemplary Methods" section according to various exemplary embodiments of the present application.

[0199] The storage unit 820 can include a readable medium in the form of volatile storage such as a random access memory (RAM) 821 and / or cache memory 822, and also can include a non-volatile storage such as a read-only memory (ROM) 823.

[0200] The storage unit 820 can further include program / utility 824 having a set of programs / modules 825, including operating systems (OS), one or more application programs, other program modules, and program data, each of or some combination of which can provide functionality for implementing network environments.

[0201] The bus 830 can represent one or more of several types of bus structures, including a storage bus or bus controller, a peripheral bus, a graphics acceleration port, a processor or local bus using any of a variety of bus architectures.

[0202] The electronic device 800 can also communicate with one or more external devices 700 such as a keyboard or pointing device, using one or more communication ports 850. Communication can also occur via a network interface device) 860 utilizing any one of a number of transfer protocols (e.g., HTTP, FTP, SMTP, etc.). The network interface device 860 can be capable of connecting to the Internet, intranet, or any other suitable network. In some embodiments, the network interface device 860 can include a wireless network interface device, a modem, or any other suitable interface device. The network interface device 860 can be connected to the bus 830 via the input / output interface 850. The electronic device 800 can also include one or more antennas 870 for transmitting and receiving wireless signals. The antennas 870 can be connected to the bus 830 via the input / output interface 850. The input / output interface 850 can also connect to one or more of the following: a display 880, a graphics processing unit 882, a storage device 884, a disk drive 886, a signal generation device 888, and a speaker 890. In some embodiments, the input / output interface 850 can connect to a transceiver 892 for communicating with one or more other electronic devices. The input / output interface 850 can also connect to a power supply 894 for powering the electronic device 800. The power supply 894 can include a rechargeable battery, a solar cell, or any other suitable power source.

[0203] Those skilled in the art can easily understand from the above description of the embodiments that the example embodiments described herein can be implemented by software or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or on a network, and includes a number of instructions to make a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) execute the methods according to the embodiments of the present disclosure.

[0204] In the example embodiments of the present disclosure, a computer readable storage medium is also provided, on which a program product capable of implementing the above-mentioned method of the present disclosure is stored. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program codes for causing a terminal device to perform the steps according to various example embodiments of the present disclosure described in the above-mentioned “example method” section of the present specification when the program product is run on the terminal device.

[0205] A program product for implementing the above-mentioned method according to the embodiments of the present disclosure is described, which can take the form of a portable compact disc read-only memory (CD-ROM) and include program codes, and can be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited to this, and in the present document, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, device, or apparatus.

[0206] The program product can take any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, be but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0207] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that can be involved in

[0208] The code can be transmitted in any form, including, but not limited to, radio frequency, optical, electrical, or the like, or any suitable combination thereof.

[0209] The program code can be implemented in any of a variety of programming languages, including, but not limited to, Java, C++, or the like, and can be executed by one or more processors of a device. The program code can execute entirely on the user's device, partly on the user's device, as a stand-alone software package, partly on the user's device and partly on a remote device or entirely on the remote device or server. In the latter scenario, the remote device can be connected to the user's device through any type of network, including, but not limited to, a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external device, such as, for example, through the Internet using an Internet Service Provider (ISP).

[0210] It should be noted that, although several modules or units for device for action execution are mentioned in the foregoing detailed description, such a division into modules or units is not mandatory. Indeed, according to an embodiment of the present disclosure, features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, features and functionalities of one module or unit described above can be further divided into several modules or units.

[0211] Moreover, although the various steps of the methods of the present disclosure are described in a particular order in the figures, this is not required or implied. Indeed, the steps can be performed in any order, some steps can be omitted, some steps can be combined into one step, one step can be divided into several steps, etc.

[0212] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the disclosure being indicated by the following claims.

Claims

1. A method of video processing, the method comprising: The method comprises the following steps: obtaining a target video segment from a video that matches a query text; determining a first video feature of the target video segment based on the target video segment; determining a first text feature of the query text and a correlation between events contained in the query text based on the query text; determining a content description of the target video segment based on the first video feature, the first text feature, and the correlation between the events.

2. The method of claim 1, wherein, The method of obtaining a target video segment from a video that matches a query text comprises the following steps: obtaining a plurality of video segments contained in the video; determining a global feature similarity and a local feature similarity between the video segment and the query text based on the video segment and the query text; determining a target video segment from the plurality of video segments based on the global feature similarity and the local feature similarity between the video segment and the query text.

3. The method of claim 2, wherein, The method of determining a global feature similarity and a local feature similarity between the video segment and the query text based on the video segment and the query text comprises the following steps: determining a second video feature of the video segment and a second text feature of the query text based on the video segment and the query text; determining a first similarity between the second text feature and the second video feature, and taking the first similarity as the global feature similarity between the video segment and the query text; determining a second similarity between entities contained in the video segment and the query text based on the second text feature and the second video feature, and taking the second similarity as the local feature similarity between the video segment and the query text.

4. The method of claim 2, wherein, The method of determining a target video segment from the plurality of video segments based on the global feature similarity and the local feature similarity between the video segment and the query text comprises the following steps: for each video segment in the plurality of video segments, performing weighted summation processing on the global feature similarity and the local feature similarity between the video segment and the query text to obtain a weighted summation result; determining the video segment with the maximum weighted summation result in the plurality of video segments as the target video segment.

5. The method of claim 1, wherein, The method of determining a first video feature of the target video segment based on the target video segment comprises the following steps: performing dependency syntax analysis processing on the query text to obtain a dependency syntax tree of the query text; determining a depth of the dependency syntax tree based on the dependency syntax tree; determining a feature extraction mode of the target video segment based on the depth of the dependency syntax tree; extracting the first video feature from the target video segment based on the feature extraction mode.

6. The method of claim 5, wherein, The method of determining a feature extraction mode of the target video segment based on the depth of the dependency syntax tree comprises the following steps: when the depth of the dependency syntax tree is less than a first preset threshold, determining to extract the feature of the target video segment through a sliding window. When the depth of the dependency syntax tree is greater than or equal to the first preset threshold, it is determined to extract the feature of the target video segment by a TimeSformer model, and the feature extraction density of the sliding window is less than the feature extraction density of the TimeSformer model.

7. The method of claim 1, wherein, The determining, based on the query text, the first text feature of the query text, and the association relationship between events contained in the query text comprises: determining, based on the query text, a syntactic structure between words contained in the query text; matching, based on a semantic dependency template, the events contained in the query text from the syntactic structure; matching, based on an event relationship template, the association relationship between events from the events contained in the query text.

8. The method of claim 1, wherein, The determining, based on the query text, the first text feature of the query text, and the association relationship between events contained in the query text comprises: extracting, based on a first language model, the first text feature from the query text.

9. The method according to any one of claims 1-8, characterized in that, The determining, based on the first video feature, the first text feature, and the association relationship between events, the content description of the target video segment comprises: determining, based on the timestamp of the target video segment, the first video feature, the first text feature, and the association relationship between events, the content description of the target video segment.

10. The method of claim 9, wherein, The determining, based on the timestamp of the target video segment, the first video feature, the first text feature, and the association relationship between events, the content description of the target video segment comprises: generating a causal graph based on the association relationship between events, wherein nodes in the causal graph are used to represent events in the association relationship, and edges in the causal graph are used to represent the association relationship between events; encoding the causal graph into an adjacency matrix; mapping the adjacency matrix, the timestamp, the first video feature, and the first text feature into a joint embedding feature; determining the content description of the target video segment based on the joint embedding feature.

11. A video processing apparatus, comprising: comprises: an acquisition module configured to acquire a target video segment matching a query text from a video; a first determination module configured to determine a first video feature of the target video segment based on the target video segment; a second determination module configured to determine, based on the query text, a first text feature of the query text, and an association relationship between events contained in the query text; a third determination module configured to determine, based on the first video feature, the first text feature, and the association relationship between events, a content description of the target video segment.

12. An electronic device, comprising: comprises: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the method of any one of claims 1-10 by executing the executable instructions.

13. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-10. The computer program is executed by the processor to implement the method of any one of claims 1-10.