Method and device for determining video clip, equipment and storage medium

By extracting feature representations of video frames, audio, subtitles and description information, and using the multimodal transformer model to determine candidate video clips, the problem of difficulty in extracting highlight clips in lengthy videos is solved, and efficient and accurate video clip extraction and customized video generation are achieved.

CN119946355AActive Publication Date: 2025-05-06BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510089967.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

When dealing with lengthy videos, it is difficult for the audience to quickly grasp the essence of the video. The existing technology has limitations in extracting highlight clips and cannot meet the customized needs of users.

Method used

By acquiring the target video and description information associated with the target user, feature representations of video frames, audio, subtitles and description information are extracted, candidate video clips matching the description information are determined using a multimodal transformer model, and target video clips related to the target user are filtered out based on the subtitle information.

Benefits of technology

The efficiency and accuracy of extracting target video clips related to target users from target videos are improved, and the user's needs for customized highlight videos are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946355A_ABST
    Figure CN119946355A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for determining a video clip, equipment, a storage medium and a program product. The method comprises the steps that a target video and description information associated with a target user are acquired, the target video comprises video frame information, audio information and subtitle information, and the description information indicates requirements related to fragment extraction of the target video; extracting a visual feature representation from the video frame information, extracting an audio feature representation from the audio information, extracting a first textual feature representation from the subtitle information, and extracting a second textual feature representation from the description information; determining at least one candidate video clip matched with the description information from the target video based on the visual feature representation, the audio feature representation, the first text feature representation and the second text feature representation; and determining at least one target video clip related to the target user from the at least one candidate video clip at least based on the subtitle information corresponding to the at least one candidate video clip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for determining video segments. Background Art

[0002] With the development of information technology, video, as a vivid means of communication, has penetrated into every aspect of human life. However, as video content becomes increasingly rich, the audience's time becomes more and more precious. Faced with lengthy videos, how to quickly grasp the essence of them has become a common problem faced by viewers and producers. Therefore, it is particularly important to extract highlight clips from long videos. Summary of the invention

[0003] In a first aspect of the present disclosure, a method for determining a video segment is provided. The method includes: obtaining a target video and description information associated with a target user, the target video including video frame information, audio information, and subtitle information, the description information indicating requirements related to segment extraction of the target video; extracting a visual feature representation from the video frame information, extracting an audio feature representation from the audio information, extracting a first text feature representation from the subtitle information, and extracting a second text feature representation from the description information; determining at least one candidate video segment matching the description information from the target video based on the visual feature representation, the audio feature representation, the first text feature representation, and the second text feature representation; and determining at least one target video segment related to the target user from at least one candidate video segment based at least on the subtitle information corresponding to each of the at least one candidate video segment.

[0004] In a second aspect of the present disclosure, a device for determining a video segment is provided. The device includes: an information acquisition module configured to acquire a target video and description information associated with a target user, the target video including video frame information, audio information and subtitle information, the description information indicating requirements related to segment extraction of the target video; a feature extraction module configured to extract a visual feature representation from the video frame information, an audio feature representation from the audio information, a first text feature representation from the subtitle information, and a second text feature representation from the description information; a candidate video segment determination module configured to determine at least one candidate video segment matching the description information from the target video based on the visual feature representation, the audio feature representation, the first text feature representation and the second text feature representation; and a target video segment determination module configured to determine at least one target video segment associated with the target user from at least one candidate video segment based at least on the subtitle information corresponding to each of the at least one candidate video segment.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor. When the instructions are executed by the at least one processor, the electronic device executes the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored on the medium, and when the computer-executable instructions are executed by a processor, the method of the first aspect is implemented.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer executable instructions, and when the computer executable instructions are executed by a processor, the method of the first aspect is implemented.

[0008] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A schematic diagram showing a process of determining a video segment according to some embodiments of the present disclosure;

[0012] Figure 3 A schematic diagram showing determining at least one candidate video segment using a multimodal transformer model according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A schematic diagram showing an example process of determining a target video segment from candidate video segments according to some embodiments of the present disclosure;

[0014] Figure 5 A schematic diagram showing an example process of determining identification information of a target user according to some embodiments of the present disclosure;

[0015] Figure 6 A schematic diagram showing an example process of determining a second set of video clips according to some embodiments of the present disclosure;

[0016] Figure 7A flowchart of a method for determining a video segment according to some embodiments of the present disclosure is shown;

[0017] Figure 8 An apparatus for determining a video segment according to some embodiments of the present disclosure is shown; and

[0018] Fig. 9 A block diagram of an electronic device is shown in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0020] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0021] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0022] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to software or hardware such as electronic devices, applications, servers or storage media that execute operations of the technical solution of the present disclosure based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information is sent to the user in a manner such as a pop-up window, in which the prompt information can be presented in text form. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0026] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multi-layer processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0027] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, and typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes input from the previous layer.

[0028] Generally, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values ​​obtained through training to determine the corresponding model output.

[0029] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In the environment 100, a segment extraction system 120 is deployed in an electronic device 110. The segment extraction system 120 is configured to determine a plurality of video segments 104-1 to 104-N (for convenience of description, may be collectively or individually referred to as video segments 104) in the target video 102 based on the target video 102.

[0030] In some embodiments, the target video 102 may include videos such as movies, news reports, sports events, game videos, etc. The video clips 104 may be relatively important clips in the overall content of the target video 102, also known as "highlight" clips. These clips may be video clips in the target video 102 that contain a large amount of information or leave a deep impression on the audience.

[0031] In some embodiments, the segment extraction system 120 may extract the video segment 104 from the target video 102 based on a machine learning model.

[0032] In environment 100, electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phone, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, media computer, multimedia tablet, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA), audio / video player, digital camera / camcorder, positioning device, television receiver, radio broadcast receiver, e-book device, game device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of interface for the user (such as "wearable" circuit, etc.).

[0033] It should be understood that the structure and function of the various elements in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0034] As mentioned above, it is very important to extract highlight clips from long videos. For videos related to interactive events (e.g., interactive videos, live videos, game videos, etc.), sharing the user's own wonderful operations in interactive events is a major demand of users on content platforms. For example, for game scenes, game screen recording is an important carrier for recording these wonderful operations. However, pure game screen recording has a low information density, lacks highlights and focus, and is not conducive to the rapid dissemination of content.

[0035] One solution to generate highlight segments in videos is to use anchor point information mounted in applications (e.g., live broadcast applications, game applications, etc.) to locate specific highlight segments. However, it is difficult and costly for content platforms to obtain anchor point information within applications, which cannot meet users' requirements for extracting highlight segments.

[0036] Another solution is to train a powerful machine learning model to cover all applications, use the model to retrieve highlight clips, and automatically synthesize a collection of wonderful clips. However, because the machine learning model needs to cover as many types of applications as possible, the model training objectives need to be as unified as possible, so only relatively single highlight content can be produced. And because the definition of highlights is subjective, a single highlight content is difficult to satisfy the preferences of all users, and it is impossible to produce more customized highlight videos. It can be seen that the solutions of related technologies have certain limitations in the generation of highlight clips.

[0037] In an embodiment of the present disclosure, a method for determining a video segment is proposed. Specifically, a target video and description information associated with a target user are obtained, the target video includes video frame information, audio information, and subtitle information, and the description information indicates requirements related to segment extraction of the target video. A visual feature representation is extracted from the video frame information, an audio feature representation is extracted from the audio information, a first text feature representation is extracted from the subtitle information, and a second text feature representation is extracted from the description information. Based on the visual feature representation, the audio feature representation, the first text feature representation, and the second text feature representation, at least one candidate video segment that matches the description information is determined from the target video. At least based on the subtitle information corresponding to each of the at least one candidate video segment, at least one target video segment related to the target user is determined from the at least one candidate video segment.

[0038] According to the scheme disclosed in the present invention, by using the multimodal information in the target video and the description information, the content included in the target video can be fully understood, and the description information can be used to determine the candidate video segments that match the segment extraction requirements of the target user. In addition, the subtitle information in the candidate video segments can filter out the video segments that are irrelevant to the target user. In this way, the efficiency and accuracy of extracting the target video segments related to the target user from the target video can be improved.

[0039] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0040] Figure 2 FIG. 2 is a schematic diagram showing a process 200 of determining a video segment according to some embodiments of the present disclosure. Figure 2 As shown, first, a target video 102 and description information 202 associated with a target user are obtained. The target video 102 includes video frame information, audio information, and subtitle information. The description information 202 indicates requirements related to segment extraction of the target video 102. In one example, the target video 102 is a movie, and the description information 202 may indicate that segments that highlight the essence of the movie and leave a deep impression on the audience are extracted from the movie. In another example, the target video 102 is a game video, and the description information 202 may indicate that segments in which the target player completes wonderful operations are extracted from the game video.

[0041] In some embodiments, video frames may be extracted from the target video 102 at a frequency corresponding to a predetermined frame rate, such as a frequency corresponding to 1 FPS (frames per second), and audio information aligned with the video frames may be extracted. Each extracted video frame may be considered as an image. For each extracted frame of the image, an optical character recognition (OCR) model may be used to extract subtitle information in the frame, and the subtitle information may include text information and a bounding box where the text information is located.

[0042] In block 204, a multimodal preprocessing operation may be performed on the target video 102 and the description information 202. Specifically, a visual feature representation is extracted from the video frame information, an audio feature representation is extracted from the audio information, a first text feature representation is extracted from the subtitle information, and a second text feature representation is extracted from the description information.

[0043] In some embodiments, different feature encoders may be used to extract feature vectors (also referred to as feature representations) from information of different modalities. The visual feature representation may be extracted using a visual encoder, the audio feature representation may be extracted using an audio encoder, and the first text feature representation and the second text feature representation may be extracted using a text encoder.

[0044] In box 206, at least one candidate video segment 208 that matches the description information can be determined from the target video 102 based on the visual feature representation, the audio feature representation, the first text feature representation, and the second text feature representation. In some embodiments, at least one candidate video segment 208 can be determined using a multimodal transformer model. According to an embodiment of the present disclosure, multimodal data can enable the multimodal transformer model to more comprehensively understand the content of the target video 102, and the multimodal data can provide additional contextual information to help the multimodal transformer model extract candidate video segments that meet the needs of the target user from the target video 102. In this way, the efficiency and accuracy of extracting the target video segment related to the target user from the target video can be improved. Reference will be made to the following. Figure 3 The process of determining at least one candidate video segment 208 using a multimodal transformer model is described.

[0045] At block 210, at least one target video segment 212 related to the target user is determined from at least one candidate video segment 208 based at least on the caption information corresponding to each of the at least one candidate video segments. In some embodiments, the at least one target video segment 212 may be merged into a complete video, and the complete video may be provided to the target user.

[0046] In some embodiments, for each candidate video segment in at least one candidate video segment, a candidate user identifier associated with the subtitle broadcast information included in the candidate video segment can be determined. It can then be determined whether the candidate user identifier matches the identification information of the target user, and in response to the candidate user identifier matching the identification information of the target user, the candidate video segment is determined as the target video segment 212. If the candidate user identifier matches the identification information of the target user, it means that the candidate video segment is related to the target user, so the candidate video segment can be determined as the target video segment 212. In some examples, the candidate user identifier associated with the subtitle broadcast information may include a user name, an avatar, etc. In one example, the subtitle broadcast information is information about a user entering a live broadcast room, and the candidate user identifier associated with the subtitle broadcast information may be a user name. In another example, the subtitle broadcast information is game status information in a game, and the candidate user identifier associated with the subtitle broadcast information may be a user avatar. According to an embodiment of the present disclosure, based on the subtitle broadcast information and identification information specific to the video related to the interactive event, the target video segment related to the target user can be accurately determined. In this way, the accuracy and recall rate of the target video segment generation can be improved.

[0047] In a candidate video clip, the candidate user identification usually appears at a specific location at a specific time point. According to the visual style of the candidate user identification or the similarity between the candidate user identification and other known user identifications, it can be determined whether the candidate user identification corresponds to the user identification of the target user.

[0048] In some embodiments, an area containing subtitle broadcast information can be determined from a candidate video clip. Then, the identification information at a predetermined position in the area is determined as a candidate user identification. The subtitle broadcast information may include the current operation of each player in the game, and the identification information of the players involved in the current operation will be displayed at a specific position of the subtitle broadcast information. Therefore, the candidate user identification can be extracted according to the area containing the subtitle broadcast information. In some examples, the area containing the subtitle broadcast information can be a bounding box. For example, the game state information in the game can be included in the bounding box, and the identification information at a predetermined position (e.g., left or right) of the bounding box can be determined as a candidate user identification. It should be noted that determining the candidate user identification by the area containing the subtitle broadcast information is only an example, and the present disclosure does not limit the method for determining the candidate user identification. The method for extracting the user identification can be configured according to the fixed way in which the user identification may appear in the specific application scenario.

[0049] In a game scenario, according to the settings of some games, the visual style (e.g., border color) of the user ID (e.g., user avatar) can be used to distinguish between the own members (i.e., teammates) and the opposing members. For example, a user ID with a blue border indicates that the user is a member of the own team, and a user ID with a red border indicates that the user is an opposing member. After determining that a user ID belongs to a member of the own team, it is also necessary to determine whether the user ID is the user ID of the target user himself or the user ID of the target user's teammate. Usually, the teammate's ID will be displayed at a specific location in the game video. By comparing the similarity between the user ID and the teammate's user ID, it can be determined whether the user ID is the player's own user ID or the teammate's user ID.

[0050] In some embodiments, it can be determined whether the candidate user identification matches the identification information of the target user based on the visual style of the candidate user identification and the similarity between the identification information of the interactive user corresponding to the target user and the candidate user identification. The visual style of the candidate user identification may include the color or shape of the border of the candidate user identification. For example, when the color of the border of the candidate user identification is a specific color (e.g., blue), it can be determined that the candidate user identification is the user identification of the party. Then, it can be determined whether the candidate user identification matches the identification information of the target user based on the similarity between the candidate user identification and the identification information of the interactive user corresponding to the target user. In a live broadcast scenario, the interactive user corresponding to the target user may include a user (e.g., anchor) who is connected to the target user. In a game scenario, the interactive user corresponding to the target user may include the teammates of the target user. In some examples, the identification information of the interactive user corresponding to the target user can be obtained in advance. If the similarity between the candidate user identification and the identification information of the interactive user is greater than a predetermined threshold (e.g., 60%), it means that the candidate user identification does not match the identification information of the target user. If the similarity between the candidate user identification and the identification information of the interactive user is less than a predetermined threshold, it means that the candidate user identification matches the identification information of the target user. In this way, video clips that are not related to the target user can be filtered, thereby improving the accuracy of generating the target video clip. Figure 4 An example process of determining the target video segment 212 from the candidate video segment 208 is described.

[0051] In addition to using at least one candidate video segment of the multimodal transformer model and determining the target video segment 212 from the at least one candidate video segment according to the subtitle information as described in the above blocks 204 to 210, the target video segment 212 related to the target user can also be directly filtered out from the target video 102 by the user identification of the target user. Therefore, the identification information of the target user needs to be extracted first.

[0052] At block 214, identification information of the target user may be extracted. In some examples, the user identification of the target user may be user-specified, or may be determined from a video associated with the interactive event by automatic detection.

[0053] In some embodiments, a reference video clip (also referred to as a first reference video clip) including identification information of the target user can be determined from a video related to an interactive event. The reference video clip has predetermined subtitle broadcast information or the reference video clip is associated with an operation performed by the target user in the interactive event. Then, the identification information of the target user can be extracted from the reference video clip. In some examples, the reference video clip has predetermined subtitle broadcast information, and the predetermined subtitle broadcast information can indicate that the reference video clip includes the identification information of the target user. For example, the predetermined subtitle broadcast information includes "You have entered the live broadcast room", "You have been defeated by the other party", etc. In some examples, the operation performed by the target user in the interactive event can indicate that the first reference video clip includes the identification information of the target user. Taking a game video as an example, the operation performed by the target user in the interactive event can include specific skills that the target user can release when manipulating the character. The following will refer to Figure 5 An example process for determining identification information of a target user.

[0054] At block 216, if the identification information of the target user is successfully extracted, the process 200 proceeds to block 218. At block 218, the identification information of the target user determined from block 214 may be used to recall the video segments including the identification information of the target user to determine the target video segments 212. At block 220, the recalled target video segments 212 may be deduplicated from the target video segments 212 determined using the operation in block 210, so that all the target video segments may be determined.

[0055] In some embodiments, the target video segment 212 determined by the operation in block 210 may be referred to as a first group of video segments. In some examples, at least one candidate video segment may be determined using a multimodal converter modality, and the first group of video segments may be determined based on the subtitle information corresponding to each of the at least one candidate video segment.

[0056] In some embodiments, the target video segments 212 determined by the operation in block 218 may be referred to as a second set of video segments. In some examples, the second set of video segments including identification information of the target user may be determined from the target video 102 based on the identification information of the target user.

[0057] In some embodiments, in order to implement the operation of recalling the target video segment 212 in block 218, at least one reference video segment (also referred to as a second reference video segment) may be determined from the target video 102, each of which has subtitle information matching the description information 202. For each reference video segment in the at least one reference video segment, it may be determined whether the candidate user identification associated with the subtitle information matches the identification information of the target user. If the candidate identification information matches the identification information of the target user, the reference video segment is determined as a video segment in the second set of target video segments.

[0058] After the second group of video clips is determined, at least one target video clip may be determined by merging the first group of video clips with the second group of video clips. The merging of the first group of video clips with the second group of video clips is mainly for performing deduplication of video clips.

[0059] In some embodiments, the second set of video segments may be filtered to be repeated with the first set of target video segments, and the filtered second set of video segments and the first set of target video segments are jointly determined as the final target video segments related to the target user. Figure 6 An example process for recalling a second set of video clips is described.

[0060] The following will continue to refer to the attached Figure 3-6 Example procedures according to some example embodiments of the present disclosure are described.

[0061] Figure 3 FIG. 3 is a schematic diagram showing a process 300 of determining at least one candidate video segment 208 using a multimodal transformer model according to some embodiments of the present disclosure. Figure 3 As shown, for information of different modalities, different encoders can be used to encode the information (also called feature extraction). First, a visual encoder 308 can be used to extract a visual feature representation 313 from the video frame information 302. In some embodiments, a visual encoder (e.g., a ViT-B visual encoder) of a comparative language-image pre-training (CLIP) model pre-trained in the field of the target video 102 can be used to extract main features from the video frame information 302, and the dimension of the main feature is 768. In addition, in order to make the multimodal transformer model 318 more scalable and the recognition ability of general highlight segments better, the ViT-B visual encoder of the original CLIP model can be used to extract auxiliary features, and the dimension of the auxiliary features is 512. Then, the main features and the auxiliary features are concatenated to obtain a visual feature representation 313 with a dimension of (N, 1280), where N is the total number of video frames of the target video 102.

[0062] Then, an audio feature representation 314 may be extracted from the audio information 304 using an audio encoder 310. In some embodiments, the audio information 304 may be encoded using an audio feature encoder pre-trained on large-scale audio data to obtain the audio feature representation 314. The dimension of the audio feature representation 314 is (N, 2048), where N is the total number of video frames of the target video 102.

[0063] Next, a text encoder 312 may be used to extract a first text feature representation 315 and a second text feature representation 316 from the subtitle information 306 and the description information 202, respectively. In some embodiments, a text encoder (e.g., a BERT text encoder) of a CLIP model pre-trained in the domain of the target video 102 may be used to extract the text feature representation. The text feature of the subtitle information 306 is a classification word unit (CLS token) feature, whose dimension is 768. For a video frame without subtitle information 306, it may be initialized to a feature that is all 0, and finally a first text feature representation 315 with a dimension of (N, 768) is obtained. The dimension of the second text feature representation 316 is (N, 768).

[0064] It should be noted that the dimensions of the visual feature representation 313, the dimensions of the audio feature representation 314, the dimensions of the first text feature representation 315, and the second text feature representation 316 described above are merely examples, and the dimensions of these feature representations may also have other values, which is not limited by the present disclosure.

[0065] Based on the extracted feature representations, the multimodal transformer model 318 may output a probability distribution 319 for different video segments in the target video 102. The post-processor 320 may determine at least one candidate video segment 208 that matches the description information 202 based on the probability distribution 319. In some examples, the at least one candidate video segment 208 may be in the form of a triple, including a start time point, an end time point, and a segment type.

[0066] In some embodiments, after extracting the above features, the features related to the target video 102, namely the visual feature representation 313, the audio feature representation 314 and the first text feature representation 315, can be spliced ​​to obtain a video feature representation with a dimension of (N, 3328). Before fusing the video feature representation and the second text feature representation 316, a linear layer can be used to map the video feature representation and the second text feature representation 316 to the same dimension (e.g., 512 dimensions) to perform feature alignment. Then, in the multimodal transformer model 318, the self-attention mechanism can be used to fuse the video feature representation and the second text feature representation 316 to obtain a fused feature. For the prediction of at least one candidate video segment 208, the multimodal transformer model 318 can initialize M 512-dimensional query features, and a cross-attention mechanism can be performed on the query features and the fused features, so that the start time point and the end time point of at least one candidate video segment 208 can be predicted.

[0067] In some embodiments, when the multimodal transformer model 318 is trained, the model parameters of the visual encoder 308, the audio encoder 310, and the text encoder 312 remain unchanged. In this way, the training speed of the multimodal transformer model 318 is accelerated, and the risk of overfitting can be reduced.

[0068] In some embodiments, the target video 102 may include a video related to the interactive event, and the subtitle information includes multiple subtitle broadcast information for prompting the progress of the interactive event. At least one candidate video segment can be determined from the video related to the interactive event, and each candidate video segment in the at least one candidate segment includes one subtitle broadcast information in multiple subtitle broadcast information. In some examples, the video related to the interactive event may include a live video, a game video, etc. In the case where the target video 102 is a live video, the subtitle broadcast information may include information about the user entering the live broadcast room, information about the user sending gifts to the anchor, user comments, etc. In the case where the target video 102 is a game video, the subtitle broadcast information may include game status prompts, information about defeating opponents, etc.

[0069] Figure 4 FIG. 4 is a schematic diagram showing an example process 400 of determining a target video segment from candidate video segments according to some embodiments of the present disclosure. Figure 4 In the example, assume that the video related to the interactive event is a game video. Figure 4As shown, in block 404, video frame data 405, subtitle broadcasting information 406, and a bounding box 407 containing the subtitle broadcasting information 406 are extracted from the candidate video segment 402. In block 408, candidate user identifications 409 and identification information 410 of interactive users can be extracted based on the video frame data 405, the subtitle broadcasting information 406, and the bounding box 407. In some examples, identification information at a predetermined position (e.g., the left side) of the bounding box 407 can be determined as the candidate user identification 409. In some examples, identification information 410 of interactive users can be obtained at a specific position of the video frame data 405.

[0070] In box 412, the border color of the candidate user identification 409 can be detected (as an example of a visual style). In some game settings, the color of the border of the user identification can be used to distinguish whether the user identification is the identification of the own user or the identification of the opponent user. For example, a blue border can be used to identify the own user, and a red border can be used to identify the opponent user. In box 413, if it is determined that the border color is not blue, it means that the candidate video segment 402 is the highlight segment of the opponent. The process 400 proceeds to box 414. In box 414, the candidate video segment 402 can be deleted. In box 413, if it is determined that the border color is blue, it means that the candidate video segment 402 is the highlight segment of the own side, and the process 400 proceeds to box 416. In box 416, it is necessary to further detect the similarity between the candidate user identification and the identification information of the interactive user. If the similarity is greater than a predetermined threshold, it means that the candidate video segment 402 is the highlight segment of the teammate, and in box 414, the candidate video segment 402 can be deleted. If the similarity is not greater than the predetermined threshold, it means that the candidate video segment 402 is the highlight segment of the video segment itself, and the candidate video segment 402 can be determined as the target video segment in block 418. It should be noted that determining whether the candidate user identification matches the identification information of the target user by judging whether the border of the user identification is blue is only an example, and other colors or other visual styles can also be used to determine whether the candidate user identification matches the identification information of the target user, and the present disclosure does not limit this.

[0071] Figure 5 FIG. 5 is a schematic diagram showing an example process 500 for determining identification information of a target user according to some embodiments of the present disclosure. Figure 5 In the example of FIG. 5 , it is assumed that the video 501 related to the interactive event is a game video. Figure 5 As shown, in block 502, it can be determined whether there is a segment in which the character controlled by the target user is defeated by the opponent according to frame data, subtitles or bounding box data in the video 501 related to the interactive event. The segment in which the character controlled by the user is defeated by the opponent can indicate that the video segment includes the identification information of the target user.

[0072] If there is no segment defeated by the opponent, an empty data set is returned in block 504. If there is a segment defeated by the opponent, in block 506, video frame data 507, subtitle broadcast information 508, and a bounding box 509 containing the subtitle broadcast information 508 are extracted from the segment. In block 511, if there are other segments defeated by the opponent, the operation in block 506 can be repeatedly performed.

[0073] In box 510, identification information 512 of the target user may be extracted based on the video frame data 507, the subtitle broadcast information 508, and the enclosing frame 509. For example, identification information with a mark of being defeated by the opponent (e.g., a backslash) in the frame data 507 may be determined as identification information 512 of the target user. The identification information 512 of the target user may then be stored in an identification information pool 513. In box 514, low-quality identification information is filtered in the identification information pool 513. In box 516, identification information clustering is performed in the identification information pool 513. In box 518, the identification information 512 of the target user may be determined.

[0074] Figure 6 FIG. 6 is a schematic diagram showing an example process 600 for determining a second set of video clips according to some embodiments of the present disclosure. Figure 6 In the example of FIG. 5 , it is assumed that the video 501 related to the interactive event is a game video. Figure 6 As shown, in box 602, according to the subtitle broadcasting information in the video 501 related to the interactive event, a reference video segment is extracted, and the reference video segment has subtitle information matching the description information 202. In box 604, it is determined whether the reference video segment is an existing video segment in the first group of video segments. If yes, in box 605, the next reference video segment is extracted. If not, in box 606, video frame data 607, subtitle broadcasting information 608, and a bounding box 609 containing the subtitle broadcasting information 608 are extracted from the reference video segment.

[0075] In block 610, a candidate user identifier 612 may be extracted based on the video frame data 607, the subtitle broadcast information 608, and the enclosing frame 609. In one example, the identification information on the left side of the enclosing frame 609 may be determined as the candidate user identifier 612. In some embodiments, the method described above may be used, that is, determining a region including the subtitle broadcast information 608, and determining the identification information at a predetermined position of the region as the candidate user identifier 612.

[0076] In block 614, the similarity between the candidate user identification 612 and the identification information 512 of the target user may be calculated. In block 615, it is determined whether the candidate user identification 612 matches the identification information 512 of the target user. If they match, the reference video segment is divided into the second group of video segments in block 616. If they do not match, in block 605, the next reference video segment is extracted and it is determined according to the above process that the reference video frame can be divided into the second group of video segments.

[0077] Figure 7 FIG. 7 is a flowchart of a method 700 for determining a video segment according to some embodiments of the present disclosure. The method 700 is implemented in Figure 1 The electronic device 110 is referred to as Figure 1 The method 700 is described with reference to the environment 100 of FIG.

[0078] In block 710 , the electronic device 110 acquires a target video and description information associated with a target user, the target video including video frame information, audio information, and subtitle information, and the description information indicates requirements related to segment extraction of the target video.

[0079] At block 720 , the electronic device 110 extracts a visual feature representation from the video frame information, an audio feature representation from the audio information, a first text feature representation from the subtitle information, and a second text feature representation from the description information.

[0080] In block 730 , the electronic device 110 determines at least one candidate video segment matching the description information from the target video based on the visual feature representation, the audio feature representation, the first text feature representation, and the second text feature representation.

[0081] In block 740 , at least one target video segment related to the target user is determined from the at least one candidate video segment based at least on the subtitle information corresponding to each of the at least one candidate video segments.

[0082] In some embodiments, the target video includes a video related to the interactive event, the subtitle information includes multiple subtitle broadcast information for prompting the progress of the interactive event, and determining at least one candidate video segment from the target video includes: determining at least one candidate video segment from the video related to the interactive event, each candidate video segment in the at least one candidate segment includes one subtitle broadcast information among the multiple subtitle broadcast information.

[0083] In some embodiments, determining at least one target video segment includes: for each candidate video segment in at least one candidate video segment, determining a candidate user identifier associated with subtitle broadcast information included in the candidate video segment; determining whether the candidate user identifier matches identification information of the target user; in response to the candidate user identifier matching the identification information of the target user, determining the candidate video segment as the target video segment.

[0084] In some embodiments, determining the candidate user identification includes: determining a region containing subtitle broadcasting information from the candidate video segment; and determining identification information located at a predetermined position in the region as the candidate user identification.

[0085] In some embodiments, determining whether the candidate user identifier matches the identification information of the target user includes: determining whether the candidate user identifier matches the identification information of the target user based on a visual style of the candidate user identifier and a similarity between the candidate user identifier and the identification information of an interactive user corresponding to the target user.

[0086] In some embodiments, the identification information of the target user is determined by: determining a first reference video segment including the identification information of the target user from a video related to the interactive event, the first reference video segment having predetermined subtitle broadcast information or the first reference video segment being associated with an operation performed by the target user in the interactive event; and extracting the identification information of the target user from the first reference video segment.

[0087] In some embodiments, determining at least one target video segment includes: determining a first group of video segments based on subtitle information corresponding to each of at least one candidate video segments; determining a second group of video segments containing identification information of the target user from the target video based on identification information of the target user; and determining at least one target video segment by merging the first group of video segments and the second group of video segments.

[0088] In some embodiments, determining the second group of video segments includes: determining at least one second reference video segment from the target video, each second reference video segment having subtitle information that matches the description information; and for each second reference video segment in the at least one second reference video segment: determining whether a candidate user identifier associated with the subtitle information matches identification information of the target user; and in response to the candidate identification information matching the identification information of the target user, determining the second reference video segment as a video segment in the second group of target video segments.

[0089] In some embodiments, determining at least one target video segment includes: filtering video segments in the second group of target video segments that are repeated in the first group of target video segments; and merging the filtered second target video segments with the first target video segments to determine at least one target video segment.

[0090] In some embodiments, a visual feature representation is extracted using a visual encoder, an audio feature representation is extracted using an audio encoder, a first text feature representation and a second text feature representation are extracted using a text encoder, at least one candidate video segment is determined using a multimodal transformer model, and wherein model parameters of the visual encoder, the audio encoder, and the text encoder remain unchanged when the multimodal transformer model is trained.

[0091] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 8 An apparatus for determining a video segment according to some embodiments of the present disclosure is shown. The apparatus 800 may be implemented as or included in the electronic device 110. Each module / component in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.

[0092] like Figure 8 As shown, the information acquisition module 810 of the device 800 is configured to acquire a target video and description information associated with a target user, the target video includes video frame information, audio information and subtitle information, and the description information indicates requirements related to segment extraction of the target video; the feature extraction module 820 is configured to extract a visual feature representation from the video frame information, an audio feature representation from the audio information, a first text feature representation from the subtitle information, and a second text feature representation from the description information; the candidate video segment determination module 830 is configured to determine at least one candidate video segment matching the description information from the target video based on the visual feature representation, the audio feature representation, the first text feature representation and the second text feature representation; and the target video segment determination module 840 is configured to determine at least one target video segment related to the target user from at least one candidate video segment based at least on the subtitle information corresponding to each of the at least one candidate video segments.

[0093] In some embodiments, the target video includes a video related to the interactive event, and the subtitle information includes multiple subtitle broadcast information for prompting the progress of the interactive event. The candidate video segment determination module 830 is further configured to determine at least one candidate video segment from the video related to the interactive event, and each candidate video segment in the at least one candidate segment includes one subtitle broadcast information in the multiple subtitle broadcast information.

[0094] In some embodiments, the target video segment determination module 840 is further configured to determine, for each candidate video segment in at least one candidate video segment, a candidate user identifier associated with the subtitle broadcast information included in the candidate video segment; determine whether the candidate user identifier matches the identification information of the target user; and in response to the candidate user identifier matching the identification information of the target user, determine the candidate video segment as the target video segment.

[0095] In some embodiments, the target video segment determination module 840 is further configured to determine a region containing subtitle broadcast information from the candidate video segment; and determine identification information located at a predetermined position in the region as a candidate user identification.

[0096] In some embodiments, the target video segment determination module 840 is further configured to determine whether the candidate user identifier matches the identification information of the target user based on the visual style of the candidate user identifier and the similarity between the candidate user identifier and the identification information of the interactive user corresponding to the target user.

[0097] In some embodiments, the device 800 also includes an identification information determination module, which is configured to determine a first reference video segment including identification information of a target user from a video related to an interactive event, the first reference video segment having predetermined subtitle broadcast information or the first reference video segment being associated with an operation performed by the target user in the interactive event; and extracting the identification information of the target user from the first reference video segment.

[0098] In some embodiments, the target video segment determination module 840 is further configured to determine a first group of video segments based on subtitle information corresponding to each of at least one candidate video segments; determine a second group of video segments containing identification information of the target user from the target video based on identification information of the target user; and determine at least one target video segment by merging the first group of video segments and the second group of video segments.

[0099] In some embodiments, the target video segment determination module 840 is further configured to determine at least one second reference video segment from the target video, each second reference video segment having subtitle information that matches the description information; and for each second reference video segment in the at least one second reference video segment: determine whether a candidate user identifier associated with the subtitle information matches the identification information of the target user; and in response to the candidate identification information matching the identification information of the target user, determine the second reference video segment as a video segment in the second group of target video segments.

[0100] In some embodiments, the target video segment determination module 840 is further configured to filter the video segments in the second group of target video segments that are repeated in the first group of target video segments; and merge the filtered second target video segments with the first target video segments to determine at least one target video segment.

[0101] In some embodiments, a visual feature representation is extracted using a visual encoder, an audio feature representation is extracted using an audio encoder, a first text feature representation and a second text feature representation are extracted using a text encoder, at least one candidate video segment is determined using a multimodal transformer model, and wherein model parameters of the visual encoder, the audio encoder, and the text encoder remain unchanged when the multimodal transformer model is trained.

[0102] The units and / or modules included in the device 800 can be implemented in various ways, including software, hardware, firmware or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 800 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0103] It should be understood that one or more steps in the above method can be performed by a suitable electronic device or a combination of electronic devices. Such an electronic device or a combination of electronic devices may include, for example, Figure 1 The electronic device 110 in.

[0104] Fig. 9 900 is a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. It should be understood that Fig. 9 The electronic device 900 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Fig. 9 The electronic device 900 shown can be used to implement Figure 1 Electronic device 110 or Figure 8 Device 800.

[0105] like Fig. 9 As shown, the electronic device 900 is in the form of a general electronic device. The components of the electronic device 900 may include, but are not limited to, one or more processing units or processors 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 920. In a multi-processor system, multiple processors execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 900.

[0106] The electronic device 900 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 920 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 930 can be a removable or non-removable medium, and can include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data and can be accessed within the electronic device 900.

[0107] The electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Fig. 9 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. Memory 920 may include a computer program product 925 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.

[0108] The communication unit 940 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 900 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 900 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0109] The input device 950 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 960 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 900 may also communicate with one or more external devices (not shown) through the communication unit 940 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 900, or communicate with any device that allows the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0110] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0111] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0112] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so as to produce a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device for implementing the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions make the computer, programmable data processing device, and / or other equipment work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0113] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0114] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as an update, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the function involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0115] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A method for determining a video segment, comprising: Acquire a target video and description information associated with a target user, wherein the target video includes video frame information, audio information, and subtitle information, and the description information indicates requirements related to segment extraction of the target video; extracting a visual feature representation from the video frame information, extracting an audio feature representation from the audio information, extracting a first text feature representation from the subtitle information, and extracting a second text feature representation from the description information; Determining at least one candidate video segment matching the description information from the target video based on the visual feature representation, the audio feature representation, the first text feature representation, and the second text feature representation; and At least based on the subtitle information corresponding to each of the at least one candidate video segment, at least one target video segment related to the target user is determined from the at least one candidate video segment.

2. The method according to claim 1, wherein the target video includes a video game related to an interactive event, the subtitle information includes a plurality of subtitle broadcast information for prompting the progress of the interactive event, and determining the at least one candidate video segment from the target video includes: At least one candidate video segment is determined from the video related to the interactive event, and each candidate video segment in the at least one candidate segment includes one subtitle broadcast information among the multiple subtitle broadcast information.

3. The method according to claim 2, wherein determining the at least one target video segment comprises: For each candidate video segment in the at least one candidate video segment, Determining a candidate user identifier associated with the subtitle broadcasting information included in the candidate video segment; Determining whether the candidate user identification matches the identification information of the target user; In response to the candidate user identification matching the identification information of the target user, the candidate video segment is determined as the target video segment.

4. The method according to claim 3, wherein determining the candidate user identification comprises: Determining a region containing the subtitle broadcasting information from the candidate video segment; as well as The identification information located at a predetermined position in the area is determined as a candidate user identification.

5. The method according to claim 3, wherein determining whether the candidate user identification matches the identification information of the target user comprises: Based on the visual style of the candidate user identifier and the similarity between the candidate user identifier and the identifier information of the interactive user corresponding to the target user, it is determined whether the candidate user identifier matches the identifier information of the target user.

6. The method according to claim 3, wherein the identification information of the target user is determined by: Determining a first reference video segment including identification information of the target user from the video related to the interactive event, wherein the first reference video segment has predetermined subtitle broadcasting information or the first reference video segment is associated with an operation performed by the target user in the interactive event; and The identification information of the target user is extracted from the first reference video segment.

7. The method according to claim 1, wherein determining the at least one target video segment comprises: Determining a first group of video segments based on the subtitle information corresponding to each of the at least one candidate video segments; Based on the identification information of the target user, determining a second group of video clips containing the identification information of the target user from the target video; The at least one target video segment is determined by merging the first group of video segments and the second group of video segments.

8. The method of claim 7, wherein determining the second set of video clips comprises: Determining at least one second reference video segment from the target video, each second reference video segment having subtitle information matching the description information; as well as For each second reference video segment of the at least one second reference video segment: determining whether the candidate user identification associated with the subtitle information matches the identification information of the target user; and In response to the candidate identification information matching the identification information of the target user, the second reference video segment is determined as a video segment in a second group of target video segments.

9. The method according to claim 7, wherein determining the at least one target video segment comprises: filtering out video segments in the second group of target video segments that are repeated in the first group of target video segments; as well as The filtered second target video segment is merged with the first target video segment to determine the at least one target video segment.

10. The method of claim 1, wherein the visual feature representation is extracted using a visual encoder, the audio feature representation is extracted using an audio encoder, the first text feature representation and the second text feature representation are extracted using a text encoder, the at least one candidate video segment is determined using a multimodal transformer model, and Wherein, when the multimodal transformer model is trained, the model parameters of the visual encoder, the audio encoder and the text encoder remain unchanged.

11. An apparatus for determining a video segment, comprising: An information acquisition module, configured to acquire a target video and description information associated with a target user, wherein the target video includes video frame information, audio information, and subtitle information, and the description information indicates requirements related to segment extraction of the target video; a feature extraction module configured to extract a visual feature representation from the video frame information, an audio feature representation from the audio information, a first text feature representation from the subtitle information, and a second text feature representation from the description information; a candidate video segment determination module configured to determine at least one candidate video segment matching the description information from the target video based on the visual feature representation, the audio feature representation, the first text feature representation, and the second text feature representation; and The target video segment determination module is configured to determine at least one target video segment related to the target user from the at least one candidate video segment based at least on the subtitle information corresponding to each of the at least one candidate video segment.

12. An electronic device comprising: at least one processor; as well as At least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the device to perform the method according to any one of claims 1 to 10 when executed by the at least one processor.

13. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.

14. A computer program product comprising computer executable instructions, which when executed by a processor implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Caption customization and editing

    CN113962200A

  • Video retrieval method, device and equipment and storage medium

    CN114282049A

  • Video clip sharing method and device, electronic equipment and readable storage medium

    CN114449327A

  • Video generation method and device, electronic equipment, storage medium and program product

    CN118433466A

  • Caption customization and editing

    US20230055421A1