Video conversation and model training method, device, equipment and storage medium
By performing characterization and extraction, time-time processing and conversion processing on video, video embedding is generated for dialogue with text embedding, which solves the problem that it is difficult for the existing technology to realize video dialogue, and improves the effect of video dialogue and the systematic reasoning ability.
Patent Information
- Application Number
- CN202311176255.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-09-12
AI Technical Summary
The prior art is difficult to apply an image dialogue system to video scenes to achieve efficient video dialogue.
By performing characterization extraction, space-time processing and conversion processing on the target video, the video embedding is obtained and dialogue with the text embedding is performed to generate the answer text.
It improves the effect of video dialogue, makes full use of the visual and temporal information of the video, and enhances the reasoning ability of the dialogue system.
Smart Images

Figure CN117351387B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically computer vision, deep learning, large models and other technical fields, which can be applied to scenarios such as AIGC, and in particular to a video dialogue and model training method, device, equipment and storage medium. Background Art
[0002] Vision-centric multimodal dialogue systems are an important research area. Such dialogue systems usually use a pre-trained large language model (LLM) combined with an image encoder and other learnable modules to conduct dialogues with users to perform image-related tasks.
[0003] The above scheme is mainly for image dialogue. How to apply the above scheme to video scenes to realize video dialogue is a problem that needs to be solved. Summary of the invention
[0004] The present invention provides a video conversation and model training method, device, equipment and storage medium.
[0005] According to one aspect of the present disclosure, a video dialogue method is provided, comprising: performing a representation extraction process on a target video to obtain an initial video representation of the target video; performing spatiotemporal processing on the initial video representation to obtain a target video representation of the target video; performing a conversion process on the target video representation to obtain a video embedding of the target video; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question text; and performing a dialogue process on the video embedding and the text embedding to obtain an answer text.
[0006] According to another aspect of the present disclosure, a training method for a video dialogue model is provided, the video dialogue model comprising: an operation timing module, the method comprising: performing representation extraction processing on a video sample to obtain an initial video representation of the video sample; using the operation timing module to perform spatiotemporal processing on the initial video representation to obtain a target video representation of the video sample; performing conversion processing on the target video representation to obtain a video embedding of the video sample; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question sample; performing dialogue processing on the video embedding and the text embedding to obtain a predicted answer; constructing a loss function based on the predicted answer; and using the loss function to adjust the parameters of the operation timing module.
[0007] According to another aspect of the present disclosure, a video dialogue device is provided, including: an extraction module for performing a representation extraction process on a target video to obtain an initial video representation of the target video; a processing module for performing a spatiotemporal process on the initial video representation to obtain a target video representation of the target video; a conversion module for performing a conversion process on the target video representation to obtain a video embedding of the target video; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question text; and a generation module for performing a dialogue process on the video embedding and the text embedding to obtain an answer text.
[0008] According to another aspect of the present disclosure, a training device for a video dialogue model is provided, wherein the video dialogue model comprises: an operation timing module, and the device comprises: an extraction module for performing a representation extraction process on a video sample to obtain an initial video representation of the video sample; a processing module for performing spatiotemporal processing on the initial video representation using the operation timing module to obtain a target video representation of the video sample; a conversion module for performing a conversion process on the target video representation to obtain a video embedding of the video sample; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question sample; a generation module for performing a dialogue process on the video embedding and the text embedding to obtain a predicted answer; a construction module for constructing a loss function based on the predicted answer; and an adjustment module for adjusting the parameters of the operation timing module using the loss function.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method as described in any of the above aspects.
[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods according to any one of the above aspects.
[0011] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the methods described in any one of the above aspects.
[0012] According to the technical solution disclosed in the present invention, the video conversation effect can be improved.
[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0015] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0016] Figure 2 is a schematic diagram of an application scenario provided according to an embodiment of the present disclosure;
[0017] Figure 3 is a schematic diagram of the overall architecture of a dialogue system provided according to an embodiment of the present disclosure;
[0018] Figure 4 is a schematic diagram according to a second embodiment of the present disclosure;
[0019] Figure 5 is a schematic diagram according to a third embodiment of the present disclosure;
[0020] Figure 6 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0021] Figure 7 is a schematic diagram according to a fifth embodiment of the present disclosure;
[0022] Figure 8 It is a schematic diagram of an electronic device used to implement the video conversation method or the video conversation model training method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0023] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0024] In the related art, for video conversations, the video content is usually converted into text, and the text is input into the LLM as a prompt, and a conversation with the user is carried out through the LLM.
[0025] However, converting video content into text will undoubtedly lead to the loss of visual information and oversimplification of spatiotemporal complexity, affecting the reasoning effect of the dialogue system.
[0026] In order to improve the video conversation effect, the present disclosure provides the following embodiments.
[0027] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a video conversation method, the method comprising:
[0028] 101. Perform representation extraction processing on a target video to obtain an initial video representation of the target video.
[0029] 102. Perform spatiotemporal processing on the initial video representation to obtain a target video representation of the target video.
[0030] 103. Perform a conversion process on the target video representation to obtain a video embedding of the target video; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question text.
[0031] 104. Perform dialog processing on the video embedding and the text embedding to obtain an answer text.
[0032] The target video refers to the video to be processed, which can be uploaded to the dialogue system by the user.
[0033] The question text refers to the question content for the target video, which can be input by the user. For example, the user uploads the target video to the dialogue system and enters the question text for the target video. Specifically, after the user uploads the target video, the question text is "What program is this video?"
[0034] The answer text refers to the answer content obtained based on the target video and the question text. For example, based on the above example, the answer text is "This video is XX program", where XX is the specific program name obtained based on the actual situation.
[0035] For the purpose of distinction, the video representation of the target video is divided into an initial video representation and a target video representation, wherein the video representation before spatiotemporal processing is called the initial video representation, and the video representation after spatiotemporal processing is called the target video representation.
[0036] The initial video representation can be obtained by extracting the representation of the target video using a visual encoder. For example, the visual encoder is a Vision Transformer (ViT) model. The ViT model is a visual task-related representation extraction model based on a Transformer Encoder. Specifically, the target video can be input into the ViT model, and the output is the initial video representation.
[0037] Spatiotemporal processing refers to processing the time dimension and the space dimension of the initial video representation. Specifically, it may include performing four arithmetic operations on the initial video representation, and the four arithmetic operations include addition, subtraction, multiplication and division operations. Furthermore, the four arithmetic operations may be implemented by a convolution module, and the convolution module is used to perform a convolution operation, and the convolution operation may include a one-dimensional (1D) convolution operation and / or a two-dimensional (2D) convolution operation. The specific structure of the convolution module may be set according to actual needs.
[0038] The target video representation is a video representation obtained after performing spatiotemporal processing on the initial video representation.
[0039] Through spatiotemporal processing, the target video representation can contain rich spatiotemporal information. Specifically, the target video representation can represent the average property (such as the intensity of the image or the cumulative state of the image), the change property (such as the passage of time, the trend of movement), the correlation, the consistency, etc. in the target video. Among them, the average property can be represented by addition operation, the change property can be represented by subtraction operation, the correlation can be represented by multiplication operation, and the consistency can be represented by division operation.
[0040] Text embedding can be obtained by embedding the question text using a text encoder.
[0041] In order to ensure the accuracy of dialogue processing, the target video representation can be mapped to obtain video embedding. The video embedding and text embedding are in the same vector space, that is, the dimension of video embedding and text embedding is the same.
[0042] After obtaining the video embedding and text embedding, you can perform dialog processing on the video embedding and text embedding to get the answer text.
[0043] In this embodiment, the target video is represented and extracted to obtain the initial video representation. Since the processing is based on the video representation, the visual information of the target video can be fully utilized. In addition, the target video representation is obtained by performing spatiotemporal processing on the initial video representation, and the rich spatiotemporal information of the target video can also be utilized. Therefore, the video dialogue effect can be improved.
[0044] In order to better understand the embodiments of the present disclosure, application scenarios to which the embodiments of the present disclosure are applicable are described.
[0045] Figure 22 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. The scenario includes: a user terminal 201 and a server 202. The user terminal 201 may include: a personal computer (PC), a mobile phone, a tablet computer, a laptop computer, a smart wearable device, etc. The server 202 may be a cloud server or a local server. The user terminal 201 and the server 202 may communicate using a communication network, and the communication network may include, for example, a wired network and / or a wireless network.
[0046] The user can send the target video and question text to the server 202 through the user terminal 201. The server 202 performs dialogue processing based on the target video and question text to obtain the answer text; the server 202 sends the answer text to the user terminal 201 for display.
[0047] Specifically, a dialogue system can be deployed on the server to conduct human-computer dialogue through the dialogue system, that is, the dialogue system is used to generate answer text based on the target video and the question text.
[0048] The dialogue system can conduct video dialogue based on a large multimodal dialogue model of images and texts.
[0049] The large text-image multimodal dialogue model may specifically be a Large Language and Vision Assistant (LLaVA) model.
[0050] The LLaVA model is a large multimodal model that connects a visual encoder and an LLM for general vision and language understanding, where the visual encoder can be a ViT model.
[0051] LLM is a hot topic in the field of artificial intelligence in recent years. LLM is a pre-trained language model that learns rich language knowledge and world knowledge through pre-training on massive text data, so that it can achieve amazing results in various natural language processing (NLP) tasks. Wenxinyiyan, ChatGPT and other applications are developed based on LLM. They can generate fluent, logical and creative text content, and can even have natural conversations with humans.
[0052] Taking the LLaVA model as an example, the LLM corresponding to the LLaVA model can be specifically selected as StableVicuna. StableVicuna is a type of LLM.
[0053] like Figure 3As shown, the ViT model includes multiple ViT blocks. In this embodiment, two adjacent ViT blocks can be selected, and an Arithmetic Temporal Module (ATM) is additionally introduced between the two ViT blocks to perform spatiotemporal processing through the ATM.
[0054] Assume that ATM is introduced between the i-th ViT block and the (i+1)-th ViT block, where the specific value of i can be set according to actual needs.
[0055] The input of the ViT model (the first ViT block) is the target video, and the output of the ith ViT block is the initial video representation X; ATM performs spatiotemporal processing on the initial video representation X to obtain the target video representation X'.
[0056] For ATM, it mainly includes four arithmetic operations modules (expressed by addition, subtraction, multiplication and division), and may also include: context generation (context spanning) module, feature extraction (feature) module and domain transformation (domain transformation) module. The feature extraction module may include a two-dimensional convolution (Conv2D) module, and the domain transformation module may include a regroup (Regroup) module and a two-dimensional convolution (Conv2D) module.
[0057] Specifically, the dimension of the initial video representation X is T*C*H*W, where T is the number of frames of the target video, C is the number of channels, and H and W are the height and width of each image in the target video.
[0058] The context generation module is used to perform context representation extraction processing on the initial video representation X, that is, to obtain a preset number of other image representations adjacent to each image. The output of the context generation module is a video representation containing context information, which can be called a context representation, and the dimension is T*Z*C*H*W, where Z is the number of images of the context, for example, Figure 3 As shown, for each image, representations of 5 adjacent images can be obtained, and Z = 5. The context generation module can be composed of a convolution module, and the specific structure can be set as needed.
[0059] The four arithmetic operations module is used to perform four arithmetic operations on the initial video representation and the context representation. During the operation, the initial video representation can be filled to be consistent with the dimension of the context representation, and then the four arithmetic operations are performed.
[0060] Among them, the four arithmetic operation module can be composed of a convolution module (1D convolution module and / or 2D convolution module) to realize the four arithmetic operation function, and the specific structure can be set according to actual needs. In the training stage, the model parameters of the convolution module in the four arithmetic operation module can be learned, so that in the inference stage, the model structure and model parameters of the four arithmetic operation module are determined, and then the determined four arithmetic operations can be performed.
[0061] The feature extraction module may be specifically composed of a two-dimensional convolution module (Conv2D), which performs feature extraction on the representation after the four arithmetic operations to obtain the representation after feature extraction. The dimensions of the representation after the four arithmetic operations and the feature extraction are both T*Z*C*H*W.
[0062] The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
[0063] The region transformation module includes a reorganization module and a 2D convolution module.
[0064] The reorganization module is used to reorganize the dimension of the representation. The input of the reorganization module is the representation after feature extraction, and the output of the reorganization module is the representation after dimension reorganization. Specifically, the dimension of the input representation of the reorganization module is T*Z*C*H*W, and the dimension of the output representation is T*ZC*H*W, where ZC represents the multiplication of Z and C.
[0065] The 2D convolution module in the region transformation module is used to perform 2D convolution processing on the dimensionally reorganized representation to obtain the target video representation X'. The dimension of the target video representation X' is T*C*H*W.
[0066] After obtaining the target video representation, the initial video representation X and the target video representation X' can be added, and the added video representation (X+X') is used as the input of the (i+1)th ViT block. The (i+1)th ViT block further performs representation extraction processing on the input representation to obtain the output representation of the (i+1)th ViT block. After the output representation is converted, the keyword embedding (K) and value embedding (V) required by the attention network can be obtained respectively.
[0067] In addition, the input of the attention network also includes query embedding (Q). In this embodiment, the Q input by QFormer can be called the target query embedding, which includes two parts, one part is the existing Q of QFormer (original query embedding), and the other part is the additional learnable Q (newly added query embedding). The original query embedding remains unchanged during the training phase and is represented by frozen; the initial value of the newly added query embedding can be a randomized vector, and the final value of the newly added query embedding is obtained through invariant learning (adjustment) during the training phase. The frozen original query embedding and the newly added query embedding constitute the target query embedding as the input of the QFormer model.
[0068] The QFormer model is based on the attention network. After obtaining K, V, and Q, the attention network can be used to process K, V, and Q to obtain the output representation of the QFormer model. After the output representation is processed by the linear layer, the video embedding is obtained.
[0069] Through the linear layer, the dimension of the representation can be adjusted. For example, the target video representation with the dimension of T*C*H*W can be converted into a video embedding with the dimension of C*N, where N is the value obtained by compressing the multiplication of the three parameters T*H*W.
[0070] During the training phase, the model parameters of the QFormer model are kept constant (frozen), while the model parameters of the linear layer are adjusted.
[0071] The model parameters of LLM also remain unchanged during the training phase. LLM can process conversations based on text embedding and video embedding to generate answer text. Among them, text embedding and video embedding are representations of the same dimension, such as the dimension is C*N.
[0072] The LLaVA model is a multimodal model of images and texts, that is, it can understand images and texts. In this embodiment, the ViT model, ATM, QFormer model, linear layer and LLM can be used to process images and texts in a similar way, and the answer text can be generated based on video embedding and text embedding, so as to transfer the understanding of images and texts to the understanding of videos and texts, and realize knowledge transfer. Using the LLaVA model as a hot start can improve the reasoning speed and reasoning effect.
[0073] In combination with the above application scenarios, the present disclosure also provides the following embodiments.
[0074] Figure 4 is a schematic diagram according to a second embodiment of the present disclosure, and this embodiment provides a video conversation method, such as Figure 4 As shown, the method includes:
[0075] 401. Use a visual encoder to perform representation extraction processing on a target video to obtain an initial video representation of the target video.
[0076] Among them, reference Figure 3 The visual encoder can be a ViT model, and the output representation of the i-th ViT block of the ViT model is used as the initial video representation. The specific value of i can be set according to actual needs.
[0077] 402. Use an arithmetic timing module (ATM) in a video dialogue model to perform spatiotemporal processing on the initial video representation to obtain a target video representation of the target video.
[0078] Among them, reference Figure 3 , an ATM can be introduced between the i-th ViT block and the (i+1)-th ViT block. The ATM performs spatiotemporal processing on the input initial video representation X, and the output is the target video representation X'.
[0079] In this embodiment, since the ATM is trainable, an ATM with better performance can be obtained through the training process. Therefore, during reasoning, the ATM with better performance can be used to perform spatiotemporal processing on the initial video representation to obtain a target video representation with better performance, thereby improving the video dialogue effect.
[0080] like Figure 3 As shown, the ATM may include: a context generation module, a four arithmetic operation module, a feature extraction module and a region transformation module.
[0081] Accordingly, the pre-trained operation timing module is used to perform spatiotemporal processing on the initial video representation to obtain the target video representation, including:
[0082] Using the context generation module, performing context representation extraction processing on the initial video representation to obtain a context representation;
[0083] Using the four arithmetic operations module, perform four arithmetic operations on the initial video representation and the context representation to obtain a representation after the four arithmetic operations;
[0084] Using the feature extraction module, performing feature extraction processing on the representation after the four arithmetic operations to obtain a representation after feature extraction;
[0085] The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
[0086] Among them, through the four arithmetic operations, rich temporal and spatial information of the target video can be extracted, such as representing the average property of the target video (such as the intensity of the image or the cumulative state of the image) through addition operation, representing the changing property of the target video (such as the passage of time, the trend of movement) through subtraction operation, representing the correlation of the target video through multiplication operation, and representing the consistency of the target video through division operation.
[0087] Therefore, based on the four arithmetic operations, a target video representation containing rich spatiotemporal information can be obtained; in addition, through context representation extraction, feature extraction, region transformation and other processing, the representation ability of the target video representation can be further enhanced, thereby improving the video conversation effect.
[0088] 403. Perform query processing on the target video representation to obtain a video representation after query processing.
[0089] 404. Perform mapping processing on the video representation after the query processing to obtain the video embedding; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question text.
[0090] In this embodiment, by performing query processing on the target video representation, a video representation after query processing is obtained, and subsequent processing is performed based on the video representation after query processing, which can further improve the representation ability of the video representation and enhance the video conversation effect; in addition, by performing mapping processing on the video representation after query processing, a video embedding in the same vector space as the text embedding can be obtained, and then conversation processing is performed based on the video embedding and the text embedding, which can improve the accuracy of conversation processing.
[0091] The query device in the video dialogue model may be used to perform query processing on the target video representation to obtain the video representation after query processing; wherein the parameters of the query device remain unchanged during the training process of the video dialogue model.
[0092] like Figure 3 As shown, the queryer may be specifically a querying transformer (QFormer) model. The parameters of the QFormer model remain unchanged (frozen) during the training process of the video dialogue model.
[0093] Specifically, the initial video representation and the target video representation can be added to obtain an added video representation, and the added video representation can be input into the (i+1)th ViT block, and the keyword embedding (K) and the value embedding (V) can be obtained based on the output representation of the (i+1)th ViT block.
[0094] The above takes the example of ATM being set between the i-th ViT module and the (i+1)-th ViT module. It can be understood that if ATM is connected after the last ViT block, the target video representation can be directly converted into K and V; or the video representation after the initial video representation and the target video representation are added is converted into K and V.
[0095] The QFormer model processes the input K, V, and query embedding (Q), and the output is the video representation after query processing.
[0096] In this embodiment, a query device is used to query the target video representation to obtain a video representation after query processing, which can further improve the representation ability of the video representation and improve the accuracy of the video dialogue. In addition, the parameters of the query device remain unchanged during the training process of the video dialogue model, which can reduce the number of parameters that need to be adjusted and improve the model training efficiency.
[0097] To make the video embedding and the text embedding in the same vector space, the video representation after query processing can be mapped to obtain a video embedding in the same vector space as the text embedding, that is, the dimensions of the video embedding and the text embedding are the same.
[0098] As Figure 3 shown, a linear layer in the video dialogue model can be used for mapping. Among them, the parameters of the linear layer are adjustable during the training process of the video dialogue model.
[0099] In this embodiment, through the linear layer with learnable parameters, the video embedding and the text embedding can be in the same vector space, thereby improving the video dialogue effect.
[0100] For the query embedding, as Figure 3 shown, the query embedding (Q) input to the QFormer model can be called the target query embedding, which includes the original query embedding (shown on the left) and the newly added query embedding (shown on the right). The original query embedding remains unchanged (frozen) during the training process, and the newly added query embedding is adjustable (learned) during the training process.
[0101] In this embodiment, by additionally introducing a new query embedding (the newly added query embedding) and this newly added query embedding is learnable, the target video can be better contextually modeled, thereby obtaining a compact LLM-compatible video embedding and improving the video dialogue effect.
[0102] 405. Use a large language model (LLM) in the video dialogue model to perform dialogue processing on the video embedding and the text embedding to obtain an answer text; among them, the parameters of the LLM remain unchanged (frozen) during the training process of the video dialogue model.
[0103] In this embodiment, since the LLM is an existing large language model with excellent text generation capabilities, therefore, using the LLM to obtain the answer text can utilize the excellent performance of the LLM to improve the accuracy of the answer text. In addition, the parameters of the LLM remain unchanged during the training process of the video dialogue model, which can reduce the number of parameters to be adjusted and improve the model training efficiency.
[0104] The above embodiment describes the video dialogue process, and specifically, a video dialogue model can be used for video dialogue. The training process of the video dialogue model can be referred to the following embodiment.
[0105] Figure 5 is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides a training method for a video dialogue model. The video dialogue model includes: an operation timing module (ATM), and this method includes:
[0106] 501. Perform representation extraction processing on a video sample to obtain an initial video representation of the video sample.
[0107] 502. Use the timing operation module to perform spatiotemporal processing on the initial video representation to obtain a target video representation of the video sample.
[0108] 503. Perform a conversion process on the target video representation to obtain a video embedding of the video sample; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question sample.
[0109] 504. Perform dialog processing on the video embedding and the text embedding to obtain a predicted answer.
[0110] 505. Construct a loss function based on the predicted answer.
[0111] 506. Use the loss function to adjust parameters of the timing calculation module.
[0112] In the training phase, existing data sources can be used or samples can be obtained through manual collection, and the samples include video samples and question samples.
[0113] The process of obtaining the predicted answer based on the video sample and the question sample can be similar to the process of obtaining the answer text during the above-mentioned video conversation.
[0114] After obtaining the predicted answer, a loss function may be constructed based on the predicted answer and the question sample, or a loss function may be constructed based on the predicted answer and the true answer. The loss function is, for example, an autoregressive loss function.
[0115] After constructing the loss function, the back propagation (BP) algorithm can be used to adjust parameters.
[0116] In this embodiment, the initial video representation is obtained by performing representation extraction processing on the video sample. Since the processing is based on the video representation, the visual information of the video sample can be fully utilized; in addition, the target video representation is obtained by performing spatiotemporal processing on the initial video representation, and the rich spatiotemporal information of the video sample can also be utilized. Therefore, the effect of the video dialogue model can be improved.
[0117] In some embodiments, the operation sequence module includes: a context generation module, a four-arithmetic operation module, a feature extraction module and a region transformation module;
[0118] The using the timing operation module to perform spatiotemporal processing on the initial video representation to obtain a target video representation of the video sample includes:
[0119] Using the context generation module, performing context representation extraction processing on the initial video representation to obtain a context representation;
[0120] Using the four arithmetic operations module, perform four arithmetic operations on the initial video representation and the context representation to obtain a representation after the four arithmetic operations;
[0121] Using the feature extraction module, performing feature extraction processing on the representation after the four arithmetic operations to obtain a representation after feature extraction;
[0122] The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
[0123] Therefore, based on the four arithmetic operations, a target video representation containing rich spatiotemporal information can be obtained; in addition, through context representation extraction, feature extraction, region transformation and other processing, the representation ability of the target video representation can be further enhanced, thereby improving the video conversation effect.
[0124] In some embodiments, the converting the target video representation to obtain the video embedding of the target video includes:
[0125] The target video representation is query-processed to obtain a video representation after query processing;
[0126] Mapping is performed on the query-processed video representation to obtain the video embedding.
[0127] In this embodiment, by performing query processing on the target video representation, a video representation after query processing is obtained, and subsequent processing is performed based on the video representation after query processing, which can further improve the representation ability of the video representation and enhance the video conversation effect; in addition, by performing mapping processing on the video representation after query processing, a video embedding in the same vector space as the text embedding can be obtained, and then conversation processing is performed based on the video embedding and the text embedding, which can improve the accuracy of conversation processing.
[0128] In some embodiments, the video conversation model further includes: a query device;
[0129] The query processing of the target video representation to obtain a video representation after the query processing includes:
[0130] The query device is used to perform query processing on the target video representation to obtain a video representation after query processing; wherein the parameters of the query device remain unchanged during the training process of the video dialogue model.
[0131] In this embodiment, a query device is used to query the target video representation to obtain a video representation after query processing, which can further improve the representation ability of the video representation and improve the accuracy of the video dialogue. In addition, the parameters of the query device remain unchanged during the training process of the video dialogue model, which can reduce the number of parameters that need to be adjusted and improve the model training efficiency.
[0132] In some embodiments, the querying device is used to perform query processing on the target video representation to obtain a video representation after query processing, including:
[0133] The query device is used to perform query processing on the target video representation based on the target query embedding to obtain the video representation after query processing; wherein the target query embedding includes: the original query embedding and the newly added query embedding, and the parameters of the original query embedding remain unchanged during the training process;
[0134] The method further comprises:
[0135] The loss function is used to adjust the parameters of the newly added query embedding.
[0136] In this embodiment, by additionally introducing a new query embedding (newly added query embedding), and the new query embedding is learnable, the context modeling of the target video can be better performed, thereby obtaining a compact LLM-compatible video embedding and improving the video conversation effect.
[0137] In some embodiments, the video conversation model further includes: a linear layer;
[0138] The mapping process is performed on the query-processed video representation to obtain the video embedding, including:
[0139] Using the linear layer, mapping the query-processed video representation to obtain the video embedding;
[0140] The method further comprises:
[0141] The loss function is used to adjust the parameters of the linear layer.
[0142] In this embodiment, through the linear layer with learnable parameters, the video embedding and the text embedding can be placed in the same vector space, thereby improving the video conversation effect.
[0143] In some embodiments, the video conversation model further includes: a large language model;
[0144] The performing dialog processing on the video embedding and the text embedding to obtain a predicted answer includes:
[0145] The large language model is used to perform dialog processing on the video embedding and the text embedding to obtain a predicted answer; wherein the parameters of the large language model remain unchanged during the training process of the video dialog model.
[0146] In this embodiment, since LLM is an existing large language model with excellent text generation capabilities, LLM is used to obtain the answer text, and the excellent performance of LLM can be used to improve the accuracy of the answer text. In addition, the parameters of LLM remain unchanged during the training process of the video dialogue model, which can reduce the number of parameters that need to be adjusted and improve the model training efficiency.
[0147] In addition, overall, since only the parameters of ATM, the parameters of the newly added query embedding, and the parameters of the linear layer are adjusted during model training, the parameters of the other models (such as the QFormer model and LLM) remain unchanged, efficient parameter adjustment can be achieved, the number of parameters that need to be adjusted can be reduced, and the amount of instruction data (question samples) can be reduced, thereby improving model training efficiency. In addition, using the large multimodal model of images and text as a hot start can transfer image and text understanding knowledge to video and text understanding, improve the training efficiency of the video dialogue model, and improve the dialogue generation effect of the video dialogue model.
[0148] Figure 6 is a schematic diagram according to the fourth embodiment of the present disclosure. This embodiment provides a video conversation device, such as Figure 6 As shown, the device 600 includes: an extraction module 601, a processing module 602, a conversion module 603 and a generation module 604.
[0149] The extraction module 601 is used to perform representation extraction processing on the target video to obtain the initial video representation of the target video; the processing module 602 is used to perform spatiotemporal processing on the initial video representation to obtain the target video representation of the target video; the conversion module 603 is used to perform conversion processing on the target video representation to obtain the video embedding of the target video; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question text; the generation module 604 is used to perform dialogue processing on the video embedding and the text embedding to obtain the answer text.
[0150] In this embodiment, the target video is represented and extracted to obtain the initial video representation. Since the processing is based on the video representation, the visual information of the target video can be fully utilized. In addition, the target video representation is obtained by performing spatiotemporal processing on the initial video representation, and the rich spatiotemporal information of the target video can also be utilized. Therefore, the video dialogue effect can be improved.
[0151] In some embodiments, the processing module 602 is further used to: use an operation timing module in the video conversation model to perform spatiotemporal processing on the initial video representation to obtain the target video representation; wherein the parameters of the operation timing module are adjustable during the training process of the video conversation model.
[0152] In this embodiment, since the ATM is trainable, an ATM with better performance can be obtained through the training process. Therefore, during reasoning, the ATM with better performance can be used to perform spatiotemporal processing on the initial video representation to obtain a target video representation with better performance, thereby improving the video dialogue effect.
[0153] In some embodiments, the operation sequence module includes: a context generation module, a four-arithmetic operation module, a feature extraction module and a region transformation module; the processing module 602 is further used to:
[0154] Using the context generation module, performing context representation extraction processing on the initial video representation to obtain a context representation;
[0155] Using the four arithmetic operations module, perform four arithmetic operations on the initial video representation and the context representation to obtain a representation after the four arithmetic operations;
[0156] Using the feature extraction module, performing feature extraction processing on the representation after the four arithmetic operations to obtain a representation after feature extraction;
[0157] The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
[0158] Therefore, based on the four arithmetic operations, a target video representation containing rich spatiotemporal information can be obtained; in addition, through context representation extraction, feature extraction, region transformation and other processing, the representation ability of the target video representation can be further enhanced, thereby improving the video conversation effect.
[0159] In some embodiments, the conversion module 603 is further used to: perform query processing on the target video representation to obtain a video representation after the query processing; and perform mapping processing on the video representation after the query processing to obtain the video embedding.
[0160] In this embodiment, by performing query processing on the target video representation, a video representation after query processing is obtained, and subsequent processing is performed based on the video representation after query processing, which can further improve the representation ability of the video representation and enhance the video conversation effect; in addition, by performing mapping processing on the video representation after query processing, a video embedding in the same vector space as the text embedding can be obtained, and then conversation processing is performed based on the video embedding and the text embedding, which can improve the accuracy of conversation processing.
[0161] In some embodiments, the conversion module 603 is further used to: use a query device in the video conversation model to query the target video representation to obtain a video representation after query processing; wherein the parameters of the query device remain unchanged during the training process of the video conversation model.
[0162] In this embodiment, a query device is used to query the target video representation to obtain a video representation after query processing, which can further improve the representation ability of the video representation and improve the accuracy of the video dialogue. In addition, the parameters of the query device remain unchanged during the training process of the video dialogue model, which can reduce the number of parameters that need to be adjusted and improve the model training efficiency.
[0163] In some embodiments, the conversion module 603 is further used to: use the query device to perform query processing on the target video representation based on the target query embedding to obtain a video representation after query processing; wherein the target query embedding includes: original query embedding and newly added query embedding, the parameters of the original query embedding remain unchanged during the training process, and the parameters of the newly added query embedding are adjustable during the training process.
[0164] In this embodiment, by additionally introducing a new query embedding (newly added query embedding), and the new query embedding is learnable, the context modeling of the target video can be better performed, thereby obtaining a compact LLM-compatible video embedding and improving the video conversation effect.
[0165] In some embodiments, the conversion module 603 is further used to: use a linear layer in the video conversation model to map the video representation after the query processing to obtain the video embedding; wherein the parameters of the linear layer are adjustable during the training process of the video conversation model.
[0166] In this embodiment, through the linear layer with learnable parameters, the video embedding and the text embedding can be placed in the same vector space, thereby improving the video conversation effect.
[0167] In some embodiments, the generation module 604 is further used to: use a large language model in the video dialogue model to perform dialogue processing on the video embedding and the text representation to obtain the answer text; wherein the parameters of the large language model remain unchanged during the training process of the video dialogue model.
[0168] In this embodiment, since LLM is an existing large language model with excellent text generation capabilities, LLM is used to obtain the answer text, and the excellent performance of LLM can be used to improve the accuracy of the answer text. In addition, the parameters of LLM remain unchanged during the training process of the video dialogue model, which can reduce the number of parameters that need to be adjusted and improve the model training efficiency.
[0169] Figure 7 is a schematic diagram according to the fifth embodiment of the present disclosure. This embodiment provides a training device for a video dialogue model, wherein the video dialogue model comprises: a timing operation module, such as Figure 7 As shown, the device 700 includes: an extraction module 701, a processing module 702, a conversion module 703, a generation module 704, a construction module 705 and an adjustment module 706.
[0170] The extraction module 701 is used to perform representation extraction processing on the video sample to obtain the initial video representation of the video sample; the processing module 702 is used to use the operation timing module to perform spatiotemporal processing on the initial video representation to obtain the target video representation of the video sample; the conversion module 703 is used to perform conversion processing on the target video representation to obtain the video embedding of the video sample; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question sample; the generation module 704 is used to perform dialogue processing on the video embedding and the text embedding to obtain a predicted answer; the construction module 705 is used to construct a loss function based on the predicted answer; the adjustment module 706 is used to use the loss function to adjust the parameters of the operation timing module.
[0171] In this embodiment, the initial video representation is obtained by performing representation extraction processing on the video sample. Since the processing is based on the video representation, the visual information of the video sample can be fully utilized; in addition, the target video representation is obtained by performing spatiotemporal processing on the initial video representation, and the rich spatiotemporal information of the video sample can also be utilized. Therefore, the effect of the video dialogue model can be improved.
[0172] In some embodiments, the operation sequence module includes: a context generation module, a four-arithmetic operation module, a feature extraction module and a region transformation module; the processing module 702 is further used to:
[0173] Using the context generation module, performing context representation extraction processing on the initial video representation to obtain a context representation;
[0174] Using the four arithmetic operations module, perform four arithmetic operations on the initial video representation and the context representation to obtain a representation after the four arithmetic operations;
[0175] Using the feature extraction module, performing feature extraction processing on the representation after the four arithmetic operations to obtain a representation after feature extraction;
[0176] The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
[0177] Therefore, based on the four arithmetic operations, a target video representation containing rich spatiotemporal information can be obtained; in addition, through context representation extraction, feature extraction, region transformation and other processing, the representation ability of the target video representation can be further enhanced, thereby improving the video conversation effect.
[0178] In some embodiments, the conversion module 703 is further used to: perform query processing on the target video representation to obtain a video representation after the query processing; and perform mapping processing on the video representation after the query processing to obtain the video embedding.
[0179] In this embodiment, by performing query processing on the target video representation, a video representation after query processing is obtained, and subsequent processing is performed based on the video representation after query processing, which can further improve the representation ability of the video representation and enhance the video conversation effect; in addition, by performing mapping processing on the video representation after query processing, a video embedding in the same vector space as the text embedding can be obtained, and then conversation processing is performed based on the video embedding and the text embedding, which can improve the accuracy of conversation processing.
[0180] In some embodiments, the video conversation model further includes: a query device; the conversion module 703 is further used to: use the query device to query the target video representation to obtain a video representation after query processing; wherein the parameters of the query device remain unchanged during the training process of the video conversation model.
[0181] In this embodiment, a query device is used to query the target video representation to obtain a video representation after query processing, which can further improve the representation ability of the video representation and improve the accuracy of the video dialogue. In addition, the parameters of the query device remain unchanged during the training process of the video dialogue model, which can reduce the number of parameters that need to be adjusted and improve the model training efficiency.
[0182] In some embodiments, the conversion module 703 is further used to: use the query device to perform query processing on the target video representation based on the target query embedding to obtain a video representation after query processing; wherein the target query embedding includes: the original query embedding and the newly added query embedding, and the parameters of the original query embedding remain unchanged during the training process; the adjustment module 706 is also used to: use the loss function to adjust the parameters of the newly added query embedding.
[0183] In this embodiment, by additionally introducing a new query embedding (newly added query embedding), and the new query embedding is learnable, the context modeling of the target video can be better performed, thereby obtaining a compact LLM-compatible video embedding and improving the video conversation effect.
[0184] In some embodiments, the video conversation model further includes: a linear layer; the conversion module 703 is further used to: use the linear layer to map the video representation after the query processing to obtain the video embedding; the adjustment module 706 is also used to: use the loss function to adjust the parameters of the linear layer.
[0185] In this embodiment, through the linear layer with learnable parameters, the video embedding and the text embedding can be placed in the same vector space, thereby improving the video conversation effect.
[0186] In some embodiments, the video conversation model further includes: a large language model; the generation module 704 is further used to: use the large language model to perform conversation processing on the video embedding and the text embedding to obtain a predicted answer; wherein the parameters of the large language model remain unchanged during the training process of the video conversation model.
[0187] In this embodiment, since LLM is an existing large language model with excellent text generation capabilities, LLM is used to obtain the answer text, and the excellent performance of LLM can be used to improve the accuracy of the answer text. In addition, the parameters of LLM remain unchanged during the training process of the video dialogue model, which can reduce the number of parameters that need to be adjusted and improve the model training efficiency.
[0188] It can be understood that in the embodiments of the present disclosure, the same or similar contents in different embodiments can be referenced to each other.
[0189] It can be understood that the “first”, “second”, etc. in the embodiments of the present disclosure are only used for distinction and do not indicate the degree of importance, time sequence, etc.
[0190] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0191] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0192] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device 800 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0193] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0194] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0195] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the training method or video conversation method of the video conversation model. For example, in some embodiments, the training method or video conversation method of the video conversation model may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the training method or video conversation method of the video conversation model described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the video conversation model training method or the video conversation method in any other suitable manner (eg, by means of firmware).
[0196] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0197] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0198] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0199] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0200] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0201] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0202] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0203] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A video conversation method, include: Performing a representation extraction process on the target video to obtain an initial video representation of the target video; Performing spatiotemporal processing on the initial video representation to obtain a target video representation of the target video; Performing a conversion process on the target video representation to obtain a video embedding of the target video; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question text; performing dialog processing on the video embedding and the text embedding to obtain an answer text; The converting the target video representation to obtain the video embedding of the target video includes: Using a queryer with unchanged parameters, query processing is performed on the target video representation based on a target query embedding to obtain a video representation after query processing; the target query embedding includes: an unchanged original query embedding and an adjustable new query embedding; An adjustable linear layer is used to perform mapping processing on the query-processed video representation to obtain the video embedding.
2. The method according to claim 1, in, The performing spatiotemporal processing on the initial video representation to obtain a target video representation of the target video includes: The initial video representation is subjected to spatiotemporal processing using a timing operation module in a video dialogue model to obtain the target video representation; wherein the parameters of the timing operation module are adjustable during the training process of the video dialogue model.
3. The method according to claim 2, in, The operation sequence module includes: a context generation module, a four-arithmetic operation module, a feature extraction module and a region transformation module; The using the operation timing module in the video dialogue model to perform spatiotemporal processing on the initial video representation to obtain the target video representation includes: Using the context generation module, performing context representation extraction processing on the initial video representation to obtain a context representation; Using the four arithmetic operations module, perform four arithmetic operations on the initial video representation and the context representation to obtain a representation after the four arithmetic operations; Using the feature extraction module, performing feature extraction processing on the representation after the four arithmetic operations to obtain a representation after feature extraction; The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
4. The method according to any one of claims 1 to 3, in, The performing dialog processing on the video embedding and the text embedding to obtain an answer text comprises: A large language model in the video dialogue model is used to perform dialogue processing on the video embedding and the text embedding to obtain the answer text; wherein the parameters of the large language model remain unchanged during the training process of the video dialogue model.
5. A method for training a video dialogue model, wherein the video dialogue model include: The timing operation module comprises: Performing a representation extraction process on the video sample to obtain an initial video representation of the video sample; Using the operation timing module, performing spatiotemporal processing on the initial video representation to obtain a target video representation of the video sample; Performing a conversion process on the target video representation to obtain a video embedding of the video sample; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question sample; Performing dialog processing on the video embedding and the text embedding to obtain a predicted answer; Constructing a loss function based on the predicted answer; Using the loss function, adjusting the parameters of the timing operation module; The converting the target video representation to obtain the video embedding of the target video includes: Using a queryer with unchanged parameters, query processing is performed on the target video representation based on a target query embedding to obtain a video representation after query processing; the target query embedding includes: an unchanged original query embedding and an adjustable new query embedding; An adjustable linear layer is used to perform mapping processing on the query-processed video representation to obtain the video embedding.
6. The method according to claim 5, in, The operation sequence module includes: a context generation module, a four-arithmetic operation module, a feature extraction module and a region transformation module; The using the timing operation module to perform spatiotemporal processing on the initial video representation to obtain a target video representation of the video sample includes: Using the context generation module, performing context representation extraction processing on the initial video representation to obtain a context representation; Using the four arithmetic operations module, perform four arithmetic operations on the initial video representation and the context representation to obtain a representation after the four arithmetic operations; Using the feature extraction module, performing feature extraction processing on the representation after the four arithmetic operations to obtain a representation after feature extraction; The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
7. The method according to any one of claims 5 to 6, in, The video dialogue model also includes: a large language model; The performing dialog processing on the video embedding and the text embedding to obtain a predicted answer includes: The large language model is used to perform dialog processing on the video embedding and the text embedding to obtain a predicted answer; wherein the parameters of the large language model remain unchanged during the training process of the video dialog model.
8. A video conversation device, include: An extraction module, used for performing a representation extraction process on a target video to obtain an initial video representation of the target video; A processing module, configured to perform spatiotemporal processing on the initial video representation to obtain a target video representation of the target video; A conversion module, configured to convert the target video representation to obtain a video embedding of the target video; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question text; A generation module, configured to perform dialog processing on the video embedding and the text embedding to obtain an answer text; The conversion module is further used for: Using a queryer with unchanged parameters, query processing is performed on the target video representation based on a target query embedding to obtain a video representation after query processing; the target query embedding includes: an unchanged original query embedding and an adjustable new query embedding; An adjustable linear layer is used to perform mapping processing on the query-processed video representation to obtain the video embedding.
9. The device according to claim 8, in, The processing module is further configured to: The initial video representation is subjected to spatiotemporal processing using a timing operation module in a video dialogue model to obtain the target video representation; wherein the parameters of the timing operation module are adjustable during the training process of the video dialogue model.
10. The device according to claim 9, in, The operation sequence module includes: a context generation module, a four-arithmetic operation module, a feature extraction module and a region transformation module; The processing module is further configured to: Using the context generation module, performing context representation extraction processing on the initial video representation to obtain a context representation; Using the four arithmetic operations module, perform four arithmetic operations on the initial video representation and the context representation to obtain a representation after the four arithmetic operations; Using the feature extraction module, performing feature extraction processing on the representation after the four arithmetic operations to obtain a representation after feature extraction; The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
11. The device according to any one of claims 8 to 10, in, The generating module is further used for: A large language model in the video dialogue model is used to perform dialogue processing on the video embedding and the text embedding to obtain the answer text; wherein the parameters of the large language model remain unchanged during the training process of the video dialogue model.
12. A training device for a video dialogue model, wherein the video dialogue model include: A timing operation module, the device comprises: An extraction module, used for performing a representation extraction process on the video sample to obtain an initial video representation of the video sample; A processing module, configured to perform spatiotemporal processing on the initial video representation using the timing operation module to obtain a target video representation of the video sample; A conversion module, configured to convert the target video representation to obtain a video embedding of the video sample; wherein the dimension of the video embedding is the same as the dimension of the text embedding of the question sample; A generation module, configured to perform dialog processing on the video embedding and the text embedding to obtain a predicted answer; A construction module, configured to construct a loss function based on the predicted answer; An adjustment module, used to adjust the parameters of the timing operation module by using the loss function; The conversion module is further used for: Using a queryer with unchanged parameters, query processing is performed on the target video representation based on a target query embedding to obtain a video representation after query processing; the target query embedding includes: an unchanged original query embedding and an adjustable new query embedding; An adjustable linear layer is used to perform mapping processing on the query-processed video representation to obtain the video embedding.
13. The device according to claim 12, in, The operation sequence module includes: a context generation module, a four-arithmetic operation module, a feature extraction module and a region transformation module; The processing module is further configured to: Using the context generation module, performing context representation extraction processing on the initial video representation to obtain a context representation; Using the four arithmetic operations module, perform four arithmetic operations on the initial video representation and the context representation to obtain a representation after the four arithmetic operations; Using the feature extraction module, performing feature extraction processing on the representation after the four arithmetic operations to obtain a representation after feature extraction; The region transformation module is used to perform region transformation processing on the representation after feature extraction to obtain the target video representation.
14. The device according to any one of claims 12 to 13, in, The video dialogue model also includes: a large language model; The generating module is further used for: The large language model is used to perform dialog processing on the video embedding and the text embedding to obtain a predicted answer; wherein the parameters of the large language model remain unchanged during the training process of the video dialog model.
15. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, in, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.