Target multi-modal model system and construction method, video processing model training method, and video processing method

WO2025186663A8PCT designated stage Publication Date: 2025-10-02CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/052098
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-08
Filing Date
2025-02-27
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing technologies lose important dynamic information in video analysis, affecting scenarios that require dynamic information analysis, such as urban traffic accidents and urban elderly people falling and fainting.

Method used

A target multimodal model system is adopted, including a temporal feature extraction encoding layer and a dynamic feature extraction encoding layer. By performing temporal feature extraction on the first modality data and dynamic feature extraction on the second modality data, the dynamic information understanding ability of the model is enhanced by combining language supervision, and the modules are deployed in a distributed system to improve processing efficiency.

Benefits of technology

It realizes the accurate extraction and analysis of dynamic information of video data, improves the accuracy and efficiency of video content analysis, adapts to real-time analysis needs, and overcomes the memory bottleneck problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025052098_02102025_PF_FP_ABST
    Figure IB2025052098_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a target multi-modal model system and construction method, a video processing model training method, and a video processing method. The video processing model training method comprises: inputting a video sample and each initial text sample into a video processing model, wherein the initial text sample is a text for performing category description on video content of the video sample; using the video processing model to perform feature extraction on the video sample to obtain a temporal motion feature and a fused image feature; using the video processing model to perform feature extraction on the initial text sample to obtain a dynamic text feature and a fused text feature; and training the video processing model on the basis of the temporal motion feature and the dynamic text feature, and the fused image feature and the fused text feature.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Target multimodal model system and construction method, video processing model training method, video processing method technical field

[0002]

[0001] The present disclosure relates to the field of computer technology, and more particularly to a target multimodal model system and construction method, a video processing model training method, and a video processing method.

[0003]

[0002] In practical applications, video analysis and processing is a very important visual analysis content. Existing solutions usually use image-level detection and tracking to obtain images of the target of interest, and then perform image content understanding, such as image classification, image feature comparison, etc.

[0004]

[0003] Image-based processing logic can ensure the efficiency of algorithm operation, but for video data, it undoubtedly loses important dynamic information, affecting scenarios that require dynamic information analysis, such as urban traffic vehicle accidents, abnormal behavior analysis of key urban individuals, and urban elderly people falling and fainting.

[0005] In view of this, embodiments of the present disclosure provide a target multimodal model system. One or more embodiments of the present disclosure also relate to a target multimodal model construction method, a video processing model training method, a video processing method, a video processing model training device, a video processing device, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0006]

[0005] According to a first aspect of an embodiment of the present disclosure, a target multimodal model system is provided, comprising a target multimodal model, wherein the target multimodal model comprises a first modal fusion feature extraction module and a second modal fusion feature extraction module, wherein the first modal fusion feature extraction module comprises a temporal feature extraction coding layer, and the second modal fusion feature extraction module comprises a dynamic feature extraction coding layer, wherein the temporal feature extraction coding layer is configured to perform temporal feature extraction on first modal data to obtain temporal motion features, and obtain first modal fusion features based on the temporal motion features and the first modal features, wherein the first modal features are obtained by performing first modal feature extraction on the first modal data; and the dynamic feature extraction coding layer is configured to perform dynamic feature extraction on the second modal data to obtain dynamic features, and obtain second modal fusion features based on the dynamic features and the static features, wherein The second modality data is obtained by enhancing the initial second modality data, the static features are obtained by extracting static features from the initial second modality data, and the first modality fusion features and the second modality fusion features are used to determine target second modality data.

[0007]

[0006] According to a second aspect of an embodiment of the present disclosure, a method for constructing a target multimodal model is provided, comprising: receiving a task request, wherein the task request includes task parameters; obtaining an input layer, a temporal feature extraction coding layer, a dynamic feature extraction coding layer, and an output layer based on the task parameters; connecting the input layer with the temporal feature extraction coding layer and the dynamic feature extraction coding layer, respectively, and connecting the temporal feature extraction coding layer and the dynamic feature extraction coding layer with the output layer to obtain a multimodal model to be trained; training the multimodal model to be trained to obtain a target multimodal model that has completed training.

[0008]

[0007] According to a third aspect of an embodiment of the present disclosure, a video processing model training method is provided, comprising: inputting a video sample and an initial text sample into a video processing model, wherein the initial text sample is a text that categorizes the video content of the video sample; performing feature extraction on the video sample using the video processing model to obtain temporal motion features and fused image features; performing feature extraction on the initial text sample using the video processing model to obtain dynamic text features and fused text features; and training the video processing model based on the temporal motion features and the dynamic text features, as well as the fused image features and the fused text features.

[0009]

[0008] According to a fourth aspect of an embodiment of the present disclosure, a video processing model training device is provided, comprising: an input module, configured to input a video sample and an initial text sample into a video processing model, wherein the initial text sample is a text that categorizes the video content of the video sample; a video extraction module, configured to perform feature extraction on the video sample using the video processing model to obtain temporal motion features and fused image features; a text extraction module, configured to perform feature extraction on the initial text sample using the video processing model to obtain dynamic text features and fused text features; and a training module, configured to train the video processing model based on the temporal motion features and the dynamic text features, as well as the fused image features and the fused text features.

[0010]

[0009] According to a fifth aspect of an embodiment of the present disclosure, a video processing method is provided, comprising: inputting a target video and a plurality of initial texts into a video processing model, wherein the plurality of initial texts are texts that categorize the video content of the target video; using the video processing model, determining a target text corresponding to the target video from the plurality of initial texts, wherein the video processing model is trained by the above-mentioned video processing model training method.

[0011]

[0010] According to the sixth aspect of the embodiment of the present disclosure, a video processing device is provided, comprising: an input module, configured to input a target video and multiple initial texts into a video processing model, wherein the multiple initial texts are texts that categorize the video content of the target video; a determination module, configured to use the video processing model to determine a target text corresponding to the target video from the multiple initial texts, wherein the video processing model is trained by the above-mentioned video processing model training method.

[0012]

[0011] According to a seventh aspect of an embodiment of the present disclosure, a video processing method is provided, which is applied to a traffic scene, comprising: inputting a target traffic video and a plurality of initial event texts into a video processing model, wherein the plurality of initial event texts are texts that categorize the video content of the target traffic video; using the video processing model, determining a target event text corresponding to the target traffic video from the plurality of initial event texts, wherein the video processing model is trained by the above-mentioned video processing model training method.

[0013]

[0012] According to an eighth aspect of an embodiment of the present disclosure, a computing device is provided, comprising: a memory and a processor; the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the above-mentioned video processing model training method or video processing method are implemented.

[0014]

[0013] According to a ninth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the above-mentioned video processing model training method or video processing method are implemented.

[0015]

[0014] According to the tenth aspect of the embodiment of the present disclosure, a computer program product is provided, including a computer program / instruction, which implements the above-mentioned video processing model training method or the steps of the video processing method when executed by a processor.

[0016]

[0015] The target multimodal model system provided by the embodiment of the present disclosure includes a target multimodal model, wherein the target multimodal model includes a first modal fusion feature extraction module and a second modal fusion feature extraction module, wherein the first modal fusion feature extraction module includes a temporal feature extraction coding layer, and the second modal fusion feature extraction module includes a dynamic feature extraction coding layer, wherein the temporal feature extraction coding layer is used to perform temporal feature extraction on the first modal data to obtain temporal motion features, and obtain first modal fusion features based on the temporal motion features and the first modal features, wherein the first modal features are obtained by performing first modal feature extraction on the first modal data; the dynamic feature extraction coding layer is used to perform dynamic feature extraction on the second modal data to obtain dynamic features, and obtain second modal fusion features based on the dynamic features and static features, wherein the second modal data is obtained by performing enhancement processing on the initial second modal data, and the static features are obtained by performing static feature extraction on the initial second modal data, and the first modal fusion features and the second modal fusion features are used to determine the target second modal data.

[0017]

[0016] Based on this, the target multimodal model system provided by the embodiment of the present disclosure uses the target multimodal model to perform time series feature extraction on the input first modal data to obtain time series motion features, thereby obtaining dynamic information of the first modal data. In order to ensure the accuracy of the dynamic information extracted by the target multimodal model, the target multimodal model is used to perform feature extraction on the input second modal data to obtain dynamic features. When the first modal data is video or image data and the second modal data is text data, language supervision is used to enhance the target multimodal model's learning ability for action details in the video or image. The first modal fusion feature extraction module and the second modal fusion feature extraction module in the target multimodal model can be deployed on different hardware nodes respectively. Each hardware node only needs to be responsible for the data and calculations of its corresponding part, effectively avoiding the memory bottleneck problem and improving the processing efficiency of each hardware node.

[0018] FIG1 is a schematic diagram of a video processing method according to an embodiment of the present disclosure;

[0019] FIG2 is a system structure diagram of a target multimodal model system provided by one embodiment of the present disclosure;

[0020] FIG3 is a flow chart of a method for constructing a target multimodal model according to an embodiment of the present disclosure;

[0021] FIG4 is a flow chart of a video processing model training method provided by one embodiment of the present disclosure;

[0022] FIG5 is a flow chart of a video processing model training method according to an embodiment of the present disclosure;

[0023] FIG6 is an overall framework diagram of a video processing model provided by one embodiment of the present disclosure;

[0024] FIG7 is a flow chart of a video processing method provided by one embodiment of the present disclosure;

[0025]

[0024] FIG8 is a flow chart of a video processing method applied to a traffic scene provided by one embodiment of the present disclosure;

[0026] FIG9 is a schematic diagram of a video processing model training device according to an embodiment of the present disclosure;

[0027] FIG10 is a schematic structural diagram of a video processing device provided by one embodiment of the present disclosure;

[0028]

[0027] FIG11 is a block diagram of a computing device provided by an embodiment of the present disclosure.

[0029] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways different from those described herein, and those skilled in the art can make similar promotions without violating the connotation of the present disclosure. Therefore, the present disclosure is not limited by the specific implementation disclosed below.

[0030]

[0029] The terms used in one or more embodiments of the present disclosure are intended only to describe specific embodiments and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a," "an," "the," and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0031] It should be understood that although the terms "first," "second," and so on may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first," without departing from the scope of one or more embodiments of the present disclosure. Depending on the context, the term "if" as used herein may be interpreted as "at the time," "when," or "in response to a determination."

[0032]

[0031] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0033] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, typically including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities. Examples include large language models (LLMs) and multimodal pre-trained models.

[0034]

[0033] In practical applications, large models only require a small number of samples to fine-tune the pre-training model and can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields, specifically in computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0035]

[0034] First, the terms involved in one or more embodiments of the present disclosure are explained.

[0036]

[0035] Video action recognition: Automatically identifying different actions or motion categories from video data; it is an important task in the field of computer vision and machine learning, aiming to understand and infer the behaviors shown in the video by analyzing the motion patterns and action sequences in the video.

[0037]

[0036] CLIP (Contrastive Language-Image Pretraining) is a large multimodal model for text and image data. It learns a unified representation space by jointly training image and text data, so that images and text can be compared and matched in this space. The core idea of ​​the CLIP model is to use contrastive learning to

[0038] (contrastive learning) is used to train the representation of images and texts, so that similar images and texts are closer in the representation space, while dissimilar images and texts are farther away.

[0039]

[0037] Traditional large multimodal models do not have the ability to analyze and understand videos. Although some technologies have explored how to introduce a timing module into the network middle layer of a large multimodal model of images and text to give the model dynamic understanding capabilities, there is a lack of exploration and utilization of language supervision. Another technology studies how to simultaneously introduce a multimodal prompt learning module into the input layer of the model to learn visual-text semantic matching. However, this method has limited ability to extract dynamic information and has not conducted in-depth research on how to model and learn dynamic video information.

[0040]

[0038] The video processing model training method provided in the embodiments of the present disclosure aims to study how to migrate a large multimodal model of images and text to video tasks, and use a lightweight dynamic information perception module to give the large multimodal model the ability to understand dynamic information in videos, thereby meeting the real-time analysis requirements in practical applications.

[0041]

[0039] In the present disclosure, a video processing model training method is provided. The present disclosure also relates to a video processing method, a video processing model training device, a video processing device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0042]

[0040] Referring to FIG1 , FIG1 shows a scene diagram of a video processing method provided according to an embodiment of the present disclosure.

[0043]

[0041] Specifically, the video processing method is implemented by applying the end-device 102 and the server 104. The end-device 102 is used to send a target video and multiple initial texts to the server 104. For example, the video content in the target video is "a scene of a car-to-car collision in an urban traffic scene", and the multiple initial texts can be description texts of different types of traffic accidents such as "car-to-car collision, car-to-pedestrian collision, electric vehicle-to-car collision, electric vehicle-to-electric vehicle collision"; in actual application, the user can input the initial text in the end-device 102 in text or voice. If voice is used, the end-device 102 will also include a corresponding voice processing part, such as voice analysis, voice-to-text, voice synthesis and other modules, which are used to convert the questions input by the user through voice into text, and the present disclosure does not limit this.

[0044]

[0042] A video processing model is trained in the server 104. When the server 104 receives the target video and multiple initial texts sent by the end device 102, the video processing model is used to extract features of the input target video to obtain temporal motion features and image features. The fused image features are obtained by fusing the temporal motion features and the image features. The video processing model is used to extract features of the input initial text to obtain dynamic text features and static text features. The fused text features are obtained by fusing the dynamic text features and the static text features. Thus, the similarity between the target video and each initial text is obtained by fusing the image features and the fused text features. The similarities are arranged in descending order, and the initial text corresponding to the first similarity is determined as the target text corresponding to the target video. The target text matches the video content of the target video, that is, the target text is "cars collide with cars", and the target text is returned to the end device 102.

[0045]

[0043] The end device 102 may include a browser, an APP (Application), or a web application such as an H5 (Hypertext Markup Languages, version 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The end device may be developed based on a software development kit (SDK) of a corresponding service provided by the server, such as a real-time communication (RTC) SDK. The end device may be deployed in an electronic device and may rely on the device to run or on certain APPs in the device to run. The electronic device may have a display and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, or a personal computer. Various other types of applications may also be configured in the electronic device, such as human-computer interaction applications, model training applications, video processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0046]

[0044] Server 104 can be understood as a server that provides various services, including physical servers and cloud servers. For example, a server that provides communication services to multiple clients, a server that supports backend training of models used on clients, or a server that processes data sent by clients. It should be noted that server 104 can be implemented as a distributed server cluster consisting of multiple servers or as a single server. Server 104 can also be a server in a distributed system or a server integrated with blockchain. Server 104 can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0047]

[0045] It is worth noting that the video processing method provided in the embodiment of the present disclosure can be executed by the server 104. In other embodiments of the present disclosure, the video processing model can be deployed in the terminal G1 device 102, so that the terminal G1 device 102 can also have similar functions as the server 104, thereby executing the video processing method provided in the embodiment of the present disclosure; in other embodiments, the data detection method provided in the embodiment of the present disclosure can also be jointly executed by the terminal G1 device 102 and the server 104.

[0048]

[0046] The video processing method provided by the embodiment of the present disclosure uses a video processing model to perform feature extraction on an input target video to obtain temporal motion features, thereby obtaining dynamic information of a video sample. To ensure the accuracy of the dynamic information extracted by the video processing model, the video processing model is used to perform feature extraction on an input initial text to obtain dynamic text features. By improving the video processing model's ability to analyze the dynamic information of the video, the video analysis capability of the video content is improved. Therefore, when the video processing model is used to select text from the initial text according to the target video, the target text that matches the video content of the target video can be accurately selected from multiple initial texts that describe the category of the video content, thereby improving the accuracy of the target text.

[0049]

[0047] Referring to FIG2 , FIG2 shows a system structure diagram of a target multimodal model system provided by an embodiment of the present disclosure.

[0050]

[0048] Specifically, the target multimodal model system includes a target multimodal model, and the target multimodal model includes a first modal fusion feature extraction module and a second modal fusion feature extraction module, the first modal fusion feature extraction module includes a temporal feature extraction coding layer, and the second modal fusion feature extraction module includes a dynamic feature extraction coding layer.

[0051]

[0049] In one embodiment of the present disclosure, the target multimodal model may be an audio processing model, and correspondingly, the first modal data may be audio data, and the second modal data may be text data, and the audio processing model is used to perform text recognition on the audio; in another embodiment of the present disclosure, the target multimodal model may be a video processing model, and correspondingly, the first modal data may be video or image data, and the second modal data may be text data, and the video processing model in this case is used to analyze and identify the video content.

[0052]

[0050] The embodiment of the present disclosure takes the target multimodal model system applied to the visual field, the target multimodal model is a visual processing model, the first modal data is visual data, the visual data includes video or image, and the second modal data is text data as an example to describe the target multimodal model system in detail.

[0053]

[0051] The temporal feature extraction coding layer is used to perform temporal feature extraction on the first modal data to obtain temporal motion features, so as to obtain first modal fusion features based on the temporal motion features and the first modal features, wherein the first modal features are obtained by performing first modal feature extraction on the first modal data; wherein, the first modal data can be understood as video data; and the initial second modal data can be understood as a descriptive text of an action category corresponding to the video data.

[0054]

[0052] Specifically, when the first modal data is input into the temporal feature extraction coding layer, the temporal motion features of the first modal data are extracted using the temporal feature extraction coding layer, and the dynamic features of the second modal data are obtained by inputting the second modal data into the dynamic feature extraction coding layer.

[0055] In practical applications, to ensure the accuracy of temporal motion feature extraction corresponding to the first modal data, the initial second modal data is enhanced to obtain more detailed and motion-related second modal data corresponding to the initial second modal data. The second modal data is used as supervision to ensure the accuracy of the temporal motion features. A specific implementation is as follows: The second modal fusion feature extraction module further includes a data preprocessing layer; the data preprocessing layer is configured to invoke a language model to obtain second modal concept data and second modal distinction data corresponding to the initial second modal data, and to determine the second modal concept data and the second modal distinction data as the second modal data.

[0056]

[0054] The language model can be understood as a large-scale language model; the second modal concept data can be understood as data obtained by using the language model to describe the primary second modal data in detail, including dynamic descriptions such as the action and state changes of the primary second modal data; the second modal distinction data can be understood as descriptive data that shows the typical distinctions and differences between the target second modal data and other initial second modal data, wherein the target second modal data can be understood as any second modal data in the initial second modal data.

[0057]

[0055] Specifically, a plurality of initial second modal data are input into the target multimodal model, and by calling the language model, any one of the initial second modal data is described in detail to obtain second modal concept data, and by querying the language model, second modal distinction data for distinguishing any one of the initial second modal data from other initial second modal data is obtained, so as to show the distinction and difference between the initial second modal data and other initial second modal data through the second modal distinction data.

[0058]

[0056] The target multimodal model system provided by the embodiment of the present disclosure calls the language model through the data preprocessing layer to obtain the second modal concept data and the second modal distinction data related to the initial second modal data. When the second modal data not only covers the basic meaning of the initial second modal data but also highlights its unique characteristics compared with other initial second modal data, the second modal data is subsequently used as language supervision to ensure the accuracy of the output results of the video processing model.

[0059]

[0057] In one or more embodiments of the present disclosure, the first modal data is sampled at different frequencies in the input layer, thereby obtaining different sampled data. By inputting the different sampled data into different structures in the target multimodal model, different features are obtained for the first modal data. A specific implementation is as follows: The first modal fusion feature extraction module further includes a first modal data encoding layer and a first modal feature fusion layer; the first modal data encoding layer is configured to perform first modal feature extraction on the first sampled data to obtain the first modal feature, wherein the first sampled data is obtained by sampling the first modal data at a first sampling frequency; the time series feature extraction encoding layer is configured to perform time series feature extraction on the second sampled data to obtain time series motion features, wherein the second sampled data is obtained by sampling the first modal data at a second sampling frequency; and the first modal feature fusion layer is configured to perform feature fusion on the time series motion features and the first modal features to obtain the first modal fusion features.

[0060]

[0058] The sampling frequency can be understood as the number of times data is collected from the first modal data per second; the first sampling frequency can be understood as a low sampling frequency, such as a frequency of collecting data once per second, and the second sampling frequency can be understood as a high sampling frequency, such as a frequency of collecting data eight times per second.

[0059] Specifically, when the first modal data is video data, when the video data is sampled according to the first sampling frequency, the obtained first video frame sequence (first sampled data) can be understood as a low-frame-rate video frame sequence. This low-frame-rate video frame sequence focuses on preserving spatial details, such as high-resolution or high-quality static image information. When the video data is sampled according to the second sampling frequency, the obtained second video frame sequence (second sampled data) can be understood as a high-frame-rate video frame sequence. This high-frame-rate video frame sequence captures more motion information or temporal variation characteristics.

[0061]

[0060] Specifically, when the first modal data is video data, the sampled data obtained by different sampling frequencies are input into different layers of the target multimodal model; the first sampled data is input into the first modal data encoding layer, and the first modal feature extraction is performed on the first sampled data to obtain image features that can reflect the content features of each frame of the image; for the second sampled data, the temporal motion features of the video are extracted using the temporal feature extraction encoding layer, that is, the motion information, dynamic changes, and temporal dependencies between frames.

[0062]

[0061] A first modal feature fusion layer of a cross-attention structure is used to integrate temporal motion features and first modal features to obtain first modal fusion features, which contain static information and dynamically changing content of the first modal data.

[0063]

[0062] The target multimodal model system provided by the embodiment of the present disclosure samples the first modal data according to different sampling frequencies. While meeting the requirements for features at different levels, it can also balance the relationship between computing resource consumption and result accuracy to a certain extent. In addition, the first modal fusion features are obtained through the first modal feature fusion layer, which can provide a comprehensive and context-dependent representation, and help improve the performance of the target multimodal model on various tasks.

[0064]

[0063] The dynamic feature extraction coding layer is used to perform dynamic feature extraction on the second modality data to obtain dynamic features, so as to obtain second modality fusion features based on the dynamic features and static features, wherein the second modality data is obtained by performing enhancement processing on the initial second modality data, and the static features are obtained by performing static feature extraction on the initial second modality data. The first modality fusion features and the second modality fusion features are used to determine the target second modality data.

[0065]

[0064] Specifically, the dynamic feature extraction coding layer is used to perform feature extraction on the second modal data, and the dynamic feature extraction coding layer includes a feature extraction unit and an adaptation unit; the feature extraction unit is used to perform feature extraction on the second modal data to obtain descriptive features; the adaptation unit is used to perform dynamic feature extraction on the descriptive features to obtain the dynamic features.

[0066]

[0065] The feature extraction unit is used to extract features from the second modal data, and the adaptation unit is a multi-head self-attention structure with updateable parameters. As training progresses, the adaptation unit gradually learns the text representation related to motion in the descriptive features. The descriptive features can be understood as features that contain basic information and meaning of the second modal data, and the dynamic features can be understood as features that contain dynamic information of the second modal data.

[0067]

[0066] Specifically, the second modal data is input into the feature extraction unit. After passing through the feature extraction unit, a preliminary and static feature expression of the second modal data is obtained, which covers the basic information of the second modal data; the description feature is input into the adaptation unit to obtain dynamic features related to motion.

[0068]

[0067] It should be noted that the adaptation unit can dynamically adjust and optimize the features based on different application scenarios or task requirements; it can obtain more accurate and efficient features by learning and highlighting those feature dimensions that are more important for the current specific task.

[0069]

[0068] The video processing model training method provided by the embodiment of the present disclosure obtains description features by utilizing a feature extraction unit, and then dynamically extracts, selects or enhances the description features through an adaptation unit to generate dynamic features related to a specific task.

[0070]

[0069] In one or more embodiments of the present disclosure, the second modality fusion feature extraction module further includes a second modality encoding layer and a second modality feature fusion layer; the second modality encoding layer is used to perform static feature extraction on the initial second modality data to obtain the static features; the second modality feature fusion layer is used to perform feature fusion on the dynamic features and the static features to obtain the second modality fusion features.

[0071]

[0070] Specifically, when the dynamic feature extraction coding layer is used to obtain dynamic features, it is also necessary to use the second modality coding layer to perform feature extraction on the initial second modality data to capture the context dependency of the initial second modality data.

[0072]

[0071] Specifically, the second modal feature fusion layer of the cross-attention structure is used to fuse dynamic features and static features. The purpose of feature fusion is to integrate features extracted from different levels and dimensions to construct more comprehensive and expressive features.

[0073]

[0072] In the case where the second modal data is text data, feature extraction is performed on the initial text (initial second modal data), and static semantic features (static features) related to the initial text are obtained; based on the target description text

[0074] Feature extraction is performed on the (second modality data) to obtain dynamic text features (dynamic features) that contain rich action and dynamic context information. The static text features extracted from the initial text are fused with the dynamic text features corresponding to the target description text, integrating the two different levels of information to form fused text features (second modality fused features) that combine static semantics and dynamic behavior information.

[0075]

[0073] According to the first modal fusion feature and the second modal fusion feature, the matching degree between the first modal fusion feature and each second modal fusion feature is determined; according to the matching degree, target second modal data is determined from the second modal data, where the target second modal data is second modal data related to the content of the first modal data; and the target second modal data is output as an output result.

[0076]

[0074] The target multimodal model system provided by the embodiment of the present disclosure realizes the feature fusion of dynamic features and static features through the second modality fusion feature, so that the target multimodal model can fully utilize the advantages of various features and improve the performance on downstream tasks.

[0077]

[0075] In practical applications, the target multimodal model can be combined with the hardware nodes of the distributed system, that is, the first modality fusion feature extraction module and the second modality fusion feature extraction module in the target multimodal model can be deployed on different hardware nodes respectively, thereby distributing large-scale and complex multimodal data processing tasks to multiple hardware nodes for parallel processing, significantly improving the model training and inference speed, and greatly shortening the response time while ensuring the accuracy of the target multimodal model. When the target multimodal model needs to process real-time streaming data or respond to high-concurrency requests, the hardware nodes of the distributed system can provide powerful real-time computing capabilities and high-concurrency processing performance.

[0078]

[0076] The hardware nodes of a distributed system can be understood as computing devices or special devices that are dispersed in a network and communicate with each other to jointly complete specific tasks. These nodes can be independent computers, servers, embedded systems, Internet of Things (IoT) devices, or hardware accelerators with special processing capabilities such as GPU acceleration cards.

[0079]

[0077] For example, in an embodiment of the present disclosure, the first modality fusion feature extraction module of the target multimodal model is deployed on an FPGA (Field-Programmable Gate Array) processor to utilize its efficient real-time signal processing capabilities to process calculations on high-dimensional data such as videos or images; the second modality fusion feature extraction module of the target multimodal model is deployed on another independent hardware acceleration device or server cluster node, such as a GPU.

[0080] Through this distributed architecture design, the different modules of the target multimodal model can fully utilize the advantageous resources of their respective nodes and exchange data and work together with the help of high-speed networks (such as InfiniBand and 100Gbps Ethernet).

[0081]

[0078] In this way, not only is concurrent processing of multimodal input data (first modal data, second modal data) achieved, but also the limitations of a single device in terms of memory, storage, and computing power are overcome. That is, when the modules in the target multimodal model are deployed on different hardware nodes, each hardware node is only responsible for the data and calculation of its corresponding part, effectively avoiding the memory bottleneck problem. Moreover, through distributed deployment, hardware nodes that are good at specific tasks can be used in a targeted manner, thereby efficiently using and deploying computing resources, improving the processing efficiency of each hardware node, and enabling large-scale multimodal training tasks to maintain high accuracy while improving the training and inference speed of the target multimodal model.

[0082]

[0079] Referring to FIG3 , FIG3 shows a flow chart of a method for constructing a target multimodal model provided by an embodiment of the present disclosure; the method specifically includes the following steps.

[0083]

[0080] Step 302: Receive a task request, wherein the task request includes task parameters.

[0084]

[0081] Among them, the task request can be understood as a task request for creating a model sent from a client or other server; the task parameters can be understood as detailed parameter information required to build a model, such as model type, data source, objective function and other parameters.

[0085]

[0082] Specifically, in machine learning, deep learning or other data processing tasks, when a task request for creating a model is received from a client or other server, the corresponding model building operation is performed according to the task parameters carried in the task request, that is, the various specific parameter configurations required for building the model.

[0086]

[0083] Step 304: Based on the task parameters, obtain a first modality fusion feature extraction module and a second modality fusion feature extraction module, wherein the first modality fusion feature extraction module includes a temporal feature extraction coding layer, and the second modality fusion feature extraction module includes a dynamic feature extraction coding layer.

[0087]

[0084] Specifically, according to the task parameters, two feature extraction modules for different modal data are constructed and configured, namely, a first modal fusion feature extraction module and a second modal fusion feature extraction module. In this case, when the first modal fusion feature extraction module includes a temporal feature extraction coding layer, the temporal feature extraction coding layer is used to perform feature extraction on the first modal data having temporal features to obtain temporal motion features corresponding to the first modal data; and when the second modal fusion feature extraction module includes a dynamic feature extraction coding layer, the dynamic feature extraction coding layer is used to perform feature extraction on the second modal data to obtain dynamic features corresponding to the second modal data.

[0088]

[0085] Step 306: Connect the first modal fusion feature extraction module and the second modal fusion feature extraction module to obtain a multimodal model to be trained.

[0089]

[0086] When the first modality fusion feature extraction module and the second modality fusion feature extraction module are connected, a complete multimodal model architecture can be formed, that is, a multimodal model to be trained is obtained.

[0090]

[0087] Step 308: Train the multimodal model to be trained to obtain a target multimodal model that has completed training.

[0091]

[0088] Specifically, the multimodal model to be trained can be jointly trained in an end-to-end manner to optimize a common objective function, such as a loss function such as cross entropy loss and mean square error loss, so as to train the multimodal model to be trained so that it can understand and utilize information from two different modalities to complete a specific task.

[0092]

[0089] The target multimodal model construction method provided by the embodiment of the present disclosure can dynamically construct a feature extraction module according to the task parameters in the task request, and when the first modality fusion feature extraction module and the second modality fusion feature extraction module are obtained, the multimodal model to be trained is trained so that the multimodal model to be trained can learn more comprehensive and richer data representations, thereby improving the overall performance.

[0093]

[0090] Referring to FIG4, FIG4 shows a flow chart of a video processing model training method provided by an embodiment of the present disclosure, which specifically includes the following steps.

[0094]

[0091] Specifically, the video processing model includes a first modality fusion feature extraction module and a second modality fusion feature extraction module, the first modality fusion feature extraction module includes a temporal feature extraction coding layer, and the second modality fusion feature extraction module includes a dynamic feature extraction coding layer.

[0095]

[0092] Step 402: Input the video sample and the initial text sample into the video processing model, wherein the initial text sample is a text that describes the category of the video content of the video sample.

[0096]

[0093] The initial text sample can be understood as a brief description of an action category corresponding to the video sample; for example, when the video content of the video sample is "a person puts a cup on the table", the corresponding initial text sample can be "the cup is put on the table"; when the video content of the video sample is "a car collides with another car", the corresponding initial text sample can be "a traffic accident in which cars collide with each other".

[0097]

[0094] The video processing model can be understood as a multimodal large model for analyzing and understanding videos, which is used to select text that matches the video content of the video from the multiple texts based on the ability to analyze and understand the video when a video and multiple texts are input.

[0098]

[0095] Specifically, the video sample and its initial text sample are input into the video processing model together. The video processing model learns the semantic correlation between the video sample and its initial text sample by understanding the visual content in the video and extracting the semantic information of the initial text sample.

[0099]

[0096] Step 404: Use the video processing model to extract features from the video sample to obtain temporal motion features and fused image features.

[0100]

[0097] Wherein, when a video is composed of video frames, the temporal motion feature is used to represent the relationship and change between video frames and to understand the continuity and dynamic characteristics of the video content. In the field of video processing and analysis, temporal motion features may include: Frame rate: The number of frames played per second in a video determines the video smoothness and motion perception. A high frame rate can provide smoother motion performance. Time interval: The time interval between adjacent video frames is constant, which reflects the natural time passage of continuous action in real scenes. Motion information: The difference between adjacent video frames, namely motion vectors or motion compensation data, describes the temporal changes in image content, which is an important basis for video encoding and compression. Inter-frame dependency: In video coding technology, frames are generally divided into I-frames (independent frames), P-frames (forward predicted frames), and B-frames (bidirectionally predicted frames). They have temporal dependencies and are used to reduce redundant information. Time series analysis: In video analysis, by analyzing the changes in a series of video frames, temporal features such as object speed, direction, and trajectory can be extracted. This is very important for video content understanding, behavior recognition, and event detection. In other words, the temporal motion features between multiple video frames are not only related to the physical properties of video playback, but also to the temporal properties of the video. Like frame rate, it also contains a lot of information about changes in video content.

[0101]

[0098] The fused image features can be understood as features that contain the overall video information of the video sample.

[0102]

[0099] Video is composed of a series of continuous static images. Therefore, when the dynamic motion features are extracted from the video sample, the image features of the video sample need to be extracted as well, so as to obtain a fused image feature based on the dynamic motion features and the image features. The fused image feature contains the overall information of the video sample.

[0103]

[0100] In one or more embodiments of the present disclosure, different video frame sequences are obtained by using different sampling frequencies, thereby obtaining features containing different video information when a video processing model is used to extract features from video frames in the video frame sequence. A specific implementation method is as follows: the first modality fusion feature extraction module further includes an image coding layer and an image feature fusion layer; using the video processing model to extract features from the video samples to obtain temporal motion features and fused image features includes: sampling the video samples according to a first sampling frequency to obtain a first video frame sequence, and sampling the video samples according to a second sampling frequency to obtain a second video frame sequence; using the image coding layer to extract image features from the video frames in the first video frame sequence to obtain image features; using the temporal feature extraction coding layer to extract temporal motion features from the video frames in the second video frame sequence to obtain the temporal motion features; and using the image feature fusion layer to fuse the image features and the temporal motion features to obtain the fused image features.

[0104]

[0101] The sampling frequency can be understood as the number of times video frames are collected from the video sample per second; the first sampling frequency can be understood as a low sampling frequency, such as a frequency of collecting 1 frame per second, and the second sampling frequency can be understood as a high sampling frequency, such as a frequency of collecting 8 frames per second.

[0105]

[0102] Accordingly, when the video samples are sampled according to the first sampling frequency, the obtained first video frame sequence can be understood as a low-frame-rate video frame sequence, which focuses on retaining spatial details, such as high-resolution or high-quality static image information; when the video samples are sampled according to the second sampling frequency, the obtained second video frame sequence can be understood as a high-frame-rate video frame sequence, which captures more motion information or time-varying characteristics.

[0106]

[0103] Specifically, a spatio-temporal disentangled video coding framework is adopted, and different video frame sequences are input into different layers of a video processing model. A first video frame sequence is input into the video processing model, and the image coding layer of the video processing model is used to extract features of the video frames in the first video frame sequence to obtain image features that can reflect the content features of each frame image. For the second video frame sequence, the second video frame sequence is input into a motion encoder of the video processing model, and the motion encoder is used to extract temporal motion features, i.e., motion information, dynamic changes, and temporal dependencies between frames.

[0107]

[0104] Specifically, a first video frame sequence obtained according to a first sampling frequency is input into an image coding layer. After being processed by the image coding layer, each video frame in the first video frame sequence is converted into a high-dimensional vector. This vector contains the main feature information of the frame image, such as object shape, texture, color distribution, etc.; the image features corresponding to all video frames in the first video frame sequence are further integrated into global image features of the video sample. These global image features can be used for subsequent tasks, such as video classification, target detection, action recognition, etc., thereby improving the ability of the video processing model to understand and analyze video content.

[0108]

[0105] Specifically, in order to integrate the image features in the spatial domain and the temporal motion features in the temporal domain, an image feature fusion layer with a cross-attention structure is adopted. This layer is designed to effectively combine the two different types of features. Through the operation of the image feature fusion layer, a fused image feature is finally generated. The fused image feature contains the static visual information and dynamic changes of the video sample. The image features obtained at two different sampling frequencies are fused with the temporal motion features to comprehensively consider the information in the spatial and temporal dimensions, thereby obtaining a fused image feature that contains rich spatial features and takes into account the temporal motion characteristics.

[0109]

[0106] The video processing model training method provided by the embodiment of the present disclosure obtains different image features and temporal motion features by adopting different sampling frequencies to obtain video frame sequences, and fuses the image features with the temporal motion features, so that the final fused image features can take into account information of both spatial dimension and time dimension at the same time, thereby obtaining a more accurate and comprehensive judgment basis when the video processing model handles complex video analysis tasks, and can flexibly adjust the two sampling frequencies according to specific task requirements, which can not only meet the needs of different levels of features, but also balance the relationship between computing resource consumption and result accuracy to a certain extent.

[0110]

[0107] In one or more embodiments of the present disclosure, the image coding layer has a multi-layer structure. Image features corresponding to the first video frame sequence are obtained by processing the image coding layers layer by layer. The specific implementation is as follows.

[0111]

[0108] Using the image coding layer, performing image feature extraction on the video frames in the first video frame sequence to obtain image features, including: using the multi-layer image coding layer to respectively perform image feature extraction on the video frames in the first video frame sequence to obtain intermediate layer image features output by each layer of the image coding layer; and determining image features based on the intermediate layer image features output by each layer of the image coding layer.

[0112]

[0109] Among them, the image coding layer usually includes multiple convolutional layers, pooling layers and other structures, which can gradually extract more and more abstract and high-level feature representations from the original pixel-level information.

[0113]

[0110] Specifically, the first image coding layer of the image coding layer is determined as the target image coding layer, and the target image coding layer is used to perform image feature extraction on the video frames in the first video frame sequence to obtain intermediate layer image features; when the target image coding layer has a next image coding layer, the next image coding layer is determined as the target image coding layer, and the target image coding layer is used to perform image feature extraction on the intermediate layer image features to obtain the next intermediate layer image features. When the target image coding layer does not have a next image coding layer, the image feature extraction is terminated to obtain the image features.

[0114]

[0111] For example, an image encoding layer comprising a multi-layer convolutional neural network (CNN) is used to process each video frame from a first video frame sequence layer by layer. Each convolutional layer extracts feature information at different levels and abstractions from the input image. For example, a low-level layer may capture basic information such as edges and colors, while a high-level layer may learn more complex semantic features of parts of objects or the entire scene.

[0115]

[0112] Depending on the actual application requirements and model design, all or part of the intermediate layer image features can be selected for combination and fusion to generate the final image feature representation; for example, deep features can be selected for high-level semantic understanding tasks, while shallow features can be combined to maintain local detail information, or features of each layer can be integrated through pooling, attention mechanisms, etc.

[0116]

[0113] The video processing model training method provided in the embodiment of the present disclosure utilizes a multi-layer image coding layer to fully mine the multi-level and multi-scale information of the video frame, and based on this constructs a comprehensive and rich image feature expression, which can improve the accuracy of subsequent video analysis tasks.

[0117]

[0114] In one or more embodiments of the present disclosure, the video processing model includes a temporal feature extraction and encoding layer, which is used to obtain temporal motion features corresponding to the second video frame sequence. The specific implementation is as follows.

[0118]

[0115] The method of using the temporal feature extraction coding layer to extract temporal motion features from the video frames in the second video frame sequence to obtain the temporal motion features includes: using the temporal feature extraction coding layer to extract temporal features from the video frames in the second video frame sequence to obtain local temporal motion features, and obtaining global temporal motion features based on the intermediate layer image features; and using the temporal feature extraction coding layer to perform feature fusion on the local temporal motion features and the global temporal motion features to obtain the temporal motion features.

[0119]

[0116] The temporal feature extraction coding layer can be understood as the motion coding layer in the above embodiment, which is a model structure with time dimension processing capabilities, used to extract dynamic temporal information of a video frame sequence. Local temporal motion features can be understood as temporal motion features of video frames and local objects in the video frames; global temporal motion features can be understood as reflecting the entire action flow, scene transitions, or other temporally and spatially related global dynamic characteristics of the video.

[0120]

[0117] Specifically, the temporal feature extraction coding layer analyzes the continuous video frames in the second video frame sequence, captures the subtle changes and motion information between each frame and adjacent frames, and thus extracts the temporal motion features of the video frames and local objects in the video frames.

[0121]

[0118] Combined with the intermediate layer image features obtained from the multi-layer image coding layer, the intermediate layer image features contain a high-level abstract representation of the entire scene, which helps to grasp the unchanging or slowly changing global content in the video frame sequence, such as the video background, scene layout, etc., so that the global temporal motion features can be further derived.

[0122]

[0119] In practical applications, high frame rate sampling can provide richer motion information, thereby more accurately capturing rapid or subtle changes in movement. By extracting local temporal features of the second video frame sequence, the dynamic behavior of the object in the temporal dimension can be better understood. In addition, a high frame rate provides better temporal resolution, but may increase computational complexity and storage requirements. Combining the intermediate layer features of low frame rate images can reduce redundancy and improve processing efficiency while maintaining a certain degree of temporal continuity.

[0123]

[0120] Specifically, when the temporal feature extraction coding layer is used for feature fusion, the intermediate layer image features obtained by each layer of the image coding layer are input into the temporal feature extraction coding layer. When the temporal feature extraction coding layer is also a multi-layer structure, the global temporal motion features obtained based on the intermediate layer image features are fused with the local temporal motion features obtained by each layer of the temporal feature extraction coding layer, and the fusion result is input into the next layer of the current temporal feature extraction coding layer. When there is no next layer of the temporal feature extraction coding layer, the temporal motion features are obtained.

[0124]

[0121] The video processing model training method provided by the embodiment of the present disclosure enables the video processing model to learn more comprehensive video content by combining local temporal features and global temporal features. It not only focuses on action details but also comprehensively considers changes in the overall environment. This multi-scale fusion helps to improve the performance and robustness of the video processing model in various visual tasks (such as video classification, object detection, and event detection).

[0125]

[0122] Step 406: Use the video processing model to extract features from the initial text sample to obtain dynamic text features and fused text features.

[0126]

[0123] Among them, dynamic text features can be understood as text features corresponding to the rich action descriptions generated by the initial text sample; fused text features can be understood as features containing the overall text information of the initial text sample.

[0127]

[0124] Specifically, the video processing model is used to extract features from the initial text sample to obtain dynamic text features and static text features, and then the fused text features are obtained by fusing the dynamic text features and the static text features.

[0128] In one or more embodiments of the present disclosure, for an initial text sample, its corresponding target description text is determined, thereby utilizing the data preprocessing layer of a video processing model to obtain static text features corresponding to the initial text sample and dynamic text features corresponding to the target description text. Specific implementation methods are described below.

[0129]

[0126] The second modality fusion feature extraction module also includes a text encoding layer and a text feature fusion layer; the use of the video processing model to extract features from the initial text sample to obtain dynamic text features and fused text features includes: determining a target description text corresponding to the initial text sample, wherein the target description text is obtained by performing action enhancement description on the initial text sample; using the text encoding layer to extract static text features from the initial text sample to obtain static text features; using the dynamic feature extraction encoding layer to extract dynamic text features from the target description text to obtain dynamic text features; using the text feature fusion layer to fuse the static text features and the dynamic text features to obtain the fused text features.

[0130]

[0127] Among them, the target description text can be understood as a text that expands or details the initial text and emphasizes dynamic elements such as actions and state changes; the static text feature can be understood as a feature that contains the semantic information of the initial text sample.

[0131]

[0128] Specifically, first, the target description text of the action enhancement description is determined based on the initial text sample, and the text encoding layer of the video processing model is used to extract features for the original initial text sample to obtain static semantic features related to the initial text sample; the target description text is input into the dynamic text feature extraction encoding layer, and the deep semantic features and context information in the target description text are captured through the dynamic text feature extraction encoding layer to obtain dynamic text features containing rich action and dynamic situational information; the static text features extracted from the initial text sample are fused with the dynamic text features corresponding to the target description text to integrate the two different levels of information to form a fused text feature that combines static semantics and dynamic behavior information.

[0132]

[0129] In practical applications, the fixed-dimensional feature representation obtained for the initial text sample through the basic text encoding layer reflects the basic semantic information of the initial text sample; the dynamic text features obtained by the adaptation unit can usually capture context-sensitive information related to a specific task or key features after weight adjustment. In the embodiment of the present disclosure, the dynamic text features obtained by the adaptation unit are dynamic information related to motion in the target description text.

[0133]

[0130] Through the text feature fusion layer of the cross attention structure, the static text features and the dynamic text features are effectively combined, the complementary information in the static text features and the dynamic text features are comprehensively considered, and the final fused text features are generated.

[0134]

[0131] The video processing model training method provided by the embodiment of the present disclosure integrates two different levels of information to form a fused text feature that combines static semantics and dynamic behavior information, which can extract richer semantic information of the initial text sample, and realize feature fusion through the text feature fusion layer, so that the video processing model can fully utilize the advantages of various features and improve the performance on downstream tasks.

[0135]

[0132] In one or more embodiments of the present disclosure, the target description text is composed of a concept description text and a difference description text, both of which are obtained by calling a language model. The specific implementation is as follows.

[0136]

[0133] The second modal fusion feature extraction module also includes a data preprocessing layer; the determination of the target description text corresponding to the initial text sample includes: determining a preset number of similar text samples corresponding to the initial text sample; calling a language model through the data preprocessing layer to obtain a concept description text corresponding to the initial text sample; inputting the initial text sample and the similar text sample into the language model to obtain a difference description text corresponding to the initial text sample, wherein the difference description text is a text that describes the difference between the initial text sample and the similar text sample; and determining the concept description text and the difference description text as the target description text.

[0137]

[0134] Among them, the language model can be understood as a large-scale language model; the concept description text can be understood as a text that provides a detailed and rich explanation of the movement details in the initial text sample; the difference description text can be understood as a description text that shows the typical differences and differences between the initial text sample and similar text samples.

[0138]

[0135] First, video samples and initial text samples are obtained from a data set. Text features are extracted from the initial text samples in the data set, and the cosine similarities between the initial text samples are calculated. For the initial text samples currently to be used for training the video processing model, the cosine similarities are arranged from high to low, and a preset number (e.g., 5) of similar text samples corresponding to the initial text samples are selected.

[0139]

[0136] By calling a pre-trained language model, a concept description text is generated for the initial text sample; and the initial text sample and similar text samples are input into the same language model to identify and extract the key differences between them, and output the difference description text.

[0140]

[0137] The obtained concept description text is combined with the difference description text to form a target description text that integrates the core content and unique characteristics of the initial text sample.

[0141]

[0138] For example, when the initial text sample is “put something in front of something” and the determined similar text sample is “throw something in front of something”; the concept description text is “put an object in front of another object or in a position in front of it”, and the difference description text is “a conscious and orderly placement behavior”.

[0142]

[0139] The video processing model training method provided by the embodiment of the present disclosure obtains a target description text of a large amount of video dynamic information by obtaining concept description text and difference description text. The target description text not only covers the basic meaning of the original text, but also highlights its unique features compared with other similar texts. Subsequently, the target description text is used as language supervision to ensure the training effect of the video processing model.

[0143] In one or more embodiments of the present disclosure, the dynamic text feature extraction coding layer includes a feature extraction unit and an adaptation unit. The feature extraction unit obtains description text features corresponding to the target description text, and the adaptation unit extracts dynamic text features from the description text features. The specific implementation method is as follows.

[0144]

[0141] The dynamic text feature extraction coding layer includes a feature extraction unit and an adaptation unit; the use of the dynamic text feature extraction coding layer to perform dynamic text feature extraction on the target description text to obtain the dynamic text feature includes: using the feature extraction unit to perform feature extraction on the target description text to obtain description text features; using the adaptation unit to perform dynamic feature extraction on the description text features to obtain the dynamic text features.

[0145]

[0142] Among them, the feature extraction unit is used to extract features from the text, and the adaptation unit is a multi-head self-attention structure with updateable parameters. As the training proceeds, the text representation related to motion in the description text features is gradually learned; the description text features can be understood as the text features obtained by the feature extraction unit of the target description text, which contain the basic information and meaning of the target description text.

[0146]

[0143] Specifically, the target description text is input into the feature extraction unit. After passing through the feature extraction unit, a preliminary and static feature expression of the target description text is obtained, which covers the basic semantic information of the target description text; the description text features are input into the adaptation unit to obtain dynamic text features related to the movement.

[0147]

[0144] It should be noted that the adaptation unit can dynamically adjust and optimize features based on different application scenarios or task requirements. It can obtain more accurate and efficient text features by learning and highlighting those feature dimensions that are more important for the current specific task.

[0148]

[0145] The video processing model training method provided by the embodiment of the present disclosure applies a feature extraction unit to the target description text to obtain description text features, and then dynamically selects, strengthens or fuses the description text features through an adaptation unit to generate dynamic text features related to a specific task.

[0149]

[0146] Step 408: Train the video processing model based on the temporal motion features and the dynamic text features, as well as the fused image features and the fused text features.

[0150]

[0147] In one or more embodiments of the present disclosure, the video processing model is trained based on the temporal motion features and the dynamic text features, and the fused image features and the fused text features, including: determining a first loss function based on the temporal motion features and the dynamic text features, and determining a second loss function based on the fused image features and the fused text features; and training the video processing model based on the first loss function and the second loss function.

[0151]

[148] Specifically, when the temporal motion feature is an action, trajectory or high-order feature representing temporal continuity extracted from a video frame sequence, and the dynamic text feature contains text description information related to motion in the target description text, by determining the loss function of the temporal motion feature-dynamic text feature, it is used to ensure that there is a high degree of matching between the temporal motion feature of the video and the corresponding text description, thereby constraining the motion encoder (i.e., the temporal feature extraction encoding layer) to extract the motion detail information in the video, and enhancing the learning ability of the video action details; and by determining the loss function of the fused image feature-fused text feature, it is possible to promote the mutual correspondence and understanding between the visual information of the video sample and the text information of the initial text sample, thereby constraining the training process of the entire video processing model.

[0152]

[149] The video processing model training method provided by the embodiment of the present disclosure, during the training process, the video processing model calculates two loss functions to achieve effective matching and understanding of video content and related text descriptions, thereby improving the overall training effect of the video processing model.

[0153]

[150] The video processing method provided by the embodiment of the present disclosure generates a rich action description for each initial text sample by querying the language model. In addition, in order to make full use of the generated target description text, a motion encoder with learnable parameters is additionally constructed, high frame rate video frames are input to the motion encoder, and the target description text is used as language supervision. Through the temporal motion feature-dynamic text feature contrast loss, the learning process of the temporal motion feature is constrained, thereby enhancing the video processing model's ability to understand the semantics of motion details.

[0154]

[151] See FIG5 , which shows a flow chart of a processing process of a video processing model training method provided by an embodiment of the present disclosure, which specifically includes the following steps.

[0155]

[152] Step 502: Input the training data into the video processing model.

[0156]

[153] The training data includes videos and video category names, and the video category names can be understood as the initial text samples in the above embodiment.

[0157]

[154] Specifically, in visual stimuli, when a video is input into a video processing model, the input video is first sampled, and different video frames are obtained according to different sampling frequencies; for example, the video is sampled at a frequency of 1 frame per second to obtain a low frame rate video frame, and the video is sampled at a frequency of 8 frames per second to obtain a high frame rate video frame.

[0158]

[155] In text 1, first, the training data is the data in the data set, so the video category names contain multiple ones. By extracting the text features of each video category name, the cosine similarity between the video category names is calculated, and the cosine similarity is arranged in order from high to low to obtain a preset number (such as 5) of similar video category names that are easily confused with the target video category name. The target video category name is the video category name currently used to train the video processing model; by calling the language model, the motion details in the target video category name are elaborated in detail to obtain the concept description text corresponding to the target video category name, and by inputting the target video category name and the similar video category name into the language model, the language model is required to output the difference description text, which can be understood as a text description that shows the typical differences and differences between the target video category name and the similar video category name. In the case that there are 5 similar video category names, the difference description text also includes 5 accordingly. The 1 concept description text corresponding to the target video category name and the 5 difference description texts are determined as the target description text, and the target description text all contains the motion information in the video category name.

[0159]

[156] Step 504: Use the video processing model to extract features from the video sample.

[0160]

[157] See FIG6 , which shows an overall framework diagram of a video processing model provided by an embodiment of the present disclosure.

[0161]

[158] According to Figure 6, the framework of the video processing model is a symmetrical extension of the CLIP-based image encoder and text encoder (i.e., the static text encoder in the figure), with the addition of a motion encoder, a video feature fusion layer, an adapter, and a text feature fusion layer.

[0162]

[159] In actual applications, during the training of the video processing model, the parameters of the image encoder and text encoder of the original CLIP model are not updated, in order to maintain the CLIP model's ability to recognize and generalize image content; the parameters of the newly added text encoder and visual encoder modules can be updated, in order to give the video processing model the ability to extract visual target motion information in the video based on the CLIP model.

[0163]

[0160] Specifically, high-frequency video frames are input into a motion encoder (i.e., the motion coding layer in the above embodiment), local temporal features in the video are extracted, and combined with the image intermediate layer features of low-frame-rate video frames to obtain global temporal features that can perceive inter-frame changes in global image content. The temporal motion features of the video are obtained by integrating the local temporal features and the global temporal features; low-frequency video frames are input into an image encoder (i.e., the image coding layer in the above embodiment) to obtain image features of the video.

[0164]

[0161] Step 506: Use the video processing model to extract features from the initial text sample.

[0165] Specifically, the target description text (concept description text and difference description text) is input into a dynamic text encoder (i.e., a feature extraction unit in the dynamic text feature extraction coding layer in the above embodiment) to extract the description text features corresponding to the target description text. When the dynamic text encoder is a text encoder of the CLIP model, considering that the graphic and text pre-training data of the CLIP model mainly include descriptions of static information, in order to enable the CLIP text encoding result to better understand the target description text, an adapter composed of a multi-head self-attention network is introduced (i.e., an adaptation unit in the dynamic text feature extraction coding layer in the above embodiment). The parameters of the adapter are updateable. As training proceeds, the adapter gradually learns and understands motion-related concepts and semantics, and extracts corresponding text features. Therefore, when the description text features corresponding to the target description text are input into the adapter, the adapter is used to understand the motion information in the target description text and extract the dynamic text features from the description text features.

[0166]

[0163] On the other hand, the target video category name is input into a static text encoder (ie, the text encoding layer in the above embodiment) to obtain static text features corresponding to the target video category name.

[0167]

[0164] Step 508: Perform feature fusion on the extracted features.

[0168]

[0165] In the visual layer 1, the image features and temporal motion features are fused through a video feature fusion layer composed of a cross-attention network to obtain the video features of the entire video, that is, the fused image features in the above embodiment.

[0169]

[0166] In the text layer 1, the static text features and the dynamic text features are fused through a text feature fusion layer composed of a cross attention network to obtain the overall text features of the video category name, that is, the fused text features in the above embodiment.

[0170]

[0167] Step 510: Train the video processing model through the loss function.

[0171]

[0168] Specifically, a first loss function is determined based on temporal motion features and dynamic text features. The first loss function is used to constrain the alignment of temporal motion features of visual examples and dynamic text features of text examples. By using dynamic text features as language supervision, the motion encoder is constrained to extract motion detail information in the video, thereby enhancing the ability to learn video action details.

[0172]

[0169] The second loss function is determined based on the fused image features and the fused text features. Considering that identifying the action category of a video requires understanding both static features and dynamic features, the fused image features are aligned with the fused text features to constrain the training process of the entire framework of the video processing model.

[0173]

[0170] When the video processing model is trained, it can be applied to urban scenes where there is a significant demand for dynamic video information, such as urban traffic vehicle accidents and abnormal behavior analysis of key urban individuals. By improving the accuracy of semantic judgment, the industry's bottleneck in vehicle accident analysis accuracy can be broken through.

[0174]

[0171] The video processing model training method provided by the embodiment of the present disclosure uses a video processing model to perform feature extraction on an input video sample to obtain temporal motion features, thereby obtaining dynamic information of the video sample. To ensure the accuracy of the dynamic information extracted by the video processing model, the video processing model is used to perform feature extraction on an input initial text sample to obtain dynamic text features, and a first loss function is determined by using temporal motion features and dynamic text features to achieve the use of language supervision to enhance the video processing model's learning ability for video action details; and a second loss function is determined by fusing image features and fused text features to ensure that the video information contained in the video sample is aligned with the text information contained in the initial text sample, so that when the video processing model is applied to select text from the initial text according to the target video, the target text that matches the video content of the target video can be accurately selected from multiple initial texts that describe the category of the video content. By improving the video processing model's ability to analyze the dynamic information of the video, the video analysis ability of the video content is improved, thereby achieving more accurate selection of descriptive text related to the video content.

[0175]

[0172] Referring to FIG. 7 , FIG. 7 shows a flow chart of a video processing method provided by an embodiment of the present disclosure, which specifically includes the following steps.

[0176]

[0173] Step 702: Input the target video and multiple initial texts into the video processing model, wherein the multiple initial texts are texts that categorize the video content of the target video.

[0177]

[0174] Step 704: Using the video processing model, determine the target text corresponding to the target video from the multiple initial texts, wherein the video processing model is trained by the above-mentioned video processing model training method.

[0178]

[0175] In one or more embodiments of the present disclosure, after determining the target text corresponding to the target video, it also includes: feeding back the target text to a front-end user; receiving feedback information sent by the front-end user, wherein the feedback information is information provided by the front-end user on the target text; constructing sample task data based on the feedback information; and training the video processing model based on the sample task data.

[0179]

[0176] In practical applications, the video processing model can be iteratively optimized based on user feedback. First, the video processing model outputs the obtained target text to the terminal corresponding to the front-end user to feed back the target text to the front-end user; after reading or experiencing the target text, the front-end user will send feedback information based on his or her understanding and feelings. The feedback information includes but is not limited to text quality assessment, accuracy correction, emotional tendency judgment or other customized evaluation indicators.

[0180]

[0177] After collecting user feedback information, sample task data is constructed based on this feedback information; for example, if the user points out an error in the text and provides a corrected version, then a training sample containing the original text and the corrected text can be formed; if the feedback is a score of overall satisfaction, it can be used as a supervisory signal to adjust the parameters of the video processing model.

[0181]

[0178] The constructed sample task data is used to update and train the video processing model. Through continuous iterative optimization, the video processing model can better output target documents that meet user needs and expectations to improve user experience.

[0182]

[0179] Taking an actual scenario as an example, the video processing model can be a video action recognition model. Accordingly, using the video processing model to determine the target text corresponding to the target video from the multiple initial texts includes: using the video action recognition model to determine the target text corresponding to the target video from the multiple initial texts, wherein the target text is a text that describes the action of the target video.

[0183]

[0180] When the video processing model is a video action recognition model, the determined target text is the text related to motion in the target video identified by the video action recognition model.

[0184]

[0181] It should be noted that, in the application phase of the video processing model, the fused features are used to identify and predict the target text corresponding to the target video. Specific implementations can be found in the above embodiments, which will not be described in detail here.

[0185]

[0182] The video processing method provided by the embodiment of the present disclosure migrates a large multimodal model of images and texts to a video task, and uses a lightweight dynamic information perception module (motion encoder, video feature fusion layer, adapter, and text feature fusion layer) to give the large multimodal model the ability to understand dynamic information in the video, thereby meeting the real-time analysis requirements in practical applications.

[0186]

[0183] Referring to FIG. 8 , FIG. 8 shows a flowchart of a video processing method applied to a traffic scene provided by an embodiment of the present disclosure, which specifically includes the following steps.

[0187]

[0184] Step 802: Input the target traffic video and multiple initial event texts into the video processing model, wherein the multiple initial event texts are texts that categorize the video content of the target traffic video; Step 804: Utilize the video processing model to determine the target event text corresponding to the target traffic video from the multiple initial event texts, wherein the video processing model is trained using the above-mentioned video processing model training method.

[0188]

[0185] For specific implementations, please refer to the above embodiments and will not be described in detail here.

[0186] The video processing method provided in the embodiments of the present disclosure utilizes a video processing model to provide an efficient video analysis and processing solution. Combined with the powerful openness and generalization effect of the large multimodal graphic and text model (CLIP model) for different scenarios, it enables analysis and understanding of video content in complex and changing urban environments.

[0189]

[0187] Corresponding to the above method embodiment, the present disclosure also provides an embodiment of a video processing model training device. FIG9 shows a structural diagram of a video processing model training device provided by an embodiment of the present disclosure.

[0190]

[0188] As shown in Figure 9, the device includes: an input module 902, configured to input a video sample and an initial text sample into a video processing model, wherein the initial text sample is a text that categorizes the video content of the video sample; a video extraction module 904, configured to use the video processing model to perform feature extraction on the video sample to obtain temporal motion features and fused image features; a text extraction module 906, configured to use the video processing model to perform feature extraction on the initial text sample to obtain dynamic text features and fused text features; a training module 908, configured to train the video processing model based on the temporal motion features and the dynamic text features, as well as the fused image features and the fused text features.

[0191]

[0189] Optionally, the video extraction module 904 is further configured to: sample the video samples according to a first sampling frequency to obtain a first video frame sequence, and sample the video samples according to a second sampling frequency to obtain a second video frame sequence; utilize the image coding layer to perform feature extraction on the video frames in the first video frame sequence to obtain image features; utilize the temporal feature extraction coding layer to perform feature extraction on the video frames in the second video frame sequence to obtain the temporal motion features; utilize the image feature fusion layer to perform feature fusion on the image features and the temporal motion features to obtain the fused image features.

[0192]

[0190] Optionally, the video extraction module 904 is further configured to: use the multi-layer image coding layer to respectively extract image features of the video frames in the first video frame sequence to obtain the intermediate layer image features output by each layer of the image coding layer; and determine the image features based on the intermediate layer image features output by each layer of the image coding layer.

[0193]

[0191] Optionally, the video extraction module 904 is further configured to: utilize the temporal feature extraction coding layer to perform temporal feature extraction on the video frames in the second video frame sequence to obtain local temporal motion features, and obtain global temporal motion features based on the intermediate layer image features; utilize the temporal feature extraction coding layer to perform feature fusion on the local temporal motion features and the global temporal motion features to obtain the temporal motion features.

[0194]

[0192] Optionally, the text extraction module 906 is further configured to: determine a target description text corresponding to the initial text sample, wherein the target description text is obtained by performing action enhancement description on the initial text sample; utilize the text encoding layer to perform static text feature extraction on the initial text sample to obtain static text features; utilize the dynamic feature extraction encoding layer to perform dynamic text feature extraction on the target description text to obtain the dynamic text features; utilize the text feature fusion layer to perform feature fusion on the static text features and the dynamic text features to obtain the fused text features.

[0195]

[0193] Optionally, the text extraction module 906 is further configured to: determine a preset number of similar text samples corresponding to the initial text sample; call the language model through the data preprocessing layer to obtain the concept description text corresponding to the initial text sample; input the initial text sample and the similar text sample into the language model to obtain the difference description text corresponding to the initial text sample, wherein the difference description text is a text that describes the difference between the initial text sample and the similar text sample; and determine the concept description text and the difference description text as the target description text.

[0196]

[0194] Optionally, the text extraction module 906 is further configured to: utilize the feature extraction unit to perform feature extraction on the target description text to obtain description text features; utilize the adaptation unit to perform dynamic feature extraction on the description text features to obtain the dynamic text features.

[0197]

[0195] Optionally, the training module 908 is further configured to: determine a first loss function based on the temporal motion features and the dynamic text features, and determine a second loss function based on the fused image features and the fused text features; and train the video processing model based on the first loss function and the second loss function.

[0196] The above is a schematic diagram of a video processing model training device according to this embodiment. It should be noted that the technical solution of the video processing model training device and the technical solution of the video processing model training method described above are based on the same concept. For details not described in detail in the technical solution of the video processing model training device, please refer to the description of the technical solution of the video processing model training method described above.

[0198]

[0197] Corresponding to the above method embodiment, the present disclosure also provides a video processing device embodiment. FIG10 shows a schematic structural diagram of a video processing device provided by an embodiment of the present disclosure.

[0199]

[0198] As shown in Figure 10, the device includes: an input module 1002, configured to input a target video and multiple initial texts into a video processing model, wherein the multiple initial texts are texts that categorize the video content of the target video; a determination module 1004, configured to use the video processing model to determine the target text corresponding to the target video from the multiple initial texts, wherein the video processing model is trained by the above-mentioned video processing model training method.

[0200]

[0199] The device further includes: an optimization module, configured to feed back the target text to a front-end user; receive feedback information sent by the front-end user, wherein the feedback information is information provided by the front-end user on the target text; construct sample task data based on the feedback information; and train the video processing model based on the sample task data.

[0201]

[0200] Optionally, the determination module 1004 is further configured to: use the video action recognition model to determine the target text corresponding to the target video from the multiple initial texts, wherein the target text is a text that describes the action of the target video.

[0202]

[0201] The above is a schematic diagram of a video processing device according to this embodiment. It should be noted that the technical solution of the video processing device and the technical solution of the video processing method described above are based on the same concept. For details not described in detail in the technical solution of the video processing device, please refer to the description of the technical solution of the video processing method described above.

[0203]

[0202] FIG11 shows a block diagram of a computing device 1100 according to one embodiment of the present disclosure. Components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0204]

[0203] The computing device 1100 also includes an access device 1140 that enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1140 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0205] In one embodiment of the present disclosure, the aforementioned components of the computing device 1100 and other components not shown in FIG. 11 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG. 11 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.

[0206] The computing device 1100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1100 may also be a mobile or stationary server.

[0207]

[0206] The processor 1120 is used to execute the following computer program / instruction, which implements the steps of the above-mentioned video processing model training method or video processing method when executed by the processor.

[0208]

[0207] The various embodiments in this disclosure are described in a progressive manner. Similar portions between the various embodiments may be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the computing device embodiment is generally similar to the video processing model training method embodiment, so the description is relatively simple. For relevant portions, refer to the partial description of the video processing model training method embodiment.

[0209]

[0208] An embodiment of the present disclosure also provides a computer-readable storage medium, which stores a computer program / instruction, and when the computer program / instruction is executed by a processor, it implements the above-mentioned video processing model training method or the steps of the video processing method.

[0210]

[0209] The various embodiments in this disclosure are described in a progressive manner. Similar portions between the various embodiments may be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the computer-readable storage medium embodiment is generally similar to the video processing model training method embodiment, so the description is relatively simple. For relevant portions, refer to the description of the video processing model training method embodiment.

[0211]

[0210] An embodiment of the present disclosure also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned video processing model training method or video processing method.

[0212] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the video processing model training method described above are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the video processing model training method described above.

[0213]

[0212] The above description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0214]

[0213] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0215]

[0214] It should be noted that, for the sake of simplicity, the aforementioned method embodiments are all described as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the embodiments of the present disclosure.

[0216]

[0215] In the above embodiments, the description of each embodiment is given with emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0217]

[0216] The preferred embodiments of the present disclosure disclosed above are intended only to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

Claims 1. A target multimodal model system, comprising a target multimodal model, the target multimodal model comprising a first modality fusion feature extraction module and a second modality fusion feature extraction module, the first modality fusion feature extraction module comprising a temporal feature extraction coding layer, the second modality fusion feature extraction module comprising a dynamic feature extraction coding layer, the temporal feature extraction coding layer being configured to perform temporal feature extraction on first modality data to obtain temporal motion features, and to obtain first modality fusion features based on the temporal motion features and first modality features, wherein: The first modal feature is obtained by performing first modal feature extraction on the first modal data; the dynamic feature extraction coding layer is used to perform dynamic feature extraction on the second modal data to obtain dynamic features, so as to obtain dynamic features based on the dynamic features and static features. A second modality fusion feature is obtained, wherein the second modality data is obtained by enhancing initial second modality data, the static feature is obtained by extracting static features from the initial second modality data, and the first modality fusion feature and the second modality fusion feature are used to determine target second modality data.

2. The target multimodal model system according to claim 1, wherein the first modal fusion feature extraction module further comprises a first modal data encoding layer and a first modal feature fusion layer; the first modal data encoding layer is configured to extract the first modal feature from the first sampled data to obtain the first modal feature, wherein: The first sampling data is obtained by sampling the first modal data according to a first sampling frequency; the time series feature extraction and encoding layer is used to extract time series features on the second sampling data to obtain time series motion features, wherein the second sampling data is obtained by sampling the first modal data according to a second sampling frequency; the first modal feature fusion layer is used to fuse the time series motion features with the first modal features to obtain first modal fusion features.

3. The target multimodal model system according to claim 1, wherein the dynamic feature extraction encoding layer includes a feature extraction unit and an adaptation unit; the feature extraction unit is used to perform feature extraction on the second modality data to obtain descriptive features; and the adaptation unit is used to perform dynamic feature extraction on the descriptive features to obtain the dynamic features.

4. The target multimodal model system according to claim 1, wherein the second modal fusion feature extraction module further includes a second modal encoding layer and a second modal feature fusion layer; the second modal encoding layer is used to extract static features from the initial second modal data to obtain the static features; and the second modal feature fusion layer is used to fuse the dynamic features and the static features to obtain the second modal fusion features.

5. The target multimodal model system according to claim 1, wherein the second modal fusion feature extraction module further includes a data preprocessing layer; the data preprocessing layer is used to call a language model to obtain second modal concept data and second modal distinction data corresponding to the initial second modal data, and determine the second modal data based on the second modal concept data and the second modal distinction data.

6. The target multimodal model system according to claim 1, wherein the target multimodal model is a visual processing model, the first modal data is visual data including video or image, and the second modal data is text data.

7. A method for constructing a target multimodal model, comprising: receiving a task request including task parameters; obtaining a first modality fusion feature extraction module and a second modality fusion feature extraction module based on the task parameters, wherein the first modality fusion feature extraction module includes a temporal feature extraction coding layer, and the second modality fusion feature extraction module includes a dynamic feature extraction coding layer; and connecting the first modality fusion feature extraction module and the second modality fusion feature extraction module to obtain a multimodal model to be trained; The multimodal model to be trained is trained to obtain a target multimodal model that has completed training.

8. A video processing model training method, comprising: Inputting a video sample and an initial text sample into a video processing model, wherein the initial text sample is a text that categorizes the video content of the video sample; using the video processing model to extract features from the video sample to obtain temporal motion features and fused image features; using the video processing model to extract features from the initial text sample to obtain dynamic text features and fused text features; and training the video processing model based on the temporal motion features, the dynamic text features, and the fused image features and the fused text features.

9. The video processing model training method according to claim 8, wherein the video processing model includes a first modality fusion feature extraction module and a second modality fusion feature extraction module, the first modality fusion feature extraction module includes a temporal feature extraction coding layer, and the second modality fusion feature extraction module includes a dynamic feature extraction coding layer.

10. The video processing model training method according to claim 9, wherein the first modality fusion feature extraction module further comprises an image encoding layer and an image feature fusion layer; wherein extracting features from the video sample using the video processing model to obtain temporal motion features and fused image features comprises: The video samples are sampled according to a first sampling frequency to obtain a first video frame sequence, and the video samples are sampled according to a second sampling frequency to obtain a second video frame sequence; using the image coding layer, image feature extraction is performed on the video frames in the first video frame sequence to obtain image features; using the temporal feature extraction coding layer, temporal motion feature extraction is performed on the video frames in the second video frame sequence to obtain the temporal motion features; using the image feature fusion layer, feature fusion is performed on the image features and the temporal motion features to obtain the fused image features.

11. The video processing model training method according to claim 10, wherein the image coding layer is a multi-layer structure; and using the image coding layer to extract image features from video frames in the first video frame sequence to obtain image features comprises: Using multiple image coding layers, image features are extracted from video frames in the first video frame sequence to obtain intermediate image features output by each image coding layer; and image features are determined based on the intermediate image features output by each image coding layer.

12. The video processing model training method according to claim 11, utilizing the temporal feature extraction coding layer, performing temporal motion feature extraction on the video frames in the second video frame sequence, and obtaining the temporal motion feature, comprising: utilizing the temporal feature extraction coding layer, performing temporal motion feature extraction on the video frames in the second video frame sequence, obtaining local temporal motion features, and obtaining global temporal motion features based on the intermediate layer image features; utilizing the temporal feature extraction coding layer, performing feature fusion on the local temporal motion features and the global temporal motion features, and obtaining the temporal motion features.

13. The video processing model training method according to claim 9, wherein the second modality fusion feature extraction module further comprises a text encoding layer and a text feature fusion layer; wherein the step of extracting features from the initial text sample using the video processing model to obtain dynamic text features and fused text features comprises: determining a target description text corresponding to the initial text sample, wherein the target description text is obtained by performing action enhancement description on the initial text sample; performing static text feature extraction on the initial text sample using the text encoding layer to obtain static text features; The dynamic feature extraction encoding layer is used to extract dynamic text features from the target description text to obtain the dynamic text features; and the text feature fusion layer is used to fuse the static text features and the dynamic text features to obtain the fused text features.

14. The video processing model training method according to claim 13, wherein the second modality fusion feature extraction module further comprises a data preprocessing layer; and determining the target description text corresponding to the initial text sample comprises: Determine a preset number of similar text samples corresponding to the initial text sample; call a language model through the data preprocessing layer to obtain concept description text corresponding to the initial text sample; input the initial text sample and the similar text sample into the language model to obtain difference description text corresponding to the initial text sample, wherein the difference description text is text that describes the difference between the initial text sample and the similar text sample; and determine the concept description text and the difference description text as the target description text.

15. The video processing model training method according to claim 13, wherein the dynamic text feature extraction coding layer comprises a feature extraction unit and an adaptation unit; and the step of extracting dynamic text features from the target description text using the dynamic text feature extraction coding layer to obtain the dynamic text features comprises: The feature extraction unit is used to perform feature extraction on the target description text to obtain description text features; and the adaptation unit is used to perform dynamic feature extraction on the description text features to obtain the dynamic text features.

16. The video processing model training method according to claim 8, wherein the step of training the video processing model based on the temporal motion feature and the dynamic text feature, and the fused image feature and the fused text feature comprises: Determining a first loss function based on the temporal motion feature and the dynamic text feature, and determining a second loss function based on the fused image feature and the fused text feature; The video processing model is trained according to the first loss function and the second loss function.

17. The video processing model training method according to claim 8, wherein the video processing model is a video action recognition model.

18. A video processing method, comprising: A target video and multiple initial texts are input into a video processing model, wherein the multiple initial texts are texts that categorize the video content of the target video; and a target text corresponding to the target video is determined from the multiple initial texts using the video processing model, wherein the video processing model is trained using the video processing model training method described in any one of claims 8 to 17.

19. The video processing method according to claim 18, after determining the target text corresponding to the target video, further comprising: Feedback the target text to the front-end user; Receive feedback information sent by the front-end user, wherein the feedback information is information provided by the front-end user regarding the target text; construct sample task data based on the feedback information; and train the video processing model based on the sample task data.

20. The video processing method according to claim 18, wherein the video processing model is a video action recognition model; and determining a target text corresponding to the target video from the plurality of initial texts using the video processing model comprises: Using the video action recognition model, determine the video corresponding to the target video from the multiple initial texts. Target text, wherein the target text is a text describing the action of the target video.

21. A video processing method, applied to traffic scenes, comprising: Inputting a target traffic video and a plurality of initial event texts into a video processing model, wherein the plurality of initial event texts are texts that categorize the video content of the target traffic video; and utilizing the video processing model to determine a target event text corresponding to the target traffic video from the plurality of initial event texts, wherein the video processing model is trained using the video processing model training method according to any one of claims 8 to 17.

22. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 7 to 21 are implemented.

23. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 7 to 21.

24. A computer program product comprising a computer program / instructions, which, when executed by a processor, implements the steps of the method according to any one of claims 7 to 21.