Video processing method and system, model training method, computing device
By mapping multimodal information to the same feature space and fusing it, the problem of insufficient flexibility and interactivity in video colorization processing is solved, resulting in better color effects and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2025-01-03
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies for video colorization processing offer limited flexibility and interactivity, resulting in poor color quality in the generated videos and a subpar user experience.
By acquiring the original video and multimodal information, including at least two modalities and cue information, mapping them to the same feature space, fusing the information, and then performing color processing, the target video is generated.
It improves the flexibility and accuracy of video colorization, resulting in better color effects in the generated videos and enhancing the user experience.
Smart Images

Figure CN122340326A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology and video data processing, specifically to a video processing method and system, a model training method, and a computing device. Background Technology
[0002] In the field of video processing, colorization or recoloring of videos allows for the adjustment, modification, or restoration of colors, making the processed video more in line with users' color preferences. It can be applied to film and television production, video restoration, advertising, and other scenarios. For example, in the post-production of movies and TV series, colorization helps producers design and adjust the color tone of the image, making the entire work more visually impactful and artistic. Colorizing old videos can improve image quality and clarity, enhancing the user's viewing experience. In advertisements and promotional videos, colorization can highlight product features or strengthen visual impact to attract a wider target audience. Colorization or recoloring allows for personalized adjustments to video colors, resulting in better visual effects and user experience.
[0003] Currently, colorization methods in related technologies are typically suitable for colorizing images or pictures. When applied to colorizing or recoloring videos, inconsistencies in timing can cause flickering and abrupt color changes in the resulting videos. Furthermore, since video colorization or recoloring methods are usually fully automated, the process lacks flexibility and interactivity. In summary, current video colorization technologies suffer from low flexibility and interactivity, and produce poor color effects, resulting in a subpar user experience.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a video processing method and system, a model training method, and a computing device to at least solve the technical problems in related technologies, such as low flexibility and interactivity in video colorization processing and poor color effects in the generated videos.
[0006] According to one aspect of the embodiments of this application, a video processing method is provided, the method comprising: acquiring an original video and multimodal information, wherein the multimodal information includes: at least two modal information and cue information, the cue information being used to indicate the color information of at least one sub-region in different video frames of the original video; mapping the at least two modal information to the same feature space to obtain fusion information; and performing color processing on the original video based on the fusion information and the cue information to obtain a target video.
[0007] According to another aspect of the embodiments of this application, a video processing method is also provided, the method comprising: responding to an input command applied to an operation interface, determining the original video corresponding to the input command; acquiring multimodal information, wherein the multimodal information includes: at least two modal information and prompt information, the prompt information being used to indicate the color information of at least one sub-region in different video frames of the original video; mapping the at least two modal information to the same feature space to obtain fusion information; performing color processing on the original video based on the fusion information and the prompt information to obtain a target video; and displaying the target video on the operation interface.
[0008] According to another aspect of the embodiments of this application, a video processing method is also provided. The method includes: obtaining an original video and multimodal information by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the original video and multimodal information, the multimodal information includes: at least two modal information and prompt information, the prompt information being used to prompt color information of at least one sub-region in different video frames of the original video; mapping the at least two modal information to the same feature space to obtain fusion information; performing color processing on the original video based on the fusion information and the prompt information to obtain a target video; and outputting the target video by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the target video.
[0009] According to another aspect of the embodiments of this application, a video processing system is also provided, the system comprising: a client for sending an original video; a server connected to the client for acquiring multimodal information, wherein the multimodal information includes: at least two modal information and cue information, the cue information being used to indicate the color information of at least one sub-region in different video frames of the original video; mapping the at least two modal information to the same feature space to obtain fusion information; performing color processing on the original video based on the fusion information and the cue information to obtain a target video; the client is also used to output the target video.
[0010] According to another aspect of the embodiments of this application, a model training method is also provided. The method includes: acquiring training data, wherein the training data includes: a training video and first modality data; mapping the first modality data and second modality data in the training video to the same feature space to obtain training fusion information; generating training prompt information based on the training video, wherein the training prompt information is used to prompt the color information of at least one sub-region in different video frames of the training video; and training a preset processing model based on the training fusion information, the training prompt information, and the training video to obtain a target processing model, wherein the target processing model includes: a multimodal processing module, a color extraction module, a video processing module, and a depth processing module.
[0011] According to another aspect of the embodiments of this application, a computing device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0012] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory storing an executable program; and a processor connected to the memory via a bus for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0013] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0014] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0015] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the methods in various embodiments of this application.
[0016] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.
[0017] In this embodiment, firstly, the original video and multimodal information can be acquired. The multimodal information includes at least two modalities and cue information, wherein the cue information is used to guide the color information of at least one sub-region in the original video. Next, the at least two modalities can be mapped to the same feature space to obtain fused information. Finally, based on the fused information and the cue information, the original video can be colorized to obtain the target video. It is noteworthy that this application, by mapping at least two modal information to the same feature space and performing colorization processing on the original video based on the obtained fusion information and prompt information, can fully utilize multiple modal information to assist in video colorization processing. By mapping at least two modal information to the same feature space to obtain fusion information, alignment of at least two modal information can be achieved, ensuring that the processed target video is consistent across different modalities. Combining multiple modal information improves the understanding and analysis capabilities of video content, providing a more comprehensive and accurate information foundation for colorization processing. This enables more accurate colorization based on user prompts, improving the accuracy and flexibility of colorization. Consequently, it solves the technical problems of low flexibility and interactivity in video colorization processing in related technologies, and poor color effects in the generated videos.
[0018] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 This is a schematic diagram illustrating an application scenario of a video processing method according to an embodiment of this application;
[0021] Figure 2 This is a flowchart of a video processing method according to an embodiment of this application;
[0022] Figure 3 This is a schematic diagram of an optional video processing procedure according to an embodiment of this application;
[0023] Figure 4 This is a schematic diagram of an optional multimodal processing model according to an embodiment of this application;
[0024] Figure 5 This is a schematic diagram of an optional color extraction model processing procedure according to an embodiment of this application;
[0025] Figure 6This is a schematic diagram of an optional splicing information generation process according to an embodiment of this application;
[0026] Figure 7 This is a flowchart of an optional video processing method according to an embodiment of this application;
[0027] Figure 8 This is a flowchart of another optional video processing method according to an embodiment of this application;
[0028] Figure 9 This is a schematic diagram of a video processing system according to an embodiment of this application;
[0029] Figure 10 This is a schematic diagram of a model training method according to an embodiment of this application;
[0030] Figure 11 This is a structural block diagram of a computing device according to an embodiment of this application;
[0031] Figure 12 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] The technical solution provided in this application is mainly implemented using large-scale model technology. Here, "large-scale model" refers to a deep learning model with a massive number of parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of parameters. Large-scale models can also be called foundation models. They are pre-trained using large-scale unlabeled corpora to produce pre-trained models with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0035] It's important to note that in practical applications, large models can be fine-tuned using a small number of samples after pre-training, allowing them to be applied to various tasks. For example, large models can be widely used in Natural Language Processing (NLP), computer vision, and speech processing. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. Therefore, the main application scenarios for large models include, but are not limited to, digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0036] Some of the terms or terms that appear in the description of the embodiments of this application herein shall be interpreted as follows:
[0037] Diffusion models are a type of generative model that can generate high-quality image and video data.
[0038] Video colorization converts a given grayscale video into a color video. This technology can be used in fields such as old video restoration and artistic creation. According to an embodiment of this application, a video processing method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0039] Considering the large number of model parameters in large models and the limited computing resources of mobile terminals, the video processing method provided in this application can be applied to, for example, Figure 1 The application scenarios shown are not limited to these. In, for example... Figure 1 In the application scenario shown, the large model is deployed on server 10. Server 10 can connect to one or more client devices 20 via a local area network (LAN), wide area network (WAN), internet connection, or other types of data network. These client devices 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users through a graphical user interface to access the large model, thereby implementing the method provided in this embodiment.
[0040] In this embodiment of the application, the system consisting of a client device and a server can perform the following steps: the client device can interact with the server; the server can obtain the original video and multimodal information; at least two types of modal information are mapped to the same feature space to obtain fused information; based on the fused information and prompt information, the original video is colorized to obtain the target video.
[0041] It should be noted that with the rapid development of high-performance computing units, the methods provided in this application embodiment can also be applied to model-in-machine systems in other application scenarios. In one optional embodiment, the model-in-machine system has multiple built-in models, and users can select one model to adjust as needed to obtain their own model. The high-performance computing unit built into the model-in-machine system can then directly call the adjusted model to execute the methods provided in this application embodiment. In another optional embodiment, the large model-in-machine system has a pre-trained model built-in, and the high-performance computing unit built into the model-in-machine system can then directly call that model to execute the methods provided in this application embodiment.
[0042] Furthermore, when users need to train their own models, they can upload their own datasets via the client. These datasets are then sent to the server, allowing the server to adjust the pre-trained model using the dataset to obtain the user's customized model, which can then be deployed to the production environment. To facilitate users' model adjustment needs, the server provides complete adjustment tools, development frameworks, and processes, supporting multiple adjustment strategies. This allows the adjusted model to better adapt to different application domains and achieve a high degree of customization.
[0043] Under the aforementioned operating environment, this application provides the following: Figure 2 The video processing method shown. Figure 2 This is a flowchart of a video processing method according to an embodiment of this application. Figure 2 As shown, the specific steps may include the following:
[0044] Step S202: Obtain the original video and multimodal information.
[0045] The multimodal information includes at least two modal information and cue information, wherein the cue information is used to indicate the color information of at least one sub-region in different video frames of the original video.
[0046] The aforementioned original video can refer to a video that needs to be colorized or recolored. The original video can be a color video that does not meet the user's color requirements or has missing colors in some areas, or it can be a grayscale video, etc. The original video can be determined according to actual needs, and there is no limitation here.
[0047] The aforementioned at least two modal information may include at least two different modal information, including but not limited to image information, text information, audio information, video information, etc., which can be determined according to actual needs and are not limited here. Among them, image information may be one or more frames of images arbitrarily selected from the original image, or one or more frames of images provided by the user, or images selected that are similar or related to the content and color of the original image. The method of determining image information can be determined according to actual needs and is not limited here. Text information may be information describing the overall content and color of the original video, or descriptions or introductory information of the original video, or text content describing the overall content and color of the original video edited or entered by the user. The method of determining text information can be determined according to actual needs and is not limited here.
[0048] The aforementioned prompts may be prompts about the local colors of the original video obtained by superpixel segmentation and prompt synthesis of the original video, or they may be prompts describing the local colors of the original video that are edited or entered by the user. The method of determining the prompts can be determined according to actual needs and is not limited here.
[0049] In one optional embodiment, obtaining the original video and multimodal information during video processing can provide data support for subsequent video colorization or recoloring. By obtaining the original video and based on the information provided by the user, at least two modal information and prompt information can be obtained. Users can directly provide the system with the original video and specify various multimodal information, or users can provide the system with the original video and specify one or more multimodal information. The system can determine the remaining multimodal information based on the default settings or the original video. Alternatively, users can only provide the system with the original video, and the system can supplement the complete multimodal information based on the default settings or the original video. Preferably, when at least two types of modal information include image information and text information, the specific process of the system supplementing multimodal information based on the original video can be as follows: First, the system can receive the original video, including each frame image in the video. Through video decoding technology, the video file can be decoded into a series of image frames. These image frames can be used as the basic data units for video processing and can be used to extract image information. The image information can be one or more frames arbitrarily selected from the original images or selected by the user. The image information can be used to indicate the color information in the video. Through image processing technology, the color features, texture features, and other information of the image can be extracted, thereby providing a reference for video colorization. At the same time, the image information can also be used to train a deep learning model to achieve the function of automatic video colorization. Simultaneously, text information can be extracted. This text information can describe the overall content and color-related information of the original video, including overall descriptions, explanations, and introductions, helping to better understand the video's color information. Through natural language processing technology, key information can be extracted from the text to provide guidance for video colorization. Furthermore, cue information can be extracted. This cue information can be local color cues obtained by superpixel segmentation and cue synthesis of the original video. Superpixel segmentation is an image segmentation technique that aggregates adjacent pixels into superpixels, making image segmentation more accurate. The resulting cue information can help indicate the color information of local areas in the video, providing local guidance for video colorization. By acquiring the original video and multimodal information, the video content and color information can be fully understood, providing more accurate guidance for video colorization. A more comprehensive understanding of video color information can be achieved, thereby improving the accuracy and quality of video colorization. By analyzing the original video and multimodal information, deep learning models can also be trained to achieve automatic video colorization, saving manpower and material resources. This provides data support for video colorization or recoloring, improving the accuracy and efficiency of video colorization, enabling automatic video colorization, and enriching the functionality of video processing.
[0050] Step S204: Map at least two modal information to the same feature space to obtain fused information.
[0051] In one optional embodiment, in video processing, at least two modalities can be mapped to the same feature space to obtain fused information, thereby achieving the fusion of at least two modalities to better understand the video content and color information for subsequent colorization or recoloring processing. Preferably, when the at least two modalities include image information and text information, the image information and text information can first be represented as feature vectors. For image information, convolutional neural networks or other image processing techniques can be used to extract image features to obtain a high-dimensional image feature vector. For text information, natural language processing techniques can be used to convert text descriptions into word vectors or sentence vectors. This allows for the representation of image information and text information in their respective feature spaces. Next, the image information and text information can be mapped to the same feature space, which can be achieved through a multimodal fusion technique. Multimodal fusion can fuse information from different modalities to obtain a more comprehensive and integrated feature representation. Multimodal fusion methods can employ concatenation fusion, summation fusion, product fusion, attention fusion, etc. For example, in the concatenation fusion method, image feature vectors and text feature vectors can be concatenated to form a longer feature vector. In summation fusion methods, two feature vectors are added together to obtain a combined feature vector. In product fusion methods, two feature vectors are multiplied element-wise to obtain a combined feature vector. In attention fusion methods, attention mechanisms can be used to dynamically fuse two feature vectors. By selecting an appropriate multimodal fusion method, a fused representation of image and text information in the same feature space can be obtained, i.e., fused information. This fused information can contain rich information from both images and text, and can better describe video content and color information. In the above process, mapping at least two modalities of information to the same feature space can improve the understanding and description of video content, thereby enabling more accurate video processing. The obtained fused information can help better capture complex relationships and structures in the video, improving the effect and performance of video processing.
[0052] Step S206: Based on the fusion information and prompt information, the original video is colorized to obtain the target video.
[0053] In one optional embodiment, the original video can be colorized based on fusion information and prompt information. This can be done manually, frame by frame, by manually selecting the corresponding color to fill or smear based on the fusion and prompt information, or by preparing a color reference table beforehand, matching the corresponding color according to the prompt and fusion information, and then performing colorization frame by frame. Alternatively, image processing software can be used, employing color fill tools to colorize the video. Based on the prompt and fusion information, the corresponding color can be selected for filling, and the automatic fill function of the tool can be used to speed up the colorization process. The specific method for colorizing the original video can be chosen according to the actual situation, and can be continuously tried and improved in practice to achieve satisfactory results.
[0054] In another alternative embodiment, an artificial intelligence model can be used to colorize the original video based on fusion information and cue information. For example, generative adversarial networks (GANs) or conditional generative adversarial networks (CANs) can be used to achieve video colorization. First, each frame of the original video can be used as input, along with fusion information and cue information. The fusion information helps the model better understand the overall color information of the video, while the cue information guides the model to colorize or adjust specific areas of the video. This ensures that the resulting target video fully utilizes the original video's color information in both overall and local color aspects, and better meets the diverse color needs of users. After training, the model can generate a target video with good color reproduction. Through the above steps, the original video can be colorized to obtain the target video. This method combines at least two modalities of information and cue information, making full use of the advantages of multimodal information, which can effectively improve the accuracy and effect of video colorization. At the same time, the application of deep learning technology can realize the automated video colorization process, improving work efficiency. In the above process, reinforcement learning technology can be applied to the video colorization task, and the reward mechanism can guide the model to learn better colorization strategies, further improving the colorization effect. When fusing at least two modalities of information, different fusion strategies can be adopted, such as layer-by-layer fusion or parallel fusion, to obtain more comprehensive and accurate fusion information. Data augmentation technology can be used to expand the training dataset, improve the model's generalization ability and robustness, and make it perform better in a wider range of video colorization tasks.
[0055] In this embodiment, firstly, the original video and multimodal information can be acquired. The multimodal information includes at least two modalities and cue information, wherein the cue information is used to guide the color information of at least one sub-region in the original video. Next, the at least two modalities can be mapped to the same feature space to obtain fused information. Finally, based on the fused information and the cue information, the original video can be colorized to obtain the target video. It is noteworthy that this application, by mapping at least two modal information to the same feature space and performing colorization processing on the original video based on the obtained fusion information and prompt information, can fully utilize multiple modal information to assist in video colorization processing. By mapping at least two modal information to the same feature space to obtain fusion information, alignment of at least two modal information can be achieved, ensuring that the processed target video is consistent across different modalities. Combining multiple modal information improves the understanding and analysis capabilities of video content, providing a more comprehensive and accurate information foundation for colorization processing. This enables more accurate colorization based on user prompts, improving the accuracy and flexibility of colorization. Consequently, it solves the technical problems of low flexibility and interactivity in video colorization processing in related technologies, and poor color effects in the generated videos.
[0056] In the above embodiments of this application, at least two modal information includes: image information and text information; mapping at least two modal information to the same feature space to obtain fused information includes: extracting features from image information to obtain the original color features of image information; extracting features from text information to obtain the original text features of text information; mapping the original color features to the feature space to obtain target color features, and mapping the original text features to the feature space to obtain target text features; fusing the target color features and target text features to obtain fused information.
[0057] In one optional embodiment, feature extraction, mapping, and fusion of image and text information can be achieved through the following steps. First, the image information can be converted into a digitized pixel matrix. Then, by calculating the color value of each pixel in the pixel matrix, the original color features of the image information can be obtained. Next, the text information is preprocessed, including word segmentation and stop word removal. Then, the text information can be converted into a vector representation using a bag-of-words model or a term frequency-inverse document frequency method to obtain the original text features. Next, the original color features can be dimensionality-reduced, such as through principal component analysis or linear discriminant analysis, mapping them to a low-dimensional feature space to obtain the target color features. Then, the original text features can be dimensionality-reduced, mapping them to a low-dimensional feature space to obtain the target text features. Finally, the target color features and target text features can be merged, either through concatenation or by using weighted averaging, to obtain the final fused information. Through these steps, feature extraction, mapping, and fusion of image and text information can be achieved, resulting in fused information.
[0058] In another alternative embodiment, when using an artificial intelligence model, feature extraction can first be performed on the image information using a color extraction model. This model can be a deep learning-based model, such as a convolutional neural network or a recurrent neural network. The color extraction model can learn color features in the image, such as color distribution and color contrast. Through the color extraction model, the original color features of the image information can be obtained. Feature extraction can then be performed on the text information, which can describe the overall content of the original video and color-related information. Natural language processing techniques can be used for feature extraction of the text information, such as the bag-of-words model, term frequency-inverse document frequency algorithm, and word vector models, to convert the text information into vector form for subsequent processing and analysis. Next, a multimodal processing model can be used to map the original color features to a feature space. This multimodal processing model can be a neural network model, such as a multimodal fusion network or a multimodal attention network. The multimodal processing model can learn the correlations between different modalities, converting the original color features into target color features. Next, the original text features can be mapped to the feature space. A multimodal processing model can be used to map the original text features to the feature space. Finally, the target color features and target text features can be fused to obtain fused information. This process can use a fusion model, such as a multimodal fusion network or an attention mechanism. The fused information can better reflect the color and content features of the video. In practical applications, by mapping image and text information to the same feature space and fusing them, more accurate and comprehensive video processing results can be achieved. By combining color features from image information with semantic information from text information, video content can be better understood, resulting in more accurate coloring or recoloring. Multimodal information fusion can improve the robustness and generalization ability of video processing. Information from different modalities can complement each other, making the video processing system more stable and reliable when processing different types of videos. It can avoid overfitting and information loss problems, and improve the applicability of the system in practical applications. Multimodal information fusion can also improve user experience and interactivity. By fusing information from different modalities, users can be provided with richer and more diverse video processing functions to meet their personalized needs. This can improve user satisfaction, interactivity, and experience in video processing, and achieve more accurate, robust, and user-friendly video processing effects.
[0059] In the above embodiments of this application, mapping the original color features to the feature space to obtain target color features, and mapping the original text features to the feature space to obtain target text features, includes: using at least one image processing module in the multimodal processing model to perform attention processing on the preset color feature sequence and the original color features in the feature space to obtain target color features; and using at least one text processing module in the multimodal processing model to perform attention processing on the preset text feature sequence and the original text features in the feature space to obtain target text features.
[0060] The aforementioned preset color feature sequence refers to a pre-determined sequence of color features during video colorization or recoloring. It can be a set of color samples or color labels used to guide the video colorization process. These color feature sequences can be learned query information or specific color values determined based on color theory or experience. They can provide guidance and constraints for the video colorization process, helping the algorithm better understand the video content and perform accurate colorization. Using preset color feature sequences can provide color samples and hue information, helping the algorithm select appropriate colors for colorization. Preset color feature sequences can be customized and adjusted according to specific application scenarios and needs to achieve better visual effects.
[0061] The aforementioned preset text feature sequences refer to pre-determined text feature sequences in video colorization or recoloring processes. These sequences can be a set of text labels or descriptions used to guide the video colorization process. These text feature sequences can be query information learned through training or specific descriptive information determined based on text analysis or professional knowledge. They can provide guidance and constraints for the video colorization process, helping the algorithm better understand the video content and perform accurate colorization. Preset text feature sequences can provide descriptive information and semantic features, helping the algorithm understand the video content and perform more precise colorization. Furthermore, preset text feature sequences can be customized and adjusted according to specific application scenarios and needs to achieve better visual effects.
[0062] In one optional embodiment, firstly, color sequences and text sequences mapped to the same feature space can be obtained to obtain preset color feature sequences and preset text feature sequences. Then, a pre-established multimodal processing model can be invoked. This model can include at least one image processing module and at least one text processing module. The image processing module can use an attention mechanism to process the color features, and the text processing module can use natural language processing techniques to process the text features. Next, the image processing module can perform attention processing on the preset color feature sequences and the original color features to obtain target color features; similarly, the text processing module can perform attention processing on the preset text feature sequences and the original text features to obtain target text features. In this process, the preset color feature sequences and preset text feature sequences allow for more precise control over the video colorization process, achieving controllability of the colorization effect. The application of the multimodal processing model can provide users with more personalized and accurate video colorization services, improving user experience and satisfaction. It can achieve more accurate, efficient, and controllable video colorization effects, thereby improving the quality of video processing and the user experience.
[0063] In the above embodiments of this application, any image processing module includes: a first attention layer and a first feedforward neural network layer; performing attention processing on a preset color feature sequence and original color features using at least one image processing module to obtain a target color feature includes: performing attention processing on the input color features of any image processing module and the preset color feature sequence using the first attention layer to obtain a first attention feature, wherein the input color feature of the first image processing module is the original color feature, and the input color features of other image processing modules besides the first image processing module are the output color features of the previous image processing module; adding the first attention feature and the preset color feature sequence to obtain a first added feature; performing feedforward processing on the first added feature using the first feedforward neural network layer to obtain a first feedforward feature; adding the first added feature and the first feedforward feature to obtain the output color feature of any image processing module; and determining the output image feature of the last image processing module as the target color feature.
[0064] In one optional embodiment, the image processing module includes a first attention layer and a first feedforward neural network layer, and the text processing module includes a second attention layer and a second feedforward neural network layer. At least one image processing module can be used to perform attention processing on a preset color feature sequence and original color features to obtain the target color feature. Specifically, firstly, the first attention layer can be used to perform attention processing on the input color features of any image processing module and the preset color feature sequence to obtain a first attention feature. Here, the input color features of the first image processing module are the original color features, and the input color features of the remaining image processing modules are the output color features of the previous image processing module. Next, the first attention feature and the preset color feature sequence can be added together to obtain a first added feature. Then, the first feedforward neural network layer can be used to perform feedforward processing on the first added feature to obtain a first feedforward feature. Finally, the first added feature and the first feedforward feature can be added together to obtain the output color feature of any image processing module, and the output image feature of the last image processing module can be determined as the target color feature. In the above process, by introducing an attention mechanism and a feedforward neural network layer, the correlation between image and text information can be better captured, thereby improving the accuracy of video processing. By processing the preset color feature sequence and the original color features, personalized video colorization can be achieved according to user needs, meeting the needs of different users.
[0065] In the above embodiments of this application, any text processing module includes: a second attention layer and a second feedforward neural network layer; using at least one text processing module to perform attention processing on a preset text feature sequence and original text features to obtain target text features includes: using the second attention layer to perform attention processing on the input text features of any text processing module and the preset text feature sequence to obtain a second attention feature, wherein the input text features of the first text processing module are the original text features, and the input text features of other text processing modules besides the first text processing module are the output text features of the previous text processing module; adding the second attention feature and the preset text feature sequence to obtain a second added feature; using the second feedforward neural network layer to perform feedforward processing on the second added feature to obtain a second feedforward feature; adding the second added feature and the second feedforward feature to obtain the output text feature of any text processing module; and determining the output text feature of the last text processing module as the target text feature.
[0066] In one optional embodiment, a preset text feature sequence and original text features are first input into a first text processing module. In this module, a second attention layer performs attention processing on these two inputs to obtain second attention features. This attention processing can involve weighting the text information, allowing the model to focus more on important information. The second attention features can be features weighted according to the importance of the input text features. Next, the second attention features are added to the preset text feature sequence to obtain second summed features. Then, a second feedforward neural network layer performs feedforward processing on the second summed features to obtain second feedforward features. This feedforward processing helps the model better learn the relationships and semantic information between text information. Finally, the second summed features and the second feedforward features are added together to obtain the output text features of any text processing module. This output text feature serves as the input text feature for the next text processing module and determines the target output text feature of the last text processing module. In the above process, by performing attention processing and multimodal fusion on text information, the model can more accurately understand the video content and guide the colorization process, improving the accuracy and efficiency of video processing. Through the fusion of multimodal information and the processing of text features, the model can learn richer information, improve its generalization ability, and perform better when processing different types of videos. Accurate video colorization can enhance the user experience, making the video content more vivid and thus improving the user's viewing experience.
[0067] In the above embodiments of this application, feature extraction of image information to obtain the original color features of the image information includes: segmenting the image information using the segmentation module in the color extraction model to obtain multiple sets of sub-image information, wherein the shapes and / or quantities of different sets of sub-image information are different; extracting features from the multiple sets of sub-image information using the residual module in the color extraction model to obtain color features in the image space; and mapping the color features in the image space to the multimodal space using the image encoding module in the color extraction model to obtain the original color features.
[0068] In one optional embodiment, the color extraction model can extract features from image information for subsequent coloring or recoloring operations. Specifically, in the color extraction model, the segmentation module can segment the image information to obtain multiple sets of sub-image information. Image segmentation algorithms can include pixel-based segmentation methods, such as thresholding and edge detection, and region-based segmentation methods, such as region growing and watershed algorithms. These algorithms can segment the image into different regions or objects, providing a foundation for subsequent feature extraction and color mapping. In practical applications, deep learning methods can be combined for image segmentation, such as using convolutional neural networks for semantic segmentation, which can more accurately extract color information from the image and improve the accuracy and stability of color features. Next, the residual module can extract features from the multiple sets of sub-image information to obtain color features in the image space. In deep learning, the residual module can use residual networks to solve the gradient vanishing and gradient exploding problems during the training process of deep neural networks by introducing residual connections, effectively improving the training effect and performance of the network. Finally, the image encoding module can map the color features in the image space to the multimodal space to obtain the original color features. In practical applications, models such as autoencoders or generative adversarial networks can be used for image encoding to compress and transform the color information of the image, resulting in a more compact and effective feature representation. Through the above steps, the color information of the video can be extracted and mapped, providing a foundation and support for subsequent coloring or recoloring operations. This can improve the accuracy and stability of color features, increase the efficiency and quality of video processing, and make the final visual effect more natural and vivid.
[0069] In the above embodiments of this application, the original video is colorized based on fusion information and prompt information to obtain a target video, including: generating noise information corresponding to the original video; splicing the noise information, prompt information and the original video to obtain splicing information; and inputting the splicing information and fusion information into a video processing model to obtain a target video generated by the video processing model.
[0070] In one optional embodiment, noise information, such as Gaussian white noise, can first be generated. The number of frames and the size of each frame in the noise information can be the same as the original video. Next, the noise information, prompt information, and the original video can be stitched together. The prompt information can include color information for different sub-regions in the video frames, which helps to more accurately colorize the video. Stitching this information provides more reference and basis for subsequent video processing. Finally, the generated stitched and fused information can be input into a video processing model for processing. The video processing model can be a deep learning model, such as a convolutional neural network or a generative adversarial network, to colorize the video, making its colors richer and its details clearer. When generating noise information, some noise removal techniques, such as denoising algorithms or deblurring algorithms, can be used to reduce the impact of noise in the video on the colorization effect. Through the above steps, the colorization effect and processing speed of the video can be improved, making the generated target video clearer and more realistic, and multimodal information can be effectively utilized to improve the accuracy and reliability of video processing.
[0071] In the above embodiments of this application, after splicing noise information, prompt information and original video to obtain spliced information, the method further includes: predicting the depth information of the original video; using a depth processing model to extract features from the depth information to obtain the depth features of the original video; adding the spliced information and the depth features to obtain the target features; and inputting the target features and fusion information into the video processing model to obtain the target video generated by the video processing model.
[0072] In one alternative embodiment, the stitching and fusion information can first be input into a video processing model. The video processing model accepts input from the stitching and fusion information and learns how to combine this information to generate a target video. Next, the video processing model can be used to generate the target video, including predicting the depth information of the original video. Depth information can refer to the distance or depth relationships between different objects or scenes in the video, enabling the prediction of depth information in the original video based on the input stitching and fusion information. Then, a depth processing model can be used to extract features from the depth information to obtain the depth features of the original video. The depth processing model can be a neural network model for processing depth information, capable of learning to extract key features from the depth information. Next, the stitching information can be added to the depth features to obtain target features, combining information from the stitching and depth features to obtain a richer feature representation. Finally, the target features and fusion information can be input into the video processing model to obtain the target video generated by the video processing model. The video processing model can comprehensively consider information from the target features and fusion information to generate the final target video. In the process of colorizing or recoloring videos, the prediction and utilization of depth information can improve the video processing effect. Depth information helps the model better understand the spatial relationships in the video, thus enabling more accurate colorizing or recoloring. By mapping splicing and fusion information to the same feature space, image and text information can be better combined, allowing the model to comprehensively consider information from different sources and improve the accuracy and effect of video processing. By adding depth features to splicing information, the model can better utilize the key features extracted from depth information when generating the target video, thereby improving the quality of video processing.
[0073] In the above embodiments of this application, the process of inputting splicing information and fusion information into a video processing model to obtain a target video generated by the video processing model includes: inputting splicing information and fusion information into a video processing model to obtain a color video generated by the video processing model; and replacing the values of preset channels in the color video with the values of preset channels in the original video to obtain the target video.
[0074] The aforementioned preset channel can be a luminance (Luma) channel in a video, where the Luma channel represents brightness, or the degree of lightness or darkness. The luminance channel can be calculated by weighting the red, green, and blue channels. In video processing, the luminance channel determines the brightness information in an image or video, controlling the image's brightness and contrast. By adjusting the value of the luminance channel, the brightness of an image or video can be changed, making the image clearer or softer, while also enhancing the image's contrast, making it more impactful and attractive. The preset channel in this application can also be determined according to actual needs, and is not limited here.
[0075] In one optional embodiment, the stitching information and fusion information can be input together into the video processing model. The video processing model can learn the features and color information of the original video and then generate the target video. After training, the video processing model can generate a color video, i.e., the target video. This color video contains the color and feature information learned by the model. However, considering that the model-generated video may damage the structural information, further processing is required. The values of preset channels in the color video can be replaced with the values of preset channels in the original video. By providing the original video as input, the user has already provided the Luma channel values. Therefore, the Luma channel of the target video can be replaced with the Luma channel of the original video. This preserves the structural information of the original video while fusing the color information generated by the model, thereby improving the image quality. In video processing, in addition to replacing the Luma channel to preserve structural information, other techniques can be applied to improve image quality. For example, adaptive enhancement algorithms can be used to enhance the contrast and sharpness of the video; the visual effect of the video can be improved by adjusting the brightness and color of pixels; denoising techniques can be applied to reduce noise in the video, making the image clearer and smoother; and super-resolution techniques can be used to increase the resolution of the video, thereby obtaining a more detailed image. In the above process, replacing the Luma channel can preserve the structural information of the original video, avoid damage to the structure caused by the model-generated video, improve the picture quality, effectively improve the video processing effect, preserve the structural information of the original video, improve the picture quality and clarity, enhance the visual experience, achieve better results in video processing, and improve the user's viewing experience.
[0076] The technical solution proposed in this application will be described below with reference to an optional embodiment. This application proposes a multimodal grayscale video colorization method based on a diffusion model. The application scenarios of this application can include monochrome video conversion or video for which the user is not satisfied with the color. This application can also be applied to the recoloring of color videos. The color video can be converted to grayscale first by applying a linear transformation, and then the grayscale video can be recolorized.
[0077] The proposed technical solution, given a color video, firstly randomly selects a frame from the original video and feeds it into a color mapper, then converts the original video to grayscale. A superpixel segmentation algorithm is then used to synthesize cue information. This cue information, along with noise and the grayscale video, is concatenated and fed into a unified network (UNet). Simultaneously, depth information can be utilized to enhance spatial-temporal consistency. For text information, the text is first encoded, and the encoded text embedding, along with the color mapper's output, is fed into a multimodal processing model (Dual QFormer) for feature extraction. Finally, the features are injected into the UNet.
[0078] Specifically, a random frame from the original video can be fed into a color mapper, which then divides it into three sub-images of different shapes and numbers. These three sub-images are then fed into a residual network and a Contrastive Language-Image Pretraining (CLIP) image encoder to extract color features. Regarding the model design, setting it to three groups is a compromise; dividing the input into three groups decouples color and structure. Therefore, four or five groups could also be used, but this might increase computational resources. Three groups are preferred. During inference, three modalities can guide the colorization of the grayscale video: text, images, and prompts. This module is for the image modality; when an image is input, it is also divided into three sub-images for subsequent processing.
[0079] To align color and text information, this application designs a multimodal processing model (DualQFormer). This module includes two sets of learnable sequences, a cross-attention layer, and a feedforward neural network. The learnable sequences can actually be a set of learnable neural network parameters. Fusing them into the same feature space involves weighted summation to efficiently extract semantic information from text and images. Specifically, the two sub-modules are used to extract image and text information, respectively. When the image representation is fed into the first sub-module, it serves as the "key (K)" and "value (V)" of the attention layer, while the learnable sequence of this sub-module serves as the "query (Q)" of the attention layer. After the attention mechanism is applied to these three sets of features (Q, K, V), they are fed into the feedforward neural network layer to obtain image features. Similarly, when the text representation is fed into the second sub-module, it serves as K and V, and the learnable sequence serves as Q. After the attention mechanism is applied, it is passed through the feedforward neural network layer. After extracting features from the text and image separately, these two features can be fused together and fed into the cross-attention layer of UNet.
[0080] This application leverages the depth map information of grayscale videos to enhance the temporal-spatial consistency of generated videos. Specifically, a depth guide can be employed, which can be a neural network composed of several convolutional layers. It accepts depth information as input and uses a depth estimation model to predict that depth information. To incorporate the depth information into the network, features of the depth information can be extracted to align it with other parts of the network. The depth guide can extract these features. A depth estimation model can be used to predict the depth information of the input grayscale video. The predicted depth information is then fed into the depth guide as input. After passing through this neural network, a tensor with depth features is obtained as output, and this output is added to the implicit code.
[0081] This application observes that color overflow can cause errors in optical flow estimation. Therefore, two frames of the generated video can be randomly selected, and the optical flow of these two frames can be calculated using an existing optical flow estimation model. Similarly, optical flow estimation is also performed on the corresponding two frames in the original video and regarded as the true value. Then, the mean square error can be calculated between the optical flow calculated from the generated video and the optical flow calculated from the original video, which is used as an additional loss to constrain the color overflow problem.
[0082] This application proposes a multimodal video colorization method, aiming to achieve high-quality, interactive video colorization by combining multiple modalities. Users provide 0-3 modal conditions: grayscale video, text, images, and prompts. Unprovided modal information is automatically set to the default value; that is, if the user provides 0 inputs, the model automatically colors the video. After providing user input, the model first randomly initializes Gaussian noise and processes the user input, continuously denoising the Gaussian noise to generate the colored video. Furthermore, previous video colorization methods suffer from temporal inconsistencies and color overflow. Therefore, a depth guide and optical flow loss are designed to enhance spatial-temporal consistency and alleviate color overflow issues.
[0083] This application proposes a multimodal processing model (Dual QFormer) for efficiently extracting color features from text and images and fusing them into the same feature space, thereby improving the diffusion model's ability to align color features. The design of directly stitching together the prompt information and grayscale video and feeding it as part of the input to the diffusion model not only allows the prompt information to be efficiently preserved, but also reduces computational resources. A depth guide and optical flow loss are introduced to improve temporal consistency, while Luma channel replacement is used to improve image quality.
[0084] Figure 3 This is a schematic diagram of an optional video processing procedure according to an embodiment of this application, such as... Figure 3 As shown, image information can be obtained by randomly selecting a frame from the original video. The image information is processed by a color extraction model to obtain original color features, and text information is processed by feature extraction to obtain original text features. The original color features and original text features are then processed by a multimodal processing model to obtain fused information. The original video is processed by a prompt generation module to obtain prompt information, and the obtained prompt information is processed by a stitching module to obtain stitched information. At the same time, depth features can be obtained based on a depth processing model. The stitched information and depth features are added together to obtain the target features. The fused information and target features are then input into the video processing model to obtain the target video.
[0085] The feature extraction described above can be achieved using a Contrastive Language-Image Pre-training (CLIP) text encoder, or other methods, which are not limited here. The video processing model described above can use contextual information and optical character recognition technology. By analyzing the contextual information in the text or image and combining it with optical character recognition technology, the text or image data can be understood and recognized more accurately, which can help improve the accuracy and efficiency of tasks such as text analysis and image recognition.
[0086] Figure 4 This is a schematic diagram of an optional multimodal processing model according to an embodiment of this application, such as... Figure 4As shown, the original color features and original text features can be processed by a multimodal processing model to obtain fused information. Specifically, the multimodal processing model includes a preset color feature sequence, a preset text feature sequence, at least one image processing module (×N), and at least one text processing module (×N). Any image processing module includes a first attention layer and a first feedforward neural network layer, and any text processing module includes a second attention layer and a second feedforward neural network layer. Based on the above process, attention processing can be performed on the input color features and preset color feature sequences of any image processing module using the first attention layer to obtain the first attention feature; the first attention feature and the preset color feature sequence are added together to obtain the first added feature; and the first added feature is fed forward using the first feedforward neural network layer to obtain the first... Feedforward features: The first additive feature and the first feedforward feature are added together to obtain the output color feature of any image processing module; the output image feature of the last image processing module is determined as the target color feature; the second attention layer is used to perform attention processing on the input text feature of any text processing module and a preset text feature sequence to obtain the second attention feature; the second attention feature and the preset text feature sequence are added together to obtain the second additive feature; the second feedforward neural network layer is used to perform feedforward processing on the second additive feature to obtain the second feedforward feature; the second additive feature and the second feedforward feature are added together to obtain the output text feature of any text processing module; the output text feature of the last text processing module is determined as the target text feature; the target color feature and the target text feature are fused to obtain fused information.
[0087] Figure 5 This is a schematic diagram of an optional color extraction model processing procedure according to an embodiment of this application, such as... Figure 5 As shown, image information can be processed by a color extraction model to obtain original color features. Specifically, the color extraction model includes a segmentation module, a residual module, and an image encoding module. Based on the above structure, the image information can be segmented using the segmentation module to obtain multiple sets of sub-image information; the residual module can be used to extract features from the multiple sets of sub-image information to obtain color features in the image space; and the image encoding module can be used to map the color features in the image space to the multimodal space to obtain the original color features.
[0088] Figure 6 This is a schematic diagram of an optional splicing information generation process according to an embodiment of this application, such as... Figure 6As shown, the original video is processed by the prompt generation module to obtain prompt information, and the obtained prompt information is processed by the stitching module to obtain stitched information. Specifically, the prompt generation module may include a hints synthesis module. Superpixel segmentation of the original video can obtain superpixel segmentation information. The hint synthesis module can generate prompt information based on the original video and superpixel segmentation information. Diffusion processing of the original video can obtain noise information. Based on the prompt information, a hints mask and canvas information can be generated. The stitching module can stitch together the noise information, the hints mask, the canvas information, and the input grayscale video to obtain stitched information.
[0089] According to an embodiment of this application, a video processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0090] Figure 7 This is a flowchart of a video processing method according to an embodiment of this application, such as... Figure 7 As shown, the specific steps may include the following:
[0091] Step S702: Respond to the input command applied to the operation interface and determine the original video corresponding to the input command.
[0092] The aforementioned user interface refers to the interface through which a user interacts with a terminal device, such as a computer, mobile phone, tablet, smartwatch, or smart bracelet; it is also known as the user interface. The user interface can take various forms, including but not limited to graphical user interfaces (GUIs) and command-line interfaces. A graphical user interface displays information graphically, allowing users to interact with the computer using input devices such as a mouse and keyboard. GUIs enable intuitive computer operation and are suitable for most users, especially those unfamiliar with command-line operations. A command-line interface displays information in a text-based manner, allowing users to interact with the computer by entering commands. The specific user interface can be determined according to user needs and is not limited here.
[0093] Step S704: Obtain multimodal information, wherein the multimodal information includes: at least two modal information and cue information, the cue information being used to indicate the color information of at least one sub-region in different video frames of the original video.
[0094] Step S706: Map at least two modal information to the same feature space to obtain fused information.
[0095] Step S708: Based on the fusion information and prompt information, the original video is colorized to obtain the target video.
[0096] Step S710: Display the target video on the operation interface.
[0097] In the above embodiments of this application, obtaining multimodal information includes: responding to an input instruction, determining at least one modal information corresponding to the input instruction; when the number of at least one modal information is less than a preset number, obtaining multimodal information based on at least one modal information and preset modal information; when the number of at least one modal information is greater than or equal to the preset number, determining at least one modal information as multimodal information.
[0098] In one optional embodiment, the system can respond to user input instructions. Based on the user input instructions, the system can determine at least one modal information corresponding to the instructions, which can be any one or more of image information, text information, or prompt information, or none of them may be specified. The system can generate the target video based on user-specified multimodal information, or user-specified multimodal information combined with the system's default multimodal information, or when the user does not specify any, the system can use only the system's default multimodal information. In the above process, complete user multimodal information is not required. The system can flexibly supplement and adjust multimodal information to generate the target video, improving the flexibility of target video generation and the user experience.
[0099] In the above embodiments of this application, color processing is performed on the original video based on fusion information and prompt information to obtain a target video, including: when the original video is a grayscale video, color processing is performed on the original video based on fusion information and prompt information to obtain a target video; when the original video is a color video, the original video is converted to a grayscale video, and color processing is performed on the grayscale video based on fusion information and prompt information to obtain a target video.
[0100] In one optional embodiment, if the original video is grayscale, color processing can be directly performed on the original video based on fusion information and prompt information; if the original video is color, it can first be converted to grayscale and then color processed. Each frame image in the original video is converted to grayscale, and then color processing is performed on the grayscale images according to the above steps to obtain the target video. Considering that grayscale videos only contain luminance information and no color information, this makes color processing easier. By converting to grayscale, the complexity of processing can be effectively reduced and the processing efficiency improved. With this setting, accurate color processing of the video can be achieved, making the target video more vivid and clear.
[0101] According to an embodiment of this application, a video processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0102] Figure 8 This is a flowchart of a video processing method according to an embodiment of this application, such as... Figure 8 As shown, the specific steps may include the following:
[0103] Step S802: Obtain the original video and multimodal information by calling the first interface. The first interface includes a first parameter, the parameter value of which includes the original video and multimodal information. The multimodal information includes at least two types of modal information and prompt information. The prompt information is used to indicate the color information of at least one sub-region in different video frames of the original video.
[0104] Step S804: Map at least two modal information to the same feature space to obtain fused information.
[0105] Step S806: Based on the fusion information and prompt information, the original video is colorized to obtain the target video.
[0106] Step S808: Output the target video by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target video.
[0107] According to embodiments of this application, a video processing system for implementing the above-described video processing method is also provided. Figure 9 This is a schematic diagram of a video processing system according to an embodiment of this application, such as... Figure 9 As shown, the system includes: client 902 and server 904.
[0108] The system comprises: a client for sending the original video; a server connected to the client for acquiring multimodal information, which includes at least two modalities and cue information, the cue information indicating the color information of at least one sub-region in different video frames of the original video; mapping the at least two modalities to the same feature space to obtain fusion information; colorizing the original video based on the fusion information and cue information to obtain the target video; and the client for outputting the target video.
[0109] According to an embodiment of this application, a model training method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0110] Figure 10 This is a flowchart of a model training method according to an embodiment of this application, such as... Figure 10 As shown, the specific steps may include the following:
[0111] Step S1002: Obtain training data, which includes training videos and first modality data.
[0112] Step S1004: Map the first modality data and the second modality data in the training video to the same feature space to obtain training fusion information.
[0113] Step S1006: Based on the training video, generate training prompt information. The training prompt information is used to indicate the color information of at least one sub-region in different video frames of the training video.
[0114] Step S1008: Based on the training fusion information, training prompt information and training video, the preset processing model is trained to obtain the target processing model, wherein the target processing model includes: a multimodal processing module, a color extraction module, a video processing module and a depth processing module.
[0115] The aforementioned multimodal processing module may be a multimodal processing model in the video processing method embodiment of this application, the aforementioned color extraction module may be a color extraction model in the video processing method embodiment of this application, the aforementioned video processing module may be a video processing model in the video processing method embodiment of this application, and the aforementioned depth processing module may be a depth processing model in the video processing method embodiment of this application.
[0116] In one optional embodiment, training data can be acquired, which may include training videos and training text. A first processing model and a preset extraction model can be used to map the training text and video frames from the training videos to the same feature space to obtain training fusion information. This can be achieved using a deep learning model, such as a convolutional neural network, to extract features from the video frames. Then, the text information is converted into vector representations using natural language processing techniques. Finally, the two feature sets are fused together to form the training fusion information. Next, based on the training videos, training prompts can be generated. These prompts may include the video's theme, emotion, scene, etc., helping the model better understand the video's content and context. This process can be achieved using a text generation model or annotation tools. Then, using training fusion information, training prompts, and training videos, the first processing model, the preset extraction model, the second processing model, and the third processing model can be trained. Specifically, a multimodal deep learning model can be used to combine different types of information, allowing the model to process multiple data types simultaneously, improving the model's generalization ability and accuracy. At the same time, the color extraction model, video processing model, and deep processing model can also be trained so that the model can better understand video content, extract key information, and perform colorization or recoloring. Through this training process, a multimodal processing model can be obtained. This model can process text and video information simultaneously, realizing the automation and intelligence of video processing. Meanwhile, the color extraction model, video processing model, and deep processing model can help to better understand video content and features, improving the efficiency and quality of video processing.
[0117] In the above embodiments of this application, generating training prompt information based on the training video includes: performing superpixel segmentation on multiple video frames of the training video to obtain multiple sets of first sub-regions, wherein different sets of first sub-regions correspond to different video frames, and different first sub-regions in the first sub-region sets correspond to different semantic information; determining the region color information of any first sub-region based on the color information of different pixels in any first sub-region; determining at least one set of second sub-regions from multiple video frames, wherein different sets of second sub-regions correspond to different semantic information, and different second sub-regions in the same set of second sub-regions correspond to different video frames; and obtaining training prompt information based on at least one set of second sub-regions and the region color information of any second sub-region.
[0118] In one optional embodiment, superpixel segmentation can be performed on multiple video frames of the training video. Superpixels refer to pixel-level clustering, grouping adjacent pixels into the same category to form larger image patches. Superpixel segmentation helps to better understand the structural and semantic information of the image. By performing superpixel segmentation on the video frames, multiple sets of first sub-regions can be obtained, each set corresponding to a video frame, and different first sub-regions correspond to different semantic information. Next, the regional color information of each first sub-region can be determined based on the color information of different pixels in the first sub-regions. This helps to better understand the color distribution and features of each sub-region, thereby enabling more accurate colorization. Then, at least one set of second sub-regions can be determined from the multiple video frames. The second sub-region sets correspond to different semantic information, and different second sub-regions in the same set correspond to different video frames. This step allows for a better understanding of the different semantic information in the video, thereby enabling better colorization or recoloring. Finally, training cues can be obtained based on at least one set of second sub-regions and the region color information of each second sub-region. These cues can include the color distribution and texture features of each second sub-region, helping the algorithm to better learn and understand the video content, thus enabling more accurate colorization or recoloring. Through the specific implementation of the above steps, the structural and semantic information of the video content can be better understood, leading to more accurate colorization or recoloring. By performing superpixel segmentation and extracting region color information from the video content, the structure and color distribution of the video content can be understood more accurately, resulting in more accurate colorization or recoloring. Generating training cues helps the algorithm better learn and understand the video content, thereby improving the algorithm's robustness and enabling it to perform well on different types of videos.
[0119] In the above embodiments of this application, a preset processing model is trained based on training fusion information, training prompt information, and training video to obtain a target processing model, including: converting the training video into a training grayscale video; extracting features from the depth information of the training grayscale video using a first processing module in the preset processing model to obtain training depth features; concatenating the training noise information, training prompt information, and training grayscale video corresponding to the training video to obtain training concatenation information; adding the training concatenation information and the training depth features to obtain target training features; inputting the target training features and training fusion information into a second processing module in the preset processing model to obtain a target video generated by the second processing module; and adjusting the model parameters of the preset processing model based on the training video and the target video to obtain the target processing model.
[0120] In one optional embodiment, different processing models can be trained using training fusion information, training cue information, and training videos to obtain more accurate results. Specifically, firstly, the training video can be converted into a training grayscale video, reducing computational load. Converting the video to grayscale simplifies the processing and makes it easier to extract depth information. Next, a third processing model can be used to extract features from the depth information of the training grayscale video, obtaining training depth features. Depth feature extraction allows for a better understanding of the video content, aiding subsequent processing. Next, the training noise information, training cue information, and training grayscale video corresponding to the training video can be concatenated to obtain training concatenation information. Noise information and cue information help to better understand the video content and characteristics; the concatenated information more comprehensively reflects the video's features. Next, the training concatenation information can be added to the training depth features to obtain the target training features. This step integrates various aspects of the video's features, resulting in a more comprehensive feature representation, which is beneficial for the accuracy and effectiveness of subsequent processing. The target training features and training fusion information are then input into a second processing model to obtain the target video generated by the second processing model. The second processing model can generate more accurate video results based on the input feature information. It can also be adjusted based on the training video and the target video to improve the performance and accuracy of the model, thus obtaining a multimodal processing model, a color extraction model, a video processing model, and a depth processing model.
[0121] In the above embodiments of this application, adjusting the model parameters of a preset processing model based on training videos and target videos to obtain a target processing model includes: constructing a first loss function value based on the difference between the training videos and the target videos; performing flow estimation on any two target video frames in the target video using a flow estimation model to obtain a first flow value in the target video, and performing flow estimation on two training video frames in the training video using the same model to obtain a second flow value in the training video, wherein the two training video frames correspond to any two target video frames; constructing a second loss function value based on the difference between the first flow value and the second flow value; summing the first loss function value and the second loss function value to obtain a total loss function value; and adjusting the model parameters of the preset processing model based on the total loss function value to obtain the target processing model.
[0122] In one optional embodiment, the differences between the training video and the target video can be analyzed. A first loss function value can be constructed by comparing pixel information, color distribution, motion trajectories, etc., between the training and target videos. By comparing the differences between the training and target videos, a value reflecting the similarity between them can be obtained. Next, a streamer estimation model can be used to perform streamer estimation on frames in both the target and training videos to obtain corresponding streamer values. By comparing the difference between the first streamer value in the target video and the second streamer value in the training video, a second loss function value can be constructed. Then, the first and second loss function values can be summed to obtain a total loss function value. The total loss function value reflects the overall difference between the training and target videos and can serve as a basis for adjusting the processing model. Finally, based on the total loss function value, the first processing model, the preset extraction model, the second processing model, and the third processing model can be adjusted to obtain a multimodal processing model, a color extraction model, a video processing model, and a depth processing model. By continuously adjusting the model parameters, the model can better adapt to the differences between the training and target videos, thereby achieving better video processing results. Through the above process, by constructing a multimodal processing model, a color extraction model, a video processing model, and a depth processing model, and adjusting the model parameters according to the loss function value, we can better adapt to the differences between the training video and the target video, thereby achieving more refined and personalized video processing effects.
[0123] Optionally, Figure 11 This is a structural block diagram of a computing device according to an embodiment of this application. Figure 11 As shown, the computing device A may include one or more (only one is shown in the figure) processors 102, memory 104 and peripheral interfaces 106, wherein the processors 102, memory 104 and peripheral interfaces 106 are interconnected via a bus 108.
[0124] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0125] The processor can access information and applications stored in the memory via a transmission device to execute the steps in each embodiment.
[0126] Embodiments of this application may provide an electronic device. Figure 12 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 12 As shown, the electronic device may include: an input / output device 1202; a memory 1204; and a processor 1206, wherein the processor 1206 is connected to the input / output device 1202 and the memory 1204 via a bus 1208.
[0127] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0128] The processor can invoke an executable program stored in memory via a transmission device to perform the following method: acquiring original video and multimodal information, wherein the multimodal information includes: at least two modal information and cue information, the cue information being used to indicate the color information of at least one sub-region in different video frames of the original video; mapping the at least two modal information to the same feature space to obtain fusion information; and performing colorization processing on the original video based on the fusion information and cue information to obtain the target video.
[0129] Optionally, at least two modalities are mapped to the same feature space to obtain fused information, including: using a color extraction model to extract features from image information to obtain the original color features of the image information; extracting features from text information to obtain the original text features of the text information; using a multimodal processing model to map the original color features to the feature space to obtain the target color features, and mapping the original text features to the feature space to obtain the target text features; and fusing the target color features and the target text features to obtain fused information.
[0130] Optionally, the multimodal processing model includes: a preset color feature sequence in the feature space, a preset text feature sequence in the feature space, at least one image processing module, and at least one text processing module, wherein the preset color feature sequence is used to represent color features in the feature space; the multimodal processing model maps the original color features to the feature space to obtain target color features, and maps the original text features to the feature space to obtain target text features, including: using at least one image processing module to perform attention processing on the preset color feature sequence and the original color features to obtain target color features; and using at least one text processing module to perform attention processing on the preset text feature sequence and the original text features to obtain target text features.
[0131] Optionally, any image processing module includes a first attention layer and a first feedforward neural network layer, and any text processing module includes a second attention layer and a second feedforward neural network layer; attention processing is performed on a preset color feature sequence and original color features using at least one image processing module to obtain target color features, including: performing attention processing on the input color features of any image processing module and the preset color feature sequence using the first attention layer to obtain a first attention feature, wherein the input color features of the first image processing module are the original color features, and the input color features of other image processing modules besides the first image processing module are the output color features of the preceding image processing module; adding the first attention feature and the preset color feature sequence to obtain a first summed feature; performing feedforward processing on the first summed feature using the first feedforward neural network layer to obtain a first feedforward feature; and adding the first summed feature and the first feedforward feature to obtain the output of any image processing module. Color features; determining the output image features of the last image processing module as the target color features; using at least one text processing module to perform attention processing on the preset text feature sequence and the original text features to obtain the target text features, including: using a second attention layer to perform attention processing on the input text features of any text processing module and the preset text feature sequence to obtain the second attention features, wherein the input text features of the first text processing module are the original text features, and the input text features of other text processing modules besides the first text processing module are the output text features of the previous text processing module; adding the second attention features and the preset text feature sequence to obtain the second added features; using a second feedforward neural network layer to perform feedforward processing on the second added features to obtain the second feedforward features; adding the second added features and the second feedforward features to obtain the output text features of any text processing module; determining the output text features of the last text processing module as the target text features.
[0132] Optionally, the color extraction model includes: a segmentation module, a residual module, and an image encoding module; the color extraction model is used to extract features from image information to obtain the original color features of the image information, including: segmenting the image information using the segmentation module to obtain multiple sets of sub-image information, wherein the shapes and / or numbers of different sets of sub-image information are different; extracting features from the multiple sets of sub-image information using the residual module to obtain color features in the image space; and mapping the color features in the image space to the multimodal space using the image encoding module to obtain the original color features.
[0133] Optionally, based on the fusion information and the prompt information, the original video is colorized to obtain the target video, including: generating noise information corresponding to the original video; splicing the noise information, prompt information and the original video to obtain spliced information; and inputting the spliced information and the fusion information into the video processing model to obtain the target video generated by the video processing model.
[0134] Optionally, the splicing information and fusion information are input into the video processing model to obtain the target video generated by the video processing model, including: predicting the depth information of the original video; using the depth processing model to extract features from the depth information to obtain the depth features of the original video; adding the splicing information and the depth features to obtain the target features; and inputting the target features and fusion information into the video processing model to obtain the target video generated by the video processing model.
[0135] Optionally, the splicing information and fusion information are input into the video processing model to obtain the target video generated by the video processing model, including: inputting the splicing information and fusion information into the video processing model to obtain the color video generated by the video processing model; replacing the values of preset channels in the color video with the values of preset channels in the original video to obtain the target video.
[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0137] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause the processing unit to execute the methods in the various embodiments of this application.
[0139] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0140] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.
[0141] Optionally, in this embodiment, the storage medium may be located in a computing device.
[0142] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, which, when the executable program is running, controls the device where the computer-readable storage medium is located to execute the method described in any of the above embodiments.
[0143] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.
[0144] The aforementioned computer program products can refer to software programs that have been written, tested, and released, and can run on computers or other devices. Computer program products can include application programs, operating systems, utility software, etc., used to achieve specific functions or solve specific problems.
[0145] Embodiments of this application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the method provided in the above embodiments.
[0146] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.
[0147] Embodiments of this application also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.
[0148] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.
[0149] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0150] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0151] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0152] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0153] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0154] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method of video processing, the method comprising: include: Acquire the original video and multimodal information, wherein the multimodal information includes: at least two modal information and cue information, the cue information being used to indicate the color information of at least one sub-region in different video frames of the original video; The at least two modal information are mapped to the same feature space to obtain fused information; Based on the fusion information and the prompt information, the original video is colorized to obtain the target video.
2. The method of claim 1, wherein, The at least two modal information includes: image information and text information; mapping the at least two modal information to the same feature space to obtain fused information includes: Feature extraction is performed on the image information to obtain the original color features of the image information; Feature extraction is performed on the text information to obtain the original text features of the text information; The original color features are mapped to the feature space to obtain the target color features, and the original text features are mapped to the feature space to obtain the target text features; The target color features and the target text features are fused to obtain the fused information.
3. The method of claim 2, wherein, The step of mapping the original color features to the feature space to obtain target color features, and mapping the original text features to the feature space to obtain target text features, includes: By using at least one image processing module in a multimodal processing model, attention processing is performed on the preset color feature sequence in the feature space and the original color features to obtain the target color features; By using at least one text processing module in the multimodal processing model, attention processing is performed on the preset text feature sequence in the feature space and the original text features to obtain the target text features.
4. The method of claim 3, wherein, Any image processing module includes: a first attention layer and a first feedforward neural network layer; using at least one image processing module in the multimodal processing model, attention processing is performed on the preset color feature sequence in the feature space and the original color features to obtain the target color features, including: The first attention layer is used to perform attention processing on the input color features of any image processing module and the preset color feature sequence to obtain the first attention feature. The input color features of the first image processing module are the original color features, and the input color features of other image processing modules besides the first image processing module are the output color features of the previous image processing module. The first attention feature and the preset color feature sequence are added together to obtain the first additive feature; The first feedforward feature is obtained by using the first feedforward neural network layer to perform feedforward processing on the first additive feature; The first additive feature and the first feedforward feature are added together to obtain the output color feature of any one of the image processing modules; The output image features of the last image processing module are determined to be the target color features.
5. The method of claim 3, any one text processing module comprising: Second attention layer and second feedforward neural network layer; Using at least one text processing module in a multimodal processing model, attention processing is performed on a preset text feature sequence in the feature space and the original text features to obtain the target text features, including: The second attention layer is used to perform attention processing on the input text features of any text processing module and the preset text feature sequence to obtain the second attention feature. The input text features of the first text processing module are the original text features, and the input text features of other text processing modules besides the first text processing module are the output text features of the previous text processing module. The second attention feature and the preset text feature sequence are added together to obtain the second additive feature; The second feedforward feature is obtained by using the second feedforward neural network layer to perform feedforward processing on the second additive feature; The second additive feature and the second feedforward feature are added together to obtain the output text feature of any one of the text processing modules; The output text features of the last text processing module are determined to be the target text features.
6. The method of claim 2, wherein, The step of extracting features from the image information to obtain the original color features of the image information includes: The image is segmented using the segmentation module in the color extraction model to obtain multiple sets of sub-image information, wherein the shapes and / or quantities of different sets of sub-image information are different; The residual module in the color extraction model is used to extract features from the multiple sets of sub-image information to obtain color features in the image space. The image encoding module in the color extraction model is used to map the color features of the image space to a multimodal space to obtain the original color features.
7. The method according to any one of claims 1 to 6, characterized in that, The step of colorizing the original video based on the fusion information and the prompt information to obtain the target video includes: Generate noise information corresponding to the original video; The noise information, the prompt information, and the original video are spliced together to obtain spliced information; The stitching information and the fusion information are input into the video processing model to obtain the target video generated by the video processing model.
8. The method of claim 7, wherein, After splicing the noise information, the prompt information, and the original video to obtain spliced information, the method further includes: Predict the depth information of the original video; The depth information is used to extract features to obtain the depth features of the original video; The stitched information is added to the depth features to obtain the target features; The target features and the fusion information are input into the video processing model to obtain the target video generated by the video processing model.
9. The method of claim 7, wherein, The step of inputting the splicing information and the fusion information into the video processing model to obtain the target video generated by the video processing model includes: The stitching information and the fusion information are input into the video processing model to obtain the color video generated by the video processing model. The values of the preset channels in the color video are replaced with the values of the preset channels in the original video to obtain the target video.
10. A method for video processing, comprising: include: In response to an input command applied to the user interface, determine the original video corresponding to the input command; Acquire multimodal information, wherein the multimodal information includes: at least two modal information and cue information, wherein the cue information is used to indicate the color information of at least one sub-region in different video frames of the original video; The at least two modal information are mapped to the same feature space to obtain fused information; Based on the fusion information and the prompt information, the original video is colorized to obtain the target video; The target video is displayed on the user interface.
11. The method according to claim 10, characterized in that, The acquisition of multimodal information includes: In response to the input command, at least one modal information corresponding to the input command is determined; When the number of the at least one modal information is less than a preset number, the multimodal information is obtained based on the at least one modal information and the preset modal information; If the number of the at least one modal information is greater than or equal to a preset number, the at least one modal information is determined to be the multimodal information.
12. The method according to claim 10 or 11, characterized in that, The step of colorizing the original video based on the fusion information and the prompt information to obtain the target video includes: If the original video is a grayscale video, the original video is colorized based on the fusion information and the prompt information to obtain the target video; If the original video is a color video, the original video is converted to a grayscale video, and based on the fusion information and the prompt information, the grayscale video is colorized to obtain the target video.
13. A method for video processing, comprising: include: The original video and multimodal information are obtained by calling a first interface. The first interface includes a first parameter, the value of which includes the original video and the multimodal information. The multimodal information includes at least two modal information and prompt information. The prompt information is used to indicate the color information of at least one sub-region in different video frames of the original video. The at least two modal information are mapped to the same feature space to obtain fused information; Based on the fusion information and the prompt information, the original video is colorized to obtain the target video; The target video is output by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target video.
14. A video processing system characterized by include: The client is used to send the raw video; The server, connected to the client, is used to acquire multimodal information, wherein the multimodal information includes: at least two modal information and prompt information, the prompt information being used to indicate the color information of at least one sub-region in different video frames of the original video; the at least two modal information are mapped to the same feature space to obtain fusion information; based on the fusion information and the prompt information, the original video is colorized to obtain the target video; The client is also used to output the target video.
15. A model training method, comprising: include: Acquire training data, wherein the training data includes: training videos and first modality data; The first modal data and the second modal data in the training video are mapped to the same feature space to obtain training fusion information; Based on the training video, training prompt information is generated, which is used to prompt the color information of at least one sub-region in different video frames of the training video; Based on the training fusion information, the training prompt information, and the training video, a preset processing model is trained to obtain a target processing model, wherein the target processing model includes: a multimodal processing module, a color extraction module, a video processing module, and a depth processing module.
16. The method of claim 15, wherein, The step of generating training prompt information based on the training video includes: Superpixel segmentation is performed on multiple video frames of the training video to obtain multiple sets of first sub-regions, wherein different sets of first sub-regions correspond to different video frames, and different first sub-regions in the first sub-region sets correspond to different semantic information. Based on the color information of different pixels in any first sub-region, determine the region color information of the arbitrary first sub-region; At least one set of second sub-regions is determined from the plurality of video frames, wherein different sets of second sub-regions correspond to different semantic information, and different second sub-regions in the same set of second sub-regions correspond to different video frames; The training prompt information is obtained based on the at least one set of second sub-regions and the region color information of any one of the second sub-regions.
17. The method according to claim 15 or 16, characterized in that, The step of training a preset processing model based on the training fusion information, the training prompt information, and the training video to obtain a target processing model includes: Convert the training video into a grayscale training video; The depth information of the training grayscale video is extracted using the first processing module in the preset processing model to obtain training depth features; The training noise information, the training prompt information, and the training grayscale video corresponding to the training video are spliced together to obtain training splicing information; The training concatenation information is added to the training depth features to obtain the target training features; The target training features and the training fusion information are input into the second processing module in the preset processing model to obtain the target video generated by the second processing module. The model parameters of the preset processing model are adjusted based on the training video and the target video to obtain the target processing model.
18. The method according to claim 17, characterized in that, The step of adjusting the model parameters of the preset processing model based on the training video and the target video to obtain the target processing model includes: Based on the difference between the training video and the target video, a first loss function value is constructed; A first stream value in the target video is obtained by performing stream estimation on any two target video frames in the target video using the stream estimation model, and a second stream value in the training video is obtained by performing stream estimation on two training video frames in the training video using the same stream estimation model, wherein the two training video frames correspond to the two target video frames. Based on the difference between the first stream value and the second stream value, a second loss function value is constructed; The first loss function value and the second loss function value are summed to obtain the total loss function value; The model parameters of the preset processing model are adjusted based on the total loss function value to obtain the target processing model.
19. A computing device, comprising: include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 18.
20. An electronic device, comprising: include: Memory, which stores executable programs; A processor, connected to the memory via a bus, is used to run the program, wherein the program, when running, executes the method according to any one of claims 1 to 18.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 18.
22. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 18.