Video conversion model determination method and device based on multiple modes
Multi-dimensional information of video is obtained through multi-modal technology, multi-modal model extracts features and adaptively matches the video conversion model, solving the problems of inefficiency and unstable effects caused by artificial dependence in the existing technology, realizing automated and efficient video conversion, presenting a more natural picture.
Patent Information
- Application Number
- CN202510696204.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-22
AI Technical Summary
The existing video conversion technology relies on manual analysis, is inefficient and subjective, making it difficult to ensure the consistency and optimal effect of the conversion.
Multimodal technology is used to obtain multi-dimensional information of video, extract multimodal features through preset multimodal models, and adaptively match the target video conversion model for automatic conversion.
Automatic video analysis is realized, reducing human subjective judgment deviations, improving processing efficiency, stable conversion effect, and presenting a more natural and realistic picture.
Smart Images

Figure CN120358389A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video processing, and particularly to a method and device for determining a video conversion model based on multi-modalities. Background Art
[0002] Videos in different formats, such as different dynamic ranges and color presentations, will present different visual experiences to users. Videos with a lower dynamic range, such as SDR videos (Standard Dynamic Range), are the most common and widely used videos at present. They usually use 8-bit color depth, but have a relatively low contrast and limited brightness range, and cannot display the details of extremely dark or extremely bright scenes. SDR video content can be compatible with the vast majority of existing display devices and playback hardware, and does not require special hardware support. High-dynamic-range videos, such as HDR videos (High Dynamic Range), support a wider color range and higher color depth (usually 10 bits or more), making the colors more rich and delicate, and can display a wide range of brightness levels from very dark to very bright. The higher contrast makes the image more layered, providing deeper blacks and brighter whites.
[0003] With the continuous upgrading of video display devices, providing richer color presentations, higher contrast, and more realistic visual experiences can greatly improve the quality of users' video viewing. Considering that the original video is a video with a lower dynamic range, it can be converted into a high-dynamic-range video to adapt to high-dynamic-range video display devices. However, in actual applications, most of the existing video conversions rely on manual analysis of the video content to be converted, and manually select a video conversion model for conversion according to the analysis results. This is not only inefficient but also has strong subjectivity, making it difficult to ensure the consistency and optimal effect of the conversion. Summary of the Invention
[0004] In view of the above problems, embodiments of this application are proposed to provide a method and device for determining a video conversion model based on multi-modalities that can overcome or at least partially solve the above problems.
[0005] According to the first aspect of the embodiments of this application, a method for determining a video conversion model based on multi-modalities is provided, which includes:
[0006] Obtain multi-dimensional video information of a first video to be converted;
[0007] Input the multi-dimensional video information into a preset multi-modal model to obtain multi-modal features including multiple features;
[0008] An adaptive match is performed between the multi-modal features and each video conversion model in the video conversion model library to obtain a target video conversion model, and the first video is converted based on the target video conversion model to obtain a second video.
[0009] Optionally, the multi-dimensional video information at least includes picture information and text information;
[0010] Further obtaining the multi-dimensional video information of the first video to be converted includes:
[0011] Performing frame sampling processing on the first video to obtain sampled frames containing picture information; the frame sampling processing includes performing frame sampling processing according to a preset playback time interval, a preset video position, and / or a preset sampling rate;
[0012] Performing conversion processing on the audio content of the first video based on speech recognition to obtain corresponding text information; the text information also includes the name, label, and / or classification of the first video.
[0013] Optionally, further obtaining the multi-modal features including multiple features by inputting the multi-dimensional video information into a preset multi-modal model includes:
[0014] Inputting the picture information and text information in the multi-dimensional video information into a preset multi-modal model, and the preset multi-modal model performs cross-modal interaction feature extraction on the picture information and text information to obtain multi-modal features including the semantic features and picture quality features of the first video.
[0015] Optionally, further obtaining the multi-modal features including the semantic features and picture quality features of the first video by inputting the picture information and text information in the multi-dimensional video information into a preset multi-modal model, and the preset multi-modal model performs cross-modal interaction feature extraction on the picture information and text information includes:
[0016] Inputting the picture information and text information into a preset multi-modal model, and the preset multi-modal model extracts multiple feature data included in the picture information and compares them with a preset feature threshold to determine the picture quality features; and determining the semantic features according to the text information in combination with the picture information; the picture quality features include brightness, contrast, color, dynamic range, and / or lens movement mode; the semantic features include video scene type and / or video content style.
[0017] Optionally, the method further includes:
[0018] Collecting multiple video samples to be converted and annotating corresponding multi-modal sample features, and training the preset multi-modal model according to the video samples and multi-modal sample features.
[0019] Optionally, the method further includes:
[0020] Collect multiple video samples to be converted, and determine the multi-modal sample features of each video sample;
[0021] Pre-train each corresponding video conversion model according to each video sample;
[0022] Establish a video conversion model library, and save each video conversion model and the applicable conditions of the video conversion model; the applicable conditions are determined according to the multi-modal sample features of the corresponding video sample; the applicable conditions include multi-modal sample features, feature data ranges, and / or feature priorities.
[0023] Optionally, adaptively match the multi-modal features with each video conversion model in the video conversion model library to obtain a target video conversion model, and further include converting the first video to the second video based on the target video conversion model:
[0024] Based on the semantic features and picture quality features of the multi-modal features, match the applicable conditions of the video conversion models in the video conversion model library based on a preset adaptive algorithm, and determine the target video conversion model according to the matching degree with the applicable conditions.
[0025] Optionally, the dynamic range of the first video is lower than that of the second video.
[0026] According to the second aspect of the embodiments of the present application, there is provided a multi-modal-based video conversion model determination device, which includes:
[0027] An information acquisition module, adapted to acquire multi-dimensional video information of a first video to be converted;
[0028] A multi-modal module, adapted to input the multi-dimensional video information into a preset multi-modal model to obtain multi-modal features including multiple features;
[0029] A model matching module, adapted to adaptively match the multi-modal features with each video conversion model in the video conversion model library to obtain a target video conversion model, and convert the first video to the second video based on the target video conversion model.
[0030] According to the third aspect of the embodiments of the present application, there is provided a computing device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0031] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above multi-modal-based video conversion model determination method.
[0032] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided. At least one executable instruction is stored in the storage medium, and the executable instruction causes a processor to perform operations corresponding to the above-described multi-modal based video conversion model determination method.
[0033] According to a fifth aspect of the embodiments of the present application, a computer program product is provided, including at least one executable instruction, and the executable instruction causes a processor to perform operations corresponding to the above-described multi-modal based video conversion model determination method.
[0034] According to the multi-modal based video conversion model determination method and device provided by the present application, for the first video to be converted, the multi-modal features of the first video can be accurately extracted and determined by using a preset multi-modal model. Furthermore, an accurate target video conversion model can be adaptively matched according to the multi-modal features, which can avoid the dependence on manual video analysis and manual selection of video conversion models, realize automated video analysis and video conversion model selection, reduce the deviation of human subjective judgment, improve the processing efficiency, make the conversion effect more stable, and make the converted video present a more natural and real picture.
[0035] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically illustrates the specific embodiments of the present application. Description of the Drawings
[0036] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0037] Figure 1 Shows a flowchart of a multi-modal based video conversion model determination method according to an embodiment of the present application;
[0038] Figure 2 Shows a flowchart of a multi-modal based video conversion model determination method according to another embodiment of the present application;
[0039] Figure 3 Shows a flowchart of the training process of a preset multi-modal model;
[0040] Figure 4 Shows a flowchart of establishing a video conversion model library;
[0041] Figure 5The structural schematic diagram of a multimodal-based video conversion model determination device according to an embodiment of the present application is shown;
[0042] Figure 6 The structural schematic diagram of a computing device according to an embodiment of the present application is shown. Detailed implementation manners
[0043] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.
[0044] First, the noun terms involved in one or more embodiments of the present application are explained.
[0045] Multimodal: MLLM (Multimodal Large Language Model) is an intelligent model that can simultaneously process and understand multiple modalities of data (such as text, images, audio, video, etc.). It realizes cross-modal interaction, generation, and reasoning by integrating the semantic understanding and reasoning capabilities of a large language model (LLM) and the encoder features of modalities such as vision and audition.
[0046] SDR video: Standard Dynamic Range, a common color display method, with a smaller information size compared to HDR (High Dynamic Range) and a high popularity.
[0047] HDR video: High Dynamic Range, a video technology used to enhance the brightness, contrast, and color range of the picture. Compared with standard dynamic range (SDR) video, HDR video can present deeper blacks, brighter highlight areas, and richer colors, making the picture closer to the real world that the human eye can perceive.
[0048] Figure 1 The flowchart of a multimodal-based video conversion model determination method according to an embodiment of the present application is shown. As Figure 1 shown, the method includes the following steps:
[0049] Step S101, obtain multi-dimensional video information of the first video to be converted.
[0050] The technological improvement of video display devices can bring users a better viewing experience. At the same time, the video also needs to be enhanced accordingly. Even on high-quality video display devices, videos with low picture quality cannot present good visual effects. Therefore, for videos with low picture quality, video conversion is required to improve the picture quality to present a better visual effect. Existing video conversions generally use manual operations, converting videos based on manual experience, relying strongly on manual labor. The videos obtained after conversion have subjective judgment biases, and the effects obtained each time are unstable.
[0051] Considering the above problems, in this embodiment, starting from multi-dimensional video information of the video to be converted, multi-modal technology is adopted to perform standard quantization processing on the video, and a corresponding target video conversion model is matched for the video, achieving the effect of reducing manual intervention, removing subjective biases, and automatically and accurately matching the target video conversion model, realizing precise and efficient video conversion.
[0052] Videos contain information such as visual content (such as pictures, shot language, editing, special effects, graphics, etc.), auditory content (such as sounds, music, sound effects, etc.), and content theme types. Different videos contain different information, and different video conversion models should be used for conversion. Therefore, for the first video to be converted, it is necessary to determine the multi-dimensional video information it contains according to the video content of the first video itself. Here, the multi-dimensional video information includes video information in different dimensions such as pictures and texts.
[0053] Step S102: Input the multi-dimensional video information into a preset multi-modal model to obtain multi-modal features containing multiple features.
[0054] The preset multi-modal model can process various types of data. The data structures of different modalities are very different. The preset multi-modal model performs cross-modal interaction on different types of data, performs data fusion and alignment on different types of data, and comprehensively analyzes the video using different types of data, and can obtain the multi-modal features of the video. The multi-modal features can contain multiple features, reflecting different features of the video from different dimensions, such as picture quality features and semantic features. The preset multi-modal model can adopt, for example, a multi-modal large language model and be trained through sample data. When training the preset multi-modal model, the sample data includes video samples and multi-modal sample features of the video samples. To improve the accuracy of the preset multi-modal model, the video samples can use multiple video samples with different picture qualities, different themes, different video types, different durations, etc., to ensure the richness and diversity of the video samples. For example, the video samples can include various types of videos such as documentaries, animations, news, and short videos, covering various video samples with different brightnesses and different colors.
[0055] Input the multi-dimensional video information of the first video into a preset multi-modal model. The multi-modal model analyzes the multi-dimensional video information and extracts multi-modal features containing multiple features from it. The multi-modal features can include, for example, picture quality features, semantic features, etc. The picture quality features are used to indicate the picture situation presented by the first video, such as various visual features in the first video, such as the brightness distribution of the picture, the contrast of the picture, the color distribution of the picture, the color saturation, the dynamic range of the brightness difference range presented simultaneously in the picture, etc., and also include, for example, the camera movement mode determined according to different video pictures. The extraction of the picture quality features can be performed by the multi-modal model using, for example, a convolutional neural network to extract visual features, analyze the video in combination with the time dimension of the video, and determine the picture quality features of the video by dividing the scoring range or evaluation metrics, etc. The semantic features are used to indicate the video scene type, video content style, etc. of the first video. Although it focuses on different aspects from the picture quality features, it needs to be comprehensively analyzed in combination with the picture quality features. For example, according to the picture content of the first video combined with the subtitle content in the audio, the picture and text are cross-modally fused to determine the video scene type, video content style, etc. Further, when the multi-modal model analyzes the multi-dimensional video information, it also considers the continuity of different pictures in the multi-dimensional video information, and determines the overall semantic features of the first video based on the continuous picture change situation. The extraction of the semantic features can be performed by the multi-modal model according to the extracted picture quality features combined with audio information such as the sound, music, intonation, etc. in the video to determine the video content style, or according to information such as each video frame, audio, title, label, etc. of the video, perform semantic alignment through a cross-modal attention mechanism, and corresponding domain knowledge (such as different fields like sports, film and television, technology, etc.) can also be injected, and the video scene type can be determined in combination with a knowledge graph.
[0056] The above is for illustrative purposes, and it is specifically set according to the implementation situation and is not limited here.
[0057] Step S103: Perform adaptive matching between the multi-modal features and each video conversion model in the video conversion model library to obtain a target video conversion model, and based on the target video conversion model, convert the first video to obtain a second video.
[0058] Multiple video conversion models can be pre-stored in the video conversion model library. The video conversion models can convert the dynamic range, color gamut, brightness, etc. of the video to be converted to obtain a video with better visual effects. The video conversion models can analyze the color gamut, dynamic range, brightness, etc. of the video to be converted by, for example, extracting frames from the video to be converted for decoding, and enhance the contrast and color gamut in the video frames through feature mapping, conversion, etc., optimize the brightness through dark part enhancement and bright part suppression, and adjust the color, hue, etc. Further, the video conversion models can also process artifacts between multiple video frames to avoid artifacts, etc., and complete the video conversion. The above is for illustrative purposes, and the specific video conversion models can be set according to the actual implementation situation and are not limited here.
[0059] Different video conversion models can be used to convert different videos. The processing of different video conversion models is different, and the converted videos obtained are also different. To ensure the effect of video conversion, for the video to be converted, the conversion effects of different video conversion models can be compared to determine the video conversion model suitable for each video to be converted. At the same time, according to the video to be converted, the applicable conditions of the video conversion model suitable for it can be determined. For example, the applicable conditions can be obtained based on the image quality characteristics and semantic characteristics of the video to be converted. Multiple video conversion models and the applicable conditions of the corresponding video conversion models can be stored in the video conversion model library.
[0060] After obtaining the multi-modal features of the first video, the target video conversion model can be adaptively matched and selected from the video conversion model library according to the multi-modal features of the first video. Specifically, the applicable conditions of each video conversion model in the video conversion model library are different. By adaptively matching the multi-modal features of the first video with the applicable conditions of the video conversion model, the target video conversion model is determined. The adaptive matching can be, for example, classifying each feature and then calculating the matching degree of each feature in the multi-modal features, determining the matching applicable conditions based on the matching degree, and then determining the target video conversion model. Using the target video conversion model to convert the first video, the brightness, color, dynamic range, etc. of the first video are adjusted specifically to obtain the second video. When the second video is displayed on the video display device, it can provide a better visual effect for the user and improve the user's viewing experience.
[0061] According to the method for determining a video conversion model based on multi-modal provided in this application, for the first video to be converted, the multi-modal features of the first video can be accurately extracted and determined using a preset multi-modal model. Furthermore, an accurate target video conversion model can be adaptively matched according to the multi-modal features, which can avoid relying on manual video analysis and manual selection of video conversion models, realize automated video analysis and video conversion model selection, reduce the deviation of human subjective judgment, improve the processing efficiency, make the conversion effect more stable, and make the converted video present a more natural and real picture.
[0062] Figure 2 The flowchart of a method for determining a multi-modal based video conversion model according to an embodiment of the present application is shown. As Figure 2 shown, the method includes the following steps:
[0063] Step S201: Perform frame sampling processing on the first video to obtain sampled frames containing picture information.
[0064] The first video is the video to be converted. Compared with the converted second video, its dynamic range is lower than that of the second video. By converting the first video into the second video, the brightness, contrast, color performance, etc. of the video can be greatly improved, the details of the video can be better presented, the colors of the video can be enriched, and the video picture transition is smoother. For example, the first video can be an SDR standard dynamic range video, and the second video can be an HDR high dynamic range video. Converting the SDR video into an HDR video can display an HDR video with a wider brightness range, richer content details, stronger light and dark layering in the picture, and more colors for the user based on the video display device. The above is for illustrative purposes, and the specific video can be determined according to the implementation situation and is not limited here.
[0065] When performing video conversion on the first video, a video conversion model is used to convert it. After the first video is converted by a suitable video conversion model, a second video with better visual effects can be obtained. After the first video is converted by an unsuitable video conversion model, the obtained second video cannot meet the conversion requirements. Therefore, it is necessary to select a suitable video conversion model for the first video for conversion. Most of the existing technologies rely on manual experience for selection, without a unified selection standard, and are strongly dependent on manual work, resulting in uneven conversion quality. In this embodiment, when performing video conversion, starting from the video content of the first video itself, according to the multi-dimensional video information contained in the video content of the first video, a multi-modal model is used to perform cross-modal fusion on different types of data in the multi-dimensional video information, and the multi-modal features of the first video are extracted. The multi-modal features standardize the data of the first video, and a suitable video conversion model can be accurately matched based on the multi-modal features.
[0066] The multi-modal features need to be determined according to the video content of the first video itself. In this embodiment, the video's picture and audio are obtained respectively, and multi-modal features are obtained through multi-dimensional video information such as pictures and audio. For the picture of the first video, frame sampling processing can be performed on the first video to obtain sampled frames containing picture information. The frame sampling processing can collect multiple video frames in the first video. According to the pictures of multiple video frames and in combination with the time continuity of the video frames, the picture information of the first video is determined. The frame sampling processing can be sampled according to a preset playback time interval. For example, according to the playback time, 1 frame or several video frames are sampled every 1 minute; or, the frame sampling processing is sampled according to a preset video position, such as the video start position, the video start X position, the X position before the video end, the video end position, etc., to collect video frames at different positions; or, the frame sampling processing is performed on the first video according to a preset sampling rate, and video frames are randomly sampled based on the preset sampling rate, etc. The above is for illustrative purposes only and is specifically set according to the actual situation and is not limited here.
[0067] Step S202: Based on speech recognition, perform conversion processing on the audio content of the first video to obtain corresponding text information.
[0068] The frame sampling processing of the first video can obtain sampled frames containing picture information. The multi-dimensional video information of the first video includes not only picture information but also text information. The text information can be obtained by performing conversion processing on the audio content of the first video through speech recognition, such as speech recognition. ASR (Automatic Speech Recognition) can be used to recognize and convert the audio content in the first video to obtain corresponding text information. Here, speech recognition can directly recognize the audio file in the first video without considering whether the first video contains subtitles, and speech recognition does not need to recognize the subtitles of each video frame frame by frame, which is faster and more convenient.
[0069] In addition to the audio content of the first video, the text information also includes text information such as the name, tags, and classification of the first video, which can help determine the video scene type, video content style, etc. of the first video.
[0070] Furthermore, if the audio content of the first video is in Chinese, the text information obtained through speech recognition is in Chinese; if the audio content of the first video is in a foreign language, during speech recognition, it can also be translated into Chinese to obtain Chinese text information, etc.
[0071] Step S203: Input the picture information and text information in the multi-dimensional video information into a preset multi-modal model, and the preset multi-modal model performs cross-modal interaction feature extraction on the picture information and text information to obtain multi-modal features including the semantic features and picture quality features of the first video.
[0072] According to steps S201 - S202, the picture information and text information of the first video are obtained, and the picture information and text information are input into a preset multimodal model together. The preset multimodal model can extract multiple feature data contained in the picture information from the picture information. For example, using visual feature extraction technology, the underlying visual features of the sampled frames can be extracted, etc., and then the corresponding feature data can be obtained. The picture quality features include, for example, brightness, contrast, color, dynamic range, lens movement mode, etc., and comprehensively reflect the picture content of the first video from different aspects. The extraction of picture quality features can be determined according to different sampled frames in combination with the time dimension of each sampled frame. When the picture quality feature values in different sampled frames are different, the average value or the highest value, etc., can be selected according to the actual situation, and there is no limitation here.
[0073] To better match the video conversion model using multimodal features, the feature data of the picture quality features can be further quantified. For example, when the feature value of brightness of the feature data is 300 and the feature value of brightness is 310, the same video conversion model can be used for conversion, which does not affect the visual effect after conversion and can reduce the computational complexity when using specific values for matching. The preset multimodal model can compare the extracted feature data with a preset feature threshold to determine the final picture quality feature. For example, for the brightness, contrast, color (color contrast, color saturation, etc.), dynamic range, etc. in the picture quality features, after specific values are extracted according to the picture information contained in the sampled frames, respective preset feature thresholds can be set for brightness, contrast, and color. The preset feature threshold can include multiple different threshold ranges. Taking brightness as an example, the brightness can be set to different feature values such as high, medium, and low according to different threshold ranges. By comparing the extracted feature data with the preset feature threshold, it can be determined that, for example, the brightness feature value is high. Further quantifying the feature values can reduce the computational complexity of matching and improve the processing efficiency when subsequently matching the video conversion model. The specific feature values can be set according to the actual situation, or more feature values can be used. Refining the feature values can more accurately correspond to the applicable conditions of the video conversion module, and there is no limitation here.
[0074] When the preset multi-modal model determines the picture quality features based on the picture information, it also determines the semantic features by combining the text information with the picture information at the same time. The semantic features include, for example, the video scene type, the video content style, etc. The video content style can include different styles such as animation, movie, documentary, etc., and the video scene type includes different types such as sports, technology, film and television, etc. Relying only on the text information may lead to deviations in the determination of semantic features. The preset multi-modal model performs cross-modal fusion of the text information and the picture information, which can more accurately determine the semantic features and ensure the accuracy of the subsequent matching video conversion model. When the preset multi-modal model performs cross-modal fusion of the picture information and the text information and aligns them, for example, by combining the rich colors contained in the picture information with the two-dimensional conversation mode, language habits, labels, etc. in the text information, the video content style can be accurately determined as animation, etc.
[0075] Further, for the preset multi-modal model, it can be pre-trained, and the training process is as Figure 3 shown:
[0076] Step S301, collect multiple video samples to be converted and label the corresponding multi-modal sample features.
[0077] The preset multi-modal model can be trained using any multi-modal large language model, and no limitation is made here.
[0078] During training, collect multiple video samples to be converted, annotate the video samples, and label their corresponding multi-modal sample features. When annotating, the multi-modal sample features can be respectively labeled as picture quality features and text features. For the picture quality features, computer vision technology and other methods can also be borrowed to determine each feature data, which will not be elaborated here.
[0079] Step S302, train the preset multi-modal model according to the video samples and the multi-modal sample features.
[0080] For the video samples, the video samples can be frame-sampled to obtain the corresponding sampled frames. Perform speech recognition on the audio content of the video samples to obtain text information. Then input the sampled frames and the text information into the preset multi-modal model for training, compare the output result with the multi-modal sample features labeled by the video samples, and adjust the training parameters of the preset multi-modal model to finally obtain the trained preset multi-modal model.
[0081] Further, when the preset multi-modal model extracts features from a video in this embodiment, it can first obtain the sampled frames and text information of the first video and input them into the preset multi-modal model, or directly input the first video into the preset multi-modal model. In the preset multi-modal model, frame sampling processing and speech recognition are directly performed on the first video to obtain text information, and the picture information and text information are cross-modally fused to extract and determine multi-modal features, which are specifically set according to the actual situation and are not limited here.
[0082] Step S204: Based on the semantic features and picture quality features of the multi-modal features, match the applicable conditions of the video conversion models in the video conversion model library according to the preset adaptive algorithm, and determine the target video conversion model according to the matching degree with the applicable conditions.
[0083] The video conversion model library can be established in advance. The video conversion model library contains multiple video conversion models, and each video conversion model corresponds to its applicable conditions, which is convenient for matching the target video conversion model based on the applicable conditions.
[0084] The establishment of the video conversion model library is as Figure 4 shown, including the following steps:
[0085] Step S401: Collect multiple video samples to be converted and determine the multi-modal sample features of each video sample.
[0086] Different video conversion models are applicable to different videos. Therefore, different video conversion models need to be pre-trained for various different video samples. When collecting video samples to be converted, different video samples can be collected based on different video content styles, different video scene types, and different picture quality features. For each video sample, the multi-modal sample features of each video sample can be determined using, for example, the preset multi-modal model, or the multi-modal sample features of the video sample can be manually marked, that is, the video samples and marked multi-modal sample features in step S301 can be reused, which is not limited here.
[0087] Step S402: Pre-train the corresponding video conversion model for each video sample.
[0088] For each collected video sample, a suitable video conversion model is trained specifically for it. During training, different video conversion models can be used for training according to the different multi-modal sample features of the video sample. For example, when the video content style in the multi-modal sample features of the video sample is anime, the video conversion model focuses on the converted video presenting richer colors. The parameters of the video conversion model during training focus on color conversion, enriching colors, paying attention to details, etc. The training of the video conversion model can be completed by, for example, detecting the converted video to determine that it meets the conversion requirements. Or, the video sample can be an existing video that has been converted and the video conversion model used during the conversion, which can be directly used to establish a video conversion model library.
[0089] Step S403, establish a video conversion model library, and save each video conversion model and the applicable conditions of the video conversion model.
[0090] When establishing the video conversion model library, in addition to saving each video conversion model, the applicable conditions of the video conversion model are also saved correspondingly for the matching of the video conversion model.
[0091] The applicable conditions can be determined according to the multi-modal sample features of the corresponding video sample. For example, for video sample b using video conversion model a, the multi-modal sample features of video sample b include high brightness, low dynamic range, video content style documentary, etc., then the applicable conditions of video conversion model a include high brightness, low dynamic range, video content style documentary. Further, the applicable conditions can include multi-modal sample features and feature data ranges. In addition, feature priorities can be set for the multi-modal sample features. For example, if video conversion model a pays more attention to the video content style, the priority of the video content style is set to high, that is, when the multi-modal sample features in the applicable conditions are not exactly the same during matching, according to the feature priorities of the matched multi-modal sample features, the matching degree is determined, and a more suitable video conversion model is selected. The above is for illustrative purposes, and it is specifically set according to the implementation situation and is not limited here.
[0092] After obtaining the semantic features and picture quality features of the multi-modal features, according to the semantic features and picture quality features, a preset adaptive algorithm is used to match with the applicable conditions of the video conversion models in the video conversion model library, and the target video conversion model is determined according to the matching degree with the applicable conditions. For example, the preset adaptive algorithm can classify the multi-modal features first, perform feature matching on the semantic features and picture quality features respectively, calculate the similarity or distance of each feature, combine the feature priorities, calculate the overall matching degree of the multi-modal features, and select the video conversion model with the highest matching degree among the applicable conditions as the target video conversion model according to the overall matching degree. The first video is converted using the target video conversion model to obtain a second video. The dynamic range of the first video is lower than that of the second video, and the target video conversion model converts the dynamic range of the first video to a higher range, and the display of the second video can provide users with richer colors, more detailed displays, etc.
[0093] According to the method for determining a video conversion model based on multi-modal provided by this application, frame sampling processing is performed on the first video, and conversion processing is performed on the audio content of the first video by voice recognition, and the picture information and text information of the first video can be obtained. The picture information and text information are input into a preset multi-modal model, and the picture information and text information are cross-modally fused to extract the multi-modal features of the first video. The multi-modal features include picture quality features and semantic features, which characterize the overall characteristics of the first video from multiple levels and provide a basis for accurate matching and comparison for subsequent matching of video conversion models. The preset multi-modal model is comprehensively evaluated based on the picture information and text information of the first video without manual intervention. The preset adaptive algorithm matches the multi-modal features of the first video with the applicable conditions of each video conversion model in the video conversion model library, determines the matching degree between the multi-modal features and each multi-modal sample feature in the applicable conditions, and the target video conversion model determined by the matching degree can select the most suitable video conversion model, and the conversion effect is more stable and more in line with the content characteristics of the first video.
[0094] Figure 5 The structural schematic diagram of a video conversion model determination device based on multi-modal provided by an embodiment of this application is shown. As Figure 5 shown, the device includes:
[0095] An information acquisition module 510, adapted to acquire multi-dimensional video information of a first video to be converted;
[0096] A multi-modal module 520, adapted to input the multi-dimensional video information into a preset multi-modal model to obtain multi-modal features including multiple features;
[0097] The model matching module 530 is adapted to adaptively match the multi-modal features with each video conversion model in the video conversion model library to obtain a target video conversion model, and based on the target video conversion model, convert the first video to obtain a second video.
[0098] Optionally, the multi-dimensional video information at least includes picture information and text information;
[0099] The information acquisition module 510 is further adapted to:
[0100] Perform frame sampling processing on the first video to obtain sampling frames containing picture information; the frame sampling processing includes frame sampling processing according to a preset playback time interval, a preset video position, and / or a preset sampling rate;
[0101] Perform conversion processing on the audio content of the first video based on speech recognition to obtain corresponding text information; the text information also includes the name, label, and / or classification of the first video.
[0102] Optionally, the multi-modal module 520 is further adapted to:
[0103] Input the picture information and text information in the multi-dimensional video information into a preset multi-modal model, and the preset multi-modal model extracts cross-modal interaction features from the picture information and text information to obtain multi-modal features including the semantic features and picture quality features of the first video.
[0104] Optionally, the multi-modal module 520 is further adapted to:
[0105] Input the picture information and text information into a preset multi-modal model, and the preset multi-modal model extracts multiple feature data included in the picture information and compares them with a preset feature threshold to determine the picture quality features; and, determine the semantic features according to the text information in combination with the picture information; the picture quality features include brightness, contrast, color, dynamic range, and / or lens movement mode; the semantic features include video scene type and / or video content style.
[0106] Optionally, the device further includes: a first training module 540, which is adapted to collect a plurality of video samples to be converted and label the corresponding multi-modal sample features, and train the preset multi-modal model according to the video samples and multi-modal sample features.
[0107] Optionally, the device further includes: a second training module 550, which is adapted to collect a plurality of video samples to be converted and determine the multi-modal sample features of each video sample; pre-train each corresponding video conversion model according to each video sample; establish a video conversion model library, and save each video conversion model and the applicable conditions of the video conversion model; the applicable conditions are determined according to the multi-modal sample features of the corresponding video sample; the applicable conditions include multi-modal sample features, feature data ranges, and / or feature priorities.
[0108] Optionally, the model matching module 530 is further adapted to:
[0109] Match the applicable conditions of the video conversion models in the video conversion model library based on the semantic features and picture quality features of the multimodal features, and determine the target video conversion model according to the matching degree with the applicable conditions.
[0110] Optionally, the dynamic range of the first video is lower than that of the second video.
[0111] The descriptions of the above modules refer to the corresponding descriptions in the method embodiments and will not be elaborated here.
[0112] According to the device for determining a video conversion model based on multimodality provided by the present application, for the first video to be converted, the preset multimodal model can accurately extract and determine the multimodal features of the first video, and then an accurate target video conversion model can be adaptively matched according to the multimodal features, which can avoid the dependence on manual video analysis and manual selection of video conversion models, realize automatic video analysis and video conversion model selection, reduce the deviation of human subjective judgment, improve the processing efficiency, make the conversion effect more stable, and make the converted video present a more natural and real picture.
[0113] The present application also provides a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction, and the executable instruction can execute the operations corresponding to the method for determining a video conversion model based on multimodality in any of the above method embodiments.
[0114] The present application also provides a computer program product, and the computer program product includes at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to execute the operations corresponding to the method for determining a video conversion model based on multimodality in any of the above method embodiments.
[0115] Figure 6 FIG. shows a schematic structural diagram of a computing device according to an embodiment of the present application, and the specific implementation of the computing device is not limited in the specific embodiments of the present application.
[0116] As Figure 6 shown, the computing device may include: a processor 602, a communications interface 604, a memory 606, and a communication bus 608.
[0117] Wherein:
[0118] The processor 602, the communications interface 604, and the memory 606 complete communication with each other through the communication bus 608.
[0119] A communication interface 604 for communicating with network elements of other devices such as clients or other servers.
[0120] A processor 602 for executing a program 610, which can specifically execute the relevant steps in the above embodiments of the method for determining a multi-modal video conversion model.
[0121] Specifically, the program 610 may include program code, and the program code includes computer operation instructions.
[0122] The processor 602 may be a central processing unit (CPU), or a specific integrated circuit (ASIC) (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0123] A memory 606 for storing the program 610. The memory 606 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0124] The program 610 can specifically be used to cause the processor 602 to execute the method for determining a multi-modal video conversion model in any of the above method embodiments. For the specific implementation of each step in the program 610, reference may be made to the corresponding steps and descriptions in the corresponding units in the above embodiments of the method for determining a multi-modal video conversion model, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.
[0125] The algorithms or displays provided herein are not inherently related to any specific computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The structure required to construct such a system is obvious from the above description. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the descriptions made above for specific languages are for disclosing the preferred embodiments of the present application.
[0126] In the specification provided herein, a large number of specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0127] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0128] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0129] In addition, those skilled in the art will be able to understand that, although some of the embodiments herein include certain features included in other embodiments but not other features, the combination of the features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0130] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the present application. The present application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0131] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
Claims
1. A method for determining a video conversion model based on multi-modalities, comprising: Obtaining multi-dimensional video information of a first video to be converted; Inputting the multi-dimensional video information into a preset multi-modal model to obtain multi-modal features including multiple features; Performing adaptive matching on the multi-modal features with each video conversion model in a video conversion model library to obtain a target video conversion model, and converting the first video based on the target video conversion model to obtain a second video.
2. The method according to claim 1, wherein, The multi-dimensional video information at least includes picture information and text information; The obtaining of the multi-dimensional video information of the first video to be converted further includes: Performing frame sampling processing on the first video to obtain sampled frames including picture information; the frame sampling processing includes performing frame sampling processing according to a preset playing time interval, a preset video position, and / or a preset sampling rate; Performing conversion processing on the audio content of the first video based on speech recognition to obtain corresponding text information; the text information further includes the name, tags, and / or classification of the first video.
3. The method according to claim 1 or 2, wherein The inputting of the multi-dimensional video information into a preset multi-modal model to obtain multi-modal features including multiple features further includes: Inputting the picture information and text information in the multi-dimensional video information into a preset multi-modal model, and the preset multi-modal model performs cross-modal interaction feature extraction on the picture information and text information to obtain multi-modal features including the semantic features and picture quality features of the first video.
4. The method according to claim 3, wherein, The inputting of the picture information and text information in the multi-dimensional video information into a preset multi-modal model, and the preset multi-modal model performs cross-modal interaction feature extraction on the picture information and text information to obtain multi-modal features including the semantic features and picture quality features of the first video further includes: Inputting the picture information and text information into a preset multi-modal model, and the preset multi-modal model extracts multiple feature data included in the picture information and compares them with a preset feature threshold to determine the picture quality features; and determining the semantic features according to the text information in combination with the picture information; the picture quality features include brightness, contrast, color, dynamic range, and / or lens movement mode; the semantic features include video scene type and / or video content style.
5. The method according to any one of claims 1 - 4, wherein, The method further includes: Collecting multiple video samples to be converted and annotating corresponding multi-modal sample features, and training the preset multi-modal model according to the video samples and multi-modal sample features.
6. The method according to any one of claims 1-5, wherein, The method further includes: Collecting multiple video samples to be converted and determining the multi-modal sample features of each video sample; Pre-training respective corresponding video conversion models according to each video sample; Establishing a video conversion model library, saving each video conversion model and the applicable conditions of the video conversion model; the applicable conditions are determined according to the multi-modal sample features of the corresponding video sample; the applicable conditions include multi-modal sample features, feature data ranges, and / or feature priorities.
7. The method according to any one of claims 1-6, wherein, The step of adaptively matching each video conversion model in the video conversion model library according to the multi-modal features to obtain a target video conversion model, and then converting the first video based on the target video conversion model to obtain a second video further includes: Based on the semantic features and picture quality features of the multi-modal features, match the applicable conditions of the video conversion models in the video conversion model library based on a preset adaptive algorithm, and determine the target video conversion model according to the matching degree with the applicable conditions.
8. The method according to any one of claims 1-7, wherein The dynamic range of the first video is lower than that of the second video.
9. A multi-modal based video conversion model determination device, comprising: An information acquisition module, adapted to acquire multi-dimensional video information of a first video to be converted; A multi-modal module, adapted to input the multi-dimensional video information into a preset multi-modal model to obtain multi-modal features including multiple features; A model matching module, adapted to adaptively match each video conversion model in the video conversion model library according to the multi-modal features to obtain a target video conversion model, and then convert the first video based on the target video conversion model to obtain a second video.
10. A computing device, comprising: A processor, a memory, a communication interface and a communication bus, where the processor, the memory and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the multi-modal based video conversion model determination method according to any one of claims 1-8.
11. A computer storage medium, in which at least one executable instruction is stored, and the executable instruction causes a processor to perform the operations corresponding to the multi-modal based video conversion model determination method according to any one of claims 1-8.
12. A computer program product, including at least one executable instruction, and the executable instruction causes a processor to perform the operations corresponding to the multi-modal based video conversion model determination method according to any one of claims 1-8.