Video processing method and device based on artificial intelligence, equipment and medium
Through the video processing method based on artificial intelligence, the video images are divided into stationary images and motion images using convolutional neural networks, and independently processed, solving the problems of low video editing efficiency and poor viewing experience in the existing technology, and achieving efficient video synthesis and smooth video playback effects.
Patent Information
- Application Number
- CN202510202294.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video processing methods have shortcomings in editing efficiency and viewing experience, especially when processing videos containing a large number of still and moving elements, which makes it difficult to efficiently separate and process, resulting in inefficient editing and poor results.
Using artificial intelligence-based video processing methods, video images are accurately divided into still images and motion images through technologies such as convolutional neural networks, and processed to form independent still video and motion videos, and finally synthesize them to improve the video synthesis efficiency, and process low-frame-rate videos through interpolated frames to solve the problem of picture lag.
It realizes automated video processing, improves editing efficiency and video quality, solves the problem of low-frame rate video screen lag, and significantly improves the viewing experience.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and specifically to a video processing method, device, equipment and medium based on artificial intelligence. Background Art
[0002] With the continuous development of video technology, videos play an increasingly important role in people's lives and work. However, traditional video processing methods have many limitations. For example, in the video editing process, it is difficult to efficiently separate and process videos containing a large number of static and moving elements, resulting in low editing efficiency and poor effects. In addition, for videos with a low frame rate, problems such as frame freezing and unsmoothness are likely to occur, affecting the viewing experience. Although there are some video processing technologies attempting to solve these problems currently, they often rely on complex algorithms and a large amount of manual intervention, and cannot meet the growing demand for efficient and automated video processing. Summary of the Invention
[0003] (1) Technical Problems to be Solved
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a video processing method, device, equipment and medium based on artificial intelligence, and solves the problems of low editing efficiency and poor viewing experience existing in the existing video processing methods.
[0005] (2) Technical Solutions
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A video processing method based on artificial intelligence includes the following steps:
[0007] Step 1: Obtain the video to be processed and convert the video to be processed into a video picture set;
[0008] Step 2: Identify each video picture in the video picture set through artificial intelligence, and divide the video pictures into static pictures and moving pictures according to the change of pixel points;
[0009] Step 3: Convert the static picture areas of multiple video pictures into multiple independent static pictures marked, and combine the multiple independent static pictures into an independent static picture set;
[0010] Step 4: Correlate each static picture in the independent static picture set with one or more moving pictures;
[0011] Step 5: When performing video synthesis, editing or analysis, convert all moving pictures into a moving video. When performing video acquisition, synthesize the independent static picture set into a static video according to time, and then synthesize the moving video and the static video.
[0012] Preferably, in the fifth step, for videos with low frame numbers, by using a convolutional neural network deep learning model, spatio-temporal features in the video can be learned, and realistic interpolation frames can be generated. Many video frame interpolation models based on deep learning have been proposed by researchers.
[0013] Preferably, in the second step, a convolutional neural network is used to extract features from each video picture in the video picture set. By analyzing the change amplitude and change pattern of pixel points between adjacent frames, it is determined whether the video picture is a static picture or a moving picture.
[0014] Preferably, in the third step, when splicing multiple static picture regions of video pictures, an image segmentation algorithm is used to accurately segment the static picture regions, and then adjacent and content-related static picture regions are seamlessly spliced according to the coherence of the image content and the spatial position relationship to form an independent static picture. For example, for a video picture containing multiple static objects in a scene, the regions where each static object is located are segmented by an image segmentation algorithm, and then they are spliced according to the actual position relationship of the objects in the scene to form a complete independent static picture.
[0015] Preferably, in the fourth step, when corresponding the independent static picture with the moving picture, by analyzing the motion trajectory and motion direction of the object in the moving picture, its position relationship and time sequence with the corresponding object in the independent static picture are determined. For example, when an object moves from left to right in the moving picture, according to its initial position and final position in the static picture, as well as the speed change during the motion process, the moving picture is associated with the corresponding independent static picture so that the moving picture and the static picture can be accurately fused in the subsequent video synthesis.
[0016] Preferably, when the convolutional neural network deep learning model learns the spatio-temporal features in the video, a multi-scale feature extraction strategy is adopted to analyze the features of video frames at different scales to capture spatio-temporal information at different levels in the video.
[0017] A video processing device based on artificial intelligence, the device includes:
[0018] Video acquisition and conversion module: responsible for acquiring the video file to be processed, supporting the input of multiple video formats, and it can decompose the video into a series of video pictures according to the set frame rate to form a video picture set;
[0019] Video picture recognition and classification module: using advanced artificial intelligence technologies, especially convolutional neural networks, to deeply analyze each picture in the video picture set. It can identify the change situation of pixel points in the picture and accurately classify the video pictures into static pictures and moving pictures according to the change amplitude and pattern;
[0020] Image Correspondence Establishment Module: Its main task is to establish the correspondence between the still images and the motion images in the independent still image set. By analyzing information such as the motion trajectory, motion direction, and speed change of the objects in the motion images, it determines the positional relationship and time sequence between the corresponding objects in the independent still images.
[0021] Video Synthesis and Editing Module: This module plays a key role in video synthesis, editing, or analysis. For motion images, it can convert them into motion videos, and through advanced video processing technologies, make the motion pictures smooth and natural. When collecting videos, this module will synthesize the independent still image set into a still video in chronological order.
[0022] Control and Management Module: As the central part of the entire device, this module is responsible for coordinating the control and management of each functional module. It rationally allocates the work processes and parameter settings of each module according to the user's input instructions and the requirements of video processing.
[0023] An electronic device, characterized in that the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above-mentioned artificial intelligence-based video processing method is implemented.
[0024] A computer-readable storage medium, the computer-readable storage medium stores a computer program, characterized in that when the computer program is executed by a processor, the above-mentioned artificial intelligence-based video processing method is implemented.
[0025] (III) Beneficial Effects
[0026] The present invention provides an artificial intelligence-based video processing method, device, equipment, and medium.
[0027] It has the following beneficial effects:
[0028] 1. Through artificial intelligence technologies such as convolutional neural networks, video pictures can be automatically and accurately divided into still images and motion images without manual frame-by-frame judgment, greatly saving time and labor costs.
[0029] 2. When synthesizing videos, the independent still image set and motion images are processed separately and then synthesized, avoiding repeated processing of the entire video and improving the synthesis efficiency. For example, when processing a video containing a large amount of still background and a small number of moving objects, only the moving object part needs to be processed intensively, and the background part is directly synthesized using the still image set, significantly accelerating the processing speed.
[0030] 3. For videos with low frame rates, a convolutional neural network deep learning model is used to learn spatio-temporal features and generate realistic interpolated frames, enabling the video to maintain a smooth visual effect even at low frame rates and solving the problem of frame freezing in traditional videos at low frame rates. Detailed implementation
[0031] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] Embodiment:
[0033] The embodiments of the present invention provide a video processing method, device, equipment and medium based on artificial intelligence. A video processing method based on artificial intelligence includes the following steps:
[0034] Step 1: Obtain a video to be processed and convert the video to be processed into a video picture set.
[0035] Step 2: Identify each video picture in the video picture set through artificial intelligence, and classify the video pictures into static pictures and moving pictures according to the change of pixel points. A convolutional neural network is used to extract features from each video picture in the video picture set. By analyzing the change amplitude and change mode of pixel points between adjacent frames, it is determined whether the video picture is a static picture or a moving picture. When the convolutional neural network deep learning model learns the spatio-temporal features in the video, a multi-scale feature extraction strategy is adopted to analyze the features of video frames at different scales to capture spatio-temporal information at different levels in the video.
[0036] Step 3: Convert the static picture regions of multiple video pictures into multiple independent static pictures marked, and combine the multiple independent static pictures into an independent static picture set. When splicing the static picture regions of multiple video pictures, an image segmentation algorithm is used to accurately segment the static picture regions, and then adjacent and content-related static picture regions are seamlessly spliced according to the coherence of image content and spatial position relationship to form an independent static picture. For example, for a video picture containing multiple static objects in a scene, the regions where each static object is located are segmented by an image segmentation algorithm, and then spliced according to the actual position relationship of the objects in the scene to form a complete independent static picture.
[0037] Step 4: Correlate each still image in the independent still image set with one or more motion images. When correlating the independent still images with the motion images, analyze the motion trajectory and direction of the objects in the motion images to determine their positional relationship and chronological order with the corresponding objects in the independent still images. For example, when an object moves from left to right in a motion image, correlate the motion image with the corresponding independent still image based on its initial and final positions in the still image and the speed changes during the motion process, so as to accurately fuse the motion image and the still image in subsequent video synthesis.
[0038] Step 5: When performing video synthesis, editing, or analysis, convert all motion images into a motion video. When collecting video, synthesize the independent still image set into a still video according to time, and then synthesize the motion video and the still video. For videos with low frame rates, by using a convolutional neural network deep learning model, the spatio-temporal features in the video can be learned and realistic interpolated frames can be generated. Many video frame interpolation models based on deep learning have been proposed by researchers.
[0039] A video processing device based on artificial intelligence, the device includes:
[0040] Video acquisition and conversion module: Responsible for acquiring the video file to be processed, supporting the input of multiple video formats. It can decompose the video into a series of video pictures according to the set frame rate to form a video picture set;
[0041] Video picture recognition and classification module: Using advanced artificial intelligence technologies, especially convolutional neural networks, deeply analyze each picture in the video picture set. It can identify the changes in pixel points in the picture and accurately classify the video pictures into still images and motion images according to the amplitude and pattern of the changes;
[0042] Image correspondence establishment module: The main task is to establish the correspondence between the still images and the motion images in the independent still image set. It determines their positional relationship and chronological order with the corresponding objects in the independent still images by analyzing information such as the motion trajectory, motion direction, and speed changes of the objects in the motion images;
[0043] Video synthesis and editing module: Plays a key role when performing video synthesis, editing, or analysis. For motion images, it can convert them into a motion video, and through advanced video processing technologies, make the motion pictures smooth and natural. When collecting video, this module will synthesize the independent still image set into a still video according to the chronological order.
[0044] Control and management module: As the center of the entire device, this module is responsible for coordinating the control and management of each functional module. It reasonably allocates the work processes and parameter settings of each module according to the user's input instructions and the requirements of video processing.
[0045] An electronic device, characterized in that the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above-mentioned video processing method based on artificial intelligence is implemented.
[0046] A computer-readable storage medium stores a computer program, characterized in that when the computer program is executed by a processor, the above-mentioned video processing method based on artificial intelligence is implemented.
[0047] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A video processing method based on artificial intelligence, characterized in that: The following steps are involved: Step 1: Obtain the video to be processed and convert it into a video picture set; Step 2: Use artificial intelligence to identify each video image in the video image set, and divide the video images into still images and motion images according to the changes in pixel points; Step 3: converting the plurality of video image still image regions into a plurality of independent still images marked as a plurality of independent still images, and combining the plurality of independent still images into an independent still image set; Step 4: Assign each still image in the independent still image set to one or more motion images; Step 5: When performing video synthesis, editing or analysis, all motion images are converted into motion videos. When performing video acquisition, independent still image sets are synthesized into still videos according to time, and then the motion videos are synthesized with the still videos.
2. The video processing method based on artificial intelligence according to claim 1, characterized in that: In step 5, for videos with low frame rates, the spatiotemporal features in the video can be learned by using a convolutional neural network deep learning model, and realistic interpolation frames can be generated. Researchers have proposed many video interpolation models based on deep learning.
3. The video processing method based on artificial intelligence according to claim 1, characterized in that: In the step 2, a convolutional neural network is used to extract features from each video picture in the video picture set, and by analyzing the change amplitude and change pattern of pixel points between adjacent frames, it is determined whether the video picture is a still picture or a moving picture.
4. The video processing method based on artificial intelligence according to claim 1, characterized in that: In the step three, when splicing still image areas of multiple video pictures, the still image areas are accurately segmented using an image segmentation algorithm, and then adjacent and content-related still image areas are seamlessly spliced according to the continuity of the image content and the spatial position relationship to form an independent still image.
5. The video processing method based on artificial intelligence according to claim 1, characterized in that: In the step 4, when the independent static image is matched with the moving image, the positional relationship and time sequence between the object in the moving image and the corresponding object in the independent static image are determined by analyzing the moving trajectory and moving direction of the object in the moving image.
6. The video processing method based on artificial intelligence according to claim 2, characterized in that: The convolutional neural network deep learning model adopts a multi-scale feature extraction strategy when learning the spatiotemporal features in the video, and performs feature analysis on the video frames at different scales to capture the spatiotemporal information at different levels in the video. For example, for a video containing a complex scene and multiple moving objects, the model extracts the overall motion trend of the scene and the approximate motion direction of the object at a coarse scale, and extracts the detailed motion features of the object at a fine scale, such as the local deformation of the object, texture changes, etc., so as to generate more realistic interpolated frames, so that the video can maintain a smooth visual effect even with a low frame rate.
7. A video processing device based on artificial intelligence, characterized in that: The device comprises: Video acquisition and conversion module: responsible for acquiring the video files to be processed, supporting the input of multiple video formats. It can decompose the video into a series of video pictures according to the set frame rate to form a video picture set; Video image recognition and classification module: Using advanced artificial intelligence technology, especially convolutional neural network, it conducts in-depth analysis on each image in the video image collection. It can identify the changes in the pixels in the image and accurately classify the video images into still images and moving images according to the magnitude and pattern of the changes. Image correspondence establishment module: The main task is to establish the correspondence between the static images and the moving images in the independent static image set. It determines the positional relationship and time sequence between the static images and the corresponding objects in the moving images by analyzing the motion trajectory, motion direction, and speed change of the objects in the moving images. Video synthesis and editing module: This module plays a key role in video synthesis, editing or analysis. For motion pictures, it can convert them into motion videos. Through advanced video processing technology, the motion pictures are smooth and natural. When performing video acquisition, this module will synthesize independent still picture sets into still videos in chronological order. Control and management module: As the center of the entire device, this module is responsible for coordinating, controlling and managing various functional modules. It rationally allocates the workflow and parameter settings of each module according to the user's input instructions and video processing requirements.
8. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the artificial intelligence-based video processing method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the artificial intelligence-based video processing method according to any one of claims 1 to 6 is implemented.