Video Image Transmission Method, Device, Equipment and Storage Medium

By decomposing video frames and extracting hierarchical similar features, the method addresses the inflexibility of fixed video compression algorithms, achieving efficient storage and transmission of diverse video content.

CN118338010BActive Publication Date: 2025-07-15CORE MICRO (SHENZHEN) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410515512.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-26
Publication Date
2025-07-15
Estimated Expiration
2044-04-26

AI Technical Summary

Technical Problem

Existing video compression methods rely on fixed algorithms and are difficult to adapt to diverse video content, resulting in waste of storage space and transmission bandwidth.

Method used

By splitting and arranging the video frames, the maximum similar features of the video image sequence are extracted, the video image construction model is constructed, and multiple levels of image splitting and encoding are performed.

Benefits of technology

Reduces video data redundancy, improves storage efficiency and transmission efficiency, and supports the reconstruction and flexible processing of high-quality videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118338010B_ABST
    Figure CN118338010B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of video processing, and discloses a method, device, equipment and storage medium for transmitting video images. The present invention splits a target video to obtain a video image sequence corresponding to the video, and circularly extracts the maximum similarity features of the video image sequence to obtain a maximum similarity feature sequence of the video image sequence. Then, based on the maximum similarity feature sequence, the video image sequence is split to obtain a video image construction model of the video image sequence. By performing corresponding encoding according to the video image construction model, an image encoding model of the target video can be obtained, solving the problem that in the prior art, video compression methods rely on fixed algorithms and are difficult to adapt to diverse videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly to a method, apparatus, device and storage medium for transmitting video images. Background Art

[0002] With the improvement of video resolution and the increase of video data volume, how to reduce the storage space and transmission bandwidth required for video data while ensuring video quality has become an urgent problem to be solved.

[0003] Currently, traditional video compression and encoding technologies often rely on some fixed algorithms and parameters, and it is difficult to adapt to the diversity of video content. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, apparatus, device and storage medium for transmitting video images, aiming to solve the problem that the video compression method in the prior art relies on fixed algorithms and is difficult to adapt to diverse videos.

[0005] The present invention is implemented as follows. In the first aspect, the present invention provides a method for transmitting video images, including:

[0006] Splitting and arranging the target video by frames to obtain a video image sequence of the target video; wherein, the video image sequence includes a plurality of video images arranged in chronological order;

[0007] Performing hierarchical cyclic extraction processing of maximum similarity features on adjacent video images in the video image sequence to obtain a maximum similarity feature sequence of the video image sequence; wherein, the maximum similarity feature sequence includes a plurality of sequence levels, the lowest sequence level is used to describe the maximum similarity features of adjacent video images in the video image sequence, and higher sequence levels are used to describe the maximum similarity features of adjacent maximum similarity features in lower sequence levels;

[0008] Performing multi-level image splitting processing on each video image in the video image sequence according to the maximum similarity feature sequence of the video image sequence to obtain a video image construction model of the video image sequence; wherein, the video image construction model is used to perform construction analysis on each video image in the video image sequence;

[0009] Performing image encoding on the target video according to the video image construction model to obtain an image encoding model of the target video.

[0010] Preferably, the step of splitting and arranging the target video by frames to obtain a video image sequence of the target video includes:

[0011] Perform frame-by-frame reading processing on the target video to sequentially obtain the image information corresponding to each frame number of the target video;

[0012] Generate sequence tags for the image information corresponding to each frame number of the target video according to the time relationship of the image information corresponding to each frame number of the target video relative to the target video, and perform binding processing on each of the sequence tags with the corresponding image information to obtain each of the video images;

[0013] Sort each of the video images according to the sequence tags of each of the video images to obtain the video image sequence of the target video.

[0014] Preferably, the step of performing hierarchical cyclic extraction processing of the maximum similarity features on adjacent video images in the video image sequence to obtain the maximum similarity feature sequence of the video image sequence includes:

[0015] Perform extraction processing of the maximum similarity features on adjacent video images in the video image sequence to obtain the maximum similarity features of adjacent video images in the video image sequence, and arrange the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence;

[0016] Perform extraction processing of the maximum similarity features on each adjacent maximum similarity feature in the lowest level of the maximum similarity feature sequence to obtain the maximum similarity features of adjacent maximum similarity features in the maximum similarity feature sequence, and arrange the maximum similarity features of adjacent maximum similarity features in the maximum similarity feature sequence to obtain a higher sequence level of the maximum similarity feature sequence;

[0017] Repeatedly and cyclically perform extraction processing of the maximum similarity features on the maximum similarity feature sequence to obtain several higher sequence levels of the maximum similarity feature sequence;

[0018] Sort each of the sequence levels according to the high-low relationship of each of the sequence levels to obtain the maximum similarity feature sequence.

[0019] Preferably, the step of performing extraction processing of the maximum similarity features on adjacent video images in the video image sequence to obtain the maximum similarity features of adjacent video images in the video image sequence, and arranging the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence includes:

[0020] Performing extraction processing of the maximum similarity features on adjacent video images in the video image sequence through a pre-trained similarity recognition AI model to obtain the maximum similarity features of adjacent video images in the video image sequence;

[0021] Performing arrangement processing on the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence.

[0022] Preferably, the training steps of the similarity recognition AI model include:

[0023] Constructing an input layer, a convolutional layer, and three fully connected layers;

[0024] Collecting several groups of similar video images and substituting each group of the similar video images into the input layer;

[0025] The input layer receives each group of the collected similar video images and transmits each group of the similar video images to the convolutional layer, and the convolutional layer is used for feature acquisition of each group of the similar video images to obtain the maximum similarity features of each group of the similar video images;

[0026] The three fully connected layers are used for performing continuous vector flattening processing on each of the maximum similarity features extracted by the convolutional layer to flatten the maximum similarity features into one-dimensional vector features; the one-dimensional vector features are the basic graphical expressions of the maximum similarity features.

[0027] Preferably, the steps of obtaining the video image construction model of the video image sequence by performing multi-level image splitting processing on each video image of the video image sequence according to the maximum similarity feature sequence of the video image sequence include:

[0028] Constructing corresponding similarity feature frameworks respectively according to each sequence level of the video image sequence;

[0029] Performing difference analysis processing on the similarity feature framework corresponding to a lower sequence level according to the similarity feature framework corresponding to the sequence level to obtain the difference feature distribution of the similarity feature framework corresponding to the lower sequence level; wherein, the difference feature distribution includes several difference features, and each of the difference features is combined with the similarity feature framework to obtain the similarity feature framework corresponding to the lower sequence level;

[0030] The repeated loop performs differential analysis processing on each of the similar feature frameworks to obtain the differential feature distributions corresponding to each of the sequence levels, combines each of the differential feature distributions with the similar feature framework to obtain each video construction level, and combines each of the video construction levels to obtain the video image construction model.

[0031] In a second aspect, the present invention provides a video image transmission device, including:

[0032] A video splitting module, configured to split and arrange a target video by frames to obtain a video image sequence of the target video; wherein, the video image sequence includes a plurality of video images arranged in chronological order;

[0033] A feature extraction module, configured to perform hierarchical cyclic extraction processing of the maximum similarity features on adjacent video images in the video image sequence to obtain a maximum similarity feature sequence of the video image sequence; wherein, the maximum similarity feature sequence includes a plurality of sequence levels, the lowest sequence level is used to describe the maximum similarity features of adjacent video images in the video image sequence, and higher sequence levels are used to describe the maximum similarity features of adjacent maximum similarity features in lower sequence levels;

[0034] An image splitting module, configured to perform multi-level image splitting processing on each of the video images in the video image sequence according to the maximum similarity feature sequence of the video image sequence to obtain a video image construction model of the video image sequence; wherein, the video image construction model is used to perform construction analysis on each of the video images in the video image sequence;

[0035] An image encoding module, configured to perform image encoding on the target video according to the video image construction model to obtain an image encoding model of the target video.

[0036] In a third aspect, the present invention provides a computer device, including a memory and a processor, the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements the video image transmission method according to any one of the first aspects.

[0037] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the processor is caused to execute the video image transmission method according to any one of the first aspects.

[0038] The present invention provides a video image transmission method, which has the following beneficial effects:

[0039] The present invention solves the problem in the prior art that video compression methods rely on fixed algorithms and are difficult to adapt to diverse videos. By splitting the target video, a video image sequence corresponding to the video is obtained, and the maximum similarity features of the video image sequence are cyclically extracted to obtain the maximum similarity feature sequence of the video image sequence. Based on the maximum similarity feature sequence, the video image sequence is split to obtain the video image construction model of the video image sequence. By performing corresponding encoding according to the video image construction model, the image encoding model of the target video can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a schematic diagram of the steps of a method for transmitting video images provided by an embodiment of the present invention;

[0041] Figure 2 is a schematic diagram of the structure of a device for transmitting video images provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0043] The implementation of the present invention will be described in detail below with reference to specific embodiments.

[0044] Refer to Figure 1 and Figure 2 shown, which are preferred embodiments provided by the present invention.

[0045] In a first aspect, the present invention provides a method for transmitting video images, including:

[0046] S1: Splitting and arranging the target video by frames to obtain the video image sequence of the target video; wherein, the video image sequence includes a plurality of video images arranged in chronological order;

[0047] S2: Performing hierarchical cyclic extraction processing on the adjacent video images in the video image sequence to obtain the maximum similarity feature sequence of the video image sequence; wherein, the maximum similarity feature sequence includes a plurality of sequence levels, and the lowest sequence level is used to describe the maximum similarity features of the adjacent video images in the video image sequence, and the higher sequence levels are used to describe the maximum similarity features of the adjacent maximum similarity features in the lower sequence levels;

[0048] S3: Perform a multi - level image splitting process on each of the video images in the video image sequence according to the maximum similarity feature sequence of the video image sequence, to obtain a video image construction model of the video image sequence; wherein, the video image construction model is used to perform construction analysis on each of the video images in the video image sequence.

[0049] S4: Perform image encoding on the target video according to the video image construction model to obtain an image encoding model of the target video.

[0050] Specifically, in step S1 of the embodiment provided by the present invention, a target video file is decoded using a video processing tool or library (such as FFmpeg, OpenCV, etc.). During the decoding process, the data in the video file is converted into a series of independent frames, and each frame is a separate image.

[0051] More specifically, the decoded video data is read frame by frame. Each time a frame is read, it is processed as an independent image. In the time order of the original video, the extracted frames are arranged in sequence to ensure the correct order of each frame, so as to preserve the temporal coherence and narrative logic of the video.

[0052] More specifically, the extracted frames can be stored on the disk in the form of image files, such as JPEG or PNG format, or the frame data can be directly used for subsequent video processing or analysis tasks without physical storage.

[0053] It can be understood that splitting the video into a single - frame image sequence enables direct access to, viewing, and analysis of specific moments and content in the video, facilitating frame - by - frame video content analysis and processing.

[0054] Specifically, starting from the starting frame of the video image sequence, adjacent video images are compared pairwise to identify the maximum similarity features between them. These features can be color distribution, texture pattern, edge information, etc. For each pair of adjacent images in the sequence, an image processing algorithm (such as a feature matching algorithm) is used to identify and extract the maximum similarity features between them.

[0055] More specifically, the identified maximum similarity features are organized into a sequence to form the lowest - level similarity feature sequence, and each element represents the maximum similarity feature between a pair of adjacent video images.

[0056] More specifically, on the basis of the lowest level, further hierarchical cyclic extraction is performed on the similarity feature sequence. That is, the similarity between adjacent features in the lowest - level sequence is compared and analyzed to form higher - level similarity features, and this process is repeated until no higher - level similarity features can be extracted.

[0057] Through the above steps, a sequence of similar features with several levels is finally formed, and each level describes different levels of image similarity information.

[0058] It can be understood that by identifying and extracting similar features in the video sequence and organizing them hierarchically, the data redundancy can be greatly reduced, which is very effective for video compression and storage optimization. During the video transmission process, the coding based on the hierarchical similar feature sequence can effectively reduce the amount of data to be transmitted. Especially in an environment with limited bandwidth, it can improve the smoothness of the video stream.

[0059] Specifically, in step S3 of the embodiment provided by the present invention, using the previously obtained maximum similar feature sequence, the similarity between each frame and its adjacent frames in the video image sequence is analyzed, and the common features and their changing parts in these frames are identified.

[0060] More specifically, according to the hierarchical structure of the similar features, the video frames are split. Each level corresponds to different degrees of similarity and change. At the lowest level, the split images only reflect minor changes; at higher levels, they reflect scene transitions or significant action changes.

[0061] More specifically, the images obtained through hierarchical splitting processing are constructed into a multi-level construction model. This model contains the detailed structural information of the video sequence, and this model can describe how each image in the video is constructed from a set of base images and changing images.

[0062] More specifically, the video image construction model is used to deeply analyze the video image sequence, which includes identifying the key frames, changing frames in the video and their relationships. During the analysis process, the features of the video content, such as dynamic changes, scene transitions, etc., can be further extracted.

[0063] It can be understood that by removing duplicate or non-critical information and compressing the video data, the storage space requirement is reduced and the transmission efficiency is improved. The construction model provides a structured representation of the video content, which is convenient for quickly locating and retrieving specific scenes or events in the video.

[0064] Specifically, in step S4 of the embodiment provided by the present invention, the video image construction model is analyzed to identify the base image elements and changing elements that make up the video. The base image elements refer to the image parts that repeatedly appear in the video sequence, while the changing elements refer to the parts that change relative to the base elements. For each frame of image, a difference image is constructed according to its difference from the base image elements. The difference image only contains the parts that have changed compared with the base image, which can greatly reduce the data volume of each frame of image.

[0065] More specifically, the encoded basic image elements and the differential images are combined into a final image coding model, which contains all the information required to reconstruct the original video, including the basic image elements, the differential images, and their position information in the video sequence.

[0066] It can be understood that by only storing the basic image elements and the differential information instead of the complete images of each frame, the storage size of the video can be significantly reduced, achieving efficient video compression. Moreover, by precisely recording the basic image elements and the changed parts, a high-quality video can be reconstructed during decoding, maximizing the preservation of the visual quality of the original video. At the same time, the structured characteristics of the coding model make the editing and processing of the video (such as cropping, scaling, adding special effects, etc.) more efficient and flexible. Additionally, during the decoding process, since the amount of data to be processed is reduced, the decoding efficiency of the video is improved, especially on devices with limited computing resources.

[0067] The present invention provides a method for transmitting video images, which has the following beneficial effects:

[0068] The present invention splits the target video to obtain a video image sequence corresponding to the video, and circularly extracts the maximum similarity features of the video image sequence to obtain a maximum similarity feature sequence of the video image sequence. Then, based on the maximum similarity feature sequence, the video image sequence is split to obtain a video image construction model of the video image sequence. By performing corresponding encoding according to the video image construction model, an image coding model of the target video can be obtained, solving the problem that the video compression method in the prior art depends on a fixed algorithm and is difficult to adapt to diverse videos.

[0069] Preferably, the steps of splitting and arranging the target video according to the number of frames to obtain the video image sequence of the target video include:

[0070] S11: Perform frame-by-frame reading processing on the target video to sequentially obtain the image information corresponding to each frame number of the target video;

[0071] S12: Generate sequence tags for the image information corresponding to each frame number of the target video according to the time relationship of the image information corresponding to each frame number of the target video with respect to the target video, and perform binding processing on each sequence tag with the corresponding image information to obtain each video image;

[0072] S13: Sort each video image according to the sequence tag of each video image to obtain the video image sequence of the target video.

[0073] Specifically, use video processing tools (such as OpenCV, FFmpeg, etc.) to open the target video file, read it frame by frame, and obtain the image information of each frame.

[0074] More specifically, for each image information read frame by frame, generate sequential tags according to its chronological order in the video. These sequential tags can be simple serial numbers (such as 1, 2, 3, …) or more complex timestamps (such as 00:00:01, 00:00:02, …).

[0075] More specifically, bind each image information to its corresponding sequential tag. This step can be achieved by storing the image information and the sequential tag in the same data structure. For example, a list where each element is a tuple or dictionary containing the image information and the sequential tag.

[0076] More specifically, sort all the images according to the sequential tags of the image information. This ensures that the video image sequence is consistent with the frame order in the original video. The sorted set of image information constitutes the video image sequence of the target video, and this sequence reflects the complete content and time flow of the video.

[0077] It can be understood that ensuring the chronological order of the image sequence is exactly the same as that of the original video is very important for video analysis, editing, and other processing. Converting the video content into an ordered image sequence facilitates frame-based content management and retrieval. The ordered video image sequence provides a basis for subsequent video processing tasks.

[0078] Preferably, the step of performing a hierarchical cyclic extraction process of the maximum similarity features on adjacent video images in the video image sequence to obtain the maximum similarity feature sequence of the video image sequence includes:

[0079] S21: Perform a maximum similarity feature extraction process on adjacent video images in the video image sequence to obtain the maximum similarity features of adjacent video images in the video image sequence. Arrange the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence;

[0080] S22: Perform a maximum similarity feature extraction process on each pair of adjacent maximum similarity features in the lowest level of the maximum similarity feature sequence to obtain the maximum similarity features of adjacent maximum similarity features in the maximum similarity feature sequence. Arrange the maximum similarity features of adjacent maximum similarity features in the maximum similarity feature sequence to obtain a higher sequence level of the maximum similarity feature sequence;

[0081] S23: Repeatedly loop to perform the extraction process of the maximum similarity features on the maximum similarity feature sequence to obtain several higher sequence levels of the maximum similarity feature sequence;

[0082] S24: Sort the respective sequence levels according to the high-low relationship of each sequence level to obtain the maximum similarity feature sequence.

[0083] Specifically, analyze adjacent video images in the video image sequence and extract the maximum similarity features between them. These features include image attributes such as color distribution, texture, and shape.

[0084] More specifically, use image processing and analysis algorithms to assist in extracting these features, and arrange the maximum similarity features of adjacent video images extracted in the order they appear in the video sequence to form the lowest level of the maximum similarity feature sequence.

[0085] More specifically, for the lowest-level similarity feature sequence, continue to analyze the similarity between adjacent features and extract higher-level maximum similarity features. This process involves identifying and extracting the similarity at a more abstract level between features. This step can adopt more advanced data analysis and pattern recognition techniques, such as clustering analysis, deep learning, etc.

[0086] More specifically, according to the higher-level similarity features extracted, arrange them according to their logical relationship in the original sequence to form a new level, and repeat this process until no higher-level similarity features can be extracted. Each iteration generates a higher-level similarity feature sequence.

[0087] More specifically, sort and organize according to the high-low relationship of each sequence level and integrate them into a complete maximum similarity feature sequence, which contains multiple similarity levels from the most specific to the most abstract.

[0088] It can be understood that by hierarchically analyzing the similarity features of the video image sequence, the details and overall structure of the video content can be understood more deeply, providing a new dimension for video content analysis and understanding; by using the similarity between video images, more efficient video coding and compression strategies can be designed to reduce redundant data and lower the costs of storage and transmission.

[0089] Preferably, the steps of performing the extraction process of the maximum similarity features on adjacent video images in the video image sequence to obtain the maximum similarity features of adjacent video images in the video image sequence, and arranging the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence include:

[0090] S221: Extract the maximum similarity features of adjacent video images in the video image sequence through a pre-trained similarity recognition AI model, and obtain the maximum similarity features of adjacent video images in the video image sequence;

[0091] S222: Arrange the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence.

[0092] Specifically, a large amount of labeled video data is used to train the AI model to enable it to learn to recognize and extract the similarity features between video images. During the training process, deep learning techniques such as convolutional neural networks (CNNs) can be used.

[0093] More specifically, input adjacent video images in the video image sequence into the trained AI model. The model will analyze and identify the maximum similarity features between these images, such as patterns, textures, colors, etc.

[0094] More specifically, according to the maximum similarity features extracted by the model, arrange these features in chronological order in the video sequence to form the lowest level of the maximum similarity feature sequence.

[0095] It can be understood that the pre-trained AI model can more accurately identify the subtle similarities between video images, thereby improving the overall accuracy of video analysis. Compared with traditional manual feature extraction methods, using the AI model can quickly process a large amount of video data, significantly improving the processing speed. The AI model can automatically extract similarity features, reducing the need for manual participation and realizing the automation of the video analysis process. Using deep learning models can not only identify basic image features but also understand the interactions between more complex scenes and objects, thus achieving a deep understanding of video content.

[0096] It should be noted that this step is an example, and a similar method is also adopted for the cyclic extraction of each subsequent maximum similarity feature.

[0097] Preferably, the training steps of the similarity recognition AI model include:

[0098] S2211: Construct an input layer, a convolutional layer, and three fully connected layers;

[0099] S2212: Collect several groups of similar video images and substitute each group of the similar video images into the input layer;

[0100] S2213: The input layer receives each group of the collected similar video images and transmits each group of the similar video images to the convolutional layer, and the convolutional layer is used to perform feature collection on each group of the similar video images to obtain the maximum similarity features of each group of the similar video images;

[0101] S2214: The three fully connected layers are used to perform continuous vector flattening processing on each of the maximum similarity features extracted by the convolutional layer to flatten the maximum similarity features into one-dimensional vector features; the one-dimensional vector features are the basic graphical expressions of the maximum similarity features.

[0102] Specifically, construct the network structure:

[0103] Input layer: Designed to receive the original pixel values of video images, and these images serve as the input to the network.

[0104] Convolutional layer: Following closely after the input layer, it is used to automatically extract features in the image. Through convolutional operations, local features of the image, such as edges and textures, can be captured.

[0105] Fully connected layer: Usually placed at the end of the network. After the feature maps extracted by the convolutional layer are flattened, they are input into the fully connected layer. Here, three fully connected layers are designed to further process and integrate the features, and finally output one-dimensional vector features.

[0106] More specifically, collect several groups of similar video images as training data, and these images need to be preprocessed, such as resizing and normalizing, to meet the requirements of the input layer.

[0107] More specifically, input the preprocessed video image dataset into the convolutional layer, and the network will automatically extract features from these images. The convolutional layer extracts the maximum similarity features of the images by applying multiple filters.

[0108] More specifically, after being processed by the convolutional layer and possibly the pooling layer, the obtained feature maps will enter the fully connected layer. Before that, the feature maps need to be flattened into one-dimensional vectors, and the three fully connected layers further process these one-dimensional vectors to form the final graphical expression, which reflects the maximum similarity features of the video images.

[0109] It can be understood that using a convolutional neural network to automatically extract the most important features from video images reduces the need for manual feature design and improves the analysis efficiency. Through multi-layer processing of the convolutional layer and the fully connected layer, the network can learn feature representations from low-level to high-level, providing a basis for in-depth understanding of video content, and converting the similar features of the images into one-dimensional vector form, which is convenient for subsequent machine learning and data analysis tasks.

[0110] Preferably, the steps of performing multi - level image splitting processing on each of the video images in the video image sequence according to the maximum similarity feature sequence of the video image sequence to obtain the video image construction model of the video image sequence include:

[0111] S31: Construct corresponding similarity feature frameworks respectively according to each of the sequence levels of the video image sequence;

[0112] S32: Perform difference analysis processing on the similarity feature framework corresponding to a lower sequence level according to the similarity feature framework corresponding to the sequence level, to obtain the difference feature distribution of the similarity feature framework corresponding to the lower sequence level; wherein, the difference feature distribution includes a number of difference features, and each of the difference features is combined with the similarity feature framework to obtain the similarity feature framework corresponding to the lower sequence level;

[0113] S33: Repeatedly perform difference analysis processing on each of the similarity feature frameworks to obtain the difference feature distributions corresponding to each of the sequence levels, and combine each of the difference feature distributions with the similarity feature framework to obtain each video construction level, and combine each of the video construction levels to obtain the video image construction model.

[0114] Specifically, for each sequence level of the video image sequence, corresponding similarity feature frameworks are respectively constructed, and these frameworks describe the similarity relationship between images based on similarity features in the sequence, such as color, shape, texture, etc.

[0115] More specifically, perform difference analysis on each similarity feature framework, comparing the differences between similarity feature frameworks at different levels. This step aims to identify the unique or different features between feature frameworks at each level.

[0116] More specifically, based on the results of the difference analysis, obtain the difference feature distribution corresponding to a lower level, which includes a set of difference features, and these features reflect the differences between similarity feature frameworks at different levels.

[0117] More specifically, combine the difference feature distribution with the corresponding similarity feature framework, that is, the difference feature distribution and the similarity feature framework together constitute the similarity feature framework of the lower level.

[0118] More specifically, repeatedly perform difference analysis and combination processing on all similarity feature frameworks until all sequence levels are covered, so as to obtain the difference feature distribution of each level and gradually construct a complete video image construction model.

[0119] More specifically, by combining all video construction levels, a comprehensive video image construction model is formed, which deeply reflects the hierarchical structure and similarity distribution of video data.

[0120] It can be understood that the constructed video image model optimizes the representation of video data by finely describing the levels and differences of video content, which helps to improve the efficiency of video processing and analysis. The video image construction model can guide the design of efficient video coding and compression algorithms by encoding only key differential features to reduce the data volume and increase the compression ratio.

[0121] Referring to Figure 2 As shown, in a second aspect, the present invention provides a video image transmission device, comprising:

[0122] A video splitting module, configured to split and arrange a target video by frames to obtain a video image sequence of the target video; wherein, the video image sequence includes a plurality of video images arranged in chronological order;

[0123] A feature extraction module, configured to perform a hierarchical cyclic extraction process on adjacent video images in the video image sequence to obtain a maximum similarity feature sequence of the video image sequence; wherein, the maximum similarity feature sequence includes a plurality of sequence levels, and the lowest sequence level is used to describe the maximum similarity features of adjacent video images in the video image sequence, and higher sequence levels are used to describe the maximum similarity features of adjacent maximum similarity features in lower sequence levels;

[0124] An image splitting module, configured to perform a multi-level image splitting process on each video image of the video image sequence according to the maximum similarity feature sequence of the video image sequence to obtain a video image construction model of the video image sequence; wherein, the video image construction model is used to perform construction analysis on each video image of the video image sequence;

[0125] An image encoding module, configured to perform image encoding on the target video according to the video image construction model to obtain an image encoding model of the target video.

[0126] In this embodiment, for the specific implementation of each module in the above device embodiment, please refer to that described in the above method embodiment, and details will not be repeated here.

[0127] In a third aspect, the present invention provides a computer device, comprising a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the computer program, it implements a video image transmission method according to any one of the first aspects.

[0128] In a fourth aspect, the present invention provides a computer-readable storage medium having stored thereon a computer program which, when run by a processor, causes the processor to execute a method for transmitting a video image according to any one of the first aspects.

[0129] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for transmitting video images, characterized in that, Including: Splitting and arranging the target video by frame number to obtain a video image sequence of the target video; wherein, the video image sequence includes a number of video images arranged in chronological order; Performing a hierarchical cyclic extraction process of the maximum similarity features on adjacent video images in the video image sequence to obtain a maximum similarity feature sequence of the video image sequence; wherein, the maximum similarity feature sequence includes a number of sequence levels, and the lowest sequence level is used to describe the maximum similarity features of adjacent video images in the video image sequence, and higher sequence levels are used to describe the maximum similarity features of adjacent maximum similarity features in lower sequence levels; Performing a multi-level image splitting process on each video image in the video image sequence according to the maximum similarity feature sequence of the video image sequence to obtain a video image construction model of the video image sequence; wherein, the video image construction model is used to perform construction analysis on each video image in the video image sequence; Performing image encoding on the target video according to the video image construction model to obtain an image encoding model of the target video; The step of performing a multi-level image splitting process on each video image in the video image sequence according to the maximum similarity feature sequence of the video image sequence to obtain a video image construction model of the video image sequence includes: Constructing corresponding similarity feature frameworks respectively according to each sequence level of the video image sequence; Performing a difference analysis process on the similarity feature framework corresponding to a lower sequence level according to the similarity feature framework corresponding to the sequence level to obtain a difference feature distribution of the similarity feature framework corresponding to the lower sequence level; wherein, the difference feature distribution includes a number of difference features, and each difference feature is combined with the similarity feature framework to obtain the similarity feature framework corresponding to the lower sequence level; Repeatedly performing a difference analysis process on each similarity feature framework to obtain the difference feature distribution corresponding to each sequence level, and combining each difference feature distribution with the similarity feature framework to obtain each video construction level, and combining each video construction level to obtain the video image construction model.

2. The transmission method of a video image according to claim 1, characterized in that, The step of splitting and arranging the target video by frame number to obtain a video image sequence of the target video includes: Performing a frame-by-frame reading process on the target video to sequentially obtain image information corresponding to each frame number of the target video; Generating a sequence tag for the image information corresponding to each frame number of the target video according to the time relationship of the image information corresponding to each frame number of the target video with respect to the target video, and binding each sequence tag to the corresponding image information respectively to obtain each video image; Sorting each video image according to the sequence tag of each video image to obtain a video image sequence of the target video.

3. The transmission method of a video image according to claim 1, wherein, The steps of performing a hierarchical cyclic extraction process of the maximum similarity features on adjacent video images in the video image sequence to obtain the maximum similarity feature sequence of the video image sequence include: Performing an extraction process of the maximum similarity features on adjacent video images in the video image sequence to obtain the maximum similarity features of adjacent video images in the video image sequence, and arranging the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence; Performing an extraction process of the maximum similarity features on each pair of adjacent maximum similarity features in the lowest level of the maximum similarity feature sequence to obtain the maximum similarity features of adjacent maximum similarity features in the maximum similarity feature sequence, and arranging the maximum similarity features of adjacent maximum similarity features in the maximum similarity feature sequence to obtain a higher sequence level of the maximum similarity feature sequence; Repeatedly and cyclically performing an extraction process of the maximum similarity features on the maximum similarity feature sequence to obtain several higher sequence levels of the maximum similarity feature sequence; Sorting each sequence level according to the high-low relationship of each sequence level to obtain the maximum similarity feature sequence.

4. The transmission method of a video image according to claim 3, characterized in that, The steps of performing an extraction process of the maximum similarity features on adjacent video images in the video image sequence to obtain the maximum similarity features of adjacent video images in the video image sequence, and arranging the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence include: Performing an extraction process of the maximum similarity features on adjacent video images in the video image sequence through a pre-trained similarity recognition AI model to obtain the maximum similarity features of adjacent video images in the video image sequence; Arranging the maximum similarity features of adjacent video images in the video image sequence to obtain the lowest sequence level of the maximum similarity feature sequence.

5. The transmission method of a video image according to claim 4, wherein The training steps of the similarity recognition AI model include: Constructing an input layer, a convolutional layer, and three fully connected layers; Collecting several groups of similar video images and substituting each group of the similar video images into the input layer; The input layer receives each group of the collected similar video images and transmits each group of the similar video images to the convolutional layer, and the convolutional layer is used to perform feature acquisition on each group of the similar video images to obtain the maximum similarity features of each group of the similar video images; The three fully connected layers are used to perform continuous vector flattening processing on each of the maximum similarity features extracted by the convolutional layer to flatten the maximum similarity features into one-dimensional vector features; the one-dimensional vector features are the basic graphical expressions of the maximum similarity features.

6. A transmission device for video images, characterized in that, A method for transmitting a video image, which is used to implement any one of claims 1 to 5, includes: A video splitting module, configured to split and arrange a target video by frames to obtain a video image sequence of the target video; wherein, the video image sequence includes a plurality of video images arranged in chronological order; A feature extraction module, configured to perform a hierarchical cyclic extraction process of maximum similarity features on adjacent video images in the video image sequence to obtain a maximum similarity feature sequence of the video image sequence; wherein, the maximum similarity feature sequence includes a plurality of sequence levels, the lowest sequence level is used to describe the maximum similarity features of adjacent video images in the video image sequence, and higher sequence levels are used to describe the maximum similarity features of adjacent maximum similarity features in lower sequence levels; An image splitting module, configured to perform a multi-level image splitting process on each video image in the video image sequence according to the maximum similarity feature sequence of the video image sequence to obtain a video image construction model of the video image sequence; wherein, the video image construction model is used to perform construction analysis on each video image in the video image sequence; An image encoding module, configured to perform image encoding on the target video according to the video image construction model to obtain an image encoding model of the target video.

7. A computer device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the computer program, it implements the method for transmitting a video image according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is run by a processor, the processor is caused to execute a method for transmitting a video image according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method and device for video coding and decoding

    CN104168482A

  • Video coding method and device and electronic equipment

    CN114173137A