Video data compression method and device, electronic equipment and storage medium
In drone inspection, based on preset text and semantic segmentation technology, differentiated compression strategies are adopted for areas of interest and non-interest in video frames, which solves the problem of poor compression quality of drone inspection video data, and achieves efficient data compression and transmission.
Patent Information
- Application Number
- CN202510306634.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-11
AI Technical Summary
There is a problem of poor data quality in the compression of drone inspection video data. The existing inter-frame prediction encoding technology leads to high redundancy and the inability to accurately identify negligible image areas, affecting the compression rate and quality.
By acquiring the target video data, the first target frame and the second target frame are extracted based on the preset text, and semantic segmentation is performed, and different compression strategies are adopted for the first and second regions respectively, to preserve the quality of the region of interest and reduce the redundancy of non-areas of interest.
It improves the quality of video data compression, reduces the need for transmission and storage resources, ensures the integrity and real-time nature of key information, and improves the efficiency and accuracy of drone inspections.
Smart Images

Figure CN120302046A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of power systems and data processing, and in particular, to a video data compression method, device, electronic device, and storage medium. Background Art
[0002] With the progress of technology, the drone inspection technology has become an important means in the field of industrial inspection. By taking advantage of the flexible and mobile characteristics of drones and their ability to operate at high altitudes, it can reach areas that are difficult for personnel to access, such as high altitudes, high-voltage facilities, complex terrains, or dangerous environments, to perform inspection tasks, greatly improving the inspection efficiency and safety. However, the challenge faced by drone inspections is the large amount of data. For the transmission, processing, and storage of a large amount of data, not only a large amount of bandwidth and storage resources are required, but also the energy consumption and computational complexity are increased.
[0003] In the prior art, for video data compression, frame-interpolation coding techniques are usually adopted to reduce the redundancy of the encoded video by predicting the differences between different video frames. However, there is an inherent problem of loss of video data quality through this prediction method. Moreover, in the compression of video data, some negligible image regions may not be accurately identified, resulting in a reduction in the compression ratio of the video data, thereby affecting the quality of the video data.
[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present invention provide a video data compression method, device, electronic device, and storage medium to at least solve the technical problem of poor data quality caused by data compression in the prior art.
[0006] According to one aspect of the embodiments of the present invention, a video data compression method is provided, including: obtaining target video data; extracting the target video data based on a preset text to obtain a first target frame and a second target frame in the target video data, where the semantic information of the first target frame matches the preset text, and the second target frame is a video frame other than the first target frame in the target video data; respectively performing semantic segmentation on the first target frame and the second target frame to obtain a first region and a second region in different video frames, where the first region matches the preset text, and the second region is a region other than the first region in different video frames; and using different compression strategies to compress the first region and the second region in different video frames respectively to obtain a compression result.
[0007] Optionally, perform semantic segmentation on the first target frame and the second target frame respectively to obtain the first region and the second region in different video frames, including: performing semantic segmentation on the first target frame based on a semantic segmentation model to obtain a first semantic segmentation result of the first target frame, where the first semantic segmentation result is used to represent the position of a preset text in the first target frame; determining the first region and the second region in the first target frame based on the first semantic segmentation result; mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a second semantic segmentation result of the second target frame, where the second semantic segmentation result is used to represent the position of the preset text in the second target frame; determining the first region and the second region in the second target frame based on the second semantic segmentation result.
[0008] Optionally, mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a second semantic segmentation result of the second target frame, including: mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a mapped target video frame and a target semantic segmentation result; determining the peak signal-to-noise ratio between the second target frame and the mapped target video frame; in response to the peak signal-to-noise ratio being greater than a signal-to-noise ratio threshold, determining the target semantic segmentation result as the second semantic segmentation result.
[0009] Optionally, mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a mapped target video frame and a target semantic segmentation result, including: determining a mapping relationship between the first target frame and the second target frame based on an optical flow algorithm; mapping the first semantic segmentation result to the second target frame according to the mapping relationship to obtain a mapped target video frame and a target semantic segmentation result.
[0010] Optionally, determining the peak signal-to-noise ratio between the second target frame and the mapped target video frame, including: determining the mean squared error between the second target frame and the mapped target video frame; determining the maximum value of the pixels in the mapped target video frame; determining the peak signal-to-noise ratio based on the mean squared error and the maximum value.
[0011] Optionally, extracting the first target frame and the second target frame in the target video data based on a preset text, including: constructing a text feature corresponding to the preset text using a word segmentation method; encoding the text feature based on an encoding model to obtain a target encoding feature; extracting the first target frame in the target video data based on the target encoding feature; determining the other frames except the first target frame as the second target frame.
[0012] Optionally, encoding the text feature based on an encoding model to obtain a target encoding feature, including: encoding the text feature based on an encoding model to obtain an initial encoding feature; extracting the target encoding feature from the initial encoding feature based on a preset condition.
[0013] Optionally, extracting a first target frame from the target video data based on the target encoding feature includes: determining the similarity between the target encoding feature and any video frame in the target video data; and in response to the similarity being greater than or equal to a similarity threshold, determining the video frame as the first target frame.
[0014] Optionally, different compression strategies are used to compress the first region and the second region in different video frames respectively to obtain a compression result, including: compressing the first region at a first magnification or not compressing the first region to obtain a first sub-compression result; compressing the second region at a second magnification to obtain a second sub-compression result; and obtaining the compression result based on the first sub-compression result and the second sub-compression result.
[0015] Optionally, obtaining the target video data includes: obtaining initial target video data; and performing data optimization on the initial target video data to obtain the target video data, where the data optimization includes at least one of the following: data denoising, contrast adjustment, brightness adjustment, and color correction.
[0016] According to another aspect of the embodiments of the present invention, there is also provided a video data compression system, including: an acquisition module for acquiring target video data; an extraction module for extracting from the target video data based on a preset text to obtain a first target frame and a second target frame in the target video data, where the semantic information of the first target frame matches the preset text, and the second target frame is a video frame other than the first target frame in the target video data; a segmentation module for respectively performing semantic segmentation on the first target frame and the second target frame to obtain a first region and a second region in different video frames, where the first region matches the preset text, and the second region is a region other than the first region in different video frames; and a compression module for using different compression strategies to compress the first region and the second region in different video frames respectively to obtain a compression result.
[0017] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including: a memory storing an executable program; and a processor for running the program, where when the program runs, it executes the methods in the various embodiments of the present invention.
[0018] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in the various embodiments of the present invention.
[0019] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the methods in the various embodiments of the present invention.
[0020] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the methods in the various embodiments of the present invention are implemented.
[0021] According to another aspect of the embodiments of the present invention, there is also provided a computer program, and when the computer program is executed by a processor, the methods in the various embodiments of the present invention are implemented.
[0022] In the embodiments of the present invention, target video data is obtained; based on a preset text, the target video data is extracted to obtain a first target frame and a second target frame in the target video data, where the semantic information of the first target frame matches the preset text, and the second target frame is a video frame in the target video data other than the first target frame; semantic segmentation is respectively performed on the first target frame and the second target frame to obtain a first region and a second region in different video frames, where the first region matches the preset text, and the second region is a region in different video frames other than the first region; different compression strategies are respectively used to compress the first region and the second region in different video frames to obtain a compression result. It is easy to notice that by adopting the adaptive compression method of semantic segmentation, the first region and the second region are distinguished semantically, and then different compression strategies are applied to different regions, which can not only retain the data quality of the region of interest, but also reduce the redundant data in the non-region of interest, achieving the purpose of improving the compression quality of video data, thus realizing the technical effect of ensuring the data compression quality, and further solving the technical problem of poor data quality caused by data compression in the prior art. Description of the Drawings
[0023] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0024] Figure 1 is a flowchart of a video data compression method according to an embodiment of the present invention;
[0025] Figure 2 is a schematic diagram of an optional video data compression method according to an embodiment of the present invention;
[0026] Figure 3 is a schematic diagram of an optional video data compression method according to an embodiment of the present invention;
[0027] Figure 4 is a schematic diagram of a video data compression device according to an embodiment of the present invention. Detailed Embodiments
[0028] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0030] According to an embodiment of the present invention, an embodiment of a video data compression method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0031] Figure 1 is a flowchart of a video data compression method according to an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:
[0032] Step S102, obtain target video data.
[0033] The above-mentioned target video data can be used to characterize the operation status of power equipment and its surrounding environment. In the drone inspection task, the target video data includes, but is not limited to, video images of key monitoring points such as power facilities, transmission lines, substation equipment, wind turbine blades, and solar panels. The target video data not only contains static object information, but may also include dynamic environmental changes, equipment operation status, and possible abnormal situations, such as equipment damage, fire warning, illegal intrusion, etc. By obtaining the target video data, the power equipment can be inspected.
[0034] The target video data can be used to identify obstacles, terrain features, and potential hazards in the inspection environment in real time, such as wires, trees, buildings, etc., to ensure the safe flight of the drone and avoid collisions. The target video data can also be used to autonomously determine the inspection path, avoid dangerous areas, and adjust the flight strategy according to the detected target status to achieve more efficient inspection coverage. The target video data is the basis for subsequent inspection analysis. Through image processing and machine learning technologies, the operating status of the equipment can be automatically diagnosed, and equipment defects, such as cracks, rust, overheating, etc., can be identified. However, the amount of data collected by the video is huge, and a large amount of data is transmitted to the receiving end for processing and storage. This method not only requires a large amount of bandwidth and storage resources, but also increases energy consumption and computational complexity. Therefore, the target video data is further compressed to improve the inspection effect.
[0035] In an alternative embodiment, the target video data can be obtained through various sensors, such as cameras, thermal imagers, infrared cameras, etc. The camera can be used to capture images and videos in the visible spectrum, the thermal imager can detect the thermal distribution of the equipment, and the infrared camera can identify targets under night or adverse weather conditions.
[0036] In another alternative embodiment, the target video data can also be obtained from the cloud platform. The video data collected by the drone can be transmitted to the cloud platform through secure network protocols, such as real-time streaming protocol, real-time messaging protocol, secure and reliable transport protocol, etc. The cloud can use edge computing technology or centralized computing resources to perform real-time analysis on the target video data, including target detection, behavior recognition, environmental monitoring, etc.
[0037] Step S104, extract the target video data based on the preset text to obtain the first target frame and the second target frame in the target video data.
[0038] Among them, the semantic information of the first target frame matches the preset text, and the second target frame is the video frame in the target video data except the first target frame.
[0039] The above-mentioned preset text can be used to identify the data related to the inspection task in the target video data. The role of the preset text is to guide video semantic segmentation, help identify the content related to the inspection task in the target video data, so as to perform adaptive compression, reduce the amount of data, and improve the efficiency of transmission and processing. Through the preset text, unnecessary data processing and analysis can also be reduced, and the operating efficiency of the algorithm can be improved. At the same time, the accurate description of the preset text helps to improve the accuracy of the inspection task analysis and reduce false alarms and missed alarms.
[0040] In an alternative embodiment, the preset text can be analyzed through natural language processing techniques to extract keywords and phrases, which may include but are not limited to inspection targets, specific behaviors, environmental characteristics, etc. Then, using video summarization generation techniques, the target video data is scanned, and through keyword matching or semantic understanding, the first target frame and the second target frame containing the preset text are located. In this way, the target frames can be quickly located, improving the retrieval efficiency of the target video data.
[0041] In another alternative embodiment, a deep learning model can be used for video frame segmentation, and then based on the target categories defined by the preset text, an object detection model is used to identify the video frames. By combining semantic segmentation and object detection, the first target frame and the second target frame can be accurately extracted. Combining semantic segmentation and object detection can accurately identify multiple video frames in the target video data and record the information of each video frame. The deep learning model can be a convolutional neural network, an encoder, a decoder, etc., and there is no limitation on the deep learning model here and it can be determined according to needs. The object detection model can be the YOLO algorithm. The YOLO algorithm can be an algorithm for object detection. YOLO is a real-time object detection framework that performs segmentation by predicting the bounding boxes and classes of objects, is fast, and is suitable for scenarios such as real-time monitoring and drone inspections that require quick responses. There is no limitation on the object detection model here and it can be determined according to needs.
[0042] In yet another alternative embodiment, natural language video understanding techniques, such as a video question answering model, can be used to convert the preset text into a natural language query. The video question answering model locates the first target frame and the second target frame related to the query and the preset text by understanding the content and text description of the target video data. Natural language video query and location provide a high degree of flexibility and user-friendliness, and video location can be directly performed using natural language.
[0043] Step S106, perform semantic segmentation on the first target frame and the second target frame respectively to obtain the first region and the second region in different video frames.
[0044] Among them, the first region matches the preset text, and the second region is the region in different video frames except the first region.
[0045] The above-mentioned first region can be a region of interest, and the first region can be a region related to the inspection task. In the video of power facility inspection, the first region may be equipment such as transformers, utility poles, or wires that need to be inspected. The region of interest refers to a region with a specific task in an image or video. In the field of drone inspection, the region of interest can be power equipment that needs to be inspected, the defect location of the power system, etc. Select whether to perform compression or not according to the actual bandwidth and inspection requirements to preserve the details and quality of the region of interest. For non-region-of-interest areas, the data size can be significantly reduced by reducing the detailed information of non-region-of-interest areas. The identification of the region of interest can focus the data, improve the algorithm efficiency and data processing speed. When compressing, transmitting, or storing, preserve the details and quality of the region of interest to ensure that key information is not degraded or lost. The precise positioning of the region of interest can improve the accuracy and reliability of subsequent analysis and avoid the interference of irrelevant information.
[0046] The above-mentioned second region can be a non-region-of-interest area, and the second region can be a region unrelated to the inspection task, which can be the background environment, redundant content, etc. The non-region-of-interest area refers to the area in an image or video that is irrelevant to the current task or target. These areas may contain background information, repetitive textures, or noise, etc. When transmitting or storing, the non-region-of-interest area can be efficiently compressed or deleted to reduce the data volume, lower the bandwidth requirement, and storage cost. The identification and compression of the second region aim to reduce the data volume and the resource consumption of transmission and storage.
[0047] In an alternative embodiment, semantic segmentation can be performed based on a deep learning semantic segmentation network. Use a deep learning model to perform pixel-level classification on the target frame, and label each pixel as a specific category. In this way, the video frame can be accurately segmented into different regions. Among them, the region labeled as a specific target category is the first region, and the remaining regions labeled as the background or other non-target categories are the second regions. The deep learning model can be constructed based on pyramid pooling, or an encoder, etc. Through the deep learning model, complex feature representations can be automatically learned from a large amount of data, providing high-precision segmentation results. This method can effectively process video frames under complex backgrounds and different lighting conditions, improving the robustness and adaptability of segmentation.
[0048] In another alternative embodiment, semantic segmentation can be achieved by fusing text information. By combining preset text with semantic segmentation technology, the text is first encoded into vectors, and then these vectors are used to guide the semantic segmentation model to pay more attention to the areas related to the inspection task, so that objects or features consistent with the preset text description are considered during the segmentation process, thereby improving the recognition accuracy of the first area. The semantic segmentation that fuses text information can achieve task-driven segmentation, making the results more in line with specific inspection or monitoring requirements, reducing the interference of the second area, and improving the pertinence and efficiency of segmentation.
[0049] In yet another alternative embodiment, semantic segmentation can be performed based on optical flow and spatio-temporal consistency constraints. The optical flow algorithm can be used to analyze the motion information between consecutive video frames, and the segmentation result of the first target frame is propagated to the second target frame according to the optical flow field. At the same time, considering the continuity and consistency in time and space, the segmentation result during the propagation process is clarified. By combining semantic constraints and optical flow information, the accuracy of segmentation can be improved, especially when dealing with scenarios with fast motion or large illumination changes. This method can utilize the similarity between consecutive frames to reduce redundant calculations and improve segmentation efficiency. At the same time, through spatio-temporal consistency constraints, the segmentation errors caused by object motion and occlusion in the target frame can be reduced, and the stability of the recognition of the first area and the second area can be improved.
[0050] Step S108, different compression strategies are adopted to compress the first area and the second area in different video frames respectively to obtain compression results.
[0051] The above compression can be to compress the area to reduce the data volume and improve the data quality. The compression can be spatial compression or temporal compression. Spatial compression can reduce the bit rate of the video frame, usually by removing redundant information, for example, the similarity between adjacent pixels. Temporal compression can refer to using the similarity between adjacent frames in the video sequence to remove redundant inter-frame information.
[0052] By adopting different compression strategies to compress the first region and the second region in the video frame, the high quality of the key region can be retained while a higher compression ratio can be applied to the non-key region, significantly reducing the resource requirements for video transmission and storage while maintaining the integrity of the key information. It can also reduce the data volume of the non-key region, thereby reducing network transmission latency, storage space occupancy, and computing resource consumption. Especially for edge computing and mobile devices, such a compression strategy can improve the processing speed and response time. In scenarios that require real-time transmission such as drone inspections and remote monitoring, the adaptive compression strategy can ensure the real-time nature of information without sacrificing real-time transmission performance due to excessive data volume. And in an environment with limited bandwidth, using different compression ratios for different regions can more effectively utilize the available bandwidth, ensuring high-quality transmission of the key region while compressing or reducing the quality of the non-key region to save transmission resources.
[0053] The above compression results can be used to characterize the results related to the inspection task. The compression results can be video data, binary code streams, etc. The binary code stream usually consists of compressed video frame data, coding parameters, metadata, etc., and can be transmitted in the network efficiently and securely. The compression results can also include a quantitative comparison of the size of the original video data and the compressed data, i.e., the compression ratio. This helps to evaluate the efficiency of the compression algorithm and its impact on video quality. The compression results can also include encoding format information, compression parameters, the duration of the video, resolution, indexes of key frames, optical flow data, semantic segmentation labels, or text feature encodings, etc. Here, the compression results are not limited and can be determined according to needs. Based on the compression results, metrics such as peak signal-to-noise ratio and structural similarity index can also be used to evaluate the quality of the decoded video to ensure that key information is not lost during the compression and decompression processes.
[0054] In an alternative embodiment, two independent encoders are constructed. One encoder is used for the first region to ensure clarity, and the other encoder is used for the second region with a compression ratio greater than that of the first encoder to reduce the data volume. The encoder can be obtained based on standard video coding techniques or a deep learning coding model. The deep learning coding model can be a generative adversarial network coding model, an end-to-end deep learning coding model, an attention mechanism coding model, etc. Here, the deep learning coding model is not limited and can be determined according to needs.
[0055] In another alternative embodiment, motion compensation technology can be used to predict the motion of the first region between frames, reducing redundant coding, while using a high compression ratio coding for the second region. Motion compensation can accurately estimate the motion vector of an object, and then encode the motion changes without encoding the first region. The optical flow algorithm can be used for motion estimation, combined with a region-adaptive encoder, using a lower compression ratio for the first region and a high compression ratio for the second region.
[0056] In yet another alternative embodiment, compression can be achieved through quantization parameter adjustment. During the video encoding process, the setting of the quantization parameter determines the intensity and quality of video compression. An intelligent system can be designed to dynamically adjust the quantization parameters of different regions based on the dynamic analysis of video frames. For example, a quantization parameter smaller than a first preset threshold is used for the first region to maintain high quality, while a quantization parameter larger than a second preset threshold is used for the second region for compression. A quantization parameter prediction model based on deep learning can also be used, combined with the adaptive adjustment of the encoding quantization table, to achieve differential compression of the first region and the second region.
[0057] In the embodiment of the present invention, target video data is obtained; based on a preset text, the target video data is extracted to obtain a first target frame and a second target frame in the target video data, wherein the semantic information of the first target frame matches the preset text, and the second target frame is a video frame in the target video data other than the first target frame; semantic segmentation is respectively performed on the first target frame and the second target frame to obtain a first region and a second region in different video frames, wherein the first region matches the preset text, and the second region is a region in different video frames other than the first region; different compression strategies are respectively used to compress the first region and the second region in different video frames to obtain a compression result. It is easy to notice that by adopting the adaptive compression method of semantic segmentation, the first region and the second region are distinguished semantically, and then different compression strategies are applied to different regions, which can not only retain the data quality of the region of interest, but also reduce the redundant data of the non-region of interest, achieving the purpose of improving the video data compression quality, thus realizing the technical effect of ensuring the data compression quality, and further solving the technical problem of poor data quality caused by data compression in the prior art.
[0058] Optionally, semantic segmentation is respectively performed on the first target frame and the second target frame to obtain a first region and a second region in different video frames, including: performing semantic segmentation on the first target frame based on a semantic segmentation model to obtain a first semantic segmentation result of the first target frame, wherein the first semantic segmentation result is used to represent the position of the preset text in the first target frame; determining the first region and the second region in the first target frame based on the first semantic segmentation result; mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a second semantic segmentation result of the second target frame, wherein the second semantic segmentation result is used to represent the position of the preset text in the second target frame; determining the first region and the second region in the second target frame based on the second semantic segmentation result.
[0059] In an alternative embodiment, a semantic segmentation model can be constructed through a fully convolutional neural network, a dilated convolutional network, multi-scale analysis, a conditional random field, etc. to perform semantic segmentation on the first target frame, thereby obtaining a first semantic segmentation result of the first target frame to determine the position of the preset text in the first target frame. Then, the first region and the second region are divided through the first semantic segmentation result. Optionally, the first region and the second region can be divided by a probability threshold. Alternatively, connected component analysis can also be performed to analyze the categories in the first semantic segmentation result and identify connected regions composed of pixels of the same category. Generally, connected regions with a larger area or meeting specific shape rules are defined as the first region, while other smaller or irregular regions are regarded as the second region. Or, the first region and the second region can also be marked through contour detection and filling algorithms.
[0060] For example, DeepLabv3+ can also be used for semantic segmentation. DeepLabv3+ is a deep learning model for semantic segmentation. DeepLabv3+ is a semantic segmentation model based on dilated convolution and multi-scale feature fusion. This model uses dilated convolution in its structure to expand the receptive field and adopts residual connections to enhance the feature representation ability. Based on the real-time response requirements of drone patrol, replacing the dilated convolution in DeepLabv3+ with dilated convolution can not only expand the receptive field without increasing the computational amount, but also increase the resolution of video frame features while maintaining the receptive field range.
[0061] Furthermore, the first semantic segmentation result can be mapped onto the second target frame through a sparse optical flow algorithm, optical flow field estimation, or a dense optical flow algorithm to obtain a second semantic segmentation result of the second target frame to determine the position of the preset text in the second target frame. For example, the position of feature points in the second target frame can be estimated by using a sparse optical flow algorithm. Using these feature points and their displacement information, through triangulation or interpolation methods, the object boundaries or pixel classification information in the first semantic segmentation result are mapped onto the second target frame to obtain a second semantic segmentation result. The displacement vector of each pixel point between adjacent second target frames can be calculated through a dense optical flow algorithm. Based on these displacement vectors, a pixel-to-pixel mapping relationship can be constructed, thereby directly mapping the semantic segmentation result on the first target frame to the corresponding position on the second target frame to generate a second semantic segmentation result. So as to further divide the first region and the second region in the second target frame by a probability threshold. Alternatively, the first region and the second region in the second target frame can also be analyzed through connected component analysis.
[0062] Optionally, mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain the second semantic segmentation result of the second target frame, including: mapping the first semantic segmentation result to the second target frame based on the optical flow algorithm to obtain a mapped target video frame and a target semantic segmentation result; determining the peak signal-to-noise ratio between the second target frame and the mapped target video frame; in response to the peak signal-to-noise ratio being greater than the signal-to-noise ratio threshold, determining the target semantic segmentation result as the second semantic segmentation result.
[0063] In an alternative embodiment, the first semantic segmentation result can be mapped to the second target frame through a sparse optical flow algorithm, optical flow field estimation, or a dense optical flow algorithm to obtain a mapped target video frame and a target semantic segmentation result. Further, the peak signal-to-noise ratio between the second target frame and the mapped target video frame can be determined through a peak signal-to-noise ratio formula or a peak signal-to-noise ratio prediction algorithm. In the case where the peak signal-to-noise ratio is greater than the signal-to-noise ratio threshold, the target semantic segmentation result is determined as the second semantic segmentation result. The peak signal-to-noise ratio prediction algorithm can be a support vector machine, linear regression, etc., and there is no limitation on the peak signal-to-noise ratio prediction algorithm here, which can be determined according to needs.
[0064] The peak signal-to-noise ratio can be calculated by the following formula:
[0065]
[0066] In the formula, PSNR is the peak signal-to-noise ratio, X is the second target frame, Y is the mapped target video frame, log is the logarithmic operation, MSE is the mean squared error solution, and MAX is the maximum value.
[0067] The application of the optical flow algorithm enables the semantic segmentation result to maintain coherence between adjacent frames, which can improve the clarity of the video and the efficient compression of video data.
[0068] Optionally, mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a mapped target video frame and a target semantic segmentation result, including: determining the mapping relationship between the first target frame and the second target frame based on the optical flow algorithm; mapping the first semantic segmentation result to the second target frame according to the mapping relationship to obtain a mapped target video frame and a target semantic segmentation result.
[0069] In an alternative embodiment, the mapping relationship between the first target frame and the second target frame can be determined by a sparse optical flow algorithm, optical flow field estimation, or a dense optical flow algorithm. Specifically, the pixel-level optical flow field from the first target frame to the second target frame can be calculated by a dense optical flow algorithm. Then, according to the optical flow vector of each pixel, each pixel point in the semantic segmentation result of the first target frame is mapped to the corresponding position in the second target frame. Alternatively, the mapping relationship between the first target frame and the second target frame can also be determined by feature point tracking and region propagation, by tracking the mapping relationship of feature points through an optical flow algorithm. Further, the pixel semantic label can be determined according to the mapping relationship, so as to map the first semantic segmentation result onto the second target frame, and the mapped target video frame and the target semantic segmentation result are recorded.
[0070] Optionally, determining the peak signal-to-noise ratio of the second target frame and the mapped target video frame includes: determining the mean square error of the second target frame and the mapped target video frame; determining the maximum value of the pixels in the mapped target video frame; and determining the peak signal-to-noise ratio based on the mean square error and the maximum value.
[0071] In an alternative embodiment, the mean square error between the second target frame and the mapped target video frame can be determined by a mean square error formula or a gradient descent method. And the maximum value can be determined by traversing each pixel in the mapped target video frame. Alternatively, the maximum value of the pixels can also be calculated by an image processing tool. Alternatively, the maximum value of the pixels can also be determined by a partition parallel computing method. For example, the mapped target video frame is divided into multiple small regions, the maximum pixel value is calculated in parallel for each small region, and finally the maximum values of each small region are compared to determine the maximum pixel value in the mapped target video frame.
[0072] Further, the peak signal-to-noise ratio can be determined based on the peak signal-to-noise ratio formula. Alternatively, the peak signal-to-noise ratio can also be determined by a peak signal-to-noise ratio prediction algorithm. The peak signal-to-noise ratio prediction algorithm can be a support vector machine, linear regression, etc. There is no limitation on the peak signal-to-noise ratio prediction algorithm here, and it can be determined according to needs.
[0073] Optionally, extracting the first target frame and the second target frame from the target video data based on a preset text includes: constructing a text feature corresponding to the preset text by a word segmentation method; encoding the text feature based on an encoding model to obtain a target encoding feature; extracting the first target frame from the target video data based on the target encoding feature; and determining the other frames except the first target frame as the second target frame.
[0074] In an alternative embodiment, text features corresponding to the preset text can be constructed first using tokenization methods such as word fragments or byte pair encoding. The word fragment tokenization method divides words into small fragments, which can be recombined to form words in the vocabulary or can independently represent unknown or rare words, thereby improving the tokenization effect. The byte pair encoding tokenization method constructs the vocabulary by counting and merging byte pairs in the text, thereby achieving encoding.
[0075] Then, the text features are encoded through an encoding model based on word embeddings, a deep learning-based encoding model, or an encoding model based on a sub-attention mechanism to obtain target encoding features. For example, the text features can be transformed into a high-dimensional vector representation through the BGE miniaturized embeddable model. This model is a bidirectional encoder model. These vectors not only contain information about individual words but also capture the dependencies between words and the semantics of the entire sentence through the multi-layer structure and attention mechanism of the model. The pre-trained miniaturized embeddable model is used to encode the text features.
[0076] Furthermore, the first target frame can be extracted based on similarity. That is, the first target frame is extracted based on the similarity between the target encoding features and the features of each target frame in the target video data. Alternatively, the first target frame can also be extracted through a target detection network. For example, the target detection network can be the YOLO algorithm, the Single Shot MultiBox Detector, etc. There is no limitation on the target detection network here and it can be determined according to needs. Alternatively, the first target frame can also be extracted based on template matching of appearance features. Thus, frames other than the first target frame can also be determined as the second target frames.
[0077] Optionally, encoding the text features based on the encoding model to obtain the target encoding features includes: encoding the text features based on the encoding model to obtain the initial encoding features; extracting the target encoding features from the initial encoding features based on preset conditions.
[0078] In an alternative embodiment, the text features can be encoded through an encoding model based on word embeddings, a deep learning-based encoding model, or an encoding model based on a sub-attention mechanism to obtain the initial encoding features.
[0079] Further, an attention mechanism can be used to extract target encoded features from the initial encoded features. That is, the initial encoded features are weighted according to preset conditions to extract the target encoded features. Alternatively, a conditional encoding network can also be used by inputting the preset conditions into the conditional encoding network. The preset conditions can be labels, key texts, or context information, etc. Thus, the conditional encoding network adjusts the encoding process according to the preset conditions. Alternatively, a feature selection algorithm or a feature filtering algorithm can also be used to extract target encoded features from the initial encoded features. The feature selection algorithm can be an information gain algorithm, a recursive feature elimination algorithm, etc. The feature selection algorithm is not limited here and can be determined according to needs. The feature filtering algorithm can be obtained based on frequency domain filtering or time domain filtering.
[0080] Optionally, extracting a first target frame from the target video data based on the target encoded features includes: determining the similarity between the target encoded features and any video frame in the target video data; in response to the similarity being greater than or equal to a similarity threshold, determining the video frame as the first target frame.
[0081] In an alternative embodiment, the similarity between the target encoded features and any video frame in the target video data can be calculated by the cosine similarity of the feature vectors. That is, the target encoded features and the encoded representation of the video frame are transformed into high-dimensional vectors, and then the cosine similarity between the two vectors is calculated. In the case where the similarity is greater than or equal to the similarity threshold, the video frame is determined as the first target frame.
[0082] In another alternative embodiment, the similarity between the target encoded features and any video frame in the target video data can be determined by a contrast network. By pre-training the contrast network, the contrast loss encourages the distance between positive samples (i.e., the target feature and the video frame features of the same target) to be close, while pulling the distance between negative samples (i.e., the target feature and the irrelevant video frame features) apart. By adjusting the contrast loss, the contrast network can learn a feature representation that can distinguish the target from the background, thereby calculating the similarity between the target feature and the video frame feature. The contrast loss function can be a triplet loss function, a multi-triplet loss function, or an information noise contrastive estimation loss. The contrast loss function is not limited here and can be determined according to needs.
[0083] Optionally, different compression strategies are used to compress the first region and the second region in different video frames respectively to obtain a compression result, including: compressing the first region at a first magnification or not compressing the first region to obtain a first sub-compression result; compressing the second region at a second magnification to obtain a second sub-compression result; obtaining the compression result based on the first sub-compression result and the second sub-compression result.
[0084] The above first magnification factor can be less than the second magnification factor to achieve adaptive compression. Thus, through this differential compression strategy, the compression efficiency can be significantly improved, reducing the storage space requirement and transmission bandwidth consumption.
[0085] In an alternative embodiment, the compression strategy for different regions can be learned through an adaptive compression network. Or, encoding can also be used to compress the first region or the second region. Alternatively, JPEG compression can also be used for compression. The data of the first region is divided into 8×8 image blocks. After color space conversion, through discrete cosine transform, quantization, and entropy coding, a binary code stream is obtained. The binary code stream is transmitted to the cloud for text feature training and image semantic segmentation processing. After processing, it is transmitted to the edge side of the drone in the form of a binary code stream for decoding to obtain the detection result. So as to obtain the compression result after compressing the first region and the second region respectively.
[0086] Optionally, obtaining the target video data includes: obtaining the initial target video data; performing data optimization on the initial target video data to obtain the target video data, where the data optimization includes at least one of the following: data denoising, contrast adjustment, brightness adjustment, and color correction.
[0087] Through the data optimization step, the visual effect of the target video data can be improved. For a video streaming platform, it can provide higher-quality video content and enhance user stickiness.
[0088] In an alternative embodiment, the initial target video data can be directly collected from on-site video data through various sensors such as cameras, thermal imagers, and infrared cameras. Or, the initial target video data can also be obtained from a cloud platform. Since the initial target video data may be affected by factors such as light changes, weather conditions, and motion blur, resulting in poor image quality. Therefore, through data optimization, the quality of the video data can be improved, making the subsequent inspection analysis and recognition more accurate. It can also reduce redundant information, such as unnecessary noise and excessive brightness or contrast changes, which makes the subsequent video compression more efficient and can reduce the data volume and transmission bandwidth requirements while maintaining the quality of the video data.
[0089] Further, data optimization includes at least one of the following: data denoising, contrast adjustment, brightness adjustment, and color correction. However, data optimization is not limited, and can be determined according to needs. Denoising can be achieved through denoising algorithms, such as median filtering, bilateral filtering, or deep learning-based denoising networks, to reduce or eliminate these noise points, making the video data clearer and the features more prominent. Contrast can be the difference in brightness between different regions of the video data. By adjusting the contrast, the visual effect of the video data can be enhanced. For inspection videos, high contrast helps to identify features such as wear and cracks on the surface of power equipment. Brightness can refer to the overall light and dark level of the image. Under different lighting conditions, the brightness of the initial target video data may not be consistent. By adjusting the brightness, the loss of key information caused by over-brightness or over-darkness can be avoided. Color correction can be used to adjust the color balance of video frames to make the colors more realistic or meet specific requirements. In inspection videos, colors help to distinguish different parts of power equipment. For example, there are obvious color differences between rusty metal and clean metal. By color correction, this difference can be enhanced to improve the accuracy of identification.
[0090] Figure 2 is a flowchart of an optional video data compression method according to an embodiment of the present invention, as Figure 2 shown, the method includes:
[0091] Step S202, obtain the initial video data of the drone inspection.
[0092] Step S204, perform preprocessing and enhancement operations on the initial video data to obtain the target video data.
[0093] Step S206, based on text features, extract key frames and non-key frames from the target video data. Among them, the text features are determined based on the inspection task.
[0094] Step S208, adopt different compression strategies to compress the first region and the second region in the key frames and non-key frames respectively to obtain the compression result.
[0095] Figure 3 is a schematic diagram of an optional video data compression method according to an embodiment of the present invention, as Figure 3 shown, the method includes:
[0096] Input the inspection requirement text. Then, based on the micro-embeddable model, convert the input inspection requirement text into a vector representation. After inputting the video frame, combine the vector representation and the video frame, perform the optical flow algorithm, and then perform inter-frame propagation to propagate the semantics.
[0097] Moreover, it is determined whether the peak signal-to-noise ratio is less than the signal-to-noise ratio threshold, that is, the peak signal-to-noise ratio of the video frame is determined. If so, inter-frame propagation is performed and then semantic propagation is carried out. Otherwise, encoding is performed, and color space conversion, discrete cosine transform, quantization, and entropy encoding operations are carried out respectively. After that, semantic segmentation is performed.
[0098] According to another aspect of the embodiments of the present invention, an embodiment of a video data compression device is further provided. It should be noted that this device can execute the video data compression method of the above embodiments. The specific implementation solutions and application scenarios in this embodiment are the same as those in the above embodiments and will not be elaborated here.
[0099] Figure 4 is a schematic diagram of a video data compression device according to an embodiment of the present application. As Figure 4 shown, the device includes the following:
[0100] An acquisition module 40, configured to acquire target video data; an extraction module 42, configured to extract the target video data based on a preset text to obtain a first target frame and a second target frame in the target video data, where the semantic information of the first target frame matches the preset text, and the second target frame is a video frame in the target video data other than the first target frame; a segmentation module 44, configured to perform semantic segmentation on the first target frame and the second target frame respectively to obtain a first region and a second region in different video frames, where the first region matches the preset text, and the second region is a region in different video frames other than the first region; a compression module 48, configured to compress the first region and the second region in different video frames respectively by using different compression strategies to obtain a compression result.
[0101] Optionally, the segmentation module is further configured to perform semantic segmentation on the first target frame based on a semantic segmentation model to obtain a first semantic segmentation result of the first target frame, where the first semantic segmentation result is used to represent the position of the preset text in the first target frame; determine the first region and the second region in the first target frame based on the first semantic segmentation result; map the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a second semantic segmentation result of the second target frame, where the second semantic segmentation result is used to represent the position of the preset text in the second target frame; determine the first region and the second region in the second target frame based on the second semantic segmentation result.
[0102] Optionally, the segmentation module is further configured to map the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a mapped target video frame and a target semantic segmentation result; determine the peak signal-to-noise ratio between the second target frame and the mapped target video frame; in response to the peak signal-to-noise ratio being greater than the signal-to-noise ratio threshold, determine the target semantic segmentation result as the second semantic segmentation result.
[0103] Optionally, the segmentation module is further configured to determine a mapping relationship between the first target frame and the second target frame based on an optical flow algorithm; map the first semantic segmentation result to the second target frame according to the mapping relationship to obtain a mapped target video frame and a target semantic segmentation result.
[0104] Optionally, the segmentation module is further configured to determine the mean square error between the second target frame and the mapped target video frame; determine the maximum value of the pixels in the mapped target video frame; determine the peak signal-to-noise ratio based on the mean square error and the maximum value.
[0105] Optionally, the extraction module is further configured to construct text features corresponding to a preset text by using a word segmentation method; encode the text features based on an encoding model to obtain target encoded features; extract a first target frame from the target video data based on the target encoded features; determine other frames except the first target frame as the second target frame.
[0106] Optionally, the extraction module is further configured to encode the text features based on an encoding model to obtain initial encoded features; extract target encoded features from the initial encoded features based on a preset condition.
[0107] Optionally, the extraction module is further configured to determine the similarity between the target encoded features and any video frame in the target video data; determine the video frame as the first target frame in response to the similarity being greater than or equal to a similarity threshold.
[0108] Optionally, the compression module is further configured to compress the first region at a first magnification or not compress the first region to obtain a first sub-compression result; compress the second region at a second magnification to obtain a second sub-compression result; obtain a compression result based on the first sub-compression result and the second sub-compression result.
[0109] Optionally, the acquisition module is further configured to acquire initial target video data; perform data optimization on the initial target video data to obtain target video data, where the data optimization includes at least one of the following: data denoising, contrast adjustment, brightness adjustment, and color correction.
[0110] An embodiment of the present application further provides an electronic device, including: a memory storing an executable program; a processor configured to run the program, where when the program runs, it executes the methods in the various embodiments of the present invention.
[0111] An embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium includes a stored executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the methods in the various embodiments of the present invention.
[0112] Embodiments of the present application further provide a computer program product, including a computer program which, when executed by a processor, implements the methods in various embodiments of the present invention.
[0113] Embodiments of the present application further provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program which, when executed by a processor, implements the methods in various embodiments of the present invention.
[0114] Embodiments of the present application further provide a computer program which, when executed by a processor, implements the methods in the above-mentioned various embodiments of the present invention.
[0115] In the above embodiments of the present invention, the descriptions of the various embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0116] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0117] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0118] In addition, the functional units in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0119] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0120] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A video data compression method, characterized in that, Including: Obtain target video data; Extract the target video data based on a preset text to obtain a first target frame and a second target frame in the target video data, where the semantic information of the first target frame matches the preset text, and the second target frame is a video frame in the target video data other than the first target frame; Perform semantic segmentation on the first target frame and the second target frame respectively to obtain a first region and a second region in different video frames, where the first region matches the preset text, and the second region is a region in different video frames other than the first region; Compress the first region and the second region in different video frames respectively using different compression strategies to obtain a compression result.
2. The video data compression method according to claim 1, wherein Performing semantic segmentation on the first target frame and the second target frame respectively to obtain a first region and a second region in different video frames, including: Perform semantic segmentation on the first target frame based on a semantic segmentation model to obtain a first semantic segmentation result of the first target frame, where the first semantic segmentation result is used to represent the position of the preset text in the first target frame; Determine the first region and the second region in the first target frame based on the first semantic segmentation result; Map the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a second semantic segmentation result of the second target frame, where the second semantic segmentation result is used to represent the position of the preset text in the second target frame; Determine the first region and the second region in the second target frame based on the second semantic segmentation result.
3. The video data compression method according to claim 2, wherein Mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a second semantic segmentation result of the second target frame, including: Map the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a mapped target video frame and a target semantic segmentation result; Determine the peak signal-to-noise ratio between the second target frame and the mapped target video frame; In response to the peak signal-to-noise ratio being greater than a signal-to-noise ratio threshold, determine the target semantic segmentation result as the second semantic segmentation result.
4. The video data compression method according to claim 3, wherein Mapping the first semantic segmentation result to the second target frame based on an optical flow algorithm to obtain a mapped target video frame and a target semantic segmentation result, including: Determine the mapping relationship between the first target frame and the second target frame based on an optical flow algorithm; Map the first semantic segmentation result to the second target frame according to the mapping relationship to obtain the mapped target video frame and the target semantic segmentation result.
5. The video data compression method according to claim 3, wherein Determining the peak signal-to-noise ratio between the second target frame and the mapped target video frame, including: Determine the mean square error between the second target frame and the mapped target video frame; Determine the maximum value of the pixels in the mapped target video frame; Determine the peak signal-to-noise ratio based on the mean square error and the maximum value.
6. The video data compression method according to claim 1, characterized in that, Extract the target video data based on a preset text to obtain a first target frame and a second target frame in the target video data, including: Construct text features corresponding to the preset text using a word segmentation method; Encode the text features based on the encoding model to obtain target encoded features; Extract the first target frame in the target video data based on the target encoded features; Determine that the other frames except the first target frame are second target frames.
7. The video data compression method according to claim 6, wherein Encoding the text features based on the encoding model to obtain target encoded features, including: Encode the text features based on the encoding model to obtain initial encoded features; Extract the target encoded features from the initial encoded features based on preset conditions.
8. The video data compression method according to claim 6, wherein Extracting the first target frame in the target video data based on the target encoded features, including: Determine the similarity between the target encoded features and any video frame in the target video data; In response to the similarity being greater than or equal to the similarity threshold, determine that the video frame is the first target frame.
9. The video data compression method according to claim 1, wherein Compress the first region and the second region in different video frames respectively using different compression strategies to obtain compression results, including: Compress the first region at a first magnification or do not compress the first region to obtain a first sub-compression result; Compress the second region at a second magnification to obtain a second sub-compression result; Obtain the compression result based on the first sub-compression result and the second sub-compression result.
10. The video data compression method according to claim 1, wherein Obtain target video data, including: Obtain initial target video data; Perform data optimization on the initial target video data to obtain target video data, where the data optimization includes at least one of the following: data denoising, contrast adjustment, brightness adjustment, and color correction.
11. A video data compression device, characterized in that, Including: An acquisition module for acquiring target video data; An extraction module for extracting based on a preset text from the target video data to obtain the first target frame and the second target frame in the target video data, where the semantic information of the first target frame matches the preset text, and the second target frame is the video frame in the target video data other than the first target frame; A segmentation module for performing semantic segmentation on the first target frame and the second target frame respectively to obtain a first region and a second region in different video frames, where the first region matches the preset text, and the second region is the region other than the first region in different video frames; A compression module for compressing the first region and the second region in different video frames respectively using different compression strategies to obtain compression results.
12. An electronic device, characterized in that, Including: A memory storing an executable program; A processor for running the program, where the program, when running, executes the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, where, when the executable program runs, it controls the device where the storage medium is located to execute the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that, Including a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.
Citation Information
Cited By
Multi-scale interaction and semantic calibration video abstraction method for motion and appearance decoupling
CN120812373A
A multi-scale interaction and semantic calibration video summarization method decoupling motion and appearance
CN120812373B
Digital asset reconstruction method and system based on AI video dynamic modification
CN120980300A
A reconstruction method and system for digital assets based on AI video dynamic modification
CN120980300B