Dynamic compression method, system and equipment for teaching video and storage medium

By detecting mutation image frames in teaching videos and adding classification tags to video clips using image recognition technology, and generating compressed parameters for dynamic compression, the problem of redundant video clips in teaching videos is solved, and resource saving and user experience improvement is achieved.

CN120186416APending Publication Date: 2025-06-20SHANDONG INSPUR ULTRA HD INTELLIGENT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510359926.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

There are a large number of redundant video clips in teaching videos, resulting in poor user experience and wasted network and storage resources.

Method used

By detecting mutational image frames in the video, the video is divided into multiple video clips, and a classification tag is added to each clip using image recognition technology, and compression parameters are generated based on the tag for dynamic compression.

Benefits of technology

Effectively separate important and invalid clips in teaching videos, realize customized compression, reduce the proportion of redundant videos, save network and storage resources, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186416A_ABST
    Figure CN120186416A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and particularly provides a dynamic compression method, system and device for a teaching video and a storage medium, and the method comprises the steps: converting the video into an image frame sequence; detecting abrupt change image frames in the image frame sequence, and marking the abrupt change image frames as node image frames; segmenting the video into a plurality of video clips according to the node image frames, and adding classification tags to the plurality of video clips by using an image recognition technology; and generating a compression parameter for the corresponding video clip according to the classification label, and compressing the video clip according to the compression parameter. According to the invention, the dynamic compression of the teaching video is realized, and the proportion of redundant video clips is reduced, so that the network and storage resources consumed by the redundant video are reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a dynamic compression method, system, device and storage medium for teaching videos. Background Art

[0002] There are some redundant video segments in teaching videos. When users watch teaching videos, they need to manually drag the progress bar to skip these video segments. Taking the assembly tutorial of a children's building block toy as an example, typical redundant phenomena include: the second-long blank screen when the teacher adjusts the camera angle, the multi-angle close-ups of the same part shown repeatedly, the silent waiting during step transitions, the repeated demonstration shots due to operation mistakes, and some promotional introductions that have little relevance to teaching. These redundant contents account for 35%-50% of the total video duration, directly resulting in an extra 12-18 minutes of progress bar operation for parents when guiding their children to learn.

[0003] These redundant video segments not only lead to poor user experience, but also waste a large amount of network resources and storage resources. Summary of the Invention

[0004] In view of the above deficiencies of the prior art, the present invention provides a dynamic compression method, system, device and storage medium for teaching videos to solve the above technical problems.

[0005] In a first aspect, the present invention provides a dynamic compression method for teaching videos, including: Converting the video into a sequence of image frames; Detecting the mutant image frames in the sequence of image frames and marking the mutant image frames as node image frames; Segmenting the video into multiple video segments according to the node image frames, and adding classification labels to the multiple video segments respectively by using image recognition technology; Generating compression parameters for the corresponding video segments according to the classification labels, and compressing the video segments according to the compression parameters.

[0006] In an optional embodiment, the method further includes: Extracting voice data from the video; Converting the voice data into text data; Extracting keywords from the text data by using keyword extraction technology and determining the video timestamps corresponding to the keywords; Generating a detection time range according to the video timestamps and a preset time fluctuation range, and the detection time range is used to determine the mutant image frames.

[0007] In an optional embodiment, generating a detection time range according to the video timestamps and a preset time fluctuation range includes: The detection time range is from (t - t1) to (t + t2), where t is the video timestamp, t1 and t2 are preset range values, (t - t1) is the lower limit value of the detection time range, and (t + t2) is the upper limit value of the detection time range.

[0008] In an alternative embodiment, detecting mutant image frames in the image frame sequence and labeling the mutant image frames as node image frames includes: Extracting a subsequence from the image frame sequence with timestamps within the detection time range; Reading the first frame of the subsequence and converting it into a grayscale image as a reference for subsequent inter-frame difference calculation; Looping through each frame of the subsequence, converting it into a grayscale image, and calculating the absolute difference image D(x, y) between the current frame and the previous frame: D(x, y)=∣I cur (x,y)−I prev (x,y)∣ where, I cur (x,y) is the grayscale image of the current frame, which is the grayscale value of the pixel at the (x, y) position; I prev (x,y) is the grayscale image of the previous frame, which is the grayscale value of the pixel at the (x, y) position.

[0009] Comparing the pixel values of the difference image with a preset threshold, marking the pixel values exceeding the threshold as 1, and marking the pixel values not exceeding the threshold as 0; Calculating the percentage of the number of non-zero pixels in the total number of pixels; If the percentage exceeds the set threshold, adding the number of the current frame to the mutant frame list; After processing all frames, returning the list of the numbers of mutant frames.

[0010] In an alternative embodiment, splitting the video into multiple video segments according to the node image frames and adding classification labels to the multiple video segments respectively using image recognition technology includes: Splitting the video at the corresponding time positions of the video according to the timestamps of the node image frames to obtain multiple video segments; Using an action recognition model to recognize one or more human actions corresponding to the video segments; Extracting a reference image from the previous video segment of the video segment and extracting multiple sample images from the video segment; Using an object recognition model to determine the newly added components in the sample images based on the multiple sample images and the reference image, and identifying the component types of the newly added components; Preset the difficulty quantization value of the human body movement type and the quantization value corresponding to the part type, calculate the weighted sum of the quantization values of the human body movement and the part type, and obtain the difficulty evaluation value; Set the part type and the difficulty evaluation value as the classification label of the video segment.

[0011] In an alternative embodiment, preset the difficulty quantization value of the human body movement type and the quantization value corresponding to the part type, calculate the weighted sum of the quantization values of the human body movement and the part type, and obtain the difficulty evaluation value, including: Convert the human body movement into a human body movement quantization value according to the preset difficulty quantization value of the human body movement type; Convert the identified part type into a part type quantization value according to the preset quantization value of the part type; Calculate the weighted sum of the human body movement quantization value and the part type quantization value; Traverse all video segments to obtain multiple weighted sums, calculate the normalization coefficient of the multiple weighted sums, and set the normalization coefficient as the difficulty evaluation value of the corresponding video segment.

[0012] In an alternative embodiment, generate compression parameters for the corresponding video segment according to the classification label, so as to compress the video segment according to the compression parameters, including: Determine the importance level of the video segment according to the classification label; Determine the compression parameters according to the pre-set correspondence between the importance level and the compression parameters and the importance level of the video segment, and the compression parameters include the bit rate and the frame skip rate.

[0013] In a second aspect, the present invention provides a dynamic compression system for teaching videos, including: A conversion module for converting a video into a sequence of image frames; A detection module for detecting mutant image frames in the sequence of image frames and marking the mutant image frames as node image frames; An identification module for dividing the video into multiple video segments according to the node image frames and adding classification labels to the multiple video segments respectively by using image recognition technology; A compression module for generating compression parameters for the corresponding video segment according to the classification label, so as to compress the video segment according to the compression parameters.

[0014] In a third aspect, a device is provided, including: A memory for storing a dynamic compression program for teaching videos; A processor for implementing the steps of the dynamic compression method for teaching videos provided in the first aspect when executing the dynamic compression program for teaching videos.

[0015] In a fourth aspect, a computer-readable storage medium is provided, on which a dynamic compression program for teaching videos is stored. When the dynamic compression program for teaching videos is executed by a processor, the steps of the dynamic compression method for teaching videos provided in the first aspect are implemented.

[0016] The beneficial effects of the present invention are as follows. The dynamic compression method, system, device and storage medium for teaching videos provided by the present invention identify the mutant image frames in the teaching video, and then divide the video into video segments based on the mutant image frames, thereby realizing the separation of effective segments and ineffective segments in the teaching video. Then, through image recognition technology, classification labels are generated for the video segments according to the content of the video segments, and adaptive compression parameters are generated for the video segments according to the classification labels. The mapping between the classification labels and the compression parameters can be customized, and thus customized compression of different types of video segments in the teaching video can be realized. The present invention realizes the dynamic compression of teaching videos, reduces the proportion of redundant video segments, thereby reducing the network and storage resources consumed by redundant videos, and improving the user experience.

[0017] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained according to these drawings without creative efforts.

[0019] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention.

[0020] Figure 2 is another schematic flowchart of the method according to an embodiment of the present invention.

[0021] Figure 3 is a flowchart of generating compression parameters of the method according to an embodiment of the present invention.

[0022] Figure 4 is a schematic diagram of the playback processing flow of the compressed video of the method according to an embodiment of the present invention.

[0023] Figure 5 is a schematic diagram of the tree structure of the compressed video in the player according to an embodiment of the present invention.

[0024] Figure 6 is a schematic block diagram of the system according to an embodiment of the present invention.

[0025] Figure 7 Schematic structural diagram of a device provided by an embodiment of the present invention. Detailed implementation manners

[0026] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0028] The following explains the key terms that appear in the present invention.

[0029] OpenCV (Open Source Computer Vision Library) is an open-source computer vision and machine learning software library initiated by Intel Corporation in 1999 and later supported by numerous developers and continuously developed. It provides rich tools and algorithms for computer vision tasks and has been widely used in both academia and industry.

[0030] OpenPose is a real-time multi-person 2D pose estimation system developed by researchers at Carnegie Mellon University (CMU). It can detect the key points of the human body in images or videos, such as joints, limbs, etc., and connect these key points into a human skeleton, thereby realizing the recognition and analysis of human actions. Its working principle includes: Feature extraction: Use a convolutional neural network (CNN) to extract features from the input image or video frame. Usually, a pre-trained model, such as VGG-19, etc., is used to convert the image into a series of feature maps.

[0031] Confidence map and part affinity field prediction: Confidence Maps: Used to predict the positions of each key point of the human body. Each key point corresponds to a confidence map, and the value of each pixel in the map represents the probability that the corresponding key point exists at that position.

[0032] Part Affinity Fields (PAFs): Used to represent the connection relationships between key points, described by a vector field for the direction and intensity between adjacent key points.

[0033] Key point detection and connection: Key point detection: Find the position with the highest probability in the confidence map as the candidate position of the key point.

[0034] Key point connection: Utilize the Part Affinity Fields information to connect the detected key points to form a human skeleton.

[0035] The target recognition model adopts YOLO (You Only Look Once), regards the target detection problem as a regression problem, directly makes predictions on the image, and outputs the category and position of the target. Its working process includes: Load the pre-trained model: Use the torch.hub.load function to load the pre-trained model of YOLOv5.

[0036] Read the image: Specify the image path for target detection.

[0037] Perform target detection: Input the image into the model for target detection.

[0038] Display and save the detection results: Use the results.show() function to display the detection results and the results.save() function to save the detection results.

[0039] The dynamic compression method of the teaching video provided by the embodiment of the present invention is executed by a computer device. Correspondingly, the dynamic compression system of the teaching video runs in the computer device.

[0040] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention. Among them, Figure 1 The execution subject can be a dynamic compression system of a teaching video. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0041] As Figure 1 shown, the method includes: S1. Convert the video into a sequence of image frames; S2. Detect the mutant image frames in the sequence of image frames and mark the mutant image frames as node image frames; S3. Segment the video into multiple video segments according to the node image frames, and use image recognition technology to add classification labels to the multiple video segments respectively; S4. Generate compression parameters for the corresponding video segments according to the classification tags, and compress the video segments according to the compression parameters.

[0042] In one embodiment of the present invention, while performing step S1, a voice processing method is given.

[0043] 1. Extract voice data from the video.

[0044] Decode the video file (such as common formats like MP4, AVI, etc.), separate the audio stream. An open-source library such as FFmpeg can be used to complete this operation. Extract the voice data from the decoded audio stream and store it in PCM (Pulse Code Modulation) format, which is an uncompressed representation of audio data.

[0045] 2. Convert the voice data into text data.

[0046] (1) Extract Mel Frequency Cepstral Coefficients (MFCC) from the voice data. It is a feature parameter widely used in speech recognition and can effectively describe the spectral characteristics of the speech signal: Framing: Divide the speech signal into multiple short frames, usually with a frame length of 20 - 30ms, and there is a certain overlap between frames; Windowing: Multiply each frame signal by a window function (such as a Hamming window) to reduce spectral leakage; Fast Fourier Transform (FFT): Convert the time-domain signal into a frequency-domain signal; Mel filtering: Pass the frequency-domain signal through a set of Mel filter banks to obtain the Mel spectrum; Logarithmic operation: Take the logarithm of the Mel spectrum; Discrete Cosine Transform (DCT): Perform a discrete cosine transform on the logarithmic Mel spectrum to obtain the MFCC coefficients.

[0047] (2) Use a recurrent neural network model to train the extracted features and establish a mapping relationship between speech features and phonemes.

[0048] (3) Combine the grammar and semantic information of the language to process the phoneme sequence output by the acoustic model and convert it into meaningful text.

[0049] 3. Use keyword extraction technology to extract keywords from the text data and determine the video timestamps corresponding to the keywords.

[0050] (1) Preset keywords, such as "next", "attention", "installation", etc. Combine the set keywords into a keyword dictionary (such as a set or dict in Python), and then query and locate these keywords from the text data. The specific steps include: Text cleaning: Remove special characters, punctuation marks, extra spaces, etc. from the text to make the text data more regular and facilitate subsequent keyword queries.

[0051] Text word segmentation (optional): If the text is in Chinese, word segmentation may be required to split the continuous text into individual words so that it can be accurately matched with keywords. Open source word segmentation libraries such as jieba can be used.

[0052] Traversal query: Traverse each word (if word segmentation is performed) or substring in the text data to check whether it exists in the keyword dictionary.

[0053] Position recording: Once a matching keyword is found, record the position information of the keyword in the text, such as the start index and end index, for subsequent determination of its corresponding video timestamp.

[0054] (2) Timestamp matching: Determine the timestamp of each keyword in the video according to the correspondence between the speech data and the text data.

[0055] 4. Generate a detection time range based on the video timestamp and a preset time fluctuation range, and the detection time range is used to determine the mutant image frames.

[0056] The detection time range is from (t - t1) to (t + t2), where t is the video timestamp, t1 and t2 are preset range values, (t - t1) is the lower limit value of the detection time range, and (t + t2) is the upper limit value of the detection time range.

[0057] By generating the detection time range, the data processing amount for detecting mutant image frames can be reduced, that is, when detecting mutant image frames, it is not necessary to traverse all the image frames of the video, but only traverse the image frames within the detection time range.

[0058] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0059] S201. Extract a subsequence from the image frame sequence whose timestamps are within the detection time range.

[0060] S202. Read the first frame of the subsequence and convert it into a grayscale image as the benchmark for subsequent inter - frame difference calculation.

[0061] Read the first frame image from the subsequence and use an image processing library to convert the color image into a grayscale image.

[0062] For each frame of the cyclic subsequence, convert it into a grayscale image, and calculate the absolute difference image D(x, y) between the current frame and the previous frame: D(x,y)=∣I cur (x,y)−I prev (x,y)∣ where I cur (x,y) is the grayscale image of the current frame, which is the grayscale value of the pixel at the (x, y) position; I prev (x,y) is the grayscale image of the previous frame, which is the grayscale value of the pixel at the (x, y) position.

[0063] S204. Compare the pixel values of the difference image with a preset threshold. Mark the pixel values that exceed the threshold as 1, and mark the pixel values that do not exceed the threshold as 0.

[0064] S205. Calculate the percentage of the number of non-zero pixels in the total number of pixels.

[0065] Count the number of non-zero pixels in the image after threshold processing, and calculate the percentage of the number of non-zero pixels in the total number of pixels.

[0066] S206. If the percentage exceeds the set threshold, add the number of the current frame to the mutant frame list.

[0067] S207. After processing all frames, return the list of the numbers of mutant frames.

[0068] After processing all frames in the subsequence, return the list of the numbers of mutant frames. Then process the subsequence corresponding to the next detection time range, and finally summarize the list of the numbers of mutant frames of all subsequences to obtain all mutant image frames of the video.

[0069] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.

[0070] S301. Split the video at the corresponding time positions of the video according to the timestamps of the node image frames to obtain a plurality of video segments.

[0071] Based on the timestamps, split the video into continuous segments, and the timestamps are provided by the video codec metadata. Specifically, it includes: Read the video file to obtain the timestamps (unit: seconds) of all frames; According to the input time range [t_start, t_end], filter out the corresponding frames; Package the filtered frames into independent video segments.

[0072] S302. Identify one or more human actions corresponding to the video clip using an action recognition model.

[0073] Use a pre-trained action recognition model (such as C3D, I3D) to extract spatio-temporal features and classify them. Specifically, it includes: Input the video clip and preprocess it to a fixed size (such as 224×224); Extract spatial features through a convolutional neural network (CNN); Extract temporal features through a recurrent neural network (RNN / LSTM) or 3D CNN; The Softmax classifier outputs the action category (such as "installation", "adjustment"): p(y∣f1,…,f T )=Softmax(RNN(f1,…,f T )) where, f T is the spatial feature of the T-th frame image, for example.

[0074] S303. Extract a reference image from the previous video clip of the video clip, and extract multiple sample images from the video clip.

[0075] Select key frames from the video clip as the reference and samples for subsequent difference analysis, including: The last frame of the previous video clip is used as the reference image I prev .

[0076] Randomly select N frames from the current video clip as the sample images I sample (1) ,…,I sample (N) .

[0077] S304. Use an object recognition model to determine the newly added components in the sample images based on the multiple sample images and the reference image, and identify the component types of the newly added components.

[0078] Detect newly added components through image difference analysis and classify them using an object detection model (such as YOLO). Specifically, it includes: Calculate the absolute difference between the reference image and the sample image: D (k) (x,y)=∣I prev (x,y)−I sample (k) (x,y)∣ Calculate the absolute difference between the reference image and each sample image respectively to obtain multiple absolute difference images.

[0079] Perform threshold processing on the difference images to extract candidate regions.

[0080] Use the object detection model to identify the part types of candidate regions (such as "screw", "gear"): Input the sample image and the corresponding candidate region into the YOLO model to obtain the part types of the candidate regions: class = YOLO(I sample (k) , candidate region).

[0081] S305. Preset the difficulty quantization values of human action types and the quantization values corresponding to part types, calculate the weighted sum of the quantization values of human actions and part types, and obtain the difficulty evaluation value.

[0082] Quantify the difficulty by weighted fusion of action difficulty and part complexity. In this way, not only the action difficulty and part difficulty are concerned, but also the number of actions can be concerned. The more the number of actions, the greater the overall difficulty evaluation value.

[0083] S305.1 Convert the human action into a human action quantization value according to the preset difficulty quantization value of the human action type.

[0084] Look up the table to obtain the quantization value w corresponding to the action type action .

[0085] S305.2 Convert the identified part type into a part type quantization value according to the preset quantization value of the part type.

[0086] Look up the table to obtain the quantization value w corresponding to the part type part .

[0087] S305.3 Calculate the weighted sum of the human action quantization value and the part type quantization value.

[0088]

[0089] Among them, is the quantization value of the i-th action, is the weight coefficient, G is the total number of actions; N is the total number of newly added parts, is the quantization value of the k-th part type.

[0090] S305.4 Traverse all video segments to obtain multiple said weighted sums, calculate the normalization coefficient of the multiple said weighted sums, and set the normalization coefficient as the difficulty evaluation value of the corresponding video segment.

[0091] Normalize all segments of S to [0, 1]:

[0092] S306. Set the component type and the difficulty evaluation value as the classification label of the video clip.

[0093] Use the component type and the normalized difficulty value as the metadata of the video clip. Count the component types that appear in all sample images to generate a label list (such as {screw, gear}). Let S norm be the difficulty label.

[0094] Label storage: label = {parts: [p1, p2,...], difficulty: S norm}.

[0095] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0096] Determine the importance level of the video clip according to the classification label; according to the pre - set correspondence between the importance level and the compression parameters, and the importance level of the video clip, determine the compression parameters, where the compression parameters include the bitrate and the frame skip rate.

[0097] Specifically, extract the difficulty evaluation value from the classification label. Set the importance level with a difficulty evaluation value of 0 - 0.2 as low level, 0.2 - 0.6 as medium level, and 0.6 - 1 as high level.

[0098] Set the corresponding bitrate and frame skip rate for the video clip according to the importance level. Among them, the video clip with a high level remains unchanged and uses the original video. For example: def get_target_bitrate(difficulty): if difficulty < 0.2: return 500 # 500kbps (low level) elif 0.2 <= difficulty < 0.6: return 1500 # 1500 kbps (medium level) else: return None # Keep the original bitrate (high level). That is, the bitrate and frame skip rate of the low - level video clip are 500 # 500kbps, the bitrate and frame skip rate of the medium - level video clip are 1500 # 1500 kbps, and the high - level video clip remains unchanged.

[0099] Compress the low - level and medium - level video clips according to the set bitrate and frame skip rate.

[0100] Use the bitrate control algorithm to adjust the bitrate of the video clip: min Q (D(Q)+λ⋅R(Q)) Among them: D(Q) is the distortion corresponding to the quantization parameter Q (such as PSNR), R(Q) is the bitrate, and λ is the Lagrange multiplier (dynamically adjusted according to the difficulty level).

[0101] For low-level segments, scalable video coding (SVC) is adopted: Base Layer: Preserve key information, with a bitrate occupancy of 60%; Enhancement Layer: Optional details, with a bitrate occupancy of 40%.

[0102] Methods for controlling the frame skip rate include: Calculate the inter-frame motion using a block matching algorithm (such as full search):

[0103] Among them: I t is the current frame, I t−1 is the previous frame, and dx, dy are the motion displacements.

[0104] def decide_frame_skip(frame_index, motion_vector): if motion_vector < 3: # Frames with low motion can be skipped return True elif frame_index % 3 == 0: # Preserve key frames return False else: return motion_vector < 5.

[0105] Protect key frames based on the results of mutation frame detection: if frame_index in mutation_frames: skip = False # Do not skip mutation frames else: skip = decide_frame_skip(frame_index, motion_vector).

[0106] Optimize the compressed video: Use a saliency detection algorithm (such as the Itti model) to locate the ROI:

[0107] Among them: S c is the color saliency map, S o is the orientation saliency map, S t is the motion saliency map.

[0108] Adopt a higher quantization parameter for the ROI region, that is, reduce the QP value by 2 - 3.

[0109] Perform motion compensation interpolation on the frame-skipped video:

[0110] Table 1 shows an exemplary parameter configuration scheme: Table 1

[0111] Please refer to Figure 2 , this embodiment provides a video dynamic compression method, including: Mark key nodes: During video recording, mark key nodes in the following ways: Manual marking: The user clicks a button or issues a voice command to trigger marking.

[0112] Automatic recognition: Automatically mark based on image recognition (such as sudden changes in the picture, text prompts) or audio features (such as keywords).

[0113] Save the original data: The video segment corresponding to the key node (such as the moment when the operation step is completed) is saved as the original high-bitrate data.

[0114] Generate a dedicated video package: Split the video into key node segments and non-key segments, and encapsulate them into a dedicated format containing the following information: Key node timestamps and position indexes; Compression parameters for non-key segments (such as bitrate, frame-skipping strategy).

[0115] Among them, the file structure of the dedicated video package includes: 1. File header (Header) Field definition: Version number (4 bytes): Identifies the video package format version (such as 1.2.0).

[0116] Compatible player list (variable-length string): Names of players that support parsing this format (such as "Player_A;Player_B").

[0117] Creation timestamp (8 bytes): Unix timestamp records the video package generation time.

[0118] Reserved field (4 bytes): For future expansion purposes (such as encryption identification).

[0119] 2. Metadata area (Metadata) Key node index table: [ {"id": 101, "name": "Install the motherboard", "timestamp": "00:01:20.500", "data_ptr": 0x000100, / / The starting address pointing to the key node data block "level": 1 / / Node level (main node = 1, sub - node = 2)},... ].

[0120] Non - key segment compression parameter table: [{"segment_id": 201, "start_time": "00:00:00", "end_time": "00:01:15", "codec": "H.265", "bitrate": "1Mbps", "frame_rate": "5fps"},... ].

[0121] 3. Key node data block Storage rule: Save the video stream at the original high bitrate (e.g., 10Mbps), without compression or only lossless compression.

[0122] Each node stores independently, supporting random access (quickly locate through the address pointer).

[0123] Data format: Starting address (4 bytes): The offset of the data block in the file.

[0124] Length (4 bytes): The size of the data block (unit: byte).

[0125] Video stream: Raw stream or encapsulated in a lightweight format (e.g., Annex B H.264).

[0126] 4. Non - key data block Storage rule: Process the low - bitrate video stream (e.g., 1Mbps H.265) according to the dynamic compression parameters.

[0127] For the skipped - frame segments, only retain I - frames and key P - frames, and the frame rate can be reduced to 1fps.

[0128] Data format: Stored in segments, each segment contains the compressed data of a continuous time interval.

[0129] Support streaming reading to reduce memory occupancy.

[0130] 5. Index area Timestamp - address mapping table: Hash table structure, with the key being the timestamp (millisecond precision) and the value being the data block address.

[0131] Example: {"00:03:45.200": 0x000500}.

[0132] Hierarchical relationship table: Records the parent - child relationships of multi - level nodes (e.g., the parent node of child node 201 is 101).

[0133] 6. File Footer Checksum (4 bytes): CRC32 checksum value, used to verify the file integrity.

[0134] End flag (2 bytes): Fixed value 0xFFFF, indicating the end of the file.

[0135] The specific video compression process includes: Step 1: Recording and manual marking Scenario: The user records an industrial equipment assembly teaching video and needs to mark key steps such as "Install the motherboard" and "Connect the power supply".

[0136] Operation process: The user manually triggers the marking (e.g., click the button after installing the motherboard) through the "Mark button" on the recording interface when the key operation is completed.

[0137] The system automatically captures the video segments 5 seconds before and after the marked point and saves them as raw high - bitrate data (1080p, 30fps, bitrate 10Mbps).

[0138] Non - key segments (such as the process of screwing) are compressed in the following ways: Frame skipping: Reduced from 30fps to 5fps, only retaining key action frames.

[0139] Bitrate compression: Using H.265 encoding, the bitrate is reduced to 1Mbps, and the resolution is adjusted to 720p.

[0140] Step 2: Automatic marking and metadata encapsulation Automatic marking trigger: Image recognition: Detect sudden changes in the picture (such as tool replacement, component appearance) through OpenCV and automatically mark the nodes.

[0141] Audio analysis: If the video contains voice explanations, detect keywords such as "Attention" and "Next step" through automatic speech recognition (ASR) to trigger marking.

[0142] Metadata generation: Key node timestamps (such as 00:02:15), spatial indices (such as video frame coordinates X / Y), node names ("Install the motherboard").

[0143] Non - critical segment compression parameter table (e.g., segment ID = 2, bitrate = 1Mbps, frame skip rate = 80%).

[0144] Encapsulation format: Pack the original segment, compressed segment, and metadata into a.kvid format video package, with the structure as follows: [Header] Version number = 1.2 | Compatible player list = Player_A, Player_B; [Key node data block] Node 1: Start time = 00:01:10, original data pointer = 0x000100; Node 2: Start time = 00:03:45, original data pointer = 0x000500. [Non - key data block] Segment 1: Start time = 00:00:00, compression parameter = bitrate 1Mbps, frame skip rate 80%.

[0145] Then generate dynamic compression parameters: Further optimize the compression strategy for non - critical segments.

[0146] Extract scene features: Calculate the motion vector of the non - critical segment's video frame (using the optical flow method). When the motion amplitude < threshold, it is determined as "low complexity".

[0147] Detect the audio energy. If it is continuously lower than the threshold and there is no speech, it is determined as a "silent segment".

[0148] Dynamically adjust compression parameters: Low - complexity segment: Compress the bitrate to 10% of the original, and increase the frame skip rate to 90% (only keep the first and last frames).

[0149] Medium - complexity segment: Compress the bitrate to 30%, and the frame skip rate to 50%.

[0150] Silent segment: Directly remove the audio track and only keep the key - frame images.

[0151] Please refer to Figure 3 , and the technical solution of this embodiment will be described below in combination with functional modules: 1. Video analysis module Motion vector detection: Technical implementation: Calculate the motion amplitude between adjacent frames through the optical flow method (such as the Farneback algorithm in OpenCV).

[0152] Scoring rule: If the average motion amplitude < 5 pixels / frame, it is determined as a low - complexity scene (e.g., statically placed tools).

[0153] Static scene recognition: Technical implementation: Extract the HSV histograms of consecutive frames. If the similarity > 95%, it is determined as a static picture.

[0154] Scoring rule: Static scene weight score + 30% (low complexity).

[0155] Text / object detection: Technical implementation: Use the YOLO model to detect text labels or specific objects (such as screwdrivers) in the picture.

[0156] Scoring rule: If a key object is detected, complexity weight score + 20% (partial details need to be retained).

[0157] 2. Audio analysis module Audio energy calculation: Technical implementation: Analyze the audio energy through FFT. If the energy is continuously < -40dB and there is no speech segment, it is determined as a silent segment.

[0158] Scoring rule: Silent segment weight score + 40% (the audio track can be significantly compressed).

[0159] Speech keyword recognition: Technical implementation: Based on ASR (such as DeepSpeech), recognize keywords such as "Attention" and "Danger".

[0160] Scoring rule: If keywords are included, complexity weight score + 25% (the complete speech needs to be retained).

[0161] 3. Scene complexity scoring Formula: Total score = Video complexity weight × 0.6 + Audio complexity weight × 0.4 Example: Low-complexity scene: Video weight 20% + Audio weight 10% → Total score 16% (bitrate compressed to 10%); High-complexity scene: Video weight 70% + Audio weight 50% → Total score 62% (bitrate compressed to 50%).

[0162] 4. Generation of dynamic compression parameters The parameter mapping is shown in Table 1.

[0163] Table 1

[0164] Technical verification example: Input segment: A 30-second non-critical video (tool preparation process), detected: Average motion amplitude = 3 pixels / frame (low motion); Static scene proportion = 80% (high static); No key objects or text; Audio energy = -45dB (silent); Scoring result: Video complexity = 15%; Audio complexity = 10%; Total score = 15% × 0.6 + 10% × 0.4 = 13%; Compression parameters: bit rate 10%, frame skip rate 90%, audio track removed.

[0165] Verification of technical effects: Test data: Processing 10 industrial operation videos (average duration 15 minutes), the results show that: Storage space reduced by 72% - 85% (traditional MP4 average 1.2GB → dedicated video package average 220MB).

[0166] The efficiency of users to locate key content is increased by 90% (average jump time reduced from 15 seconds of manual dragging to 1.5 seconds).

[0167] The above embodiments can be implemented alone or in combination, and the specific parameters can be flexibly adjusted according to the hardware performance (such as the differences between edge devices and servers).

[0168] Please refer to Figure 4 , this embodiment provides a method for playing compressed videos, including: Scenario: When the user is watching the operation instructions, the "learning mode" and "quick preview mode" can be switched in real time.

[0169] Learning mode: The player displays key nodes at the original bit rate, and non - key segments are filled at a low bit rate (such as 20%) to ensure the continuity of the process.

[0170] It supports displaying compression tips for non - key segments when paused (such as "This segment has been compressed, original duration 30 seconds").

[0171] Quick preview mode: Only key nodes are played, non - key segments are completely skipped, and the nodes are quickly connected through a progress bar.

[0172] Mode switching response: When the user clicks the switch button, the player immediately adjusts the decoding strategy according to the new mode without re - loading the video package.

[0173] Verification of storage medium and compatibility: Storage medium: Write the dedicated video package (.kvid format) into the ROM of a solid - state drive, cloud server or embedded device.

[0174] Compatibility verification: The player reads the version number and compatibility description in the video packet header. If the versions do not match, the decoder adaptation module is automatically called. The adaptation module dynamically loads the corresponding decoding algorithm according to the compression parameter table in the metadata (for example, a lightweight decoder is used for low-bitrate segments).

[0175] Please refer to Figure 5 , the tree structure of multi-level nodes in the player, including: Root node: Represents the entire video title (such as "Automobile Repair Teaching Video"), serving as the starting point of the tree structure.

[0176] Metadata: Total video duration, version number, author information.

[0177] Main node (Level 1): Corresponds to the core operation stages of the video (such as "Changing a tire", "Checking the engine").

[0178] Interactive function: Clicking on the main node directly jumps to the starting position of that stage.

[0179] Supports expanding / collapsing the sub-node list.

[0180] Display style: Bold text + icon (such as a gear symbol).

[0181] Sub-node (Level 2): Details the specific steps under the main node (such as "Removing the screws", "Installing the spare tire").

[0182] Metadata: Timestamp (such as 00:05:20), associated parent node ID (such as parent node = main node 1).

[0183] Interactive function: Clicking on the sub-node plays its corresponding original high-bitrate segment.

[0184] When the mouse hovers, a step summary is displayed (such as "A 10mm wrench is required").

[0185] Additional elements on the player interface: Hierarchical navigation bar: A tree directory on the left, supporting scrolling to view multi-level nodes.

[0186] Progress bar markers: Main nodes are marked with red vertical lines, and sub-nodes are marked with blue dots.

[0187] Play mode switching: Supports "Playing by main node" or "Expanding all sub-nodes".

[0188] The technical implementation logic includes: The player parses the metadata of the dedicated video packet and extracts the hierarchical relationship (such as in JSON format): {"nodes": [{"id": "L1_1", "name": "Replace the tire", "type": "Main node", "children": ["L2_1", "L2_2", "L2_3"], "timestamp": "00:03:00"}, {"id": "L2_1", "name": "Remove the screws", "type": "Sub-node", "parent": "L1_1", "timestamp": "00:03:15"}]} Dynamic rendering: Recursively generate a tree-like DOM structure according to the hierarchical relationship, with the main node and sub-nodes indented and aligned (e.g., the main node is indented 0px and the sub-node is indented 20px).

[0189] Node status management: The played nodes are displayed in green, and the currently playing node is highlighted in orange. In the collapsed state, the sub-nodes are hidden and a "+" button is displayed.

[0190] User interaction response: Click on the main node: Jump to the start time of the main node and play at the original bitrate, while automatically collapsing the sub-nodes.

[0191] Click on the sub-node: Precisely jump to the timestamp of the sub-node, and automatically connect to the next sub-node or return to the parent node after playback is complete.

[0192] Drag the node: Support custom adjustment of the node order (the metadata needs to be updated and the video package needs to be repackaged).

[0193] In some embodiments, the dynamic compression system of the teaching video may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the dynamic compression system of the teaching video can be stored in the memory of the computer device and executed by at least one processor to perform (see Figure 1 description) the functions of dynamic compression of the teaching video.

[0194] In this embodiment, the dynamic compression system of the teaching video can be divided into multiple functional modules according to the functions it performs, as Figure 6 shown. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0195] A conversion module for converting a video into a sequence of image frames; A detection module for detecting mutant image frames in the sequence of image frames and marking the mutant image frames as node image frames; An identification module for segmenting the video into multiple video segments according to the node image frames and adding classification labels to the multiple video segments respectively by using image recognition technology; A compression module for generating compression parameters for corresponding video segments according to the classification labels to compress the video segments according to the compression parameters.

[0196] Figure 7 The dynamic compression method for teaching videos provided by the embodiments of the present application can be applied to a device. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown in the figure, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.

[0197] Among them, the device 700 may include: a processor 710, a memory 720, and a communication unit 730. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0198] Among them, the memory 720 may be used to store the execution instructions of the processor 710. The memory 720 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. When the execution instructions in the memory 720 are executed by the processor 710, the device 700 is enabled to execute some or all of the steps in the above method embodiments.

[0199] The processor 710 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 720, and by invoking data stored in the memory, it performs various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC), for example, it may be composed of a single packaged IC, or it may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 710 may only include a central processing unit (CPU). In the embodiments of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.

[0200] The communication unit 730 is used to establish a communication channel so that the storage device can communicate with other devices. It receives user data sent by other devices or sends user data to other devices.

[0201] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it may include some or all of the steps in the embodiments provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0202] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions to enable a computer device (which may be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0203] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0204] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or module can be in electrical, mechanical or other forms.

[0205] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0206] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0207] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention.

Claims

1. A dynamic compression method for teaching videos, characterized in that: include: Convert video into a sequence of image frames; Detecting a mutation image frame in an image frame sequence, and marking the mutation image frame as a node image frame; Dividing the video into a plurality of video segments according to the node image frames, and adding classification labels to the plurality of video segments respectively by using image recognition technology; Compression parameters are generated for corresponding video segments according to the classification labels, so as to compress the video segments according to the compression parameters.

2. The method according to claim 1, characterized in that The method further comprises: extracting voice data from the video; Converting the voice data into text data; Extracting keywords from the text data using keyword extraction technology, and determining video timestamps corresponding to the keywords; A detection time range is generated according to the video timestamp and a preset time fluctuation range, and the detection time range is used to determine a sudden change image frame.

3. The method according to claim 2, characterized in that Generating a detection time range according to the video timestamp and a preset time fluctuation range includes: The detection time range is (t-t1) to (t+t2), where t is the video timestamp, t1 and t2 are pre-set range values, (t-t1) is the lower limit of the detection time range, and (t+t2) is the upper limit of the detection time range.

4. The method according to claim 2, characterized in that: Detecting a mutation image frame in an image frame sequence and marking the mutation image frame as a node image frame, including: extracting a subsequence with a timestamp within the detection time range from the image frame sequence; Reading the first frame of the subsequence and converting it into a grayscale image as a reference for subsequent frame difference calculation; Loop through each frame of the subsequence, convert it to a grayscale image, and calculate the absolute difference image D(x,y) between the current frame and the previous frame: D(x,y)=∣I cur (x,y)−I prev (x,y)∣ Among them, I cur (x, y) is the grayscale image of the current frame, which is the grayscale value of the pixel at the (x, y) position; I prev (x, y) is the grayscale image of the previous frame, which is the grayscale value of the pixel at the (x, y) position; Compare the pixel value of the difference image with a preset threshold, record the pixel value exceeding the threshold as 1, and record the pixel value not exceeding the threshold as 0; Calculate the percentage of non-zero pixels to the total number of pixels; If the percentage exceeds the set threshold, the number of the current frame is added to the mutation frame list; After processing all frames, return a numbered list of mutation frames.

5. The method according to claim 1, characterized in that The video is divided into a plurality of video segments according to the node image frames, and classification labels are added to the plurality of video segments respectively by using image recognition technology, including: Segment the video at corresponding time positions of the video according to timestamps of node image frames to obtain multiple video segments; Using an action recognition model to identify one or more human actions corresponding to the video clip; extracting a reference image from a video segment preceding the video segment, and extracting a plurality of sample images from the video segment; Determine the newly added components in the sample images and identify the component types of the newly added components based on the plurality of sample images and the reference image using the object recognition model; Presetting the difficulty quantization value of the human action type and the quantization value corresponding to the component type, calculating the weighted sum of the quantization values ​​of the human action and the component type, and obtaining the difficulty evaluation value; The component type and the difficulty evaluation value are set as classification labels of the video clip.

6. The method according to claim 5, characterized in that Preset the difficulty quantization value of the human action type and the quantization value corresponding to the component type, calculate the weighted sum of the quantization values ​​of the human action and the component type, and obtain the difficulty evaluation value, including: According to the difficulty quantization value of the preset human action type, the human action is converted into a human action quantization value; According to the preset quantitative value of the component type, the identified component type is converted into the component type quantitative value; Calculate the weighted sum of the human body motion quantization value and the component type quantization value; All video clips are traversed to obtain a plurality of the weighted sums, normalization coefficients of the plurality of the weighted sums are calculated, and the normalization coefficients are set as difficulty evaluation values ​​of the corresponding video clips.

7. The method according to claim 1, characterized in that Generating compression parameters for corresponding video segments according to the classification labels, so as to compress the video segments according to the compression parameters, comprises: Determining the importance level of the video clip according to the classification label; The compression parameters are determined according to the preset corresponding relationship between the importance level and the compression parameters and the importance level of the video clip, and the compression parameters include a bit rate and a frame skipping rate.

8. A dynamic compression system for teaching videos, characterized in that: include: A conversion module, used for converting a video into an image frame sequence; A detection module, used for detecting a mutation image frame in an image frame sequence, and marking the mutation image frame as a node image frame; An identification module, used to divide the video into multiple video segments according to the node image frames, and add classification labels to the multiple video segments respectively by using image recognition technology; A compression module is used to generate compression parameters for corresponding video segments according to the classification labels, so as to compress the video segments according to the compression parameters.

9. A device, characterized in that: include: A memory for storing a dynamic compression program of the teaching video; A processor is used to implement the steps of the dynamic compression method of the teaching video as described in any one of claims 1-7 when executing the dynamic compression program of the teaching video.

10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a dynamic compression program for teaching videos, and when the dynamic compression program for teaching videos is executed by a processor, the steps of the dynamic compression method for teaching videos as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Method for detecting intelligent identification operation performance of security inspection equipment, detection system and computing equipment

    CN120803768A

  • A method and a detection system and a computing device for detecting intelligent recognition operation performance of security inspection equipment

    CN120803768B