Video processing method and system based on multi-thread synchronization

By extracting the motion law characteristics and dynamic relationship discrete representation values of the moving target from the industrial monitoring video data, setting discrete category labels, and adapting to the transmission method, the transmission efficiency and quality problems caused by the unconsidered relationship of the moving target are solved, and efficient video data transmission is achieved.

CN120281920AActive Publication Date: 2025-07-08HEBEI BOTU COMMUNICATION TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510642623.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-08
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The prior art fails to effectively consider the dynamic relationship of moving targets in video data in the field of industrial monitoring, especially the situation where the movement situation of the sub-motion targets is relatively discrete and deviates from the overall movement situation, resulting in large errors in inter-frame prediction results and low transmission efficiency and quality.

Method used

The processing end obtains the video segment every predetermined period, extracts the motion pattern characteristics of the moving target, calculates the dynamic relationship discrete representation value, sets the dynamic relationship discrete category label for the video segment, controls the data acquisition source for adaptive transmission, including the transmission or complete transmission of keyframes and prediction information, and performs reflow verification to adjust the extraction interval.

Benefits of technology

While reducing the amount of video transmission, the video transmission efficiency and quality are improved, and the transmission method is adaptively adjusted to reduce inter-frame prediction errors, ensuring the actual reflection effect of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281920A_ABST
    Figure CN120281920A_ABST
Patent Text Reader

Abstract

The invention relates to the field of video processing, in particular to a video processing method and system based on multi-thread synchronization. Setting a dynamic relation discrete category label for the video segment based on the motion inclusion relation of each moving target in the video segment and the dynamic relation discrete representation value, and adaptively transmitting the video segment based on the dynamic relation discrete category label of the video segment subsequently, including adaptively determining an extraction interval and extracting a demand key frame; transmitting each demand key frame and prediction information to a receiving end, or completely transmitting a video segment to the receiving end, and subsequently, performing backflow verification and adjusting an extraction interval by the receiving end; according to the method, when a large number of video data are transmitted, the influence on inter-frame prediction under the condition that the motion situation of the sub-motion target in the video data is relatively discrete and deviates from the overall motion situation of the motion target is considered, the transmission mode is adaptively adjusted, and the video transmission efficiency and the quality of the transmitted video data are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing, and in particular, to a video processing method and system based on multi-thread synchronization. Background Art

[0002] Multi-threaded video transmission refers to the transmission of video data through multiple threads, which is applicable to situations where the volume of video data to be transmitted is large. In the field of video surveillance technology, due to the large volume of surveillance videos, in order to avoid congestion, multi-threaded transmission is usually adopted to transmit videos. At the same time, a processing end can be introduced to perform specific video data analysis on the video data to be transmitted, optimize the transmission method, and thus minimize the amount of data during the video data transmission process as much as possible.

[0003] For example, Chinese Patent Publication No.: CN119865606A discloses an inter-frame prediction optimization method based on template matching and multi-reference prediction block fusion, belonging to the field of video coding, including: obtaining the current prediction block of a video frame; searching for a prediction block matching the current prediction block in the current reference frame based on the encoder through template matching technology; obtaining an additional reference frame, searching for a prediction block matching the current prediction block in the additional reference frame to obtain candidate prediction blocks; performing weighted fusion processing on the current prediction block, the prediction block matching the current prediction block searched in the current frame, and the candidate prediction blocks to generate a new prediction result; comparing and analyzing the original current prediction block result and the new prediction result, and selecting the final prediction result; based on the decoder, reproducing the generation process of the final prediction result of the encoder, and adding the final prediction result and the residual data for reconstruction to obtain the decoded video frame.

[0004] However, the following problems still exist in the prior art;

[0005] In the field of industrial monitoring, if a large number of data acquisition sources are deployed, a huge amount of video data will be generated, and the data transmission volume is large. Moreover, the video data in the industrial monitoring field usually contains moving targets. In the prior art, the dynamic relationship of moving targets in a video segment is not considered. In particular, for moving targets with a movement inclusion relationship, there may be sub-moving targets inside whose movement trend is different from the overall movement trend of the moving target. When the movement trend of the sub-moving target is relatively discrete and deviates from the overall movement trend of the moving target, it may affect the inter-frame prediction result, and the transmission method is not adaptively changed, which easily reduces the quality of the transmitted video and results in low transmission efficiency. Summary of the Invention

[0006] To this end, the present invention provides a video processing method and system based on multi-thread synchronization, which is used to overcome the problems in the prior art that when transmitting video data containing moving targets, the dynamic relationship of the moving targets in the video data is not considered, the inter-frame prediction result is affected under the condition that the motion postures of the sub-moving targets are relatively discrete and deviate from the overall motion posture of the moving targets, the transmission method is not adaptively changed, the quality of the transmitted video is easily reduced, and the transmission efficiency is relatively low.

[0007] To achieve the above object, on the one hand, the present invention provides a video processing method based on multi-thread synchronization, including:

[0008] The processing end obtains video segments in the video data collected by the multi-thread data acquisition source at each predetermined period;

[0009] The processing end extracts the moving targets in the video frames corresponding to the video segments, and analyzes the motion laws for each moving target, including determining the motion inclusion relationship of each moving target, and only determining the motion law characteristics of the sub-moving targets in the moving targets with motion inclusion relationships;

[0010] The processing end calculates the dynamic relationship discrete representation value for the video segment based on the motion law characteristics, and sets a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship and the dynamic relationship discrete representation value of each moving target in the video segment;

[0011] The processing end controls the data acquisition source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including,

[0012] Determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data acquisition source to extract the required key frames at the corresponding extraction interval, and transmitting each of the required key frames and the prediction information to the receiving end;

[0013] Or, controlling the data acquisition source to transmit the video data completely to the receiving end;

[0014] The receiving end generates a predicted video based on the required key frames and the prediction information, intercepts a predetermined proportion of the predicted video sub-segments and then performs return flow verification, so that the processing end controls the data acquisition source to adjust the extraction interval;

[0015] Wherein, the motion law characteristics include the number of sub-moving targets and the multi-dimensional discrete values of the motion vectors corresponding to the sub-moving targets, and the return flow verification includes that the receiving end transmits the predicted video sub-segment back to the processing end, and the processing end performs similarity verification on the predicted video sub-segment.

[0016] Further, the process of determining the motion inclusion relationship of each moving target includes,

[0017] Labeling all the moving targets in the video segment;

[0018] If there are other moving objects included in a moving object, it is determined that there is a moving inclusion relationship for the moving object, and each of the other moving objects is determined as a sub-moving object.

[0019] Further, the process of determining the motion law characteristics of the sub-moving objects among the moving objects with a moving inclusion relationship includes:

[0020] Determine the number of sub-moving objects;

[0021] Taking the center of the sub-moving object as the starting point and the moving direction as the vector direction, construct a motion vector for the sub-moving object;

[0022] Determine the average angular acceleration and average velocity of the sub-moving object according to the motion vector;

[0023] Determine the first absolute variance of the average angular acceleration corresponding to each sub-moving object;

[0024] Determine the second absolute variance of the average velocity corresponding to each sub-moving object;

[0025] Weighted sum the first absolute variance and the second absolute variance to obtain the multi-dimensional discrete value.

[0026] Further, the process by which the processing end calculates the dynamic relationship discrete representation value for the video segment based on the motion law characteristics includes:

[0027] Determine the ratio of the number of sub-moving objects to a preset threshold of the number of sub-moving objects as the first dynamic relationship discrete representation factor;

[0028] Determine the ratio of the multi-dimensional discrete value to a preset multi-dimensional discrete threshold as the second dynamic relationship discrete representation factor;

[0029] Weighted sum the first dynamic relationship discrete representation factor and the second dynamic relationship discrete representation factor to obtain the dynamic relationship discrete representation value.

[0030] Further, the process of setting a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship and the dynamic relationship discrete representation value of each moving object in the video segment includes:

[0031] If the dynamic relationship discrete representation value is greater than or equal to a preset dynamic relationship discrete threshold, set a strong dynamic relationship discrete category label for the video segment;

[0032] If the dynamic relationship discrete representation value of the video segment is less than the preset dynamic relationship discrete threshold, set a weak dynamic relationship discrete category label for the video segment;

[0033] Among them, if there is no motion inclusion relationship among the moving targets in the video segment, the discrete representation value of the dynamic relationship corresponding to the video segment is set to 0.

[0034] Further, the process of the processing end transmitting the video segment based on the discrete category label of the dynamic relationship of the video segment includes,

[0035] If the video segment has a weak discrete category label, the extraction interval and the corresponding prediction information are determined based on the discrete representation value of the dynamic relationship, and the data acquisition source is controlled to extract the required key frames at the corresponding extraction interval, and each of the required key frames and the prediction information is transmitted to the receiving end;

[0036] If the video segment has a strong discrete category label, the video segment is completely transmitted to the receiving end.

[0037] Further, determining the extraction interval and the corresponding prediction information based on the discrete representation value of the dynamic relationship, where,

[0038] The determined extraction interval is negatively correlated with the discrete representation value of the dynamic relationship;

[0039] The prediction information includes the number of frames to be predicted, and the number of frames corresponds to the extraction interval.

[0040] Further, the process of the receiving end generating a predicted video based on the required key frames and the prediction information includes determining the number of frames to be predicted based on the prediction information;

[0041] Predicting the predicted frames of the number of frames based on the required key frames and the adjacent required key frames;

[0042] Inserting the predicted frames between the required key frames and the adjacent required key frames based on the time sequence to obtain the predicted video.

[0043] Further, the process of the processing end performing similarity verification on the predicted video sub-segment includes,

[0044] Determining the video sub-segment corresponding to the predicted video sub-segment in the original video data;

[0045] Respectively extracting several groups of video frames corresponding to the time sequence from the predicted video sub-segment and the video sub-segment for comparison;

[0046] Determining the similarity of the video frames within the video frame group, and solving the average value of the similarities corresponding to several video frame groups;

[0047] If the average value of the similarities is less than the preset feature similarity standard value, it is determined that the similarity verification fails, and the extraction interval needs to be reduced;

[0048] If the average value of the similarities is greater than or equal to the preset feature similarity standard value, it is determined that the similarity verification passes;

[0049] Among them, two video frames corresponding to the time sequence are included in the video frame group.

[0050] On the other hand, a system applying a video processing method based on multi-thread synchronization is also provided, including

[0051] A data receiving module, which is used to obtain video segments in the video data collected by the multi-thread data acquisition source every predetermined period;

[0052] A data extraction module, which is used to extract moving targets in the video frames corresponding to the video segments, analyze the motion laws of each moving target, including determining the motion inclusion relationship of each moving target, and only determining the motion law characteristics of the sub-moving targets among the moving targets with a motion inclusion relationship;

[0053] A category label setting module, which is used to calculate the dynamic relationship discrete representation value for the video segment based on the motion law characteristics, and set a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship and the dynamic relationship discrete representation value of each moving target in the video segment;

[0054] A transmission module, which is used to control the data acquisition source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including

[0055] Determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data acquisition source to extract the required key frames at the corresponding extraction interval, and transmitting each of the required key frames and the prediction information to the receiving end;

[0056] Or, controlling the data acquisition source to transmit the video data completely to the receiving end;

[0057] A receiving and verification module, which is used to receive the predicted video sub-segments sent by the corresponding receiving end for similarity verification.

[0058] Compared with the prior art, the beneficial effects of the present invention are as follows. The present invention obtains video segments in the video data collected by the multi-threaded data acquisition source at each predetermined period through the processing end. The processing end extracts moving targets in the video frames corresponding to the video segments, analyzes the motion laws of each moving target, and sets a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship and the dynamic relationship discrete representation value among the moving targets in the video segment. Subsequently, the video segment is adaptively transmitted based on the dynamic relationship discrete category label of the video segment, including determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data acquisition source to extract the required key frames at the corresponding extraction interval, and transmitting each of the required key frames and the prediction information to the receiving end, or transmitting the video segment completely to the receiving end. Subsequently, the receiving end performs a return flow verification and adjusts the extraction interval. The present invention takes into account the influence of the relatively discrete motion trend of sub-moving targets in the video data and the deviation from the overall motion trend of the moving target on the inter-frame prediction when transmitting a large amount of video data, and adaptively adjusts the transmission method to ensure the video transmission efficiency and the quality of the transmitted video data.

[0059] In particular, the present invention analyzes the motion laws of each moving target, determines the motion inclusion relationship of the moving target and the motion law characteristics of the sub-moving targets in the moving targets with the motion inclusion relationship. In actual situations, the video data collected in the industrial monitoring field usually contains moving targets, and the corresponding motion inclusion relationships and sub-moving targets of the moving targets may be different. In actual motion, the moving target is generally in a motion trend, for example, moving in a certain direction as a whole. However, there may be multiple sub-moving targets in the moving target, and these sub-moving targets may be different from the overall motion trend of the moving target. For example, although the moving target moves translationally in a certain direction as a whole, each sub-moving target moves in different directions or has different moving methods and is relatively discrete. In this case, the regularity of the entire moving target is poor, and the probability of inter-frame prediction error will increase. When there are more prediction frames, the image features corresponding to the sub-moving targets in the prediction frames are likely to have errors and are not easy to reflect the actual motion situation. On the contrary, if there are no sub-moving targets inside the moving target or the overall motion trend of the sub-moving targets is relatively unified, this type of situation has stronger regularity and a lower inter-frame prediction error rate. Based on this, the present invention takes into account analyzing the motion laws, calculating the dynamic relationship discrete representation value, representing the situation where the motion trend of the sub-moving targets in the moving target is relatively discrete and deviates from the overall motion trend of the moving target, and setting the discrete category label, thereby facilitating the subsequent adaptive transmission of the video segment, and being able to ensure the video transmission efficiency and the quality of the transmitted video data on the premise of reducing the video transmission volume.

[0060] In particular, for the case of setting a strongly discrete category label with a dynamic relationship for a video segment, it characterizes that the motion trend of sub-moving objects in the video data is relatively discrete and deviates significantly from the overall motion trend of the moving object. In this case, the inter-frame prediction error is large. Therefore, for the case of transmitting a video segment in a complete transmission manner, for the case of setting a weakly discrete category label with a dynamic relationship for a video segment, inter-frame prediction is performed, and at the same time, the corresponding extraction interval is set through the matching of the discrete characterization value of the dynamic relationship, so that the inter-frame prediction can match the discrete difference between the sub-moving object and the overall motion trend in the moving object in the video segment, and the prediction information is synchronously sent to the receiving end, enabling the receiving end to change the inter-frame prediction method in a matching manner. Furthermore, while reducing the data transmission volume, the error caused by inter-frame prediction is reduced, ensuring the video transmission efficiency and the quality of the transmitted video data.

[0061] In particular, the receiving end of the present invention intercepts a predetermined proportion of the predicted video sub-segment for feedback verification, so that the processing end controls the data acquisition source to adjust the extraction interval, and the predicted video sub-segment is transmitted back to the processing end with less bandwidth, considering whether the adjusted inter-frame prediction method adapts to the motion trend of the moving object and the sub-moving object included therein in the current video end. Furthermore, the error caused by inter-frame prediction is reduced, ensuring the video transmission efficiency and the quality of the transmitted video data. Description of the Drawings

[0062] Figure 1 It is a step diagram of the multi-threaded synchronous video processing method according to an embodiment of the present invention;

[0063] Figure 2 It is a logic block diagram for setting a discrete category label with a dynamic relationship for a video segment according to an embodiment of the present invention;

[0064] Figure 3 It is a logic block diagram of the process for adjusting the extraction interval according to an embodiment of the present invention;

[0065] Figure 4 It is a logic block diagram of the similarity verification process according to an embodiment of the present invention Detailed Embodiments

[0066] In order to make the purpose and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0067] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0068] It should be noted that in the description of the present invention, in addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0069] Please refer to Figure 1 as shown in the figure, which is the step diagram of the multi-threaded synchronous video processing method according to the embodiment of the present invention. The video processing method based on multi-threaded synchronization according to the embodiment of the present invention includes:

[0070] Step S1, the processing end obtains video segments in the video data collected by the multi-threaded data acquisition source at every predetermined period;

[0071] Step S2, the processing end extracts moving targets in the video frames corresponding to the video segments, and analyzes the motion laws for each moving target, including determining the motion inclusion relationship of each moving target, and only determining the motion law characteristics of the sub-moving targets in the moving targets with motion inclusion relationship;

[0072] Step S3, the processing end calculates the dynamic relationship discrete representation value for the video segment based on the motion law characteristics, and sets a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship of each moving target in the video segment and the dynamic relationship discrete representation value;

[0073] Step S4, the processing end controls the data acquisition source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including,

[0074] determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data acquisition source to extract the required key frames at the corresponding extraction interval, and transmitting each of the required key frames and the prediction information to the receiving end;

[0075] or, controlling the data acquisition source to transmit the video data completely to the receiving end;

[0076] Step S5, the receiving end generates a predicted video based on the required key frames and the prediction information, intercepts a predetermined proportion of the predicted video sub-segments and then performs a feedback verification to enable the processing end to control the data acquisition source to adjust the extraction interval;

[0077] wherein, the motion law characteristics include the number of sub-moving targets and the multi-dimensional discrete values of the motion vectors corresponding to the sub-moving targets, and the feedback verification includes that the receiving end transmits the predicted video sub-segment back to the processing end, and the processing end performs a similarity verification on the predicted video sub-segment.

[0078] Specifically, the form of the multi-threaded data acquisition source is not limited. It can be an image acquisition module with multi-threaded transmission capabilities. Preferably, the image acquisition module may include an image acquisition device with image acquisition functions and a processor to respond to the information sent by the processing end. Those skilled in the art can choose by themselves, and this will not be elaborated here.

[0079] Specifically, the specific form of the processing end is not limited. It only needs to be able to receive the multi-threaded video data sent by the data acquisition source, analyze it, and control the data acquisition source. It can be composed of logic components to process the video data. The logic components include a field programmable processor, a computer, or a microprocessor in the computer.

[0080] Specifically, when the processing end acquires video segments in the video data collected by the multi-threaded data acquisition source at each predetermined period, to reduce the data transmission volume, and the video segments have a certain data representativeness and can reflect the movement laws of moving targets in the video data, the acquisition ratio is selected between 10% and 20% of the predetermined period, and the predetermined period is selected within the interval [1.5s, 3s].

[0081] Specifically, the process of determining the movement inclusion relationship of each moving target includes,

[0082] Label all moving targets in the video segment;

[0083] If there are moving targets that contain other moving targets, it is determined that there is a movement inclusion relationship for the moving targets, and each of the other moving targets is determined as a sub-moving target.

[0084] Specifically, the method for determining moving targets in the video segment is not limited. For example, the inter-frame difference method can be used to compare the pixel differences between consecutive frames, detect the changed areas, and then identify the moving targets, or an existing model based on object detection can be used, such as the YOLOv8 video processing model. Of course, other methods can also be used, and this will not be elaborated here.

[0085] Specifically, the process of determining the movement law characteristics of sub-moving targets among moving targets with a movement inclusion relationship includes,

[0086] Determine the number of sub-moving targets;

[0087] Taking the center of the sub-moving target as the starting point and the movement direction as the vector direction, construct a movement vector for the sub-moving target;

[0088] Determine the average angular acceleration and average velocity of the sub-moving target according to the movement vector;

[0089] Determine the first absolute variance of the mean angular acceleration corresponding to each sub - motion target;

[0090] Determine the second absolute variance of the mean velocity corresponding to each sub - motion target;

[0091] Weightedly sum the first absolute variance and the second absolute variance to obtain the multi - dimensional discrete value.

[0092] Specifically, in implementation, the weight of the first absolute variance is 0.6, and the weight of the second absolute variance is 0.4.

[0093] It can be understood that the purpose of determining the motion vector is to determine the motion speed and motion direction of a certain point. Those skilled in the art can determine the mean angular acceleration and mean velocity of that point once they know the motion speed and motion direction of that point.

[0094] Specifically, the process by which the processing end calculates the discrete representation value of the dynamic relationship for the video segment based on the motion law characteristics includes,

[0095] Determine the ratio of the number of sub - motion targets to the preset threshold of the number of sub - motion targets as the first discrete representation factor of the dynamic relationship;

[0096] Determine the ratio of the multi - dimensional discrete value to the preset multi - dimensional discrete threshold as the second discrete representation factor of the dynamic relationship;

[0097] Weightedly sum the first discrete representation factor of the dynamic relationship and the second discrete representation factor of the dynamic relationship to obtain the discrete representation value of the dynamic relationship.

[0098] In implementation, the multi - dimensional discrete threshold is determined in advance. Among them, several video segments in the video data corresponding to the motion targets are collected in advance, the motion laws are analyzed to determine the multi - dimensional discrete values corresponding to each video segment, the mean value of the multi - dimensional discrete values is solved, and the multi - dimensional discrete threshold is selected between 1.3 times and 1.5 times the mean value of the multi - dimensional discrete values to represent the case where the multi - dimensional discrete value is relatively large.

[0099] In implementation, the threshold of the number of sub - motion targets is determined in advance. Among them, several video segments in the video data corresponding to the motion targets are collected in advance, the motion laws are analyzed to determine the number of sub - motion targets corresponding to each video segment, the mean value of the number of sub - motion targets is solved, and the mean value of the number of sub - motion targets is determined as the threshold of the number of sub - motion targets.

[0100] In implementation, the weight of the first discrete representation factor of the dynamic relationship is 0.3, and the weight of the second discrete representation factor of the dynamic relationship is 0.7.

[0101] Specifically, the present invention analyzes the motion laws of each moving target, determines the motion inclusion relationship of the moving target and the motion law characteristics of the sub-moving targets in the moving target with the motion inclusion relationship. In actual situations, the video data collected in the industrial monitoring field usually contains moving targets. The motion inclusion relationship corresponding to the moving target and the sub-moving targets may be different. In actual motion, the moving target as a whole is in a motion state. For example, it moves in a certain direction as a whole. However, there may be multiple sub-moving targets in the moving target, and these sub-moving targets may be different from the overall motion state of the moving target. For example, although the moving target moves translationally in a certain direction as a whole, each sub-moving target moves in different directions or has a relatively discrete moving manner. In this case, the regularity of the entire moving target is poor, and the probability of frame-inter prediction error will increase. When there are more prediction frames, the image features corresponding to the sub-moving targets in the prediction frames are likely to have errors and are not easy to reflect the actual motion situation. On the contrary, if there are no sub-moving targets inside the moving target or the overall motion states of the sub-moving targets are relatively unified, the regularity of this type of situation is strong and the frame-inter prediction error rate is low. Based on this, the present invention considers performing motion law analysis, calculating the discrete representation value of the dynamic relationship, representing the situation where the motion states of the sub-moving targets in the moving target are relatively discrete and deviate from the overall motion state of the moving target, setting discrete category labels, and then facilitating the subsequent adaptive transmission of video segments, and being able to ensure the video transmission efficiency and the quality of the transmitted video data on the premise of reducing the video transmission volume.

[0102] Please refer to Figure 2 as shown, which is the logic block diagram for setting the discrete category label of the dynamic relationship for the video segment in the embodiment of the present invention. Based on the motion inclusion relationship and the discrete representation value of the dynamic relationship of each moving target in the video segment, the process of setting the discrete category label of the dynamic relationship for the video segment includes

[0103] If the discrete representation value of the dynamic relationship is greater than or equal to the preset discrete threshold of the dynamic relationship, set the strong discrete category label of the dynamic relationship for the video segment;

[0104] If the discrete representation value of the dynamic relationship of the video segment is less than the preset discrete threshold of the dynamic relationship, set the weak discrete category label of the dynamic relationship for the video segment;

[0105] Among them, if there is no motion inclusion relationship for each moving target in the video segment, set the discrete representation value of the dynamic relationship corresponding to the video segment to 0.

[0106] Specifically, the discrete threshold of the dynamic relationship is selected within the interval [1.12, 1.24].

[0107] Specifically, for the case of setting a strongly discrete category label for a video segment, it characterizes that the motion trends of sub-moving objects in the video data are relatively discrete and deviate significantly from the overall motion trend of the moving object. In this case, the inter-frame prediction error is relatively large. Therefore, the video segment is transmitted in a complete transmission manner. For the case of setting a weakly discrete category label for the video segment, inter-frame prediction is performed, and at the same time, the corresponding extraction interval is set through the matching of the discrete representation value of the dynamic relationship, so that the inter-frame prediction can match the discrete difference between the sub-moving object and the overall motion trend in the moving object of the video segment. The prediction information is synchronously sent to the receiving end, enabling the receiving end to change the inter-frame prediction method in a matching manner. Furthermore, while reducing the data transmission volume, the error caused by inter-frame prediction is reduced, ensuring the video transmission efficiency and the quality of the transmitted video data.

[0108] Specifically, the process of the processing end transmitting the video segment based on the discrete category label of the dynamic relationship of the video segment includes

[0109] If the video segment has a weakly discrete category label, the extraction interval and the corresponding prediction information are determined based on the discrete representation value of the dynamic relationship. The data acquisition source is controlled to extract the required key frames at the corresponding extraction interval, and each of the required key frames and the prediction information are transmitted to the receiving end;

[0110] If the video segment has a strongly discrete category label, the video segment is completely transmitted to the receiving end.

[0111] It can be understood that the discrete category label of the dynamic relationship is a virtual label, and its purpose is to classify the video segment, which will not be elaborated here.

[0112] Specifically, the extraction interval and the corresponding prediction information are determined based on the discrete representation value of the dynamic relationship, where

[0113] The determined extraction interval is negatively correlated with the discrete representation value of the dynamic relationship;

[0114] The prediction information includes the number of frames to be predicted, and the number of frames corresponds to the extraction interval.

[0115] In implementation, optionally,

[0116] The discrete representation value of the dynamic relationship is compared with the first preset discrete reference value of the dynamic relationship and the second preset discrete reference value of the dynamic relationship.

[0117] If the discrete representation value of the dynamic relationship is greater than or equal to the second preset discrete reference value of the dynamic relationship, the extraction interval is set to the integer part of 0.65 times the reference extraction interval;

[0118] If the discrete representation value of the dynamic relationship is less than the second preset discrete reference value of the dynamic relationship and greater than the first preset discrete reference value of the dynamic relationship, the extraction interval is set to the reference extraction interval;

[0119] If the discrete representation value of the dynamic relationship is less than or equal to the first preset discrete reference value of the dynamic relationship, the extraction interval is set to the integer obtained by multiplying the reference extraction interval by 1.35;

[0120] The first preset discrete reference value of the dynamic relationship is 0.55 times the discrete threshold of the dynamic relationship, and the second preset discrete reference value of the dynamic relationship is 0.75 times the discrete threshold of the dynamic relationship.

[0121] Specifically, the reference extraction interval is selected within the interval [5, 15].

[0122] Specifically, the process by which the receiving end generates a predicted video based on the required key frames and prediction information includes

[0123] Determining the number of frames to be predicted based on the prediction information;

[0124] Predicting the predicted frames of the number of frames based on the required key frames and adjacent required key frames;

[0125] Inserting the predicted frames into the required key frames and adjacent required key frames in chronological order to obtain the predicted video.

[0126] Specifically, there is no limitation on the method of predicting the predicted frames of the number of frames based on the required key frames and adjacent required key frames. For example, it can be implemented by using an inter-frame prediction method. A corresponding inter-frame prediction model or algorithm is deployed at the receiving end to predict the intermediate frames based on the front and back key frames, which will not be elaborated here.

[0127] It can be understood that for inter-frame prediction, the more frames need to be predicted, the greater the possibility of prediction error, which will not be elaborated here.

[0128] It can be understood that the processing end can send the prediction information and the extraction interval to the data acquisition source to inform the data acquisition source of the required extraction interval, so that the data acquisition source extracts the required key frames at the corresponding extraction interval and synchronously transmits the prediction information and the required key frames to the receiving end.

[0129] Please refer to Figure 3 and Figure 4 as shown Figure 3 is the logic block diagram of the process of adjusting the extraction interval according to the embodiment of the present invention, Figure 4 is the logic block diagram of the similarity verification process according to the embodiment of the present invention. Specifically, the process by which the processing end performs similarity verification on the predicted video segment includes

[0130] Determine the video segment in the original video data corresponding to the predicted video segment. It can be understood that the predicted video segment corresponds to the video segment in the time domain dimension to ensure comparability;

[0131] Extract several groups of video frames corresponding in time sequence from the predicted video segment and the video segment respectively for comparison;

[0132] Determine the similarity of the video frames within the group of video frames, and solve the average value of the similarities corresponding to several groups of video frames;

[0133] If the average similarity value is less than the preset feature similarity standard value, it is determined that the similarity verification fails, and the extraction interval needs to be reduced;

[0134] If the average similarity value is greater than or equal to the preset feature similarity standard value, it is determined that the similarity verification passes;

[0135] Among them, the group of video frames contains two video frames corresponding in time sequence.

[0136] Specifically, the method for determining the image similarity is not limited. For example, the method of calculating the SSIM (Structural Similarity) can be used for determination. The SSIM value between two video frames can be calculated, and the SSIM value is used as the image similarity. The closer the SSIM value is to 1, the more similar the two video frames are. Of course, other methods can also be used, which will not be elaborated here.

[0137] The preset feature similarity standard value is obtained through pre-detection. The video data within several predetermined periods is transmitted to the receiving end in advance to generate a predicted video, and a feedback verification is carried out. The average similarity value during the feedback verification is determined, and the average value of the average similarity values is solved. The product of this average value and the error coefficient is determined as the preset feature similarity standard value, and the error coefficient is selected within the interval [0.85, 0.95].

[0138] Specifically, the operation of adjusting the extraction interval is to take the integer after halving the original extraction interval.

[0139] Specifically, the receiving end of the present invention intercepts a predetermined proportion of the predicted video segment and then conducts a feedback verification, so that the processing end controls the data acquisition source to adjust the extraction interval, and the predicted video segment is transmitted back to the processing end with less bandwidth. Consider whether the adjusted inter-frame prediction method adapts to the motion trend of moving targets and sub-moving targets included inside in the current video end. Furthermore, reduce the error caused by inter-frame prediction, and ensure the video transmission efficiency and the quality of the transmitted video data.

[0140] On the other hand, a system applying the video processing method based on multi-thread synchronization is also provided, including,

[0141] A data receiving module, which is used to obtain video segments in the video data collected by a multi-threaded data acquisition source at each predetermined period;

[0142] A data extraction module, which is used to extract moving targets in the video frames corresponding to the video segments, and analyze the motion laws for each moving target, including determining the motion inclusion relationships of the moving targets, and only determining the motion law characteristics of the sub-moving targets among the moving targets with motion inclusion relationships;

[0143] A category label setting module, which is used to calculate the dynamic relationship discrete representation value for the video segment based on the motion law characteristics, and set a dynamic relationship discrete category label for the video segment based on the motion inclusion relationships of the moving targets in the video segment and the dynamic relationship discrete representation value;

[0144] A transmission module, which is used to control the data acquisition source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including,

[0145] Determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data acquisition source to extract the required key frames at the corresponding extraction interval, and transmitting each of the required key frames and the prediction information to the receiving end;

[0146] Or, controlling the data acquisition source to transmit the video data completely to the receiving end;

[0147] A receiving and verification module, which is used to receive the predicted video sub-segments sent by the corresponding receiving end for similarity verification.

[0148] Specifically, the form of the data receiving module is not limited, and it can be a data storage device capable of realizing data interaction and storage, which will not be elaborated here.

[0149] Specifically, the specific forms of the data extraction module, the category label setting module, the transmission module, and the receiving and verification module are not limited, and they can be composed of logical components, which will not be elaborated here.

[0150] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. A video processing method based on multithread synchronization, characterized in that Including: The processing end obtains video segments in the video data collected by the multi-threaded data acquisition source at each predetermined period; The processing end extracts moving targets in the video frames corresponding to the video segments, and analyzes the motion laws for each moving target, including determining the motion inclusion relationship of each moving target, and only determining the motion law characteristics of the sub-moving targets among the moving targets with motion inclusion relationships; The processing end calculates the dynamic relationship discrete representation value for the video segment based on the motion law characteristics, and sets a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship of each moving target in the video segment and the dynamic relationship discrete representation value; The processing end controls the data acquisition source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including, Determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data acquisition source to extract the required key frames at the corresponding extraction interval, and transmitting each of the required key frames and the prediction information to the receiving end; Or, controlling the data acquisition source to transmit the video data completely to the receiving end; The receiving end generates a predicted video based on the required key frames and the prediction information, intercepts a predetermined proportion of the predicted video sub-segments and then performs a return flow verification, so that the processing end controls the data acquisition source to adjust the extraction interval; Wherein, the motion law characteristics include the number of sub-moving targets and the multi-dimensional discrete values of the motion vectors corresponding to the sub-moving targets, and the return flow verification includes that the receiving end transmits the predicted video sub-segment back to the processing end, and the processing end performs a similarity verification on the predicted video sub-segment.

2. The video processing method based on multi-thread synchronization according to claim 1, wherein, The process of determining the motion inclusion relationship of each moving target includes, Labeling all moving targets in the video segment; If there are moving targets that contain other moving targets, it is determined that the moving targets have a motion inclusion relationship, and each of the other moving targets is determined as a sub-moving target.

3. The video processing method based on multi-thread synchronization according to claim 2, wherein The process of determining the motion law characteristics of the sub-moving targets among the moving targets with motion inclusion relationships includes, Determining the number of sub-moving targets; Taking the center of the sub-moving target as the starting point and the motion direction as the vector direction to construct a motion vector for the sub-moving target; Determining the average angular acceleration and the average speed of the sub-moving target according to the motion vector; Determining the first absolute variance of the average angular acceleration corresponding to each sub-moving target; Determining the second absolute variance of the average speed corresponding to each sub-moving target; Weighted summing the first absolute variance and the second absolute variance to obtain the multi-dimensional discrete value.

4. The video processing method based on multithread synchronization according to claim 1, wherein, The process by which the processing end calculates the dynamic relationship discrete representation value for the video segment based on the motion law characteristics includes, Determining the ratio of the number of sub-moving targets to the preset threshold of the number of sub-moving targets as the first dynamic relationship discrete representation factor; Determining the ratio of the multi-dimensional discrete value to the preset multi-dimensional discrete threshold as the second dynamic relationship discrete representation factor; Weighted summing the first dynamic relationship discrete representation factor and the second dynamic relationship discrete representation factor to obtain the dynamic relationship discrete representation value.

5. The video processing method based on multithread synchronization according to claim 1, characterized in that, The process of setting a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship of each moving target in the video segment and the dynamic relationship discrete representation value includes, If the dynamic relationship discrete representation value is greater than or equal to the preset dynamic relationship discrete threshold, set a strong discrete category label for the video segment; If the dynamic relationship discrete representation value of the video segment is less than the preset dynamic relationship discrete threshold, set a weak discrete category label for the video segment; Among them, if there is no motion inclusion relationship among the moving targets in the video segment, set the dynamic relationship discrete representation value corresponding to the video segment to 0.

6. The video processing method based on multi-thread synchronization according to claim 1, characterized in that The process of the processing end transmitting the video segment based on the dynamic relationship discrete category label of the video segment includes If the video segment has a weak discrete category label, determine the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, control the data acquisition source to extract the required key frames at the corresponding extraction interval, and transmit each of the required key frames and the prediction information to the receiving end; If the video segment has a strong discrete category label, transmit the video segment completely to the receiving end.

7. The video processing method based on multi-thread synchronization according to claim 1, characterized in that Determine the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, where The determined extraction interval is negatively correlated with the dynamic relationship discrete representation value; The prediction information includes the number of frames to be predicted, and the number of frames corresponds to the extraction interval.

8. The video processing method based on multi-thread synchronization according to claim 1, wherein The process of the receiving end generating a predicted video based on the required key frames and the prediction information includes Determine the number of frames to be predicted based on the prediction information; Predict the predicted frames of the number of frames based on the required key frames and the adjacent required key frames; Insert the predicted frames into the required key frames and the adjacent required key frames based on the time sequence to obtain the predicted video.

9. The video processing method based on multithread synchronization according to claim 1, wherein The process of the processing end performing similarity verification on the predicted video sub-segment includes Determine the video sub-segment corresponding to the predicted video sub-segment in the original video data; Extract several video frame groups corresponding to the time sequence from the predicted video sub-segment and the video sub-segment respectively for comparison; Determine the similarity of the video frames within the video frame group, and solve the average similarity corresponding to several video frame groups; If the average similarity is less than the preset feature similarity standard value, it is determined that the similarity verification fails, and the extraction interval needs to be reduced; If the average similarity is greater than or equal to the preset feature similarity standard value, it is determined that the similarity verification passes; Among them, the video frame group contains two video frames corresponding to the time sequence.

10. A system applying the video processing method based on multi-thread synchronization according to any one of claims 1-9, characterized in that, Include A data receiving module, which is used to obtain video segments in the video data collected by the multi-threaded data acquisition source at a predetermined period; A data extraction module, which is used to extract the moving targets in the video frames corresponding to the video segment, analyze the motion laws of each moving target, including determining the motion inclusion relationship of each moving target, and only determining the motion law characteristics of the sub-moving targets among the moving targets with motion inclusion relationships; A category label setting module, which is used to calculate the dynamic relationship discrete representation value for the video segment based on the motion law characteristics, and set the dynamic relationship discrete category label for the video segment based on the motion inclusion relationship and the dynamic relationship discrete representation value of each moving target in the video segment; A transmission module, which is used to control the data acquisition source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including Determine the extraction interval and corresponding prediction information based on the discrete representation value of the dynamic relationship, control the data acquisition source to extract the required key frames at the corresponding extraction interval, and transmit each of the required key frames and the prediction information to the receiving end; Alternatively, control the data acquisition source to transmit the video data completely to the receiving end; A receiving verification module, which is used to receive the predicted video sub-segment sent by the corresponding receiving end for similarity verification.

Citation Information

Patent Citations

  • Inter-frame prediction optimization method based on template matching and multi-reference prediction block fusion

    CN119865606A

  • Video splicing method based on adaptive key frame sampling

    CN105957017A

  • Safe compression method applied to unattended equipment video

    CN116980613A

  • Remote monitoring video transmission method for 5G Internet of Things

    CN118200490A

  • Method for optimizing route of bus

    KR1020240039705A