A video processing method and system based on multi-thread synchronization

By extracting motion pattern features of moving targets from industrial surveillance video data and setting discrete category labels for dynamic relationships, the problem of inter-frame prediction error caused by the failure to consider the dynamic relationships of moving targets is solved, thus achieving efficient and high-quality video transmission.

CN120281920BActive Publication Date: 2025-12-09HEBEI BOTU COMMUNICATION TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510642623.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-12-09
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

In the field of industrial monitoring, the dynamic relationships of moving targets in video data are not taken into account, resulting in large errors in inter-frame prediction results and low transmission efficiency and quality.

Method used

The processing unit acquires video data at predetermined intervals, extracts motion pattern features of moving targets, sets dynamic relationship discrete category labels, controls the data acquisition source for adaptive transmission, including synchronous or complete transmission of keyframes and prediction information, and performs backflow verification to adjust the extraction interval.

Benefits of technology

While reducing the amount of data transmitted, it improves the efficiency and quality of video transmission, adapts to the dynamic changes of moving targets in video data, and reduces inter-frame prediction errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281920B_ABST
    Figure CN120281920B_ABST
Patent Text Reader

Abstract

The present application relates to the field of video processing, especially to a kind of video processing method and system based on multi-thread synchronization, the method of the present application includes that motion law analysis is carried out for each moving target, dynamic relationship discrete category label is set for video segment based on the motion containing relationship and dynamic relationship discrete representation value of each moving target in video segment, subsequent adaptive dynamic relationship discrete category label based on video segment is used to transmit video segment, including adaptively determining extraction interval extraction demand key frame, each demand key frame and prediction information are transmitted to receiving end, or video segment is completely transmitted to receiving end, subsequently, receiving end carries out backflow verification, and extraction interval is adjusted;The present application considers the influence of the motion trend of the sub-motion target in the video data on the inter-frame prediction when a large amount of video data is transmitted, and adaptively adjusts the transmission mode to ensure the video transmission efficiency and the quality of the transmitted video data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of video processing, in particular to a video processing method and system based on multi-thread synchronization. BACKGROUND

[0002] Multi-thread video transmission refers to transmitting video data through multiple threads, which is suitable for cases where the amount of video data to be transmitted is large. In the field of video monitoring technology, due to the large amount of monitoring video, to avoid congestion, a multi-thread transmission method is usually used to transmit video. At the same time, a processing end can be introduced to perform specific video data analysis on the video data to be transmitted, optimize the transmission method, and thus reduce the amount of data in the video data transmission process as much as possible.

[0003] For example, Chinese Patent Publication No. CN119865606A discloses an inter-frame prediction optimization method based on template matching and multi-reference prediction block fusion, which belongs to the field of video coding. The method includes: obtaining a current prediction block of a video frame; searching for a prediction block matching the current prediction block in a current reference frame based on a template matching technique by an encoder; obtaining an additional reference frame, searching for a prediction block matching the current prediction block in the additional reference frame, and obtaining a candidate prediction block; performing weighted fusion processing on the current prediction block, the prediction block matching the current prediction block searched in the current frame, and the candidate prediction block to generate a new prediction result; comparing and analyzing the original current prediction block result and the new prediction result to select a final prediction result; and based on a decoder, reproducing the generation process of the final prediction result of the encoder, adding the final prediction result and residual data to reconstruct a decoded video frame.

[0004] However, the prior art still has the following problems:

[0005] In the field of industrial monitoring, if a large number of data acquisition sources are deployed, a large amount of video data will be generated, and the amount of data transmission is large. Moreover, the video data in the field of industrial monitoring usually contains moving targets. In the prior art, the dynamic relationship of the moving targets in the video segment is not considered, especially the moving targets that have a motion-included relationship, which may have sub-moving targets that have different motion trends from the overall motion trend of the moving target. When the motion trend of the sub-moving target is relatively discrete and deviates from the overall motion trend of the moving target, it may affect the inter-frame prediction result, and the transmission method is not adapted to change, which may reduce the quality of the transmitted video and result in low transmission efficiency. SUMMARY

[0006] To this end, the application provides a video processing method and system based on multi-thread synchronization to overcome the problem that the prior art does not consider the dynamic relationship of the moving target in the video data when transmitting the video data containing the moving target, which affects the inter-frame prediction result when the motion trend of the sub-moving target is relatively discrete and deviates from the overall motion trend of the moving target, and the transmission mode is not adapted to change, which easily reduces the quality of the transmitted video and results in low transmission efficiency.

[0007] To achieve the above-mentioned purpose, in one aspect, the application provides a video processing method based on multi-thread synchronization, comprising:

[0008] The processing end acquires a video segment in the video data collected by the multi-thread data collection source every predetermined period;

[0009] The processing end extracts the moving target in the video frame corresponding to the video segment, analyzes the motion law of each moving target, including determining the motion inclusion relationship of each moving target, and determining only the motion law feature of the sub-moving target in the moving target with the motion inclusion relationship;

[0010] The processing end calculates a dynamic relationship discrete representation value for the video segment based on the motion law feature, sets a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship of each moving target in the video segment and the dynamic relationship discrete representation value;

[0011] The processing end controls the data collection source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including,

[0012] determining an extraction interval and corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data collection source to extract the required key frame with the corresponding extraction interval, and transmitting each required key frame and the prediction information to the receiving end;

[0013] or, controlling the data collection source to transmit the video data to the receiving end completely;

[0014] The receiving end generates a predicted video based on the required key frame and the prediction information, intercepts a predetermined proportion of the predicted video segment, and then performs a backflow verification to enable the processing end to control the data collection source to adjust the extraction interval;

[0015] wherein the motion law feature includes the number of sub-moving targets and the multi-dimensional discrete value of the motion vector corresponding to the sub-moving target, and the backflow verification includes the receiving end returning the predicted video segment to the processing end, and the processing end performing a similarity verification on the predicted video segment.

[0016] Further, the process of determining the motion inclusion relationship of each moving target includes,

[0017] labeling all moving targets in the video segment;

[0018] If the motion target contains other motion targets, it is determined that the motion target has a motion containing relationship, and each of the other motion targets is determined as a sub-motion target.

[0019] Further, the process of determining the motion regularity characteristics of the sub-motion targets in the motion target having the motion containing relationship includes,

[0020] determining the number of sub-motion targets;

[0021] constructing a motion vector for the sub-motion target with the center of the sub-motion target as the starting point and the motion direction as the vector direction;

[0022] determining the average angular acceleration and the average speed of the sub-motion target according to the motion vector;

[0023] determining the first absolute deviation of the average angular acceleration corresponding to each sub-motion target;

[0024] determining the second absolute deviation of the average speed corresponding to each sub-motion target;

[0025] weighting and summing the first absolute deviation and the second absolute deviation to obtain the multi-dimensional discrete value.

[0026] Further, the process of calculating the dynamic relationship discrete representation value for the video segment based on the motion regularity characteristics by the processing end includes,

[0027] determining the ratio of the number of sub-motion targets to the preset threshold number of sub-motion targets as a first dynamic relationship discrete representation factor;

[0028] determining the ratio of the multi-dimensional discrete value to the preset multi-dimensional discrete threshold value as a second dynamic relationship discrete representation factor;

[0029] weighting and summing the first dynamic relationship discrete representation factor and the second dynamic relationship discrete representation factor to obtain the dynamic relationship discrete representation value.

[0030] Further, the process of setting a dynamic relationship discrete category label for the video segment based on the motion containing relationship and the dynamic relationship discrete representation value of each motion target in the video segment includes,

[0031] if the dynamic relationship discrete representation value is greater than or equal to the preset dynamic relationship discrete threshold value, setting a strong dynamic relationship discrete category label for the video segment;

[0032] if the dynamic relationship discrete representation value of the video segment is less than the preset dynamic relationship discrete threshold value, setting a weak dynamic relationship discrete category label for the video segment;

[0033] If each of the motion targets in the video segment does not have a motion containment relationship, the dynamic relationship discrete representation value corresponding to the video segment is set to 0.

[0034] Further, the process of transmitting the video segment by the processing end based on the dynamic relationship discrete category label of the video segment comprises,

[0035] If the video segment has a weak discrete category label, the extraction interval and the corresponding prediction information are determined based on the dynamic relationship discrete representation value, the data acquisition source is controlled to extract the required key frames at the extraction interval, and the required key frames and the prediction information are transmitted to the receiving end;

[0036] If the video segment has a strong discrete category label, the video segment is completely transmitted to the receiving end.

[0037] Further, the extraction interval and the corresponding prediction information are determined based on the dynamic relationship discrete representation value, wherein,

[0038] The determined extraction interval is negatively correlated with the dynamic relationship discrete representation value;

[0039] The prediction information includes the number of required predicted frames, and the number of frames corresponds to the extraction interval.

[0040] Further, the process of generating a predicted video by the receiving end based on the required key frames and the prediction information comprises determining the number of required predicted frames based on the prediction information.

[0041] The prediction frames of the number of required predicted frames are predicted based on the required key frames and the adjacent required key frames;

[0042] The prediction frames are inserted between the required key frames and the adjacent required key frames based on the time sequence to obtain the predicted video.

[0043] Further, the process of the processing end verifying the similarity of the predicted video sub-segment comprises,

[0044] Determining the corresponding video sub-segment of the predicted video sub-segment in the original video data;

[0045] Comparing a plurality of video frame groups corresponding to the extraction time sequence in the predicted video sub-segment and the video sub-segment, respectively;

[0046] Determining the similarity of the video frames in the video frame group, and solving the average similarity of the plurality of video frame groups;

[0047] If the average similarity is less than a preset feature similarity standard value, it is determined that the similarity verification fails, and the extraction interval needs to be reduced;

[0048] If the average similarity is greater than or equal to the preset feature similarity standard value, it is determined that the similarity verification passes;

[0049] wherein the video frame group comprises two video frames corresponding in time.

[0050] In another aspect, a system applying the video processing method based on multi-thread synchronization is also provided, comprising,

[0051] a data receiving module configured to acquire a video segment in video data collected by a data collection source of the multi-thread at a predetermined period;

[0052] a data extracting module configured to extract a moving object in a video frame corresponding to the video segment, and analyze a motion rule of each moving object, comprising determining a motion inclusion relationship of each moving object, and determining a motion rule feature of a sub-moving object in the moving object only when the motion inclusion relationship exists;

[0053] a category label setting module configured to calculate a dynamic relationship discrete representation value for the video segment based on the motion rule feature, and set a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship of each moving object in the video segment and the dynamic relationship discrete representation value;

[0054] a transmission module configured to control the data collection source to transmit the video data based on the dynamic relationship discrete category label of the video segment, comprising,

[0055] determining an extraction interval and corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data collection source to extract a required key frame at the corresponding extraction interval, and transmitting each required key frame and the prediction information to a receiving end;

[0056] or, controlling the data collection source to transmit the video data to the receiving end completely;

[0057] a receiving verification module configured to receive a predicted video segment corresponding to the receiving end to perform a similarity verification.

[0058] Compared with the prior art, the present application has the beneficial effects that, by setting the processing end to acquire video segments in the video data collected by the multi-threaded data collection source every predetermined period, the processing end extracts the moving targets in the video frames corresponding to the video segments, analyzes the motion law of each moving target, sets the dynamic relationship discrete category label for the video segments based on the motion inclusion relationship of each moving target in the video segments and the dynamic relationship discrete representation value, and subsequently transmits the video segments based on the adaptive dynamic relationship discrete category label of the video segments, including determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data collection source to extract the required key frames with the corresponding extraction interval, and transmitting each of the required key frames and the prediction information to the receiving end, or transmitting the video segments to the receiving end in their entirety, and subsequently, the receiving end performs backflow verification and adjusts the extraction interval; the present application takes into account the influence of the relatively discrete motion state of the sub-movement targets in the video data and the deviation of the motion state from the overall motion state of the movement targets on inter-frame prediction when a large amount of video data is transmitted, adaptively adjusts the transmission mode, and ensures the video transmission efficiency and the quality of the transmitted video data.

[0059] Especially, the present application analyzes the motion law of each moving target, determines the motion inclusion relationship of the moving target and the motion law characteristics of the sub-movement targets in the moving target with the motion inclusion relationship, in actual situations, the video data collected in the industrial monitoring field usually contains moving targets, and the motion inclusion relationship and the sub-movement targets corresponding to the moving targets may be different, in actual motion, the overall moving target is in a motion state, for example, the overall moving target moves in a certain direction, however, there may be multiple sub-movement targets in the moving target, and these sub-movement targets may have differences with the motion state of the overall moving target, for example, although the overall moving target moves in a certain direction, each sub-movement target moves in a different direction or moves in a different way, which is relatively discrete, in this case, the regularity of the overall moving target is poor, and the probability of inter-frame prediction error will increase, if there are many predicted frames, the image features corresponding to the sub-movement targets in the predicted frames are prone to errors, and it is not easy to reflect the actual motion situation, on the contrary, if there are no sub-movement targets in the moving target or the overall motion state of the sub-movement targets is relatively uniform, the regularity of this kind of situation is strong, and the inter-frame prediction error rate is low, based on this, the present application analyzes the motion law, calculates the dynamic relationship discrete representation value, represents the case that the motion state of the sub-movement targets in the moving target is relatively discrete and deviates from the overall motion state of the moving target, sets the discrete category label, and then facilitates the subsequent adaptive transmission of the video segments, which can reduce the video transmission volume while ensuring the video transmission efficiency and the quality of the transmitted video data.

[0060] Especially, for the case that the dynamic relationship strong discrete category label is set for the video segment, it represents that the motion state of the sub-motion target in the video data is relatively discrete and deviates from the overall motion state of the motion target. In this case, the inter-frame prediction error is large. Therefore, the complete transmission mode is used to transmit the video segment. For the case that the dynamic relationship weak discrete category label is set for the video segment, the inter-frame prediction is performed, and the corresponding extraction interval is matched by setting the dynamic relationship discrete representation value, so that the inter-frame prediction can match the discrete difference of the relative overall motion state of the sub-motion target in the motion target in the video segment. The prediction information is synchronously sent to the receiving end, so that the receiving end can matchingly change the inter-frame prediction mode, thereby reducing the error caused by the inter-frame prediction while reducing the data transmission amount, and ensuring the video transmission efficiency and the quality of the transmitted video data.

[0061] Especially, the receiving end of the present application intercepts a predetermined proportion of the predicted video sub-segment for backflow verification, so that the processing end controls the data acquisition source to adjust the extraction interval, and the predicted video sub-segment is returned to the processing end by occupying less bandwidth. Whether the adjusted inter-frame prediction mode is suitable for the motion state of the motion target and the internal sub-motion target in the current video end is considered, thereby reducing the error caused by the inter-frame prediction, ensuring the video transmission efficiency and the quality of the transmitted video data. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 The step diagram of the multi-thread synchronous video processing method of the embodiment of the present application;

[0063] Figure 2 The logic block diagram of setting the dynamic relationship discrete category label for the video segment of the embodiment of the present application;

[0064] Figure 3 The logic block diagram of the process of adjusting the extraction interval of the embodiment of the present application;

[0065] Figure 4 The logic block diagram of the similarity verification process of the embodiment of the present application DETAILED DESCRIPTION

[0066] In order to make the purpose and advantages of the present application more clear and obvious, the present application is further described below in combination with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the present application.

[0067] The preferred embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application, and are not used to limit the protection scope of the present application.

[0068] It should be noted that in the description of the present application, in addition, it should be noted that in the description of the present application, unless otherwise specified and limited, the terms "mounting", "connection", "connection" should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected, it can be mechanically connected, or it can be electrically connected, it can be directly connected, or indirectly connected through an intermediate medium, it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0069] Please refer to Figure 1 As shown in the step diagram of the multi-thread synchronization video processing method of the embodiment of the present application, the multi-thread synchronization based video processing method of the embodiment of the present application comprises:

[0070] Step S1, the processing end acquires a video segment in the video data collected by the multi-thread data collection source every predetermined period;

[0071] Step S2, the processing end extracts a moving target in a video frame corresponding to the video segment, and analyzes the motion law of each moving target, including determining the motion inclusion relationship of each moving target, and only determining the motion law feature of the sub-motion target in the moving target with the motion inclusion relationship;

[0072] Step S3, the processing end calculates a dynamic relationship discrete representation value for the video segment based on the motion law feature, sets a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship of each moving target in the video segment and the dynamic relationship discrete representation value;

[0073] Step S4, the processing end controls the data collection source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including,

[0074] determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data collection source to extract the required key frame with the corresponding extraction interval, and transmitting each required key frame and the prediction information to the receiving end;

[0075] Or, the data collection source transmits the video data to the receiving end in its entirety;

[0076] Step S5, the receiving end generates a predicted video based on the required key frame and the prediction information, and performs a playback verification after cutting a predetermined proportion of the predicted video subsegment, so that the processing end controls the data collection source to adjust the extraction interval;

[0077] Wherein, the motion law feature includes the number of sub-motion targets and the multi-dimensional discrete value of the motion vector corresponding to the sub-motion target, and the playback verification includes that the receiving end returns the predicted video subsegment to the processing end, and the processing end performs similarity verification on the predicted video subsegment.

[0078] Specifically, the form of the multi-threaded data acquisition source is not limited, which can be an image acquisition module with multi-thread transmission capability. Preferably, the image acquisition module can include an image acquisition device with image acquisition function and a processor to respond to information sent by the processing end. Those skilled in the art can choose by themselves, which will not be repeated here.

[0079] Specifically, the specific form of the processing end is not limited, which only needs to receive the multi-threaded video data sent by the data acquisition source and analyze and control the data acquisition source. It can be composed of a logic component to process the video data, including a field programmable processor, a computer or a microprocessor in a computer.

[0080] Specifically, when the processing end acquires the video segment in the multi-threaded video data collected by the data acquisition source every predetermined period, in order to reduce the data transmission amount, and the video segment has certain data representation, which can reflect the motion law of the moving target in the video data, the acquisition ratio is selected between 10% and 20% of the predetermined period, and the predetermined period is selected within the interval [1.5s, 3s].

[0081] Specifically, the process of determining the motion inclusion relationship of each moving target includes,

[0082] Labeling all moving targets in the video segment;

[0083] If there are moving targets containing other moving targets, it is determined that the moving target has a motion inclusion relationship, and each of the other moving targets is determined as a sub-moving target.

[0084] Specifically, the way to determine the moving target in the video segment is not limited, for example, using inter-frame difference method to compare the pixel difference between consecutive frames, detecting the change area, and then identifying the moving target, or using existing target detection based model, such as YOLOv8 video processing model, of course, other ways can also be used, which will not be repeated here.

[0085] Specifically, the process of determining the motion law feature of the sub-moving target in the moving target with motion inclusion relationship includes,

[0086] Determine the number of sub-moving targets;

[0087] Taking the center of the sub-moving target as the starting point and the motion direction as the vector direction to construct the motion vector of the sub-moving target;

[0088] According to the motion vector, determine the average angular acceleration and the average speed of the sub-moving target;

[0089] determining a first absolute difference of each sub-motion target corresponding angular acceleration mean value;

[0090] determining a second absolute difference of each sub-motion target corresponding speed mean value;

[0091] weighting and summing the first absolute difference and the second absolute difference to obtain the multi-dimensional discrete value.

[0092] Specifically, in the implementation, the weight of the first absolute difference is 0.6, and the weight of the second absolute difference is 0.4.

[0093] It can be understood that the purpose of determining the motion vector is to determine the motion speed and the motion direction of a point, and the motion speed and the motion direction of the point can be determined by the person skilled in the art.

[0094] Specifically, the process of calculating the dynamic relationship discrete representation value of the video segment based on the motion law characteristics by the processing end includes,

[0095] determining the ratio of the number of sub-motion targets to the preset number of sub-motion target threshold as the first dynamic relationship discrete representation factor;

[0096] determining the ratio of the multi-dimensional discrete value to the preset multi-dimensional discrete threshold as the second dynamic relationship discrete representation factor;

[0097] weighting and summing the first dynamic relationship discrete representation factor and the second dynamic relationship discrete representation factor to obtain the dynamic relationship discrete representation value.

[0098] In the implementation, the multi-dimensional discrete threshold is predetermined, wherein a plurality of video segments in video data corresponding to a plurality of motion targets are pre-collected, motion law analysis is performed to determine the multi-dimensional discrete value corresponding to each video segment, the mean value of the multi-dimensional discrete value is solved, and the multi-dimensional discrete threshold is selected between 1.3 times and 1.5 times of the mean value of the multi-dimensional discrete value to represent the case that the multi-dimensional discrete value is larger.

[0099] In the implementation, the number of sub-motion target threshold is predetermined, wherein a plurality of video segments in video data corresponding to a plurality of motion targets are pre-collected, motion law analysis is performed to determine the number of sub-motion targets corresponding to each video segment, the mean value of the number of sub-motion targets is solved, and the mean value of the number of sub-motion targets is determined as the number of sub-motion target threshold.

[0100] In the implementation, the weight of the first dynamic relationship discrete representation factor is 0.3, and the weight of the second dynamic relationship discrete representation factor is 0.7.

[0101] Specifically, the present application is directed to each moving target for motion law analysis, determine the motion of the moving target contains the relationship and the motion law characteristics of the sub motion target in the motion target of the motion containing relationship, in the actual situation, the video data collected in the industrial monitoring field usually contains the motion target, the motion containing relationship corresponding to the motion target and the sub motion target may be different, in the actual motion, the whole motion target is in a motion trend, for example, the whole moves in a certain direction, however, there may be multiple sub motion targets in the motion target, these sub motion targets may be different from the motion trend of the whole motion target, for example, although the whole motion target moves in a certain direction, each sub motion target moves in a different direction or moves in a different way, in this case, the whole motion target has poor regularity, the probability of inter-frame prediction error will increase, if there are more predicted frames, the image features corresponding to the sub motion targets in the predicted frames are prone to errors, and it is not easy to reflect the actual motion situation, on the contrary, if there is no sub motion target inside the motion target or the whole motion trend of the sub motion target is relatively uniform, the regularity of this kind of situation is strong and the inter-frame prediction error rate is low, based on this, the present application considers to analyze the motion law, calculates the dynamic relationship discrete representation value, represents the case that the motion trend of the sub motion target in the motion target is relatively discrete and deviates from the whole motion trend of the motion target, sets the discrete category label, and then facilitates the subsequent adaptive transmission of the video segment, which can reduce the video transmission amount, ensure the video transmission efficiency and the quality of the transmitted video data.

[0102] Please refer to Figure 2 Fig. 1 is a logic block diagram of setting a dynamic relationship discrete category label for a video segment according to an embodiment of the present application, and Fig. 2 is a flow chart of setting a dynamic relationship discrete category label for a video segment according to an embodiment of the present application.

[0103] If the dynamic relationship discrete representation value is greater than or equal to a preset dynamic relationship discrete threshold value, a dynamic relationship strong discrete category label is set for the video segment.

[0104] If the dynamic relationship discrete representation value of the video segment is less than the preset dynamic relationship discrete threshold value, a dynamic relationship weak discrete category label is set for the video segment.

[0105] If each motion target in the video segment does not have a motion containing relationship, the dynamic relationship discrete representation value corresponding to the video segment is set to 0.

[0106] Specifically, the dynamic relationship discrete threshold value is selected in the interval [1.12, 1.24].

[0107] Specifically, for the case that the dynamic relationship strong discrete category label is set for the video segment, it represents that the motion state of the sub-motion target in the video data is relatively discrete and deviates from the overall motion state of the motion target. In this case, the inter-frame prediction error is large. Therefore, the video segment is transmitted in a complete transmission manner. For the case that the dynamic relationship weak discrete category label is set for the video segment, inter-frame prediction is performed, and the corresponding extraction interval is set by matching the dynamic relationship discrete representation value, so that the inter-frame prediction can match the discrete difference of the relative overall motion state of the sub-motion target in the motion target in the video segment. The prediction information is synchronously sent to the receiving end, so that the receiving end can match the change of the inter-frame prediction mode, thereby reducing the error caused by the inter-frame prediction while reducing the data transmission amount, and ensuring the video transmission efficiency and the quality of the transmitted video data.

[0108] Specifically, the process of transmitting the video segment by the processing end based on the dynamic relationship discrete category label of the video segment includes,

[0109] If the video segment has a weak discrete category label, the extraction interval and the corresponding prediction information are determined based on the dynamic relationship discrete representation value, and the data acquisition source is controlled to extract the required key frames at the corresponding extraction interval. The required key frames and the prediction information are transmitted to the receiving end.

[0110] If the video segment has a strong discrete category label, the video segment is completely transmitted to the receiving end.

[0111] It can be understood that the dynamic relationship discrete category label is a virtual label, and the purpose is to classify the video segment, which will not be repeated here.

[0112] Specifically, the extraction interval and the corresponding prediction information are determined based on the dynamic relationship discrete representation value, wherein,

[0113] The determined extraction interval is negatively correlated with the dynamic relationship discrete representation value;

[0114] The prediction information includes the number of required frames, and the number of frames corresponds to the extraction interval.

[0115] In implementation, optionally,

[0116] The dynamic relationship discrete representation value is compared with a first preset dynamic relationship discrete reference value and a second preset dynamic relationship discrete reference value,

[0117] If the dynamic relationship discrete representation value is greater than or equal to the second preset dynamic relationship discrete reference value, the extraction interval is set to 0.65 times the integral of the reference extraction interval;

[0118] If the dynamic relationship discrete representation value is less than the second preset dynamic relationship discrete reference value and greater than the first preset dynamic relationship discrete reference value, the extraction interval is set to the reference extraction interval;

[0119] If the dynamic relationship discrete representation value is less than or equal to the first preset dynamic relationship discrete reference value, the extraction interval is set to 1.35 times the reference extraction interval.

[0120] The first preset dynamic relationship discrete reference value is 0.55 times the dynamic relationship discrete threshold value, and the second preset dynamic relationship discrete reference value is 0.75 times the dynamic relationship discrete threshold value.

[0121] Specifically, the reference extraction interval is selected in the interval [5, 15].

[0122] Specifically, the process of generating a predicted video based on the demand key frame and the prediction information at the receiving end includes,

[0123] Determining the number of frames to be predicted based on the prediction information;

[0124] Predicting the number of frames based on the demand key frame and the adjacent demand key frame;

[0125] Inserting the predicted frames between the demand key frame and the adjacent demand key frame based on the time sequence to obtain the predicted video.

[0126] Specifically, the way of predicting the number of frames based on the demand key frame and the adjacent demand key frame is not limited, for example, the inter-frame prediction method can be used to realize, and the corresponding inter-frame prediction model or algorithm is deployed at the receiving end to predict the intermediate frame based on the front and rear key frames, which will not be repeated here.

[0127] It can be understood that for inter-frame prediction, the more the number of frames to be predicted, the greater the possibility of prediction error, which will not be repeated here.

[0128] It can be understood that the processing end can send the prediction information and the extraction interval to the data acquisition source to inform the data acquisition source of the required extraction interval, so that the data acquisition source extracts the demand key frame at the corresponding extraction interval, and synchronously transmits the prediction information and the demand key frame to the receiving end.

[0129] Please refer to Figure 3 and Figure 4 , Figure 3 is a logic block diagram of the process of adjusting the extraction interval of the embodiment of the present application, Figure 4 is a logic block diagram of the similarity verification process of the embodiment of the present application, and specifically, the process of the processing end for similarity verification of the predicted video sub-section includes,

[0130] The corresponding video sub-segment of the predicted video sub-segment in the original video data is determined, and it can be understood that the predicted video sub-segment corresponds to the video sub-segment in the time domain to ensure comparability.

[0131] The corresponding video frame groups are compared by the predicted video sub-segment and the video sub-segment, respectively.

[0132] The similarity of the video frames in the video frame group is determined, and the average similarity of the corresponding video frame groups is solved.

[0133] If the average similarity is less than the preset feature similarity standard value, it is determined that the similarity verification fails, and the extraction interval needs to be reduced.

[0134] If the average similarity is greater than or equal to the preset feature similarity standard value, it is determined that the similarity verification passes.

[0135] The video frame group includes two video frames corresponding in time sequence.

[0136] Specifically, the way of determining the image similarity is not limited, for example, the SSIM structural similarity can be used to determine, the SSIM value between two video frames can be calculated, and the SSIM value is used as the image similarity. The closer the SSIM value is to 1, the more similar the two video frames are. Of course, other ways can also be used, which will not be repeated here.

[0137] The preset feature similarity standard value is obtained by pre-detection, the video data in a plurality of predetermined periods is transmitted to the receiving end to generate a predicted video, and the similarity average in the backflow verification process is determined. The average of the average similarity is solved, and the product of the average and the error coefficient is determined as the preset feature similarity standard value. The error coefficient is selected in the interval [0.85, 0.95].

[0138] Specifically, the operation of adjusting the extraction interval is to halve the original extraction interval and take the integer.

[0139] Specifically, the receiving end of the application intercepts a predetermined proportion of the predicted video sub-segment for backflow verification, so that the processing end controls the data acquisition source to adjust the extraction interval, occupies less bandwidth, and returns the predicted video sub-segment to the processing end. Whether the adjusted inter-frame prediction method is suitable for the motion target and the motion trend of the internal sub-motion target in the current video end is considered, and then the error caused by inter-frame prediction is reduced, the video transmission efficiency and the quality of the transmitted video data are ensured.

[0140] On the other hand, a system applying the video processing method based on multi-thread synchronization is also provided, which comprises,

[0141] a data receiving module configured to acquire a video segment in video data collected by a multi-threaded data collection source at a predetermined period;

[0142] a data extracting module configured to extract a moving object in a video frame corresponding to the video segment, and analyze a motion rule of each moving object, including determining a motion inclusion relationship of each moving object, and determining a motion rule feature of a sub-moving object in the moving object only when the motion inclusion relationship exists;

[0143] a category label setting module configured to calculate a dynamic relationship discrete representation value for the video segment based on the motion rule feature, and set a dynamic relationship discrete category label for the video segment based on the motion inclusion relationship of each moving object in the video segment and the dynamic relationship discrete representation value;

[0144] a transmission module configured to control the data collection source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including,

[0145] determining an extraction interval and corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data collection source to extract a required key frame at the corresponding extraction interval, and transmitting each required key frame and the prediction information to a receiving end;

[0146] alternatively, controlling the data collection source to transmit the video data to the receiving end completely;

[0147] a receiving verification module configured to receive a predicted video sub-segment corresponding to the receiving end to perform a similarity verification.

[0148] Specifically, the form of the data receiving module is not limited, which can be a data storage device capable of realizing data interaction and storage, and details are not repeated.

[0149] Specifically, the specific form of the data extracting module, the category label setting module, the transmission module, and the receiving verification module is not limited, which can be composed of logical components, and details are not repeated.

[0150] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Without deviating from the principles of the present application, those skilled in the art can make equivalent changes or replacements to related technical features, and the technical solutions after these changes or replacements will fall within the protection scope of the present application.

Claims

1. A method for video processing based on multi-thread synchronization, characterized in that, The method comprises the following steps: The processing end acquires a video segment in the video data collected by the multi-threaded data collection source every predetermined period; The processing end extracts the moving targets in the video frames corresponding to the video segment, and analyzes the motion law of each moving target, including determining the motion inclusion relationship of each moving target, and determining the motion law features of only the sub-moving targets in the moving targets with motion inclusion relationship; The processing end calculates the dynamic relationship discrete representation value for the video segment based on the motion law features, sets the dynamic relationship discrete category label for the video segment based on the motion inclusion relationship of each moving target in the video segment and the dynamic relationship discrete representation value; The processing end controls the data collection source to transmit the video data based on the dynamic relationship discrete category label of the video segment, including, if the dynamic relationship discrete representation value is greater than or equal to a preset dynamic relationship discrete threshold, setting a dynamic relationship strong discrete category label for the video segment; if the dynamic relationship discrete representation value of the video segment is less than the preset dynamic relationship discrete threshold, setting a dynamic relationship weak discrete category label for the video segment; if the video segment has a weak discrete category label, determining the extraction interval and the corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data collection source to extract the required key frames with the corresponding extraction interval, and transmitting each required key frame and the prediction information to the receiving end; if the video segment has a strong discrete category label, controlling the data collection source to transmit the video data to the receiving end completely; The receiving end generates a predicted video based on the required key frames and the prediction information, intercepts a predetermined proportion of the predicted video segment, and performs backflow verification to enable the processing end to control the data collection source to adjust the extraction interval; wherein the motion law features include the number of sub-moving targets and the multi-dimensional discrete value of the motion vector corresponding to the sub-moving target; the determination of the motion inclusion relationship of each moving target includes labeling all moving targets in the video segment; if there is a moving target containing other moving targets, it is determined that the moving target has a motion inclusion relationship, and each of the other moving targets is determined as a sub-moving target; the determination of the motion law features of the sub-moving targets in the moving targets with motion inclusion relationship includes determining the number of sub-moving targets; taking the center of the sub-moving target as the starting point and the motion direction as the vector direction to construct the motion vector for the sub-moving target; determining the average angular acceleration and the average speed of the sub-moving target according to the motion vector; determining the first absolute deviation of the average angular acceleration corresponding to each sub-moving target; determining the second absolute deviation of the average speed corresponding to each sub-moving target; and weighting and summing the first absolute deviation and the second absolute deviation to obtain the multi-dimensional discrete value; the backflow verification includes that the receiving end returns the predicted video segment to the processing end, and the processing end performs similarity verification on the predicted video segment and the corresponding video segment in the original video data. The dynamic relationship discrete representation value reflects a motion discrete degree between the motion target and the sub-motion target in the video segment. The smaller the dynamic relationship discrete representation value is, the more uniform the motion trend of the sub-motion target is, and the higher the accuracy of the inter-frame prediction is. The extraction interval is negatively correlated with the dynamic relationship discrete representation value. The prediction information includes a number of frames to be predicted, and the number of frames corresponds to the extraction interval.

2. The method for video processing based on multi-thread synchronization according to claim 1, wherein, The process of calculating, by the processing end, the dynamic relationship discrete representation value for the video segment based on the motion rule features includes, a ratio of the number of sub-motion targets to a preset sub-motion target number threshold is determined as a first dynamic relationship discrete representation factor; a ratio of the multi-dimensional discrete value to a preset multi-dimensional discrete threshold is determined as a second dynamic relationship discrete representation factor; the first dynamic relationship discrete representation factor and the second dynamic relationship discrete representation factor are weighted and summed to obtain the dynamic relationship discrete representation value.

3. The method for video processing based on multi-thread synchronization according to claim 1, wherein, The process of generating, by the receiving end, the predicted video based on the demand key frame and the prediction information includes, determining the number of frames to be predicted based on the prediction information; predicting the prediction frames of the number of frames based on the demand key frame and the adjacent demand key frame; inserting the prediction frames into the demand key frame and the adjacent demand key frame based on the time sequence to obtain the predicted video.

4. The method for video processing based on multi-thread synchronization according to claim 1, characterized in that, The process of the processing end performing similarity verification on the predicted video sub-segment and the corresponding video sub-segment in the original video data includes, determining the corresponding video sub-segment of the predicted video sub-segment in the original video data; comparing a plurality of video frame groups corresponding to the extraction time sequence in the predicted video sub-segment and the video sub-segment, respectively; determining the similarity of the video frames in the video frame group, and solving the similarity average value corresponding to the plurality of video frame groups; if the similarity average value is less than a preset feature similarity standard value, it is determined that the similarity verification fails, and the extraction interval needs to be reduced; if the similarity average value is greater than or equal to the preset feature similarity standard value, it is determined that the similarity verification passes; wherein the video frame group includes two video frames corresponding in time sequence.

5. A system for applying the method for video processing based on multithread synchronization according to any one of claims 1-4, characterized in that, It includes, a data receiving module configured to obtain a video segment in video data collected by a multi-thread data collection source every predetermined period; a data extraction module configured to extract motion targets in video frames corresponding to the video segment, and analyze motion rules of each motion target, including determining motion inclusion relationships of each motion target, and determining motion rule features of only sub-motion targets in motion targets with motion inclusion relationships; a category label setting module configured to calculate a dynamic relationship discrete representation value for the video segment based on the motion rule features, and set a dynamic relationship discrete category label for the video segment based on the motion inclusion relationships of each motion target in the video segment and the dynamic relationship discrete representation value; a transmission module configured to control a data collection source to transmit video data based on a dynamic relationship discrete category label of the video segment, including, determining an extraction interval and corresponding prediction information based on the dynamic relationship discrete representation value, controlling the data collection source to extract demand key frames at the corresponding extraction interval, and transmitting each of the demand key frames and the prediction information to a receiving end; or, controlling the data collection source to transmit the video data to the receiving end in its entirety. a receiving verification module, configured to receive the predicted video sub-segment sent by the corresponding receiving end for similarity verification.

Citation Information

Patent Citations

  • Inter-frame prediction optimization method based on template matching and multi-reference prediction block fusion

    CN119865606A

  • Safe compression method applied to unattended equipment video

    CN116980613A

  • Remote monitoring video transmission method for 5G Internet of Things

    CN118200490A