Method and system for dynamic desensitization of sensitive information in video stream

By employing a multimodal fusion recognition algorithm and a dynamic desensitization strategy, the problems of low accuracy and insufficient adaptability in the traditional video stream sensitive information recognition are solved. Real-time dynamic desensitization processing of high frame rate and high resolution video streams is achieved, improving recognition accuracy and adaptability.

CN121262428BActive Publication Date: 2026-04-07GUANGZHOU TURINGIT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional video stream sensitive information desensitization technology has low recognition accuracy in complex scenarios, insufficient adaptability of desensitization strategies and insufficient real-time performance, and cannot effectively handle multimodal data and high frame rate video.

Method used

A multimodal fusion recognition algorithm is adopted to extract features by acquiring real-time frame data of video streams. Combined with visual, semantic and audio feature data, the desensitization strategy is dynamically generated and updated to achieve real-time dynamic desensitization processing of high frame rate and high resolution video streams.

Benefits of technology

It improves the accuracy of sensitive information identification in video streams and the adaptability of desensitization strategies, meets the needs of low-latency scenarios, and achieves efficient processing of multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121262428B_ABST
    Figure CN121262428B_ABST
Patent Text Reader

Abstract

This application provides a method and system for dynamic desensitization of sensitive information in video streams. The method includes: extracting features from real-time video frame data of the video stream to obtain dynamic desensitization identification feature data and sensitive information category feature data; weighting and fusing the dynamic desensitization identification feature data according to the sensitive information category feature data, and inputting it into a preset video sensitive information identification model for processing to obtain a sensitive information identification index; determining whether to perform desensitization processing through threshold comparison; and finally, querying a preset desensitization strategy database based on the sensitive information category feature data to obtain the corresponding desensitization strategy, performing dynamic desensitization processing, and obtaining a desensitized video stream. This application achieves real-time dynamic desensitization processing of high frame rate and high resolution video streams through a multimodal fusion identification algorithm, dynamic desensitization strategy generation, and adaptive dynamic updating.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video data processing, in particular to a sensitive information dynamic de-sensitization processing method and system in a video stream. BACKGROUND

[0002] Video stream is the main carrier of scene restoration in the digital era, and sensitive information such as license plate number, identity information or case information needs to be effectively protected. The traditional video stream sensitive information de-sensitization technology has the following problems: first, the recognition dimension is single, and the traditional technology only relies on visual features for sensitive target detection, such as HOG features, which reduces the recognition accuracy in complex scenes and cannot accurately identify; second, the de-sensitization strategy is not adaptive enough, and the traditional technology mainly uses fixed algorithms such as global Gaussian blur and uniform character replacement, which cannot be dynamically adjusted according to scene attributes, sensitive levels and de-sensitization needs, resulting in over-de-sensitization or insufficient de-sensitization; third, the real-time performance is not enough, and the existing technology mainly uses CPU serial processing, which has high data delay and cannot meet the needs of low delay scenes such as real-time monitoring (≤50ms). The traditional technology such as CN115495772A relies on pixel high-low bit replacement to realize de-sensitization processing, which is only suitable for simple visual sensitive information and cannot process audio and semantic multi-modal data. Although CN202410909416.5 uses continuous frame tracking, it still cannot solve the problems of dynamic adjustment of de-sensitization strategy and parallel processing of high frame rate video. Therefore, there is an urgent need for a video stream intelligent dynamic de-sensitization method that covers multi-modal recognition, intelligent strategy and high-speed execution.

[0003] In view of the above problems, an effective technical solution is urgently needed. SUMMARY

[0004] The purpose of the present application is to provide a sensitive information dynamic de-sensitization processing method and system in a video stream, which can realize real-time dynamic de-sensitization processing of high frame rate and high resolution video stream through multi-modal fusion recognition algorithm, dynamic de-sensitization strategy generation and adaptive dynamic update.

[0005] The present application also provides a sensitive information dynamic de-sensitization processing method in a video stream, comprising the following steps:

[0006] Obtaining real-time video frame data of the video stream, extracting features according to the real-time video frame data, and obtaining dynamic de-sensitization recognition feature data and sensitive information category feature data;

[0007] Performing weighted fusion processing on the dynamic de-sensitization recognition feature data according to the sensitive information category feature data, and inputting it into a preset video sensitive information recognition model for processing to obtain a sensitive information recognition index;

[0008] Comparing the sensitive information recognition index with a preset sensitive information recognition threshold.

[0009] If the sensitive information identification index is less than the preset sensitive information identification threshold, no desensitization processing will be performed;

[0010] If the sensitive information identification index is greater than or equal to the preset sensitive information identification threshold, then the preset desensitization strategy database is queried according to the sensitive information category feature data to obtain the corresponding desensitization strategy;

[0011] The video stream is dynamically desensitized according to the desensitization strategy to obtain a desensitized video stream.

[0012] Optionally, in the dynamic desensitization processing method for sensitive information in a video stream described in this application, the step of acquiring real-time video frame data of the video stream and performing feature extraction based on the real-time video frame data to obtain dynamic desensitization identification feature data and sensitive information category feature data includes:

[0013] Acquire real-time video frame data from the video stream;

[0014] Feature extraction is performed on the real-time video frame data to obtain dynamic desensitized identification feature data, including visual feature data, semantic feature data and audio feature data;

[0015] The visual feature data, semantic feature data, and audio feature data are input into a preset video sensitive information type recognition model for processing to obtain sensitive information category feature data.

[0016] The sensitive information category feature data includes security monitoring category feature data, medical imaging category feature data, traffic monitoring category feature data, remote office category feature data, live video streaming category feature data, or home monitoring category feature data.

[0017] Optionally, in the dynamic desensitization processing method for sensitive information in a video stream described in this application, the step of weightedly fusing the dynamic desensitization identification feature data according to the sensitive information category feature data and inputting it into a preset video sensitive information identification model for processing to obtain a sensitive information identification index includes:

[0018] Based on the sensitive information category feature data, query the preset scene type and weight value mapping table to obtain the weight values ​​corresponding to visual feature data, semantic feature data and audio feature data, including visual feature weight values, semantic feature weight values ​​and audio feature weight values;

[0019] The visual feature data, semantic feature data, and audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain fused feature data.

[0020] The fused feature data is input into a preset video sensitive information recognition model for processing to obtain a sensitive information recognition index.

[0021] Optionally, the method for dynamically desensitizing sensitive information in a video stream as described in this application further includes:

[0022] Obtain the preset sensitive information exposure risk value, preset data availability requirement value, and preset desensitization processing timeliness requirement value corresponding to the sensitive information category feature data;

[0023] The motion evaluation data and density probability value of sensitive information in the real-time video frame data are obtained, wherein the motion evaluation data of sensitive information includes motion speed and motion trajectory data;

[0024] The motion speed and trajectory data are input into a preset sensitive information motion evaluation model for processing to obtain a sensitive information motion evaluation score.

[0025] The sensitive information density probability value is determined based on the threshold range to which the sensitive information density probability value belongs;

[0026] The video stream desensitization requirement parameters are obtained by weighting and summing the preset sensitive information exposure risk value, preset data availability requirement value, preset desensitization processing timeliness requirement value, and the sensitive information motion evaluation score and sensitive information density probability value.

[0027] The video stream desensitization requirement parameters are compared with the preset desensitization requirement thresholds, and the video stream desensitization requirement level is obtained according to the threshold range it falls into, including high level, medium level or low level.

[0028] Optionally, the method for dynamically desensitizing sensitive information in a video stream as described in this application further includes:

[0029] The visual feature data, semantic feature data, and audio feature data are input into a preset multimodal cross-domain correlation evaluation model for processing to obtain data correlation, including visual semantic correlation, visual audio correlation, and semantic audio correlation.

[0030] The visual-semantic correlation, visual-audio correlation, and semantic-audio correlation are weighted and summed to obtain the cross-modal fusion correction factor.

[0031] The sensitive information identification index is corrected based on the cross-modal fusion correction factor to obtain the video sensitive information identification correction index.

[0032] Optionally, in the dynamic desensitization processing method for sensitive information in a video stream described in this application, the step of dynamically desensitizing the video stream according to the desensitization strategy to obtain a desensitized video stream includes:

[0033] The video stream is split according to the sensitive information category feature data to obtain a video stream set;

[0034] The video stream set is extracted according to a preset duration to obtain a pre-desensitized video stream set;

[0035] The pre-desensitized video stream set is dynamically desensitized according to the desensitization strategy to obtain a pre-desensitized video stream.

[0036] The pre-desensitized video stream is analyzed and processed to obtain a qualified pre-desensitization status.

[0037] If the pre-desensitization qualification status is unqualified, then the desensitization strategy is adjusted;

[0038] If the pre-desensitization qualification status is qualified, dynamic desensitization processing is performed according to the desensitization strategy and the video stream desensitization requirement level to obtain a desensitized video stream.

[0039] Optionally, in the dynamic desensitization processing method for sensitive information in a video stream described in this application, the step of analyzing and processing the pre-desensitized video stream to obtain a pre-desensitization qualification index includes:

[0040] Obtain the average user rating of the pre-de-identified video stream;

[0041] The pre-desensitized video stream is subjected to feature extraction to obtain pre-desensitized identification feature data, including pre-desensitized visual feature data, pre-desensitized semantic feature data and pre-desensitized audio feature data;

[0042] The pre-desensitized visual feature data, pre-desensitized semantic feature data, and pre-desensitized audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain pre-desensitized fused feature data.

[0043] The pre-desensitized fusion feature data is input into a preset video sensitive information recognition model for processing to obtain the pre-desensitized information recognition index.

[0044] The average user rating is compared with a preset user rating threshold to obtain the user rating evaluation result, including pass or fail.

[0045] The pre-desensitized information identification index is compared with the preset sensitive information identification threshold to obtain the index comparison result, including pass or fail.

[0046] The user rating evaluation result is compared with the index result by AND operation. If the result is satisfactory, the pre-desensitization status is determined to be satisfactory.

[0047] Conversely, if the pre-desensitization status is not qualified, it is determined to be unqualified.

[0048] Secondly, this application provides a dynamic desensitization processing system for sensitive information in a video stream. The system includes a memory and a processor. The memory includes a program for a dynamic desensitization processing method for sensitive information in a video stream. When the program for the dynamic desensitization processing method for sensitive information in a video stream is executed by the processor, it performs the following steps:

[0049] Acquire real-time video frame data from the video stream, extract features based on the real-time video frame data, and obtain dynamic desensitization recognition feature data and sensitive information category feature data;

[0050] The dynamic desensitization and identification feature data are weighted and fused according to the sensitive information category feature data, and then input into a preset video sensitive information identification model for processing to obtain a sensitive information identification index.

[0051] The sensitive information identification index is compared with the preset sensitive information identification threshold.

[0052] If the sensitive information identification index is less than the preset sensitive information identification threshold, no desensitization processing will be performed;

[0053] If the sensitive information identification index is greater than or equal to the preset sensitive information identification threshold, then the preset desensitization strategy database is queried according to the sensitive information category feature data to obtain the corresponding desensitization strategy;

[0054] The video stream is dynamically desensitized according to the desensitization strategy to obtain a desensitized video stream.

[0055] Optionally, in the video stream dynamic desensitization processing system described in this application, the step of acquiring real-time video frame data of the video stream and performing feature extraction based on the real-time video frame data to obtain dynamic desensitization identification feature data and sensitive information category feature data includes:

[0056] Acquire real-time video frame data from the video stream;

[0057] Feature extraction is performed on the real-time video frame data to obtain dynamic desensitized identification feature data, including visual feature data, semantic feature data and audio feature data;

[0058] The visual feature data, semantic feature data, and audio feature data are input into a preset video sensitive information type recognition model for processing to obtain sensitive information category feature data.

[0059] The sensitive information category feature data includes security monitoring category feature data, medical imaging category feature data, traffic monitoring category feature data, remote office category feature data, live video streaming category feature data, or home monitoring category feature data.

[0060] Optionally, in the video stream dynamic desensitization processing system described in this application, the step of weightedly fusing the dynamic desensitization identification feature data according to the sensitive information category feature data and inputting it into a preset video sensitive information identification model for processing to obtain a sensitive information identification index includes:

[0061] Based on the sensitive information category feature data, query the preset scene type and weight value mapping table to obtain the weight values ​​corresponding to visual feature data, semantic feature data and audio feature data, including visual feature weight values, semantic feature weight values ​​and audio feature weight values;

[0062] The visual feature data, semantic feature data, and audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain fused feature data.

[0063] The fused feature data is input into a preset video sensitive information recognition model for processing to obtain a sensitive information recognition index.

[0064] As can be seen from the above, the method and system for dynamic desensitization of sensitive information in video streams provided in this application achieve real-time dynamic desensitization of high frame rate and high resolution video streams through multimodal fusion recognition algorithms, dynamic desensitization strategy generation and adaptive dynamic updates.

[0065] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing embodiments of this application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description

[0066] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0067] Figure 1 A flowchart of a method for dynamically desensitizing sensitive information in a video stream provided in an embodiment of this application;

[0068] Figure 2A flowchart illustrating the process of obtaining dynamic desensitization identification feature data and sensitive information category feature data in a video stream dynamic desensitization processing method provided in this application embodiment;

[0069] Figure 3 A flowchart illustrating the process of obtaining the sensitive information identification index in the dynamic desensitization processing method for sensitive information in a video stream provided in this application embodiment;

[0070] Figure 4 This is a high-level flowchart of the methods of various embodiments of this application. Detailed Implementation

[0071] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0072] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0073] Please refer to Figure 1 , Figure 1 This is a flowchart of a method for dynamically de-identifying sensitive information in a video stream, as described in some embodiments of this application. This method is used in terminal devices, such as computers and mobile terminals. The method includes the following steps:

[0074] S11. Obtain real-time video frame data from the video stream, extract features based on the real-time video frame data, and obtain dynamic desensitization identification feature data and sensitive information category feature data.

[0075] S12. The dynamic desensitization identification feature data is weighted and fused according to the sensitive information category feature data, and then input into the preset video sensitive information identification model for processing to obtain the sensitive information identification index.

[0076] S13. Compare the sensitive information identification index with the preset sensitive information identification threshold.

[0077] S141. If the sensitive information identification index is less than the preset sensitive information identification threshold, no desensitization processing will be performed.

[0078] S142. If the sensitive information identification index is greater than or equal to the preset sensitive information identification threshold, then query the preset desensitization strategy database according to the sensitive information category feature data to obtain the corresponding desensitization strategy.

[0079] S15. Perform dynamic desensitization processing on the video stream according to the desensitization strategy to obtain a desensitized video stream.

[0080] It should be noted that, in order to achieve intelligent dynamic desensitization processing of video streams, the video stream is first subjected to feature extraction and dynamic weighted fusion. Then, sensitive information is intelligently identified based on a preset video sensitive information identification model. After sensitive information is identified, a dynamic desensitization strategy is matched based on different categories. Finally, desensitization processing is performed based on the matched dynamic desensitization strategy, and the matching between the desensitization strategy and the desensitization effect is evaluated. At the same time, the parameters corresponding to the desensitization strategy are adaptively updated based on the evaluation results.

[0081] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating the process of obtaining dynamic desensitization identification feature data and sensitive information category feature data in a video stream dynamic desensitization processing method according to some embodiments of this application. According to embodiments of the present invention, the step of acquiring real-time video frame data of the video stream and performing feature extraction based on the real-time video frame data to obtain dynamic desensitization identification feature data and sensitive information category feature data includes:

[0082] S21. Obtain real-time video frame data from the video stream;

[0083] S22. Based on the real-time video frame data, feature extraction is performed to obtain dynamic desensitization recognition feature data, including visual feature data, semantic feature data and audio feature data;

[0084] S23. Input the visual feature data, semantic feature data and audio feature data into a preset video sensitive information type recognition model for processing to obtain sensitive information category feature data;

[0085] S24. The sensitive information category feature data includes security monitoring category feature data, medical imaging category feature data, traffic monitoring category feature data, remote office category feature data, live video streaming category feature data, or home monitoring category feature data.

[0086] It should be noted that, in order to achieve accurate desensitization processing, firstly, the video stream is frame-by-frame extracted to obtain real-time video frame data. Then, visual feature data, semantic feature data, and audio feature data are identified. The visual feature data includes color histograms, texture features, and shape contours, while the audio feature data includes speech spectrum features and speech content keywords. Then, the extracted visual feature data, semantic feature data, and audio feature data are processed through a preset video sensitive information type recognition model to obtain sensitive information category feature data. The sensitive information category feature data is represented by different identifiers. The preset video sensitive information type recognition model is obtained by training a large number of historical samples of visual feature data, semantic feature data, and audio feature data, as well as the corresponding sensitive information category feature data.

[0087] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating the process of obtaining a sensitive information identification index using a dynamic desensitization method for sensitive information in a video stream, as described in some embodiments of this application. According to an embodiment of the present invention, the step of weightedly fusing the dynamic desensitization identification feature data based on the sensitive information category feature data and inputting it into a preset video sensitive information identification model for processing to obtain a sensitive information identification index includes:

[0088] S31. Based on the sensitive information category feature data, query the preset scene type and weight value mapping table to obtain the weight values ​​corresponding to the visual feature data, semantic feature data and audio feature data, including visual feature weight values, semantic feature weight values ​​and audio feature weight values.

[0089] S32. Perform weighted fusion processing on the visual feature data, semantic feature data, and audio feature data, as well as the corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, to obtain fused feature data;

[0090] S33. Input the fused feature data into a preset video sensitive information recognition model for processing to obtain a sensitive information recognition index.

[0091] It should be noted that, in order to achieve a quantitative representation of sensitive information identification, the extracted visual feature data, semantic feature data, and audio feature data are queried from a preset scene type and weight value mapping table to determine the corresponding weight values. Then, weighted fusion is performed to obtain fused feature data, which is then processed by a preset video sensitive information identification model to obtain a sensitive information identification index, which is used to assess whether sensitive information exists. The preset scene type and weight value mapping table is constructed by those skilled in the art based on a large number of historical samples and can be dynamically adjusted. The preset video sensitive information identification model is obtained by training with fused feature data of a large number of historical samples and the corresponding sensitive information identification index.

[0092] According to an embodiment of the present invention, it further includes:

[0093] Obtain the preset sensitive information exposure risk value, preset data availability requirement value, and preset desensitization processing timeliness requirement value corresponding to the sensitive information category feature data;

[0094] The motion evaluation data and density probability value of sensitive information in the real-time video frame data are obtained, wherein the motion evaluation data of sensitive information includes motion speed and motion trajectory data;

[0095] The motion speed and trajectory data are input into a preset sensitive information motion evaluation model for processing to obtain a sensitive information motion evaluation score.

[0096] The sensitive information density probability value is determined based on the threshold range to which the sensitive information density probability value belongs;

[0097] The video stream desensitization requirement parameters are obtained by weighting and summing the preset sensitive information exposure risk value, preset data availability requirement value, preset desensitization processing timeliness requirement value, and the sensitive information motion evaluation score and sensitive information density probability value.

[0098] The video stream desensitization requirement parameters are compared with the preset desensitization requirement thresholds, and the video stream desensitization requirement level is obtained according to the threshold range it falls into, including high level, medium level or low level.

[0099] It should be noted that, to improve the efficiency and adaptability of desensitization processing, the desensitization requirements are assessed from five aspects: information exposure risk, availability requirements, processing timeliness, sensitive target movement, and sensitive information density probability value. Specifically, the preset sensitive information exposure risk value represents the degree of harm caused by the leakage of sensitive information in a video stream scenario. This value is constructed by those skilled in the art based on the sensitive information category and can be dynamically adjusted. The preset data availability requirement value refers to the proportion of effective information that must be retained in the desensitized video data; the higher the proportion, the greater the preset data availability requirement value. The preset desensitization processing timeliness requirement value is quantified based on the maximum allowable delay from video stream acquisition to desensitized output; the smaller the maximum allowable delay, the greater the preset desensitization processing timeliness requirement value. The sensitive information density probability value is pre-constructed by those skilled in the art based on the video stream scenario and can be dynamically adjusted. Simultaneously, the motion speed and trajectory data of sensitive targets in the video stream scenario are processed through a preset sensitive information motion evaluation model to obtain a sensitive information motion evaluation score. The sensitive information density probability value is then used to determine the appropriate value for each target. The threshold range determines the probability value of sensitive information density. The threshold range is pre-constructed by those skilled in the art and can be dynamically adjusted. The preset sensitive information exposure risk value, preset data availability requirement value, preset desensitization processing timeliness requirement value, sensitive information motion evaluation score, and sensitive information density probability value are represented by values ​​in [0, 1]. The preset sensitive information motion evaluation model is obtained by training a large number of historical samples with motion speed and trajectory data and corresponding sensitive information motion evaluation scores. Finally, the threshold is compared with the preset desensitization requirement threshold. The desensitization requirement level of the video stream is obtained according to the threshold range it falls into. The preset desensitization requirement threshold includes a first preset desensitization requirement threshold and a second preset desensitization requirement threshold. If it is less than the first preset desensitization requirement threshold, the desensitization requirement level of the video stream is determined to be low. If it is greater than or equal to the first preset desensitization requirement threshold and less than the second preset desensitization requirement threshold, the desensitization requirement level of the video stream is determined to be medium. If it is greater than or equal to the second preset desensitization requirement threshold and less than the third preset desensitization requirement threshold, the desensitization requirement level of the video stream is determined to be high.

[0100] According to an embodiment of the present invention, it further includes:

[0101] The visual feature data, semantic feature data, and audio feature data are input into a preset multimodal cross-domain correlation evaluation model for processing to obtain data correlation, including visual semantic correlation, visual audio correlation, and semantic audio correlation.

[0102] The visual-semantic correlation, visual-audio correlation, and semantic-audio correlation are weighted and summed to obtain the cross-modal fusion correction factor.

[0103] The sensitive information identification index is corrected based on the cross-modal fusion correction factor to obtain the video sensitive information identification correction index.

[0104] It should be noted that, in order to more accurately identify sensitive information, a cross-domain attention fusion architecture of visual, semantic and audio three-modal is constructed. The cross-attention module of the Transformer in the pre-set multimodal cross-domain correlation evaluation model is used to capture the correlation between modalities. The determined visual-semantic correlation, visual-audio correlation and semantic-audio correlation are normalized and weighted and summed to obtain the cross-modal fusion correction factor. Then, the sensitive information identification index is corrected to obtain the video sensitive information identification correction index. For example, if the sensitive information identification index is x and the cross-modal fusion correction factor is y, then (1+y)*x is the video sensitive information identification correction index.

[0105] According to an embodiment of the present invention, the step of performing dynamic desensitization processing on the video stream according to the desensitization strategy to obtain a desensitized video stream includes:

[0106] The video stream is split according to the sensitive information category feature data to obtain a video stream set;

[0107] The video stream set is extracted according to a preset duration to obtain a pre-desensitized video stream set;

[0108] The pre-desensitized video stream set is dynamically desensitized according to the desensitization strategy to obtain a pre-desensitized video stream.

[0109] The pre-desensitized video stream is analyzed and processed to obtain a qualified pre-desensitization status.

[0110] If the pre-desensitization qualification status is unqualified, then the desensitization strategy is adjusted;

[0111] If the pre-desensitization qualification status is qualified, dynamic desensitization processing is performed according to the desensitization strategy and the video stream desensitization requirement level to obtain a desensitized video stream.

[0112] It should be noted that, in order to improve the accuracy and efficiency of the desensitization process and to prevent the direct desensitization process from failing to meet the requirements, the video stream is first split according to the sensitive information category feature data to obtain the video stream set corresponding to different categories. Then, the video stream of the preset duration is extracted for pre-desensitization processing, and then it is determined whether it is qualified. Based on the processing status, the desensitization strategy is adaptively adjusted or normal desensitization processing is continued. Table 1 is a sample database of preset desensitization strategies.

[0113] Table 1

[0114]

[0115] According to an embodiment of the present invention, the step of analyzing and processing the pre-desensitized video stream to obtain a pre-desensitization qualification index includes:

[0116] Obtain the average user rating of the pre-de-identified video stream;

[0117] The pre-desensitized video stream is subjected to feature extraction to obtain pre-desensitized identification feature data, including pre-desensitized visual feature data, pre-desensitized semantic feature data and pre-desensitized audio feature data;

[0118] The pre-desensitized visual feature data, pre-desensitized semantic feature data, and pre-desensitized audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain pre-desensitized fused feature data.

[0119] The pre-desensitized fusion feature data is input into a preset video sensitive information recognition model for processing to obtain the pre-desensitized information recognition index.

[0120] The average user rating is compared with a preset user rating threshold to obtain the user rating evaluation result, including pass or fail.

[0121] The pre-desensitized information identification index is compared with the preset sensitive information identification threshold to obtain the index comparison result, including pass or fail.

[0122] The user rating evaluation result is compared with the index result by AND operation. If the result is satisfactory, the pre-desensitization status is determined to be satisfactory.

[0123] Conversely, if the pre-desensitization status is not qualified, it is determined to be unqualified.

[0124] It should be noted that this embodiment evaluates whether the desensitization strategy meets the requirements from two aspects: user rating and desensitization recognition. First, the average user rating within a preset time period is obtained and compared with the corresponding threshold. If it is greater than or equal to the preset user rating threshold, the user rating evaluation result is determined to be passed; otherwise, it is deemed unsuccessful. Then, based on the pre-desensitized video stream, pre-desensitized visual feature data, pre-desensitized semantic feature data, and pre-desensitized audio feature data are extracted and input into a preset video sensitive information recognition model for processing to obtain a pre-desensitized information recognition index. This index is then compared with a preset sensitive information recognition threshold. If it is greater than or equal to the preset sensitive information recognition threshold, the index comparison result is determined to be passed; otherwise, it is deemed unsuccessful. Finally, the two are ANDed. Only if the result is passed is the pre-desensitization qualified state deemed qualified; otherwise, the pre-desensitization qualified state is deemed unqualified, and the parameter settings corresponding to the desensitization strategy need to be adjusted.

[0125] Please refer to Figure 4 , Figure 4This is a flowchart illustrating the process of obtaining sensitive information identification in a video stream through a dynamic desensitization processing method in some embodiments of this application.

[0126] It is worth mentioning that, according to embodiments of the present invention, it further includes:

[0127] Emergency events are identified based on the real-time video frame data;

[0128] If an emergency event is detected, the desensitization strategy adjustment mechanism is triggered to obtain a temporary desensitization strategy;

[0129] Desensitization is performed according to the temporary desensitization strategy described above.

[0130] It should be noted that in special scenarios, if an emergency occurs, such as a fire, the original desensitization strategy may result in the loss of some information. Therefore, if a preset emergency event is identified, the desensitization strategy adjustment mechanism is triggered according to the identified emergency event to obtain a temporary desensitization strategy. For example, when a fire event is identified in a security monitoring scenario, the desensitization intensity of the background personnel's faces is automatically reduced, changing from cartoonish to slightly blurred.

[0131] It is worth mentioning that, according to embodiments of the present invention, it further includes:

[0132] Based on the sensitive information category feature data in the real-time video frame data, a preset sensitive information and video stream priority mapping table is queried to obtain the corresponding video stream priority;

[0133] The frame complexity of obtaining sensitive information in the real-time video frame data;

[0134] The frame payload is obtained by processing the video stream based on its priority and frame complexity.

[0135] Frame scheduling is performed based on the frame payload.

[0136] It should be noted that, in order to achieve resource coordination for desensitization processing, the input video stream is split into frames, with each frame treated as an independent task unit. The video stream priority and frame complexity are identified. A pre-built mapping table between sensitive information and video stream priority is constructed by those skilled in the art and can be dynamically adjusted; for example, the priority of a medical image stream is 6, and the priority of a home monitoring stream is 1. Frame complexity is determined based on the number of sensitive areas in the video frame. The frame load is determined based on the video stream priority and frame complexity, and then, combined with the real-time load of the GPU nodes, the video frames are scheduled to the optimal nodes. For example... For example, the system contains two GPU nodes (A100-1 and A100-2), which simultaneously process medical flow A, traffic flow B, and home flow C. The frame complexities are 7 for A, 6 for B, and 4 for C, and the priorities are 6 for A, 3 for B, and 1 for C. The load on frame A is 7 x 6 = 42, the load on frame B is 6 x 3 = 18, and the load on frame C is 4 x 1 = 4. The scheduling ratio is 42:18:4. A100-1 prioritizes processing M-frames (accounting for approximately 66%), while A100-2 processes T-frames and L-frames (accounting for approximately 28% + 6%), ensuring that resources are allocated on demand.

[0137] It is worth mentioning that, according to embodiments of the present invention, it further includes:

[0138] The real-time video frame data is divided into sensitive regions to obtain sensitive region subtasks;

[0139] Desensitization processing nodes are assigned according to the preset sensitivity level corresponding to the sub-tasks in the sensitive areas.

[0140] It should be noted that, in order to achieve parallel processing of sensitive areas, the sensitive areas in the real-time video frame data are split into sensitive area sub-tasks, such as lesion images, faces, and medical record text. Among them, the sensitivity level of lesion images > faces = medical record text. Then, Tensor Cores are assigned to process lesion images (GAN generation), and two CUDA Cores are assigned to process faces (blurring) and medical record text (replacement) in parallel. After processing, the results of each region are summarized into the frame buffer. The preset sensitivity level corresponding to the sensitive area sub-task is pre-constructed by those skilled in the art and can be dynamically adjusted.

[0141] This invention also discloses a dynamic desensitization processing system for sensitive information in a video stream, comprising a memory and a processor. The memory includes a program for a dynamic desensitization processing method for sensitive information in a video stream. When the processor executes the program for the dynamic desensitization processing method for sensitive information in a video stream, it performs the following steps:

[0142] Acquire real-time video frame data from the video stream, extract features based on the real-time video frame data, and obtain dynamic desensitization recognition feature data and sensitive information category feature data;

[0143] The dynamic desensitization and identification feature data are weighted and fused according to the sensitive information category feature data, and then input into a preset video sensitive information identification model for processing to obtain a sensitive information identification index.

[0144] The sensitive information identification index is compared with the preset sensitive information identification threshold.

[0145] If the sensitive information identification index is less than the preset sensitive information identification threshold, no desensitization processing will be performed;

[0146] If the sensitive information identification index is greater than or equal to the preset sensitive information identification threshold, then the preset desensitization strategy database is queried according to the sensitive information category feature data to obtain the corresponding desensitization strategy;

[0147] The video stream is dynamically desensitized according to the desensitization strategy to obtain a desensitized video stream.

[0148] It should be noted that, in order to achieve intelligent dynamic desensitization processing of video streams, the video stream is first subjected to feature extraction and dynamic weighted fusion. Then, sensitive information is intelligently identified based on a preset video sensitive information identification model. After sensitive information is identified, a dynamic desensitization strategy is matched based on different categories. Finally, desensitization processing is performed based on the matched dynamic desensitization strategy, and the matching between the desensitization strategy and the desensitization effect is evaluated. At the same time, the parameters corresponding to the desensitization strategy are adaptively updated based on the evaluation results.

[0149] According to an embodiment of the present invention, the step of acquiring real-time video frame data of a video stream, and performing feature extraction based on the real-time video frame data to obtain dynamic desensitization identification feature data and sensitive information category feature data includes:

[0150] Acquire real-time video frame data from the video stream;

[0151] Feature extraction is performed on the real-time video frame data to obtain dynamic desensitized identification feature data, including visual feature data, semantic feature data and audio feature data;

[0152] The visual feature data, semantic feature data, and audio feature data are input into a preset video sensitive information type recognition model for processing to obtain sensitive information category feature data.

[0153] The sensitive information category feature data includes security monitoring category feature data, medical imaging category feature data, traffic monitoring category feature data, remote office category feature data, live video streaming category feature data, or home monitoring category feature data.

[0154] It should be noted that, in order to achieve accurate desensitization processing, firstly, the video stream is frame-by-frame extracted to obtain real-time video frame data. Then, visual feature data, semantic feature data, and audio feature data are identified. The visual feature data includes color histograms, texture features, and shape contours, while the audio feature data includes speech spectrum features and speech content keywords. Then, the extracted visual feature data, semantic feature data, and audio feature data are processed through a preset video sensitive information type recognition model to obtain sensitive information category feature data. The sensitive information category feature data is represented by different identifiers. The preset video sensitive information type recognition model is obtained by training a large number of historical samples of visual feature data, semantic feature data, and audio feature data, as well as the corresponding sensitive information category feature data.

[0155] According to an embodiment of the present invention, the step of weightedly fusing the dynamic desensitization identification feature data based on the sensitive information category feature data and inputting it into a preset video sensitive information identification model for processing to obtain a sensitive information identification index includes:

[0156] Based on the sensitive information category feature data, query the preset scene type and weight value mapping table to obtain the weight values ​​corresponding to visual feature data, semantic feature data and audio feature data, including visual feature weight values, semantic feature weight values ​​and audio feature weight values;

[0157] The visual feature data, semantic feature data, and audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain fused feature data.

[0158] The fused feature data is input into a preset video sensitive information recognition model for processing to obtain a sensitive information recognition index.

[0159] It should be noted that, in order to achieve a quantitative representation of sensitive information identification, the extracted visual feature data, semantic feature data, and audio feature data are queried from a preset scene type and weight value mapping table to determine the corresponding weight values. Then, weighted fusion is performed to obtain fused feature data, which is then processed by a preset video sensitive information identification model to obtain a sensitive information identification index, which is used to assess whether sensitive information exists. The preset scene type and weight value mapping table is constructed by those skilled in the art based on a large number of historical samples and can be dynamically adjusted. The preset video sensitive information identification model is obtained by training with fused feature data of a large number of historical samples and the corresponding sensitive information identification index.

[0160] According to an embodiment of the present invention, it further includes:

[0161] Obtain the preset sensitive information exposure risk value, preset data availability requirement value, and preset desensitization processing timeliness requirement value corresponding to the sensitive information category feature data;

[0162] The motion evaluation data and density probability value of sensitive information in the real-time video frame data are obtained, wherein the motion evaluation data of sensitive information includes motion speed and motion trajectory data;

[0163] The motion speed and trajectory data are input into a preset sensitive information motion evaluation model for processing to obtain a sensitive information motion evaluation score.

[0164] The sensitive information density probability value is determined based on the threshold range to which the sensitive information density probability value belongs;

[0165] The video stream desensitization requirement parameters are obtained by weighting and summing the preset sensitive information exposure risk value, preset data availability requirement value, preset desensitization processing timeliness requirement value, and the sensitive information motion evaluation score and sensitive information density probability value.

[0166] The video stream desensitization requirement parameters are compared with the preset desensitization requirement thresholds, and the video stream desensitization requirement level is obtained according to the threshold range it falls into, including high level, medium level or low level.

[0167] It should be noted that, to improve the efficiency and adaptability of desensitization processing, the desensitization requirements are assessed from five aspects: information exposure risk, availability requirements, processing timeliness, sensitive target movement, and sensitive information density probability value. Specifically, the preset sensitive information exposure risk value represents the degree of harm caused by the leakage of sensitive information in a video stream scenario. This value is constructed by those skilled in the art based on the sensitive information category and can be dynamically adjusted. The preset data availability requirement value refers to the proportion of effective information that must be retained in the desensitized video data; the higher the proportion, the greater the preset data availability requirement value. The preset desensitization processing timeliness requirement value is quantified based on the maximum allowable delay from video stream acquisition to desensitized output; the smaller the maximum allowable delay, the greater the preset desensitization processing timeliness requirement value. The sensitive information density probability value is pre-constructed by those skilled in the art based on the video stream scenario and can be dynamically adjusted. Simultaneously, the motion speed and trajectory data of sensitive targets in the video stream scenario are processed through a preset sensitive information motion evaluation model to obtain a sensitive information motion evaluation score. The sensitive information density probability value is then used to determine the appropriate value for each target. The threshold range determines the probability value of sensitive information density. The threshold range is pre-constructed by those skilled in the art and can be dynamically adjusted. The preset sensitive information exposure risk value, preset data availability requirement value, preset desensitization processing timeliness requirement value, sensitive information motion evaluation score, and sensitive information density probability value are represented by values ​​in [0, 1]. The preset sensitive information motion evaluation model is obtained by training a large number of historical samples with motion speed and trajectory data and corresponding sensitive information motion evaluation scores. Finally, the threshold is compared with the preset desensitization requirement threshold. The desensitization requirement level of the video stream is obtained according to the threshold range it falls into. The preset desensitization requirement threshold includes a first preset desensitization requirement threshold and a second preset desensitization requirement threshold. If it is less than the first preset desensitization requirement threshold, the desensitization requirement level of the video stream is determined to be low. If it is greater than or equal to the first preset desensitization requirement threshold and less than the second preset desensitization requirement threshold, the desensitization requirement level of the video stream is determined to be medium. If it is greater than or equal to the second preset desensitization requirement threshold and less than the third preset desensitization requirement threshold, the desensitization requirement level of the video stream is determined to be high.

[0168] According to an embodiment of the present invention, it further includes:

[0169] The visual feature data, semantic feature data, and audio feature data are input into a preset multimodal cross-domain correlation evaluation model for processing to obtain data correlation, including visual semantic correlation, visual audio correlation, and semantic audio correlation.

[0170] The visual-semantic correlation, visual-audio correlation, and semantic-audio correlation are weighted and summed to obtain the cross-modal fusion correction factor.

[0171] The sensitive information identification index is corrected based on the cross-modal fusion correction factor to obtain the video sensitive information identification correction index.

[0172] It should be noted that, in order to more accurately identify sensitive information, a cross-domain attention fusion architecture of visual, semantic and audio three-modal is constructed. The cross-attention module of the Transformer in the pre-set multimodal cross-domain correlation evaluation model is used to capture the correlation between modalities. The determined visual-semantic correlation, visual-audio correlation and semantic-audio correlation are normalized and weighted and summed to obtain the cross-modal fusion correction factor. Then, the sensitive information identification index is corrected to obtain the video sensitive information identification correction index. For example, if the sensitive information identification index is x and the cross-modal fusion correction factor is y, then (1+y)*x is the video sensitive information identification correction index.

[0173] According to an embodiment of the present invention, the step of performing dynamic desensitization processing on the video stream according to the desensitization strategy to obtain a desensitized video stream includes:

[0174] The video stream is split according to the sensitive information category feature data to obtain a video stream set;

[0175] The video stream set is extracted according to a preset duration to obtain a pre-desensitized video stream set;

[0176] The pre-desensitized video stream set is dynamically desensitized according to the desensitization strategy to obtain a pre-desensitized video stream.

[0177] The pre-desensitized video stream is analyzed and processed to obtain a qualified pre-desensitization status.

[0178] If the pre-desensitization qualification status is unqualified, then the desensitization strategy is adjusted;

[0179] If the pre-desensitization qualification status is qualified, dynamic desensitization processing is performed according to the desensitization strategy and the video stream desensitization requirement level to obtain a desensitized video stream.

[0180] It should be noted that, in order to improve the accuracy and efficiency of the desensitization process and to prevent the direct desensitization process from failing to meet the requirements, the video stream is first split according to the sensitive information category feature data to obtain video stream sets corresponding to different categories. Then, the video streams of a preset duration are extracted for pre-desensitization processing, and then it is determined whether they are qualified. Based on the processing status, the desensitization strategy is adaptively adjusted or normal desensitization processing is continued.

[0181] According to an embodiment of the present invention, the step of analyzing and processing the pre-desensitized video stream to obtain a pre-desensitization qualification index includes:

[0182] Obtain the average user rating of the pre-de-identified video stream;

[0183] The pre-desensitized video stream is subjected to feature extraction to obtain pre-desensitized identification feature data, including pre-desensitized visual feature data, pre-desensitized semantic feature data and pre-desensitized audio feature data;

[0184] The pre-desensitized visual feature data, pre-desensitized semantic feature data, and pre-desensitized audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain pre-desensitized fused feature data.

[0185] The pre-desensitized fusion feature data is input into a preset video sensitive information recognition model for processing to obtain the pre-desensitized information recognition index.

[0186] The average user rating is compared with a preset user rating threshold to obtain the user rating evaluation result, including pass or fail.

[0187] The pre-desensitized information identification index is compared with the preset sensitive information identification threshold to obtain the index comparison result, including pass or fail.

[0188] The user rating evaluation result is compared with the index result by AND operation. If the result is satisfactory, the pre-desensitization status is determined to be satisfactory.

[0189] Conversely, if the pre-desensitization status is not qualified, it is determined to be unqualified.

[0190] It should be noted that this embodiment evaluates whether the desensitization strategy meets the requirements from two aspects: user rating and desensitization recognition. First, the average user rating within a preset time period is obtained and compared with the corresponding threshold. If it is greater than or equal to the preset user rating threshold, the user rating evaluation result is determined to be passed; otherwise, it is deemed unsuccessful. Then, based on the pre-desensitized video stream, pre-desensitized visual feature data, pre-desensitized semantic feature data, and pre-desensitized audio feature data are extracted and input into a preset video sensitive information recognition model for processing to obtain a pre-desensitized information recognition index. This index is then compared with a preset sensitive information recognition threshold. If it is greater than or equal to the preset sensitive information recognition threshold, the index comparison result is determined to be passed; otherwise, it is deemed unsuccessful. Finally, the two are ANDed. Only if the result is passed is the pre-desensitization qualified state deemed qualified; otherwise, the pre-desensitization qualified state is deemed unqualified, and the parameter settings corresponding to the desensitization strategy need to be adjusted.

[0191] It is worth mentioning that, according to embodiments of the present invention, it further includes:

[0192] Emergency events are identified based on the real-time video frame data;

[0193] If an emergency event is detected, the desensitization strategy adjustment mechanism is triggered to obtain a temporary desensitization strategy;

[0194] Desensitization is performed according to the temporary desensitization strategy described above.

[0195] It should be noted that in special scenarios, if an emergency occurs, such as a fire, the original desensitization strategy may result in the loss of some information. Therefore, if a preset emergency event is identified, the desensitization strategy adjustment mechanism is triggered according to the identified emergency event to obtain a temporary desensitization strategy. For example, when a fire event is identified in a security monitoring scenario, the desensitization intensity of the background personnel's faces is automatically reduced, changing from cartoonish to slightly blurred.

[0196] It is worth mentioning that, according to embodiments of the present invention, it further includes:

[0197] Based on the sensitive information category feature data in the real-time video frame data, a preset sensitive information and video stream priority mapping table is queried to obtain the corresponding video stream priority;

[0198] The frame complexity of obtaining sensitive information in the real-time video frame data;

[0199] The frame payload is obtained by processing the video stream based on its priority and frame complexity.

[0200] Frame scheduling is performed based on the frame payload.

[0201] It should be noted that, in order to achieve resource coordination for desensitization processing, the input video stream is split into frames, with each frame treated as an independent task unit. The video stream priority and frame complexity are identified. A pre-built mapping table between sensitive information and video stream priority is constructed by those skilled in the art and can be dynamically adjusted; for example, the priority of a medical image stream is 6, and the priority of a home monitoring stream is 1. Frame complexity is determined based on the number of sensitive areas in the video frame. The frame load is determined based on the video stream priority and frame complexity, and then, combined with the real-time load of the GPU nodes, the video frames are scheduled to the optimal nodes. For example... For example, the system contains two GPU nodes (A100-1 and A100-2), which simultaneously process medical flow A, traffic flow B, and home flow C. The frame complexities are 7 for A, 6 for B, and 4 for C, and the priorities are 6 for A, 3 for B, and 1 for C. The load on frame A is 7 x 6 = 42, the load on frame B is 6 x 3 = 18, and the load on frame C is 4 x 1 = 4. The scheduling ratio is 42:18:4. A100-1 prioritizes processing M-frames (accounting for approximately 66%), while A100-2 processes T-frames and L-frames (accounting for approximately 28% + 6%), ensuring that resources are allocated on demand.

[0202] It is worth mentioning that, according to embodiments of the present invention, it further includes:

[0203] The real-time video frame data is divided into sensitive regions to obtain sensitive region subtasks;

[0204] Desensitization processing nodes are assigned according to the preset sensitivity level corresponding to the sub-tasks in the sensitive areas.

[0205] It should be noted that, in order to achieve parallel processing of sensitive areas, the sensitive areas in the real-time video frame data are split into sensitive area sub-tasks, such as lesion images, faces, and medical record text. Among them, the sensitivity level of lesion images > faces = medical record text. Then, Tensor Cores are assigned to process lesion images (GAN generation), and two CUDA Cores are assigned to process faces (blurring) and medical record text (replacement) in parallel. After processing, the results of each region are summarized into the frame buffer. The preset sensitivity level corresponding to the sensitive area sub-task is pre-constructed by those skilled in the art and can be dynamically adjusted.

[0206] The present invention discloses a method and system for dynamic desensitization of sensitive information in video streams. Through multimodal fusion recognition algorithm, dynamic desensitization strategy generation and adaptive dynamic update, it achieves real-time dynamic desensitization of high frame rate and high resolution video streams.

[0207] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0208] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0209] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0210] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0211] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

Claims

1. A method for dynamically desensitizing sensitive information in a video stream, characterized in that, Includes the following steps: Acquire real-time video frame data from the video stream, extract features based on the real-time video frame data, and obtain dynamic desensitization recognition feature data and sensitive information category feature data; The dynamic desensitization and identification feature data are weighted and fused according to the sensitive information category feature data, and then input into a preset video sensitive information identification model for processing to obtain a sensitive information identification index. The sensitive information identification index is compared with the preset sensitive information identification threshold. If the sensitive information identification index is less than the preset sensitive information identification threshold, no desensitization processing will be performed; If the sensitive information identification index is greater than or equal to the preset sensitive information identification threshold, then the preset desensitization strategy database is queried according to the sensitive information category feature data to obtain the corresponding desensitization strategy; The video stream is dynamically desensitized according to the desensitization strategy to obtain a desensitized video stream. The process involves acquiring real-time video frame data from the video stream, extracting features from the real-time video frame data, and obtaining dynamic desensitization identification feature data and sensitive information category feature data, including: Acquire real-time video frame data from the video stream; Feature extraction is performed on the real-time video frame data to obtain dynamic desensitization recognition feature data, including visual feature data, semantic feature data and audio feature data; The visual feature data, semantic feature data, and audio feature data are input into a preset video sensitive information type recognition model for processing to obtain sensitive information category feature data. The sensitive information category feature data includes security monitoring category feature data, medical imaging category feature data, traffic monitoring category feature data, remote office category feature data, live video streaming category feature data, or home monitoring category feature data; Also includes: Obtain the preset sensitive information exposure risk value, preset data availability requirement value, and preset desensitization processing timeliness requirement value corresponding to the sensitive information category feature data; The motion evaluation data and density probability value of sensitive information in the real-time video frame data are obtained, wherein the motion evaluation data of sensitive information includes motion speed and motion trajectory data; The motion speed and trajectory data are input into a preset sensitive information motion evaluation model for processing to obtain a sensitive information motion evaluation score. The sensitive information density probability value is determined based on the threshold range to which the sensitive information density probability value belongs; The video stream desensitization requirement parameters are obtained by weighting and summing the preset sensitive information exposure risk value, preset data availability requirement value, preset desensitization processing timeliness requirement value, and the sensitive information motion evaluation score and sensitive information density probability value. The video stream desensitization requirement parameters are compared with the preset desensitization requirement thresholds, and the video stream desensitization requirement level is obtained according to the threshold range it falls into, including high level, medium level or low level. Also includes: The visual feature data, semantic feature data, and audio feature data are input into a preset multimodal cross-domain correlation evaluation model for processing to obtain data correlation, including visual semantic correlation, visual audio correlation, and semantic audio correlation. The visual-semantic correlation, visual-audio correlation, and semantic-audio correlation are weighted and summed to obtain the cross-modal fusion correction factor. The sensitive information identification index is corrected according to the cross-modal fusion correction factor to obtain the video sensitive information identification correction index; Also includes: Emergency events are identified based on the real-time video frame data; If an emergency event is detected, the desensitization strategy adjustment mechanism is triggered to obtain a temporary desensitization strategy; Desensitization is performed according to the temporary desensitization strategy described above.

2. The method for dynamically desensitizing sensitive information in a video stream according to claim 1, characterized in that, The step of weightedly fusing the dynamic desensitization and identification feature data according to the sensitive information category feature data, and inputting it into a preset video sensitive information identification model for processing to obtain a sensitive information identification index includes: Based on the sensitive information category feature data, query the preset scene type and weight value mapping table to obtain the weight values ​​corresponding to visual feature data, semantic feature data and audio feature data, including visual feature weight values, semantic feature weight values ​​and audio feature weight values; The visual feature data, semantic feature data, and audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain fused feature data. The fused feature data is input into a preset video sensitive information recognition model for processing to obtain a sensitive information recognition index.

3. The method for dynamically desensitizing sensitive information in a video stream according to claim 2, characterized in that, The step of dynamically desensitizing the video stream according to the desensitization strategy to obtain a desensitized video stream includes: The video stream is split according to the sensitive information category feature data to obtain a video stream set; The video stream set is extracted according to a preset duration to obtain a pre-desensitized video stream set; The pre-desensitized video stream set is dynamically desensitized according to the desensitization strategy to obtain a pre-desensitized video stream. The pre-desensitized video stream is analyzed and processed to obtain a qualified pre-desensitization status. If the pre-desensitization qualification status is unqualified, then the desensitization strategy is adjusted; If the pre-desensitization qualification status is qualified, dynamic desensitization processing is performed according to the desensitization strategy and the video stream desensitization requirement level to obtain a desensitized video stream.

4. The method for dynamically desensitizing sensitive information in a video stream according to claim 3, characterized in that, The step of analyzing and processing the pre-desensitized video stream to obtain a pre-desensitization qualified status includes: Obtain the average user rating of the pre-de-identified video stream; The pre-desensitized video stream is subjected to feature extraction to obtain pre-desensitized identification feature data, including pre-desensitized visual feature data, pre-desensitized semantic feature data and pre-desensitized audio feature data; The pre-desensitized visual feature data, pre-desensitized semantic feature data, and pre-desensitized audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain pre-desensitized fused feature data. The pre-desensitized fusion feature data is input into a preset video sensitive information recognition model for processing to obtain the pre-desensitized information recognition index. The average user rating is compared with a preset user rating threshold to obtain the user rating evaluation result, including pass or fail. The pre-desensitized information identification index is compared with the preset sensitive information identification threshold to obtain the index comparison result, including pass or fail. The user rating evaluation result is compared with the index result by AND operation. If the result is satisfactory, the pre-desensitization status is determined to be satisfactory. Conversely, if the pre-desensitization status is not qualified, it is determined to be unqualified.

5. A dynamic desensitization system for sensitive information in a video stream, characterized in that, The system includes a memory and a processor. The memory contains a program for a method of dynamically de-identifying sensitive information in a video stream. When the program for the method of dynamically de-identifying sensitive information in a video stream is executed by the processor, it performs the following steps: Acquire real-time video frame data from the video stream, extract features based on the real-time video frame data, and obtain dynamic desensitization recognition feature data and sensitive information category feature data; The dynamic desensitization and identification feature data are weighted and fused according to the sensitive information category feature data, and then input into a preset video sensitive information identification model for processing to obtain a sensitive information identification index. The sensitive information identification index is compared with the preset sensitive information identification threshold. If the sensitive information identification index is less than the preset sensitive information identification threshold, no desensitization processing will be performed; If the sensitive information identification index is greater than or equal to the preset sensitive information identification threshold, then the preset desensitization strategy database is queried according to the sensitive information category feature data to obtain the corresponding desensitization strategy; The video stream is dynamically desensitized according to the desensitization strategy to obtain a desensitized video stream. The process involves acquiring real-time video frame data from the video stream, extracting features from the real-time video frame data, and obtaining dynamic desensitization identification feature data and sensitive information category feature data, including: Acquire real-time video frame data from the video stream; Feature extraction is performed on the real-time video frame data to obtain dynamic desensitization recognition feature data, including visual feature data, semantic feature data and audio feature data; The visual feature data, semantic feature data, and audio feature data are input into a preset video sensitive information type recognition model for processing to obtain sensitive information category feature data. The sensitive information category feature data includes security monitoring category feature data, medical imaging category feature data, traffic monitoring category feature data, remote office category feature data, live video streaming category feature data, or home monitoring category feature data; Also includes: Obtain the preset sensitive information exposure risk value, preset data availability requirement value, and preset desensitization processing timeliness requirement value corresponding to the sensitive information category feature data; The motion evaluation data and density probability value of sensitive information in the real-time video frame data are obtained, wherein the motion evaluation data of sensitive information includes motion speed and motion trajectory data; The motion speed and trajectory data are input into a preset sensitive information motion evaluation model for processing to obtain a sensitive information motion evaluation score. The sensitive information density probability value is determined based on the threshold range to which the sensitive information density probability value belongs; The video stream desensitization requirement parameters are obtained by weighting and summing the preset sensitive information exposure risk value, preset data availability requirement value, preset desensitization processing timeliness requirement value, and the sensitive information motion evaluation score and sensitive information density probability value. The video stream desensitization requirement parameters are compared with the preset desensitization requirement thresholds, and the video stream desensitization requirement level is obtained according to the threshold range it falls into, including high level, medium level or low level. Also includes: The visual feature data, semantic feature data, and audio feature data are input into a preset multimodal cross-domain correlation evaluation model for processing to obtain data correlation, including visual semantic correlation, visual audio correlation, and semantic audio correlation. The visual-semantic correlation, visual-audio correlation, and semantic-audio correlation are weighted and summed to obtain the cross-modal fusion correction factor. The sensitive information identification index is corrected according to the cross-modal fusion correction factor to obtain the video sensitive information identification correction index; Also includes: Emergency events are identified based on the real-time video frame data; If an emergency event is detected, the desensitization strategy adjustment mechanism is triggered to obtain a temporary desensitization strategy; Desensitization is performed according to the temporary desensitization strategy described above.

6. The dynamic desensitization system for sensitive information in a video stream according to claim 5, characterized in that, The step of weightedly fusing the dynamic desensitization and identification feature data according to the sensitive information category feature data, and inputting it into a preset video sensitive information identification model for processing to obtain a sensitive information identification index includes: Based on the sensitive information category feature data, query the preset scene type and weight value mapping table to obtain the weight values ​​corresponding to visual feature data, semantic feature data and audio feature data, including visual feature weight values, semantic feature weight values ​​and audio feature weight values; The visual feature data, semantic feature data, and audio feature data, along with their corresponding visual feature weight values, semantic feature weight values, and audio feature weight values, are weighted and fused to obtain fused feature data. The fused feature data is input into a preset video sensitive information recognition model for processing to obtain a sensitive information recognition index.

Citation Information

Patent Citations

  • Video sensitive information processing method and device, equipment and medium

    CN115495772A

  • Video desensitization method, device and equipment based on continuous frame tracking and readable storage medium

    CN118779899A

  • Sensitive information leakage protection system for video transmission

    CN119031182A