Multimodal Large Model-Driven Data Desensitization and Video Data Protection Method and Device

Through the multimodal large model-driven data desensitization method, the video data is intelligently desensitized, solving the problem of poor desensitization effect in the existing technology and achieving high-quality video data protection.

CN119808161BActive Publication Date: 2025-05-30HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510286062.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-30
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

When the prior art desensitizes video data, the desensitization effect is poor, resulting in poor quality of desensitized videos and inability to effectively protect the original information of the video data.

Method used

The multimodal large model-driven data desensitization method is adopted. By determining the target desensitization strategy corresponding to the target to be desensitized data, the target to be desensitized data is desensitized based on the multimodal large model until the desensitization effect parameters meet the desensitization quality conditions.

Benefits of technology

The quality and desensitization effect of desensitized videos are improved, so that desensitized videos can better reflect the original information of the video data, and enhance data security and usability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808161B_ABST
    Figure CN119808161B_ABST
Patent Text Reader

Abstract

The present application provides a multi-modal large model-driven data desensitization and video data protection method and device. The method includes: determining target data to be desensitized based on the video to be desensitized; determining a target desensitization strategy corresponding to the target data to be desensitized, and performing desensitization processing on the target data to be desensitized through a multi-modal large model based on the target desensitization strategy to obtain desensitized data; determining a desensitization effect parameter based on the target data to be desensitized and the desensitized data; if the desensitization effect parameter meets the desensitization quality condition, generating a desensitized video corresponding to the video to be desensitized based on the desensitized data, and sending the desensitized video; if the desensitization effect parameter does not meet the desensitization quality condition, adjusting the target desensitization strategy to obtain an adjusted desensitization strategy, using the adjusted desensitization strategy as the target desensitization strategy, and performing desensitization processing on the target data to be desensitized based on the target desensitization strategy to obtain desensitized data. Through the solution of the present application, the desensitization strategy can be adaptively adjusted to improve the desensitization effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security, and in particular, to a data desensitization and video data protection method and device driven by a multimodal large model. Background Art

[0002] Data desensitization is a data security technology aimed at deforming sensitive information contained in the original data through preset rules and algorithms to protect private data. The main purpose of data desensitization is to ensure the privacy and security of sensitive information without affecting the value of data use. By desensitizing the original data, data security can be protected, and the risk of data leakage can be effectively reduced. However, when desensitizing video data to obtain the desensitized video, the quality of the desensitized video is usually poor, that is, the desensitization effect is poor, and the desensitized video cannot reflect the original information of the video data. Summary of the Invention

[0003] This application provides a data desensitization and video data protection method driven by a multimodal large model. The method includes: determining target data to be desensitized based on the acquired video to be desensitized; determining a target desensitization strategy corresponding to the target data to be desensitized, and desensitizing the target data to be desensitized based on the target desensitization strategy through a multimodal large model to obtain desensitized data; determining a desensitization effect parameter based on the target data to be desensitized and the desensitized data; if the desensitization effect parameter meets the desensitization quality condition, generating a desensitized video corresponding to the video to be desensitized based on the desensitized data, and sending the desensitized video; if the desensitization effect parameter does not meet the desensitization quality condition, adjusting the target desensitization strategy to obtain an adjusted desensitization strategy, using the adjusted desensitization strategy as the target desensitization strategy, and returning to execute the operation of desensitizing the target data to be desensitized based on the target desensitization strategy to obtain desensitized data.

[0004] This application provides a data desensitization and video data protection device driven by a multi-modal large model. The device includes: a determination module, configured to determine target data to be desensitized based on the acquired video to be desensitized, and determine a target desensitization strategy corresponding to the target data to be desensitized; a processing module, configured to desensitize the target data to be desensitized through a multi-modal large model based on the target desensitization strategy to obtain desensitized data; the determination module is further configured to determine a desensitization effect parameter based on the target data to be desensitized and the desensitized data, and the desensitization effect parameter is used to reflect the desensitization quality of the desensitized data; a generation module, configured to generate a desensitized video corresponding to the video to be desensitized based on the desensitized data and send the desensitized video if the desensitization effect parameter meets the desensitization quality condition; the processing module is configured to adjust the target desensitization strategy to obtain an adjusted desensitization strategy if the desensitization effect parameter does not meet the desensitization quality condition, use the adjusted desensitization strategy as the target desensitization strategy, and desensitize the target data to be desensitized based on the target desensitization strategy to obtain desensitized data.

[0005] This application provides an electronic device, a processor, and a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor is configured to execute the machine-executable instructions to implement the multi-modal large model-driven data desensitization and video data protection method in the above examples of this application.

[0006] This application provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the multi-modal large model-driven data desensitization and video data protection method in the above examples of this application.

[0007] This application provides a machine-readable storage medium that stores machine-executable instructions that can be executed by a processor. Wherein, the processor is configured to execute the machine-executable instructions, and when the machine-executable instructions are executed, implement the multi-modal large model-driven data desensitization and video data protection method in the above examples of this application.

[0008] As can be seen from the above technical solutions, in the embodiments of the present application, after desensitizing the target data to be desensitized to obtain the desensitized data, the desensitization effect parameter is determined based on the target data to be desensitized and the desensitized data. If the desensitization effect parameter does not meet the desensitization quality condition, the target desensitization strategy is adjusted to obtain the adjusted desensitization strategy, and the target data to be desensitized is desensitized again based on the adjusted desensitization strategy until the desensitization effect parameter meets the desensitization quality condition. When the desensitization effect parameter meets the desensitization quality condition, the desensitized video is generated, so that the quality of the desensitized video is good, the desensitization effect is good, and the desensitized video can reflect the original information of the video data. In this way, the desensitization effect of the desensitized video can be evaluated, the desensitization strategy can be adaptively adjusted to improve the desensitization effect, and the desensitized video has usability. While improving the data security and usability of the desensitized video, it can make the image quality evaluation more accurately adapt to the dynamic changes of the video data, evaluate whether the desensitization operation affects the value of the video data in subsequent applications, and can automatically adjust the desensitization strategy to achieve the best balance between security and usability. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a schematic flowchart of a multi-modal large model-driven data desensitization and video data protection method in the present application;

[0010] Figure 2 is a schematic flowchart of a multi-modal large model-driven data desensitization and video data protection method in the present application;

[0011] Figure 3 is a schematic diagram of the sensitive information recognition process in an embodiment of the present application;

[0012] Figure 4 is a schematic structural diagram of a multi-modal large model-driven data desensitization and video data protection device in the present application;

[0013] Figure 5 is a hardware structure diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] In the embodiments of the present application, a multi-modal large model-driven data desensitization and video data protection method is proposed. This method can be applied to an electronic device. Refer to Figure 1 shown, which is a schematic flowchart of the multi-modal large model-driven data desensitization and video data protection method. The method may include:

[0015] Step 101: Determine the target data to be desensitized based on the acquired video to be desensitized.

[0016] Step 102: Determine the target desensitization strategy corresponding to the target data to be desensitized, and desensitize the target data to be desensitized through a multimodal large model based on the target desensitization strategy to obtain the desensitized data.

[0017] Step 103: Determine the desensitization effect parameters based on the target data to be desensitized and the desensitized data.

[0018] Step 104: Determine whether the desensitization effect parameters have met the desensitization quality conditions.

[0019] If so, that is, the desensitization effect parameters have met the desensitization quality conditions, then execute Step 105.

[0020] If not, that is, the desensitization effect parameters have not met the desensitization quality conditions, then execute Step 106.

[0021] Step 105: Generate a desensitized video corresponding to the video to be desensitized based on the desensitized data, and send the desensitized video, that is, send the desensitized video to the receiving end of the video to be desensitized.

[0022] Step 106: Adjust the target desensitization strategy to obtain an adjusted desensitization strategy, use the adjusted desensitization strategy as the target desensitization strategy, and return to execute Step 102, that is, re-desensitize the target data to be desensitized based on the target desensitization strategy (i.e., the adjusted desensitization strategy) to obtain the desensitized data.

[0023] Exemplarily, the target data to be desensitized may include at least two of the data to be desensitized image data, the data to be desensitized audio data, and the data to be desensitized text data. If the target data to be desensitized includes the data to be desensitized image data, the data to be desensitized audio data, and the data to be desensitized text data, then the desensitized data includes the desensitized image data, the desensitized audio data, and the desensitized text data. Based on this, determining the desensitization effect parameters based on the target data to be desensitized and the desensitized data may include, but is not limited to: determining the feature matching degree based on the image matching degree between the desensitized image data and the data to be desensitized image data, the audio matching degree between the desensitized audio data and the data to be desensitized audio data, and the text matching degree between the desensitized text data and the data to be desensitized text data. Determine the image quality based on the desensitized image data and the data to be desensitized image data. Obtain the desensitization effect parameters, which may include the feature matching degree and the image quality. For example, for Step 104, when determining whether the desensitization effect parameters have met the desensitization quality conditions, if the feature matching degree is greater than the first threshold and the image quality is greater than the second threshold, then the desensitization effect parameters have met the desensitization quality conditions; if the feature matching degree is not greater than the first threshold, and / or the image quality is not greater than the second threshold, then the desensitization effect parameters have not met the desensitization quality conditions.

[0024] Exemplarily, determining the image quality based on the desensitized image data and the image data to be desensitized may include, but is not limited to: determining a first similarity value based on a target brightness weight and a brightness parameter value, determining a second similarity value based on a target contrast weight and a contrast parameter value, and determining a third similarity value based on a target structure weight and a structure parameter value; determining the image quality based on the first similarity value, the second similarity value, and the third similarity value. The brightness parameter value is determined based on the mean pixel value of the desensitized image data and the mean pixel value of the image data to be desensitized, the contrast parameter value is determined based on the pixel value variance of the desensitized image data and the pixel value variance of the image data to be desensitized, and the structure parameter value is determined based on the pixel value variance of the desensitized image data, the pixel value variance of the image data to be desensitized, and the covariance of the pixel values of the desensitized image data and the image data to be desensitized.

[0025] The process of determining the target brightness weight may include: determining an initial brightness weight corresponding to the frame rate of the video to be desensitized; determining the target brightness weight based on the initial brightness weight. Among them, if the frame rate is less than a first value, the initial brightness weight is a first brightness weight value; if the frame rate is greater than a second value, the initial brightness weight is a second brightness weight value; if the frame rate is not less than the first value and not greater than the second value, the initial brightness weight is between the second brightness weight value and the first brightness weight value, and the initial brightness weight is inversely proportional to the frame rate.

[0026] The process of determining the target contrast weight may include: determining an initial contrast weight corresponding to the frame rate of the video to be desensitized; determining the target contrast weight based on the initial contrast weight. Among them, if the frame rate is less than a third value, the initial contrast weight may be a first contrast weight value; if the frame rate is greater than a fourth value, the initial contrast weight may be a second contrast weight value; if the frame rate is not less than the third value and not greater than the fourth value, the initial contrast weight may be between the second contrast weight value and the first contrast weight value, and the initial contrast weight may be inversely proportional to the frame rate.

[0027] The process of determining the target structure weight may include: determining an initial structure weight corresponding to the frame rate; determining the target structure weight based on the initial structure weight. Among them, if the frame rate is less than a fifth value, the initial structure weight is a first structure weight value; if the frame rate is greater than a sixth value, the initial structure weight is a second structure weight value; if the frame rate is not less than the fifth value and not greater than the sixth value, the initial structure weight is between the first structure weight value and the second structure weight value, and the initial structure weight is directly proportional to the frame rate.

[0028] Exemplarily, determining the target brightness weight based on the initial brightness weight includes: determining the initial brightness weight as the target brightness weight; or, adjusting the initial brightness weight based on the motion degree of the video to be desensitized, and determining the adjusted brightness weight as the target brightness weight; wherein, the motion degree is determined based on the motion vectors between adjacent frames in the video to be desensitized; if the motion degree is less than the first motion threshold, the initial brightness weight is increased, and if the motion degree is greater than the second motion threshold, the initial brightness weight remains unchanged. Determining the target contrast weight based on the initial contrast weight includes: determining the initial contrast weight as the target contrast weight; or, adjusting the initial contrast weight based on the motion degree, and determining the adjusted contrast weight as the target contrast weight; if the motion degree is less than the first motion threshold, the initial contrast weight remains unchanged; if the motion degree is greater than the second motion threshold, the initial contrast weight is increased. Determining the target structure weight based on the initial structure weight may include: determining the initial structure weight as the target structure weight; or, adjusting the initial structure weight based on the motion degree, and determining the adjusted structure weight as the target structure weight; wherein, if the motion degree is less than the first motion threshold, the initial structure weight remains unchanged; if the motion degree is greater than the second motion threshold, the initial structure weight is increased.

[0029] Exemplarily, if the motion degree is not less than the first motion threshold and not greater than the second motion threshold, the initial brightness weight, the initial contrast weight, and the initial structure weight can be adjusted so that the difference between the adjusted brightness weight, the adjusted contrast weight, and the adjusted structure weight is not greater than the threshold.

[0030] Exemplarily, if the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, the desensitization process of the target data to be desensitized based on the target desensitization strategy to obtain the desensitized data may include, but is not limited to: extracting image features from the image data to be desensitized through a multimodal large model, extracting audio features from the audio data to be desensitized through a multimodal large model, extracting text features from the text data to be desensitized through a multimodal large model, and fusing the image features, audio features, and text features through a multimodal large model to obtain multimodal features; determining the image sensitive positions of the image data to be desensitized, the audio sensitive positions of the audio data to be desensitized, and the text sensitive positions of the text data to be desensitized based on the multimodal features through a multimodal large model; desensitizing the image sensitive positions of the image data to be desensitized based on the image desensitization strategy corresponding to the image data to be desensitized to obtain the desensitized image data; desensitizing the audio sensitive positions of the audio data to be desensitized based on the audio desensitization strategy corresponding to the audio data to be desensitized to obtain the desensitized audio data; desensitizing the text sensitive positions of the text data to be desensitized based on the text desensitization strategy corresponding to the text data to be desensitized to obtain the desensitized text data.

[0031] Exemplarily, if the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, the target desensitization strategy includes an image desensitization strategy, an audio desensitization strategy, and a text desensitization strategy. Determining the target desensitization strategy corresponding to the target data to be desensitized may include, but is not limited to: querying the configured desensitization rule library through the multi-modal attributes corresponding to the video to be desensitized to obtain the image desensitization strategy corresponding to the multi-modal attributes; wherein, the multi-modal attributes may include at least one of image attributes, audio attributes, and text attributes, and the desensitization rule library may include the corresponding relationship between the multi-modal attributes and the image desensitization strategy; wherein, the image desensitization strategy includes an image desensitization network model. Determining the audio desensitization strategy based on the audio attributes corresponding to the video to be desensitized, the audio desensitization strategy may include an audio desensitization network model and a vocoder, and the audio desensitization network model is used to convert the audio data to be desensitized into acoustic features, and the vocoder is used to convert the acoustic features into audio waveforms. Determining the configured text desensitization strategy, the text desensitization strategy includes a text encryption algorithm or a text replacement algorithm, the text encryption algorithm is used to perform an encryption operation on the text data to be desensitized, and the text replacement algorithm is used to perform a replacement operation or a fuzzification operation on the text data to be desensitized.

[0032] Among them, when adjusting the target desensitization strategy to obtain the adjusted desensitization strategy, if it is necessary to adjust the image desensitization strategy, the network parameters in the image desensitization network model are adjusted; if it is necessary to adjust the audio desensitization strategy, the network parameters in the audio desensitization network model are adjusted; if it is necessary to adjust the text desensitization strategy, the algorithm type adopted by the text encryption algorithm is adjusted, or the algorithm type adopted by the text replacement algorithm can be adjusted.

[0033] As can be seen from the above technical solutions, in the embodiments of the present application, after desensitizing the target data to be desensitized to obtain the desensitized data, the desensitization effect parameter is determined based on the target data to be desensitized and the desensitized data. If the desensitization effect parameter does not meet the desensitization quality condition, the target desensitization strategy is adjusted to obtain the adjusted desensitization strategy, and the target data to be desensitized is desensitized again based on the adjusted desensitization strategy until the desensitization effect parameter meets the desensitization quality condition. When the desensitization effect parameter meets the desensitization quality condition, a desensitized video is generated, so that the quality of the desensitized video is good, the desensitization effect is good, and the desensitized video can reflect the original information of the video data. In this way, the desensitization effect of the desensitized video can be evaluated, the desensitization strategy can be adaptively adjusted to improve the desensitization effect, and the desensitized video has usability. While improving the data security and usability of the desensitized video, it can make the image quality assessment more accurately adapt to the dynamic changes of the video data, can evaluate whether the desensitization operation affects the value of the video data in subsequent applications, and can automatically adjust the desensitization strategy to achieve the best balance between security and usability.

[0034] The above technical solutions of the embodiments of the present application will be described below in combination with specific application scenarios.

[0035] In the embodiments of the present application, a data desensitization and video data protection method driven by a multi-modal large model is proposed, which can be applied to an electronic device. The electronic device can be the video sending end itself, or it can be a device between the video sending end and the video receiving end (such as a network device, a security device, etc.). The type of this electronic device is not limited.

[0036] See Figure 2 As shown in the figure, it is a schematic flowchart of the data desensitization and video data protection method driven by a multi-modal large model. The method may include:

[0037] Step 201, determine target data to be desensitized based on the acquired video to be desensitized.

[0038] Exemplarily, the target data to be desensitized may include at least two of image data to be desensitized, audio data to be desensitized, and text data to be desensitized. For example, the image data to be desensitized may be an image to be desensitized, the audio data to be desensitized may be an audio to be desensitized, and the text data to be desensitized may be a text to be desensitized.

[0039] After obtaining the video to be desensitized, since the video to be desensitized consists of a large number of images, all the images of the video to be desensitized can be used as the image data to be desensitized, or some of the images of the video to be desensitized can be used as the image data to be desensitized. In this way, the image data to be desensitized can be obtained from the video to be desensitized.

[0040] After obtaining the video to be desensitized, since the video to be desensitized includes sound data, and the sound data is the audio data to be desensitized, the audio data to be desensitized can be obtained from the video to be desensitized.

[0041] After obtaining the video to be desensitized, since the video to be desensitized includes text data (such as subtitle information, bullet screen information, text description information, etc., and the type of this text data is not limited), and the text data is the text data to be desensitized, the text data to be desensitized can be obtained from the video to be desensitized. For example, natural language processing technology can be used to extract the text data to be desensitized from the video to be desensitized.

[0042] In summary, it can be seen that the image data to be desensitized, the audio data to be desensitized, and the text data to be desensitized can be obtained from the video to be desensitized, and the target data to be desensitized may include at least two of the image data to be desensitized, the audio data to be desensitized, and the text data to be desensitized. For example, the target data to be desensitized may include the image data to be desensitized, the audio data to be desensitized, and the text data to be desensitized at the same time.

[0043] Step 202: Identify sensitive information based on the target data to be desensitized to obtain the sensitive positions of the target data to be desensitized. For example, if the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, then identify sensitive information based on the image data to be desensitized, audio data to be desensitized, and text data to be desensitized to obtain the sensitive positions of the image data to be desensitized, the sensitive positions of the audio data to be desensitized, and the sensitive positions of the text data to be desensitized. If the target data to be desensitized includes image data to be desensitized and audio data to be desensitized, then identify sensitive information based on the image data to be desensitized and the audio data to be desensitized to obtain the sensitive positions of the image data to be desensitized and the sensitive positions of the audio data to be desensitized, and so on.

[0044] For the sake of convenience of description, hereinafter, an example in which the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized is used. In this way, the sensitive information identification process may include the following steps:

[0045] Step S11: Extract features from the image data to be desensitized to obtain image features.

[0046] For example, the image data to be desensitized may include multiple images to be desensitized. The image data to be desensitized can be input into the image feature extraction network of the multi-modal large model, and the image feature extraction network extracts features from the image data to be desensitized to obtain image features. The image features may include, but are not limited to, texture features, shape features, color features, and spatial layout features. The type of these image features is not limited in this embodiment.

[0047] For example, the image feature extraction network can be a ResNet (Residual Network) network, or other types of networks can also be used as long as they have the function of image feature extraction. Taking ResNet as an example for illustration. ResNet can adopt a custom activation function (such as LeakyReLU) and add a Batch Normalization layer to accelerate network convergence and improve generalization ability. In the network structure of ResNet, the convolution kernel size and stride of some convolutional layers are adjusted to adapt to the resolution and detail features of different video frames. In the network structure of ResNet, through multi-layer convolutional operations, images are subjected to multi-scale feature extraction using convolutional kernels of different sizes (such as 3*3, 5*5, etc.), and then through pooling operations (such as max pooling or average pooling, etc.) and residual connections, the image features of video frames are effectively extracted. In this way, the image features of the image data to be desensitized can be obtained through feature extraction. In the network structure of ResNet, for high-resolution videos, the depth of the convolutional layer can be increased and smaller convolutional kernels can be used to better capture the fine textures of the image data to be desensitized. For low-resolution videos, the convolutional kernel and stride can be appropriately increased to improve the feature extraction efficiency of the image data to be desensitized.

[0048] Step S12: Extract audio features from the audio data to be desensitized.

[0049] For example, the audio data to be desensitized can be input into the audio feature extraction network of the multi-modal large model, and the audio feature extraction network extracts audio features (which can also be called audio feature vectors) from the audio data to be desensitized. The audio features can include but are not limited to frequency features, amplitude features, harmonic structure features, and speech prosody features. The type of audio features is not limited in this embodiment.

[0050] For example, the audio feature extraction network can adopt an MFCC (Mel Frequency Cepstrum Coefficient) network and a DNN (Deep Neural Networks), or other types of networks, as long as it has the function of audio feature extraction. Taking the MFCC network and the DNN network as examples for illustration. Based on this, the audio data to be desensitized can be input into the MFCC network, and the audio data to be desensitized (i.e., audio signal) is decomposed into multiple Mel frequency bands by the MFCC network to obtain frequency domain features. Then, the frequency domain features are input into the DNN network, and the DNN network performs feature extraction based on the frequency domain features to obtain audio features, thereby converting the audio signal into audio features with semantic and structural information. In the network structure of the DNN network, 5 fully connected layers can be adopted, and the number of nodes in each fully connected layer is 256, 128, 64, 32, and 16 in sequence. For the training process of the DNN network, the DNN network can be trained using the stochastic gradient descent algorithm, and the Dropout technique is adopted to prevent overfitting. Through the training of the DNN network, the DNN network can capture the frequency features, amplitude features, harmonic structure features, and speech prosody features of the audio. In this way, audio features can be obtained through feature extraction. When the audio feature extraction network performs feature extraction on the audio data to be desensitized to obtain audio features, for the audio data to be desensitized in different audio coding formats (such as MP3, WAV, etc.), the audio data to be desensitized can be decoded into a unified audio signal format, and then feature extraction is performed through the audio feature extraction network. Before the audio data to be desensitized is input into the audio feature extraction network, a noise reduction algorithm (such as Wiener filtering) can also be used to process the noise interference of the audio data to be desensitized to ensure the stability of feature extraction.

[0051] Step S13: Perform feature extraction on the text data to be desensitized to obtain text features.

[0052] For example, the text data to be desensitized can be input into the text feature extraction network of the multi-modal large model, and the text feature extraction network performs feature extraction on the text data to be desensitized to obtain text features (which can also be called text feature vectors). These text features can be high-dimensional semantic features, and the type of these text features is not restricted.

[0053] For example, the text feature extraction network can adopt a Transformer network (such as a Transformer-based encoder). The core of the Transformer network is the multi-head attention mechanism. Other types of networks can also be used as long as they have the function of text feature extraction. Taking the Transformer network as an example for illustration. When using the Transformer network to process the text data to be desensitized, the multi-head attention mechanism divides the text data to be desensitized into multiple heads, and each head focuses on different semantic dimensions. For example, one head focuses on lexical semantics, and another head processes syntactic structures. Through parallel processing and aggregation, deep semantic encoding of the text is achieved. At the same time, position encoding is used to process the sequential information of the text, convert the text information into word vector representations, and mine multi-dimensional information such as the lexical semantics, syntactic structures, semantic relationships, and discourse semantics of the text. Finally, the text is converted into a high-dimensional semantic feature representation, that is, the text feature. When using the Transformer network to process the text data to be desensitized, for the text data to be desensitized in different languages, pre-trained word embeddings (such as Word2Vec, etc.) can also be used to adjust the embedding matrix according to language characteristics. In addition, for the text data to be desensitized with long texts, methods such as truncation or truncation plus padding can also be used to ensure that the input lengths of the text data to be desensitized are consistent.

[0054] Step S14: Perform feature fusion on the image feature, the audio feature, and the text feature to obtain a multi-modal feature, and determine the image sensitive position (i.e., the sensitive area) of the image data to be desensitized, the audio sensitive position of the audio data to be desensitized, and the text sensitive position of the text data to be desensitized based on the multi-modal feature.

[0055] For example, input the image feature, the audio feature, and the text feature into a multi-modal large model (such as the fusion network of the multi-modal large model, which can also be called the fusion layer of the multi-modal large model). Through the multi-modal large model, perform feature fusion on the image feature, the audio feature, and the text feature to obtain a multi-modal feature, and determine the image sensitive position of the image data to be desensitized, the audio sensitive position of the audio data to be desensitized, and the text sensitive position of the text data to be desensitized based on the multi-modal feature through the multi-modal large model.

[0056] For example, the multi-modal large model can include a fusion layer. The fusion layer of the multi-modal large model adopts an Attention network with an attention mechanism (such as an Attention network with a multi-head cross-modal attention mechanism). The image feature, the audio feature, and the text feature can be input into the fusion layer, and the fusion layer fuses these features to obtain a multi-modal feature. The multi-modal feature is also called a multi-modal joint feature or a multi-modal fusion feature.

[0057] For the fusion layer of the multimodal large model, by calculating the similarity matrix between different modal features (such as image features, audio features, and text features), weights are dynamically assigned according to the similarity matrix, and the dynamic association weights between different modal features are adaptively learned in a data-driven manner. Then, based on the dynamic association weights between different modal features, image features, audio features, and text features can be fused.

[0058] In the processing of the fusion layer of the multimodal large model, the following formula can also be used for processing: . In the above formula, Q is the query vector, K is the key vector, V is the value vector, is the dimension of the key vector. In this embodiment, there is no limitation on the processing method of this fusion layer, as long as the fusion layer can fuse image features, audio features, and text features.

[0059] For example, the multimodal large model can include a decision layer (which can also be called the decision network of the multimodal large model). Multimodal features can be input into the decision layer of the multimodal large model, and the decision layer can accurately locate the image sensitive positions of the image data to be desensitized, the audio sensitive positions of the audio data to be desensitized, and the text sensitive positions of the text data to be desensitized.

[0060] For example, the decision layer of the multimodal large model can adopt an object detection network (such as the Faster R-CNN network based on deep learning), a speech recognition network (such as the end-to-end speech recognition network based on deep learning), and a text classification network (such as the text classifier based on convolutional neural network). The input data of the object detection network is multimodal features. The object detection network filters and refines the candidate regions of the image data to be desensitized (i.e., video images) based on the multimodal features, so as to accurately locate the sensitive information region in the image data to be desensitized, and the sensitive information region is the image sensitive position of the image data to be desensitized. In this embodiment, there is no limitation on the processing process of the object detection network.

[0061] The input data of the speech recognition network is multimodal features. The speech recognition network recognizes the sensitive speech content in the audio data to be desensitized (i.e., audio) based on the multimodal features, so as to accurately locate the sensitive speech content in the audio data to be desensitized, and the position where the sensitive speech content is located is the audio sensitive position of the audio data to be desensitized. In this embodiment, there is no limitation on the processing process of the speech recognition network. Based on the audio sensitive position, the object detection network can also locate the corresponding image position in the image data to be desensitized, and this image position is also the image sensitive position of the image data to be desensitized.

[0062] The input data of the text classification network is multi-modal features. The text classification network identifies sensitive information in the text data to be de-sensitized (i.e., the text) based on the multi-modal features, so as to accurately locate the sensitive information in the text data to be de-sensitized. The position where the sensitive information is located is the text sensitive position of the text data to be de-sensitized. In this embodiment, the processing process of the text classification network is not limited. Based on the text sensitive position, the target detection network can also locate the corresponding image position in the image data to be de-sensitized from the text sensitive position. This image position is also the image sensitive position of the image data to be de-sensitized.

[0063] For example, when it is recognized that a specific sensitive word is mentioned in the dialogue of a person in a video, combined with the facial expression of the person in the image and the intonation of the audio, the image sensitive position of the image data to be de-sensitized, the audio sensitive position of the audio data to be de-sensitized, and the text sensitive position of the text data to be de-sensitized are comprehensively determined.

[0064] In a possible implementation manner, as shown in Figure 3 shown, it is a schematic diagram of the sensitive information recognition process. The image data to be de-sensitized is input to the image processing module (image feature extraction network), and the image processing module extracts features from the image data to be de-sensitized to obtain image features. The audio data to be de-sensitized is input to the audio processing module (audio feature extraction network), and the audio processing module extracts features from the audio data to be de-sensitized to obtain audio features. The text data to be de-sensitized is input to the text processing module (text feature extraction network), and the text processing module extracts features from the text data to be de-sensitized to obtain text features.

[0065] Then, the image features, audio features, and text features are input to the fusion layer of the multi-modal large model, and the fusion layer performs feature fusion on the image features, audio features, and text features to obtain multi-modal features.

[0066] Then, the multi-modal features are input to the decision layer of the multi-modal large model, and the decision layer locates the sensitive positions (such as the image sensitive position of the image data to be de-sensitized, the audio sensitive position of the audio data to be de-sensitized, and the text sensitive position of the text data to be de-sensitized), and the sensitive positions are used for subsequent de-sensitization processing.

[0067] So far, step 202 is completed, and the sensitive positions of the target data to be de-sensitized are obtained.

[0068] Step 203, determine the target de-sensitization strategy corresponding to the target data to be de-sensitized, and perform de-sensitization processing on the target data to be de-sensitized based on the target de-sensitization strategy to obtain the de-sensitized data.

[0069] For example, if the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, then the image desensitization strategy corresponding to the image data to be desensitized, the audio desensitization strategy corresponding to the audio data to be desensitized, and the text desensitization strategy corresponding to the text data to be desensitized can be determined. Based on this, the image data to be desensitized is desensitized based on the image desensitization strategy to obtain desensitized image data. The audio data to be desensitized is desensitized based on the audio desensitization strategy to obtain desensitized audio data. The text data to be desensitized is desensitized based on the text desensitization strategy to obtain desensitized text data.

[0070] If the target data to be desensitized includes image data to be desensitized and audio data to be desensitized, then the image desensitization strategy corresponding to the image data to be desensitized and the audio desensitization strategy corresponding to the audio data to be desensitized are determined. The image data to be desensitized is desensitized based on the image desensitization strategy to obtain desensitized image data. The audio data to be desensitized is desensitized based on the audio desensitization strategy to obtain desensitized audio data, and so on.

[0071] For the sake of convenience in description, hereinafter, taking the target data to be desensitized including image data to be desensitized, audio data to be desensitized, and text data to be desensitized as an example, in this way, the desensitization process may include the following steps:

[0072] Step S21: Determine the image desensitization strategy corresponding to the image data to be desensitized.

[0073] Exemplarily, the image desensitization strategy can be pre-configured. After the image data to be desensitized is obtained, the configured image desensitization strategy can be determined, that is, all the image data to be desensitized corresponds to the same image desensitization strategy.

[0074] Exemplarily, a desensitization rule library can be pre-configured. The desensitization rule library can include the correspondence between multi-modal attributes and image desensitization strategies. As shown in Table 1, it is an example of the desensitization rule library.

[0075] Table 1

[0076]

[0077] Regarding the correspondence between multi-modal attributes and image desensitization strategies, this correspondence can be configured according to experience or determined by a certain algorithm, and there is no limitation on this. After the desensitization rule library is configured, the desensitization rule library can remain unchanged, or the desensitization rule library can be updated regularly.

[0078] After obtaining the image data to be desensitized, the desensitization rule library shown in Table 1 can be queried through the multi-modal attributes corresponding to the video to be desensitized to obtain the image desensitization strategy corresponding to the multi-modal attributes. For example, if the multi-modal attributes corresponding to the video to be desensitized are multi-modal attribute a1, then by querying the desensitization rule library shown in Table 1, the image desensitization strategy corresponding to multi-modal attribute a1 is image desensitization strategy b1.

[0079] Regarding the multi-modal attributes corresponding to the video to be desensitized, the multi-modal attributes can include, but are not limited to, at least one of image attributes, audio attributes, and text attributes. For example, the image attribute is an attribute used to describe the image data to be desensitized in the video to be desensitized, such as image resolution, color depth, video frame rate, etc., and the type of this image attribute is not restricted. The audio attribute is an attribute used to describe the audio data to be desensitized in the video to be desensitized, such as audio sampling rate, etc., and the type of this audio attribute is not restricted. The text attribute is an attribute used to describe the text data to be desensitized in the video to be desensitized, and the type of this text attribute is not restricted.

[0080] Based on the multi-modal attributes corresponding to the video to be desensitized, the desensitization rule library can be queried to obtain the image desensitization strategy corresponding to the image data to be desensitized, thereby generating an image desensitization strategy that is precisely adapted.

[0081] Exemplarily, the image desensitization strategy can include an image desensitization network model (such as using a replacement model based on deep learning as the image desensitization network model, which can implement the region replacement function, that is, replace a sensitive region in the image to achieve the desensitization function). That is, a certain network model is used as the image desensitization strategy. Since the network model is used to implement image desensitization, it can also be called an image desensitization network model. Of course, the image desensitization strategy can also be other types of desensitization strategies, and this is not restricted.

[0082] For example, for multi-modal attribute a1, a network model of type 1 can be used as image desensitization strategy b1, and for multi-modal attribute a2, a network model of type 2 can be used as image desensitization strategy b2, and so on. On this basis, for each multi-modal attribute, a certain type of network model can be configured as the image desensitization strategy, and different multi-modal attributes can correspond to the same or different types of network models.

[0083] For example, if it is determined based on the multi-modal attributes that the image resolution is higher than 1080p and the background of the facial area is simple, the image desensitization strategy (i.e., the image desensitization network model) corresponding to the multi-modal attributes can be the StyleGAN model in the generative adversarial network (GAN) technology. Of course, the StyleGAN model is just an example.

[0084] Step S22: After obtaining the image desensitization strategy corresponding to the image data to be desensitized, perform desensitization processing on the image sensitive positions of the image data to be desensitized based on the image desensitization strategy to obtain the desensitized image data.

[0085] For example, the sensitive positions in the image data to be desensitized can be referred to as image sensitive positions, and the image sensitive positions can be sensitive regions, such as rectangular regions, circular regions, irregular regions, etc. On this basis, desensitization processing can be performed on the image sensitive positions. For example, an image can be generated for the image sensitive positions (the image desensitization network model corresponding to the image data to be desensitized is used to generate this image). In this way, the image sensitive positions of the image data to be desensitized can be replaced with this image, thereby obtaining the desensitized image data. For example, taking the image desensitization network model as the StyleGAN model, the StyleGAN model can generate an image for the image sensitive positions, and then use this image to replace the image sensitive positions of the image data to be desensitized.

[0086] For example, when using the StyleGAN model, according to features such as age, gender, and expression, screening and matching are performed from a pre-trained rich database to generate an image for the image sensitive positions. First, calculate the cosine similarity between the feature vectors of the image sensitive positions and each feature vector in the database, and select the group of feature vectors with the highest similarity. Then, based on the feature vectors with the highest similarity, use the generator to generate a new image, and this image is used to replace the image sensitive positions of the image data to be desensitized. During the replacement process, use the Poisson fusion algorithm or multi-band fusion algorithm in image fusion technology to perform detail adjustments such as lighting and shadows on the replaced image. The Poisson fusion algorithm solves the gradient field according to the Laplace equation to ensure a smooth transition at the fusion boundary. The multi-band fusion algorithm decomposes the image into different frequency bands and performs fusion according to the weights of different frequencies, making the replaced image blend naturally with the background and achieving the desensitization effect.

[0087] Step S23: Determine the audio desensitization strategy corresponding to the audio data to be desensitized.

[0088] Exemplarily, the audio desensitization strategy can be pre-configured. After obtaining the audio data to be desensitized, the configured audio desensitization strategy can be determined, that is, all audio data to be desensitized correspond to the same audio desensitization strategy.

[0089] Exemplarily, a desensitization database can be pre-configured, and the desensitization database can include the correspondence between audio attributes and audio desensitization strategies. As shown in Table 2, it is an example of the desensitization database.

[0090] Table 2

[0091]

[0092] Regarding the correspondence between audio attributes and audio desensitization strategies, this correspondence can be configured based on experience or determined using a certain algorithm, without any restrictions in this regard. After the desensitization database is configured, the desensitization database can remain unchanged or be updated regularly.

[0093] After obtaining the audio data to be desensitized, an audio desensitization strategy can be determined based on the audio attributes corresponding to the video to be desensitized (i.e., the audio attributes corresponding to the audio data to be desensitized). For example, by querying the desensitization database shown in Table 2 using the audio attributes corresponding to the audio data to be desensitized, the audio desensitization strategy corresponding to these audio attributes can be obtained. For instance, if the audio attributes corresponding to the audio data to be desensitized are audio attribute c1, then by querying the desensitization database shown in Table 2, the audio desensitization strategy corresponding to audio attribute c1 is audio desensitization strategy d1. Regarding the audio attributes corresponding to the audio data to be desensitized, the audio attributes are used to describe the audio data to be desensitized in the video to be desensitized. For example, the audio attributes are used to indicate that the audio is a single voice and the speaking speed is slow, the audio attributes are used to indicate that the audio is a single voice and the speaking speed is fast, the audio attributes are used to indicate that the audio is multiple voices and the speaking speed is slow, and the audio attributes are used to indicate that the audio is multiple voices and the speaking speed is fast.

[0094] Exemplarily, the audio desensitization strategy can include an audio desensitization network model and a vocoder. The audio desensitization network model is used to convert the audio data to be desensitized into acoustic features, and the vocoder is used to convert the acoustic features into an audio waveform, thereby desensitizing the audio data to be desensitized through the audio waveform. Based on this, a certain network model and a vocoder can be used as the audio desensitization strategy. Since the network model is used to implement audio desensitization, it can also be called an audio desensitization network model. For example, for audio attribute c1, a type 3 network model (audio desensitization network model) and a vocoder are used as audio desensitization strategy d1, and for audio attribute c2, a type 4 network model and a vocoder are used as audio desensitization strategy d2, and so on. On this basis, for each audio attribute, a certain type of network model (audio desensitization network model) and a vocoder can be configured as the audio desensitization strategy, and different audio attributes can correspond to the same or different types of network models.

[0095] For example, if it is determined based on the audio attributes that the audio (audio data to be desensitized) is a single voice and the speaking speed is slow, the audio desensitization strategy corresponding to the audio attributes can be the Tacotron model based on deep learning and the WaveNet vocoder, and the Tacotron model can represent the audio desensitization network model.

[0096] Step S24: After obtaining the audio desensitization strategy corresponding to the audio data to be desensitized, desensitize the audio sensitive positions of the audio data to be desensitized based on the audio desensitization strategy to obtain the desensitized audio data.

[0097] For example, the sensitive positions in the audio data to be desensitized can be referred to as audio sensitive positions. On this basis, desensitization processing can be performed on the audio sensitive positions. For example, an audio waveform is synthesized for the audio sensitive positions (the audio desensitization network model and the vocoder are used to generate this audio waveform). In this way, the audio sensitive positions of the audio data to be desensitized can be replaced with the audio waveform, thereby obtaining the desensitized audio data.

[0098] For example, taking the Tacotron model and the WaveNet vocoder as an example of the audio desensitization strategy, the Tacotron model is used to convert the audio data to be desensitized into acoustic features. The audio features of the audio data to be desensitized (see step S12) and the text features of the text data to be desensitized (see step S13) are used as the input data of the Tacotron model. The Tacotron model processes the audio features and text features through an encoder, a decoder, and an attention mechanism network to obtain the acoustic features corresponding to the audio data to be desensitized, and inputs the acoustic features into the WaveNet vocoder. The WaveNet vocoder converts the acoustic features into an audio waveform. Then, the audio waveform can be used to replace the audio sensitive positions of the audio data to be desensitized to obtain the desensitized audio data.

[0099] During the processing of the Tacotron model, according to the audio features of the audio data to be desensitized (such as fundamental frequency, harmonic structure, and speech prosody, etc.), key parameters such as the timbre, intonation, and speech rate of the synthesized audio (i.e., the audio waveform) can be precisely controlled through parameter adjustment and optimization of the Tacotron model. The mean squared error (MSE) loss function and the perceptual loss function can be used to optimize the Tacotron model to ensure a high degree of coordination and consistency of the synthesized audio with the video picture and text information in terms of semantics, emotion, and time axis, and to achieve seamless docking and natural transition. For different languages and accents, by adding training data with specific languages and accents, the parameters and structure of the Tacotron model are adjusted to improve the accuracy of speech synthesis.

[0100] Step S25: Determine the text desensitization strategy corresponding to the text data to be desensitized.

[0101] Exemplarily, the text desensitization strategy can be pre-configured. After obtaining the text data to be desensitized, the configured text desensitization strategy can be determined, that is, all the text data to be desensitized correspond to the same text desensitization strategy.

[0102] For example, the text desensitization strategy can be a text encryption algorithm, and the text encryption algorithm is used to perform an encryption operation on the text data to be desensitized. For example, configure the text encryption algorithm as the text desensitization strategy. In this way, it is determined that the text desensitization strategy corresponding to the text data to be desensitized is the text encryption algorithm. The algorithm type of the text encryption algorithm can be the block encryption mode of the AES-256 encryption algorithm (such as the CBC mode), the padding mode of the AES-256 encryption algorithm (such as the PKCS7 mode), the DES encryption algorithm, etc., and there is no limitation on this.

[0103] For example, the text desensitization strategy can be a text replacement algorithm, and the text replacement algorithm is used to perform a replacement operation or a fuzzification operation on the text data to be desensitized. For example, configure the text replacement algorithm as the text desensitization strategy. In this way, it is determined that the text desensitization strategy corresponding to the text data to be desensitized is the text replacement algorithm. The algorithm type of the text replacement algorithm can be a specific symbol sequence replacement algorithm, a fuzzification replacement algorithm, etc.

[0104] Step S26: After obtaining the text desensitization strategy corresponding to the text data to be desensitized, perform desensitization processing on the text sensitive positions of the text data to be desensitized to obtain the desensitized text data.

[0105] For example, the sensitive positions in the text data to be desensitized can be called text sensitive positions, and desensitization processing can be performed on the text sensitive positions. If the text desensitization strategy is a text encryption algorithm, the text at the text sensitive positions can be encrypted using the text encryption algorithm to obtain the encrypted content, and the text at the text sensitive positions is replaced with the encrypted content to obtain the desensitized text data. Or, if the text desensitization strategy is a text replacement algorithm, replacement content can be generated based on the text replacement algorithm, and the text at the text sensitive positions is replaced with the replacement content to obtain the desensitized text data.

[0106] For example, taking the block encryption mode of the AES-256 encryption algorithm as the text encryption algorithm as an example, a hardware random number generator can be used to generate an initial key, and then the final 256-bit key is generated through a hash function and a key expansion algorithm to ensure the randomness, uniqueness, and security of the key. During the encryption process, based on the 256-bit key, the block encryption mode is used to encrypt the text at the text sensitive positions. For example, taking the text replacement algorithm (rule-based replacement algorithm) as an example, the text at the text sensitive positions can be replaced with a specific symbol sequence or fuzzified. The replacement rules are dynamically and intelligently adjusted according to the video scene where the text is located and the degree of association with other modality information. For example, in a video related to the medical field, for sensitive medical terms, according to the medical knowledge graph and multi-modal information, domain-specific synonym replacement or fuzzification strategies are adopted to retain the text semantic information and context relevance to the greatest extent.

[0107] So far, step 203 is completed, and the target data to be desensitized can be desensitized to obtain the desensitized data.

[0108] Step 204: Determine the desensitization effect parameter based on the target data to be desensitized and the desensitized data.

[0109] For example, if the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, then determine the desensitization effect parameter based on the image data to be desensitized and the desensitized image data, the audio data to be desensitized and the desensitized audio data, and the text data to be desensitized and the desensitized text data. If the target data to be desensitized includes image data to be desensitized and audio data to be desensitized, then determine the desensitization effect parameter based on the image data to be desensitized and the desensitized image data, and the audio data to be desensitized and the desensitized audio data, and so on. For the sake of convenient description, take the target data to be desensitized including image data to be desensitized, audio data to be desensitized, and text data to be desensitized as an example. In this way, the process of determining the desensitization effect parameter can include:

[0110] Step S31: Determine the image matching degree between the desensitized image data and the image data to be desensitized.

[0111] Exemplarily, the image matching degree represents the similarity degree between the desensitized image data and the image data to be desensitized. If the desensitized image data and the image data to be desensitized are more similar, the image matching degree is greater. Based on this, adopt an image matching degree determination algorithm to determine the image matching degree between the desensitized image data and the image data to be desensitized. For example, input the desensitized image data and the image data to be desensitized into the image matching degree determination algorithm, and the image matching degree determination algorithm outputs the image matching degree.

[0112] For the image matching degree determination algorithm, a deep learning-based image feature matching algorithm such as the SIFT (Scale Invariant Feature Transform) algorithm can be adopted, or other types of algorithms can also be adopted, which is not limited here. Take the SIFT algorithm as an example. Based on the SIFT algorithm, input the image feature 1 of the desensitized image data and the image feature 2 of the image data to be desensitized into the SIFT algorithm. The SIFT algorithm matches the desensitized image data and the image data to be desensitized based on the image feature 1 and the image feature 2, and obtains the image matching degree between the desensitized image data and the image data to be desensitized.

[0113] Step S32: Determine the audio matching degree between the desensitized audio data and the audio data to be desensitized.

[0114] Exemplarily, the audio matching degree represents the similarity between the de-identified audio data and the audio data to be de-identified. The more similar the de-identified audio data and the audio data to be de-identified are, the greater the audio matching degree. Based on this, an audio matching degree determination algorithm is used to determine the audio matching degree between the de-identified audio data and the audio data to be de-identified. For example, the de-identified audio data and the audio data to be de-identified are input into the audio matching degree determination algorithm, and the audio matching degree is output by the audio matching degree determination algorithm.

[0115] For the audio matching degree determination algorithm, an image feature matching algorithm based on deep learning, such as the SIFT algorithm, can also be used. That is to say, the determination of the audio matching degree is achieved by using the SIFT algorithm. Other types of algorithms can also be used, and there is no limitation in this regard. Taking the SIFT algorithm as an example, based on the SIFT algorithm, the de-identified audio data and the audio data to be de-identified can be matched to obtain the audio matching degree.

[0116] In order to determine the audio matching degree using the SIFT algorithm, the audio feature 1 of the de-identified audio data and the audio feature 2 of the audio data to be de-identified can be determined. The audio feature 1 is mapped to a vector space with the same dimension as the image feature descriptor to obtain the audio feature 3, and the audio feature 2 is mapped to a vector space with the same dimension as the image feature descriptor to obtain the audio feature 4. For example, if the input feature of the SIFT algorithm is a feature with dimensions A * B * C (i.e., the vector space of the image feature descriptor is a feature with dimensions A * B * C), A can represent the feature width, B can represent the feature height, and C can represent the number of feature channels. Then, the audio feature 1 can be mapped to the audio feature 3 with dimensions A * B * C, and the audio feature 2 can be mapped to the audio feature 4 with dimensions A * B * C. On this basis, the audio feature 3 and the audio feature 4 can be input into the SIFT algorithm, and the SIFT algorithm matches the de-identified audio data and the audio data to be de-identified based on the audio feature 3 and the audio feature 4 to obtain the audio matching degree between the de-identified audio data and the audio data to be de-identified.

[0117] For example, regarding how to map the audio feature 1 (audio feature 2) to the audio feature 3 (audio feature 4), a feature mapping method based on deep learning can be used. This feature mapping method is used to map a feature of one dimension to a feature of another dimension, and there is no limitation on this feature mapping method.

[0118] Step S33: Determine the text matching degree between the de-identified text data and the text data to be de-identified.

[0119] Exemplarily, the text matching degree represents the similarity between the de-sensitized text data and the text data to be de-sensitized. The more similar the de-sensitized text data is to the text data to be de-sensitized, the greater the text matching degree. Based on this, a text matching degree determination algorithm can be used to determine the text matching degree between the de-sensitized text data and the text data to be de-sensitized. For example, the de-sensitized text data and the text data to be de-sensitized can be input into the text matching degree determination algorithm, and the text matching degree is output by the text matching degree determination algorithm.

[0120] For the text matching degree determination algorithm, an image feature matching algorithm based on deep learning, such as the SIFT algorithm, can also be used. That is to say, the determination of the text matching degree is achieved by using the SIFT algorithm. Other types of algorithms can also be used, which is not restricted here. Taking the SIFT algorithm as an example, based on the SIFT algorithm, the de-sensitized text data and the text data to be de-sensitized can be matched to obtain the text matching degree.

[0121] In order to use the SIFT algorithm to determine the text matching degree, the text feature 1 of the de-sensitized text data and the text feature 2 of the text data to be de-sensitized can be determined. Through word embedding and semantic encoding techniques, the text feature 1 is converted into a feature vector 1, and through linear transformation, the feature vector 1 is mapped to a vector space with the same dimension as the image feature descriptor (i.e., the SIFT feature descriptor) to obtain the text feature 3. Through word embedding and semantic encoding techniques, the text feature 2 is converted into a feature vector 2, and through linear transformation, the feature vector 2 is mapped to a vector space with the same dimension as the image feature descriptor to obtain the text feature 4. For example, if the input feature of the SIFT algorithm is a feature with dimensions A*B*C, then the text features 3 and 4 are of dimensions A*B*C. On this basis, the text features 3 and 4 can be input into the SIFT algorithm, and the SIFT algorithm matches the de-sensitized text data and the text data to be de-sensitized based on the text features 3 and 4 to obtain the text matching degree between the de-sensitized text data and the text data to be de-sensitized.

[0122] Step S34: Determine the feature matching degree based on the image matching degree, audio matching degree, and text matching degree. For example, perform a weighted operation on the image matching degree, audio matching degree, and text matching degree to obtain the feature matching degree.

[0123] For example, the following formula (1) can be used to determine the feature matching degree:

[0124] Formula (1)

[0125] In formula (1), SIFT new can represent the feature matching degree, SIFT imageIt can represent the image matching degree. Here, an example is given where the weighting coefficient of the image matching degree is 1. The weighting coefficient of the image matching degree can also be other values, which can be configured according to experience and are not restricted herein. SIFT audio It can represent the audio matching degree. The weighting coefficient of the audio matching degree is , and the weighting coefficient can be configured according to experience or determined according to the importance and relevance of the audio matching degree, and is not restricted herein. SIFT text It can represent the text matching degree. The weighting coefficient of the text matching degree is , and the weighting coefficient can be configured according to experience or determined according to the importance and relevance of the audio matching degree, and is not restricted herein.

[0126] It can be seen from formula (1) that on the basis of focusing on the local features of the image (image matching degree), the feature matching degree can also focus on audio features (audio matching degree) and text features (text matching degree). It can comprehensively consider the impact of desensitization on key information from multiple dimensions and integrate multi-modal information into the SIFT feature descriptor (feature matching degree), that is, the SIFT feature descriptor comprehensively considers the local features of the image, audio description, and text information. By using the feature matching degree as the desensitization effect parameter, the feature matching degree can be used as the degree of retention of key information, and the retention and removal effects of key information can be used as the desensitization effect evaluation index.

[0127] The weighting coefficient and the weighting coefficient are weights determined according to the importance and relevance of the modality. Through weighted operations on the image matching degree, audio matching degree, and text matching degree, the integration of multi-modal information at the feature descriptor level is realized, so that the feature descriptor not only contains the local information of the image, but also contains audio information and text information related to the image, and more comprehensively reflects the key information.

[0128] Step S35: Determine the image quality based on the desensitized image data and the image data to be desensitized.

[0129] Exemplarily, the image quality represents the image quality of the desensitized image data. The more similar the desensitized image data is to the image data to be desensitized, the better the image quality of the desensitized image data. Based on this, an image quality evaluation algorithm can be used to determine the image quality of the desensitized image data. For example, the desensitized image data and the image data to be desensitized can be input to the image quality evaluation algorithm, and the image quality evaluation algorithm determines the image quality based on the desensitized image data and the image data to be desensitized and outputs the image quality.

[0130] For the image quality assessment algorithm, the SSIM (Structural Similarity) algorithm can be adopted. The SSIM algorithm is an index for measuring the similarity between two images, and it can compare images in terms of brightness, contrast, and structural information, etc. Of course, other types of algorithms can also be used for the image quality assessment algorithm, and there is no restriction on this. Hereinafter, the SSIM algorithm will be taken as an example for illustration.

[0131] For example, the following formula (2) can be used to determine the image quality of the desensitized image data:

[0132] Formula (2)

[0133] In formula (2), SSIM improved represents the image quality of the desensitized image data, x represents the image data to be desensitized, y represents the desensitized image data, fps represents the frame rate of the video to be desensitized, that is, the image quality can be determined based on the image data to be desensitized x , the desensitized image data y and the frame rate fps to determine the image quality SSIM improved .

[0134] (fps) represents the target brightness weight, that is, the frame rate fps corresponding target brightness weight, represents the brightness parameter value, and the product value of the target brightness weight and the brightness parameter value represents the first similarity value. Among the brightness parameter values, μ x represents the average pixel value of the image data to be desensitized (that is, the average of the pixel values of all pixel points), μ y represents the average pixel value of the desensitized image data, c 1 represents the configured constant value, which is used to avoid the numerator or denominator of the brightness parameter value being 0. Obviously, the brightness parameter value is based on the average pixel value of the desensitized image data μ y and the average pixel value of the image data to be desensitized μ x to determine.

[0135] (fps) represents the frame rate fps corresponding target contrast weight, represents the contrast parameter value, and the product value of the target contrast weight and the contrast parameter value represents the second similarity value. Among the contrast parameter values, x represents the variance of the pixel values of the image data to be desensitized (i.e., the variance of the pixel values of all pixel points), y represents the variance of the pixel values of the desensitized image data, c 2 represents the configured constant value, which is used to avoid the numerator or denominator of the contrast parameter value being 0. Obviously, the contrast parameter value is determined based on the variance of the pixel values of the desensitized image data and the variance of the pixel values of the image data to be desensitized.

[0136] (fps) represents the frame rate fps the corresponding target structure weight, represents the structure parameter value, and the product value of the target structure weight and the structure parameter value represents the third similarity value. Among the structure parameter values, x represents the variance of the pixel values of the image data to be desensitized, y represents the variance of the pixel values of the desensitized image data, xy represents the covariance of the pixel values of the desensitized image data and the image data to be desensitized, c 2 represents the configured constant value. Obviously, the structure parameter value is determined based on the variance of the pixel values of the desensitized image data, the variance of the pixel values of the image data to be desensitized, and the covariance of the pixel values of the desensitized image data and the image data to be desensitized.

[0137] It can be seen from formula (2) that the image quality can be determined based on the first similarity value, the second similarity value, and the third similarity value. For example, the image quality is the sum value of the first similarity value, the second similarity value, and the third similarity value.

[0138] In a possible implementation manner, in order to obtain accurate target brightness weight, target contrast weight, and target structure weight, the video to be desensitized can be divided into different categories according to the frame rate of the video to be desensitized. For example, the video to be desensitized with a frame rate less than the first value (such as 15fps) is defined as a low-frame-rate video, the video to be desensitized with a frame rate greater than the second value (such as 30fps) is defined as a high-frame-rate video, and the video to be desensitized with a frame rate not less than the first value and not greater than the second value is defined as a medium-frame-rate video. For videos with different frame rates, initial weights are set for each component (brightness, contrast, structure) of the SSIM algorithm.

[0139] High-frame-rate videos can give higher weights in terms of structural similarity due to containing more dynamic detail information to highlight the consideration of image details and dynamic changes. Medium-frame-rate videos are relatively balanced, and low-frame-rate videos pay more attention to the overall perceived quality. On this basis, the initial weights can be set in the following ways:

[0140] For the initial brightness weight, if the frame rate is less than the first value, the initial brightness weight is the first brightness weight value, and the first brightness weight value can be configured according to experience, such as 0.4, etc. If the frame rate is greater than the second value, the initial brightness weight is the second brightness weight value, and the second brightness weight value can be configured according to experience, such as 0.2, etc., and the second brightness weight value needs to be less than the first brightness weight value. If the frame rate is not less than the first value and not greater than the second value, the initial brightness weight is between the second brightness weight value and the first brightness weight value, and the initial brightness weight is inversely proportional to the frame rate, such as determining the initial brightness weight based on the second brightness weight value, the first brightness weight value, and the frame rate, as long as the initial brightness weight is between the second brightness weight value and the first brightness weight value. See the example shown in formula (3) for determining the initial brightness weight (fps) of the example.

[0141] Formula (3)

[0142] For the initial contrast weight, if the frame rate is less than the third value (the third value can be the same as or different from the first value, taking 15 as an example), the initial contrast weight is the first contrast weight value, and the first contrast weight value can be configured according to experience, such as 0.4, etc. If the frame rate is greater than the fourth value (the fourth value can be the same as or different from the second value, taking 30 as an example), the initial contrast weight is the second contrast weight value, and the second contrast weight value can be configured according to experience, such as 0.2, etc., and the second contrast weight value needs to be less than the first contrast weight value. If the frame rate is not less than the third value and not greater than the fourth value, the initial contrast weight is between the second contrast weight value and the first contrast weight value, and the initial contrast weight is inversely proportional to the frame rate, such as determining the initial contrast weight based on the second contrast weight value, the first contrast weight value, and the frame rate, as long as the initial contrast weight is between the second contrast weight value and the first contrast weight value. See the example shown in formula (4) for determining the initial contrast weight (fps) of the example.

[0143] Formula (4)

[0144] For the initial structural weight, if the frame rate is less than the fifth value (the fifth value may be the same as or different from the first value, taking 15 as an example), the initial structural weight is the first structural weight value, and the first structural weight value is configured according to experience, such as 0.2, etc. If the frame rate is greater than the sixth value (the sixth value may be the same as or different from the second value, taking 30 as an example), the initial structural weight is the second structural weight value, and the second structural weight value is configured according to experience, such as 0.6, etc., and the second structural weight value needs to be greater than the first structural weight value. If the frame rate is not less than the fifth value and not greater than the sixth value, the initial structural weight is between the first structural weight value and the second structural weight value, and the initial structural weight is proportional to the frame rate. For example, the initial structural weight is determined based on the second structural weight value, the first structural weight value, and the frame rate, as long as the initial structural weight is between the first structural weight value and the second structural weight value. Refer to formula (5) as shown for determining the initial structural weight (fps) for an example.

[0145] Formula (5)

[0146] After obtaining the initial brightness weight, the initial brightness weight can be determined as the target brightness weight. After obtaining the initial contrast weight, the initial contrast weight can be determined as the target contrast weight. After obtaining the initial structural weight, the initial structural weight can be determined as the target structural weight. In summary, the target brightness weight, the target contrast weight, and the target structural weight can be dynamically adjusted according to the frame rate of the video to be desensitized, so as to more accurately measure the image quality of the desensitized image data.

[0147] In a possible implementation, after obtaining the initial brightness weight, the initial contrast weight, and the initial structural weight, the initial brightness weight, the initial contrast weight, and the initial structural weight can also be dynamically adjusted based on the motion degree of the video to be desensitized, so as to obtain more accurate target brightness weight, target contrast weight, and target structural weight, and improve the accuracy of image quality evaluation based on the motion state.

[0148] Exemplarily, the initial brightness weight is adjusted based on the motion degree of the video to be desensitized, and the adjusted brightness weight is determined as the target brightness weight. The initial contrast weight is adjusted based on the motion degree of the video to be desensitized, and the adjusted contrast weight is determined as the target contrast weight. The initial structural weight is adjusted based on the motion degree of the video to be desensitized, and the adjusted structural weight is determined as the target structural weight.

[0149] For the motion degree of the video to be desensitized, the motion degree is determined based on the motion vectors between adjacent frames in the video to be desensitized. For example, the optical flow method (such as the Lucas-Kanade optical flow method or the Farneback optical flow method) can be used to calculate the motion vectors between adjacent frames of the video to be desensitized. For example, a reference frame (the reference frame can be one frame or multiple frames) can be selected from the video to be desensitized. For each pixel point of the reference frame, the motion vector of this pixel point (i.e., the displacement vector of this pixel point between the reference frame and the adjacent frame) is calculated through the optical flow method, and the motion vector of this pixel point can reflect the motion condition of this pixel point. By calculating the average value or variance of the motion vectors of all pixel points, the motion degree of the video to be desensitized is quantified.

[0150] For example, the following formula (6) can be used to determine the motion degree of the video to be desensitized:

[0151] Formula (6)

[0152] In formula (6), Motion intensity represents the motion degree of the video to be desensitized, N represents the total number of pixel points in the reference frame, represents the motion vector of the i-th pixel point, and the value range of i is from 1 to N. In this way, the average value of the motion vectors of all pixel points can be used as a measure of the motion degree of the video to be desensitized.

[0153] If the motion degree of the video to be desensitized is less than the first motion threshold (the first motion threshold can be configured according to experience, and there is no limit to this), it means that under the condition of low motion intensity, the detail and structure information of the image are more important. Based on this, the initial brightness weight can be increased, and the increased brightness weight is used as the target brightness weight. The initial contrast weight can be kept unchanged, that is, the initial contrast weight is used as the target contrast weight. The initial structure weight can be kept unchanged, that is, the initial structure weight is used as the target structure weight.

[0154] If the motion degree of the video to be desensitized is greater than the second motion threshold (the second motion threshold can be configured according to experience, and there is no limit to this), it means that under the condition of high motion intensity, more attention is paid to the overall coherence and smoothness of the image. Based on this, the initial brightness weight can be kept unchanged, that is, the initial brightness weight is used as the target brightness weight. The initial contrast weight can be increased, and the increased contrast weight is used as the target contrast weight. The initial structure weight can be increased, and the increased structure weight is used as the target structure weight.

[0155] If the motion degree of the video to be desensitized is not less than the first motion threshold and not greater than the second motion threshold, it indicates that it is in the medium motion intensity, and more attention should be paid to maintaining relative balance. Based on this, the initial brightness weight, initial contrast weight, and initial structure weight can be adjusted so that the difference between the adjusted brightness weight, adjusted contrast weight, and adjusted structure weight is not greater than the threshold. The adjusted brightness weight can be used as the target brightness weight, the adjusted contrast weight can be used as the target contrast weight, and the adjusted structure weight can be used as the target structure weight. Among them, when the difference between the adjusted brightness weight and the adjusted contrast weight is not greater than the threshold, the adjusted brightness weight and the adjusted contrast weight are the same or approximately the same. When the difference between the adjusted brightness weight and the adjusted structure weight is not greater than the threshold, the adjusted brightness weight and the adjusted structure weight are the same or approximately the same. When the difference between the adjusted contrast weight and the adjusted structure weight is not greater than the threshold, the adjusted contrast weight and the adjusted structure weight are the same or approximately the same.

[0156] In summary, the target brightness weight, target contrast weight, and target structure weight can be dynamically adjusted according to the motion degree of the video to be desensitized, and the image quality of the desensitized image data can be measured more accurately. By improving the SSIM algorithm through the inter-frame weighting method based on the frame rate and motion degree, the image quality evaluation can more accurately adapt to the dynamic changes of video data, further improving the video data desensitization effect.

[0157] In a possible implementation, during the process of desensitizing video data based on the desensitization strategy, the effect of the desensitization operation can be monitored in real time. For example, the hash values of the original data and the desensitized data are calculated using a hash function, and the difference in the hash values is used as an integrity index. Through a combination of manual annotation and automatic evaluation, the accuracy of the desensitized data is evaluated, such as checking whether sensitive information is completely desensitized, and using timestamps and operation logs to ensure the consistency of the desensitization process. According to the evaluation results, a reinforcement learning algorithm (such as Q-learning) is used to adjust the desensitization strategy to ensure the stable and reliable operation of the desensitization process.

[0158] Step S36: Obtain desensitization effect parameters, where the desensitization effect parameters include feature matching degree and image quality.

[0159] So far, step 204 is completed, and desensitization effect parameters can be obtained.

[0160] Step 205: Determine whether the desensitization effect parameters have met the desensitization quality conditions.

[0161] If so, that is, the desensitization effect parameters have met the desensitization quality conditions, then step 206 is executed.

[0162] Otherwise, if the desensitization effect parameter does not meet the desensitization quality condition, then step 207 is executed.

[0163] Exemplarily, the desensitization effect parameter may include a feature matching degree and an image quality. If the feature matching degree is greater than a first threshold and the image quality is greater than a second threshold, then the desensitization effect parameter has met the desensitization quality condition. If the feature matching degree is not greater than the first threshold, and / or the image quality is not greater than the second threshold, then the desensitization effect parameter does not meet the desensitization quality condition. For example, the first threshold can be configured according to experience and is used to represent the threshold of the feature matching degree, such as 0.8, 0.85, etc., without limitation. The second threshold can be configured according to experience and is used to represent the threshold of the image quality, such as 0.8, 0.85, etc., without limitation.

[0164] Step 206: Generate a desensitized video corresponding to the video to be desensitized based on the desensitized data, and send the desensitized video, that is, send the desensitized video to the receiving end of the video to be desensitized.

[0165] For example, a desensitized video corresponding to the video to be desensitized can be generated based on the desensitized image data, desensitized audio data, and desensitized text data. There is no limitation on the generation method of this desensitized video. The desensitized video may include desensitized image data, desensitized audio data, and desensitized text data.

[0166] Step 207: Adjust the target desensitization strategy to obtain an adjusted desensitization strategy, use the adjusted desensitization strategy as the target desensitization strategy, and return to execute step 203, that is, perform desensitization processing on the target data to be desensitized based on the target desensitization strategy (i.e., the adjusted desensitization strategy) to obtain desensitized data.

[0167] For example, if the target desensitization strategy includes an image desensitization strategy, an audio desensitization strategy, and a text desensitization strategy, then simultaneously adjust the image desensitization strategy, the audio desensitization strategy, and the text desensitization strategy to obtain an adjusted image desensitization strategy, an adjusted audio desensitization strategy, and an adjusted text desensitization strategy. Or, adjust the image desensitization strategy to obtain an adjusted image desensitization strategy, and keep the audio desensitization strategy and the text desensitization strategy unchanged. Or, adjust the audio desensitization strategy to obtain an adjusted audio desensitization strategy, and keep the image desensitization strategy and the text desensitization strategy unchanged. Or, adjust the text desensitization strategy to obtain an adjusted text desensitization strategy, and keep the image desensitization strategy and the audio desensitization strategy unchanged.

[0168] In a possible implementation, the adjustment priority of the image desensitization strategy is higher than that of the audio desensitization strategy, and the adjustment priority of the audio desensitization strategy is higher than that of the text desensitization strategy.

[0169] During the first desensitization strategy adjustment process, adjust the image desensitization strategy to obtain the adjusted image desensitization strategy, and keep the audio desensitization strategy and the text desensitization strategy unchanged. Then, repeat steps 203 - 205. If the desensitization effect parameter has met the desensitization quality condition, end the adjustment process. If the desensitization effect parameter does not meet the desensitization quality condition, perform the second desensitization strategy adjustment process.

[0170] During the second desensitization strategy adjustment process, adjust the audio desensitization strategy to obtain the adjusted audio desensitization strategy, and keep the image desensitization strategy and the text desensitization strategy unchanged. Repeat steps 203 - 205. If the desensitization effect parameter does not meet the desensitization quality condition, perform the third desensitization strategy adjustment process.

[0171] During the third desensitization strategy adjustment process, adjust the text desensitization strategy to obtain the adjusted text desensitization strategy, and keep the image desensitization strategy and the audio desensitization strategy unchanged. Repeat steps 203 - 205. If the desensitization effect parameter does not meet the desensitization quality condition, perform the fourth desensitization strategy adjustment process.

[0172] During the fourth desensitization strategy adjustment process, adjust the image desensitization strategy, and so on.

[0173] Exemplarily, when adjusting the image desensitization strategy to obtain the adjusted image desensitization strategy, if the image desensitization strategy is an image desensitization network model, the network parameters in the image desensitization network model can be adjusted. Of course, this is just an example of adjustment, as long as the image desensitization strategy can be adjusted.

[0174] For example, when using an image desensitization network model (StyleGAN model) to desensitize the image sensitive positions of the image data to be desensitized, the desensitization effect depends on the network parameters in the image desensitization network model. Therefore, by adjusting the network parameters in the image desensitization network model, the desensitization effect can be optimized so that the desensitization effect parameter can meet the desensitization quality condition. In addition, regarding the adjustment method of the network parameters, no limitation is made in this embodiment, as long as the network parameters can be adjusted.

[0175] Exemplarily, when adjusting the audio desensitization strategy to obtain the adjusted audio desensitization strategy, if the audio desensitization strategy includes an audio desensitization network model (such as Tacotron model) and a vocoder (such as WaveNet vocoder), the network parameters in the audio desensitization network model can be adjusted. For example, when using an audio desensitization network model to desensitize the audio sensitive positions of the audio data to be desensitized, the desensitization effect depends on the network parameters in the audio desensitization network model. By adjusting the network parameters in the audio desensitization network model, the desensitization effect can be optimized so that the desensitization effect parameter meets the desensitization quality condition.

[0176] Exemplarily, when adjusting the text desensitization strategy to obtain the adjusted text desensitization strategy, if the text desensitization strategy is a text encryption algorithm, the algorithm type adopted by the text encryption algorithm can be adjusted. For example, if the algorithm type of the text encryption algorithm is the block encryption mode of the AES-256 encryption algorithm, it can be adjusted to the padding mode of the AES-256 encryption algorithm or the DES encryption algorithm. Or, if the text desensitization strategy is a text replacement algorithm, the algorithm type adopted by the text replacement algorithm can be adjusted.

[0177] As can be seen from the above technical solutions, in the embodiments of the present application, a video data desensitization method driven by a multimodal large model is provided. The multimodal large model is used to perform intelligent desensitization on video data, giving full play to the advantages of the multimodal large model. When the desensitization effect parameter does not meet the desensitization quality condition, the target desensitization strategy can be adjusted until the desensitization effect parameter meets the desensitization quality condition, so that the quality of the desensitized video is good, the desensitization effect is good, and the original information of the video data can be reflected. It is possible to evaluate the desensitization effect of the desensitized video, adaptively adjust the desensitization strategy to improve the desensitization effect, and make the desensitized video usable. While improving the data security and usability of the desensitized video, the image quality assessment can be more accurately adapted to the dynamic changes of the video data. During the process of the multimodal large model processing video data, the integrity and usability of the desensitized data are verified. Through the multimodal large model, the video data before and after desensitization is compared and analyzed, and the desensitization operation is evaluated from multiple modalities (such as image quality, key information retention degree, etc., and the key information retention degree can be the above-mentioned feature matching degree), and the desensitization strategy can be automatically adjusted to achieve the best balance between security and usability.

[0178] Based on the same application concept as the above method, in the embodiments of the present application, a data desensitization and video data protection device driven by a multimodal large model is proposed. Refer to Figure 4 As shown, it is a schematic structural diagram of the device, and the device may include:

[0179] A determination module 41 is configured to determine target data to be desensitized based on the acquired video to be desensitized, and determine a target desensitization strategy corresponding to the target data to be desensitized. A processing module 42 is configured to desensitize the target data to be desensitized through a multi-modal large model based on the target desensitization strategy to obtain desensitized data. The determination module 41 is further configured to determine a desensitization effect parameter based on the target data to be desensitized and the desensitized data, where the desensitization effect parameter is used to reflect the desensitization quality of the desensitized data. A generation module 43 is configured to, if the desensitization effect parameter meets the desensitization quality condition, generate a desensitized video corresponding to the video to be desensitized based on the desensitized data, and send the desensitized video. The processing module 42 is further configured to, if the desensitization effect parameter does not meet the desensitization quality condition, adjust the target desensitization strategy to obtain an adjusted desensitization strategy, use the adjusted desensitization strategy as the target desensitization strategy, and desensitize the target data to be desensitized based on the target desensitization strategy to obtain desensitized data.

[0180] Exemplarily, the target data to be desensitized includes at least two of image data to be desensitized, audio data to be desensitized, and text data to be desensitized. If the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, the desensitized data includes desensitized image data, desensitized audio data, and desensitized text data. When the determination module 41 determines the desensitization effect parameter based on the target data to be desensitized and the desensitized data, it is specifically configured to: determine a feature matching degree based on the image matching degree between the desensitized image data and the image data to be desensitized, the audio matching degree between the desensitized audio data and the audio data to be desensitized, and the text matching degree between the desensitized text data and the text data to be desensitized; determine the image quality based on the desensitized image data and the image data to be desensitized; obtain a desensitization effect parameter, where the desensitization effect parameter includes the feature matching degree and the image quality; if the feature matching degree is greater than a first threshold and the image quality is greater than a second threshold, the desensitization effect parameter meets the desensitization quality condition; if the feature matching degree is not greater than the first threshold, and / or the image quality is not greater than the second threshold, the desensitization effect parameter does not meet the desensitization quality condition.

[0181] Exemplarily, when the determining module 41 determines the image quality based on the desensitized image data and the to-be-desensitized image data, it is specifically configured to: determine a first similarity value based on a target brightness weight and a brightness parameter value, determine a second similarity value based on a target contrast weight and a contrast parameter value, and determine a third similarity value based on a target structure weight and a structure parameter value; the brightness parameter value is determined based on the mean pixel value of the desensitized image data and the mean pixel value of the to-be-desensitized image data, the contrast parameter value is determined based on the pixel value variance of the desensitized image data and the pixel value variance of the to-be-desensitized image data, and the structure parameter value is determined based on the pixel value variance of the desensitized image data, the pixel value variance of the to-be-desensitized image data, and the covariance of the pixel values of the desensitized image data and the to-be-desensitized image data; determine the image quality based on the first similarity value, the second similarity value, and the third similarity value.

[0182] Exemplarily, when the determining module 41 determines the target brightness weight, it is specifically configured to: determine an initial brightness weight corresponding to the frame rate of the to-be-desensitized video; determine the target brightness weight based on the initial brightness weight; wherein, if the frame rate is less than a first value, the initial brightness weight is a first brightness weight value; if the frame rate is greater than a second value, the initial brightness weight is a second brightness weight value; if the frame rate is not less than the first value and not greater than the second value, the initial brightness weight is between the second brightness weight value and the first brightness weight value, and the initial brightness weight is inversely proportional to the frame rate.

[0183] Exemplarily, when the determining module 41 determines the target contrast weight, it is specifically configured to: determine an initial contrast weight corresponding to the frame rate; determine the target contrast weight based on the initial contrast weight; wherein, if the frame rate is less than a third value, the initial contrast weight is a first contrast weight value; if the frame rate is greater than a fourth value, the initial contrast weight is a second contrast weight value; if the frame rate is not less than the third value and not greater than the fourth value, the initial contrast weight is between the second contrast weight value and the first contrast weight value, and the initial contrast weight is inversely proportional to the frame rate.

[0184] Exemplarily, when the determining module 41 determines the target structure weight, it is specifically configured to: determine an initial structure weight corresponding to the frame rate; determine the target structure weight based on the initial structure weight; wherein, if the frame rate is less than a fifth value, the initial structure weight is a first structure weight value; if the frame rate is greater than a sixth value, the initial structure weight is a second structure weight value; if the frame rate is not less than the fifth value and not greater than the sixth value, the initial structure weight is between the first structure weight value and the second structure weight value, and the initial structure weight is directly proportional to the frame rate.

[0185] Exemplarily, when determining the target brightness weight based on the initial brightness weight, the determining module 41 is specifically configured to: determine the initial brightness weight as the target brightness weight; or, adjust the initial brightness weight based on the motion degree of the video to be desensitized, and determine the adjusted brightness weight as the target brightness weight; wherein, the motion degree is determined based on the motion vector between adjacent frames in the video to be desensitized; if the motion degree is less than the first motion threshold, increase the initial brightness weight, and if the motion degree is greater than the second motion threshold, keep the initial brightness weight unchanged.

[0186] Exemplarily, when determining the target contrast weight based on the initial contrast weight, the determining module 41 is specifically configured to: determine the initial contrast weight as the target contrast weight; or, adjust the initial contrast weight based on the motion degree, and determine the adjusted contrast weight as the target contrast weight; if the motion degree is less than the first motion threshold, keep the initial contrast weight unchanged; if the motion degree is greater than the second motion threshold, increase the initial contrast weight.

[0187] Exemplarily, when determining the target structure weight based on the initial structure weight, the determining module 41 is specifically configured to: determine the initial structure weight as the target structure weight; or, adjust the initial structure weight based on the motion degree, and determine the adjusted structure weight as the target structure weight; wherein, if the motion degree is less than the first motion threshold, keep the initial structure weight unchanged; if the motion degree is greater than the second motion threshold, increase the initial structure weight.

[0188] Exemplarily, if the motion degree is not less than the first motion threshold and not greater than the second motion threshold, adjust the initial brightness weight, the initial contrast weight, and the initial structure weight so that the difference between the adjusted brightness weight, the adjusted contrast weight, and the adjusted structure weight is not greater than the threshold.

[0189] Exemplarily, if the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, when the processing module 42 performs desensitization processing on the target data to be desensitized based on the target desensitization strategy to obtain desensitized data, it is specifically used for: extracting image features from the image data to be desensitized through a multimodal large model, extracting audio features from the audio data to be desensitized through a multimodal large model, extracting text features from the text data to be desensitized through a multimodal large model, and performing feature fusion on the image features, the audio features, and the text features through a multimodal large model to obtain multimodal features; determining the image sensitive positions of the image data to be desensitized, the audio sensitive positions of the audio data to be desensitized, and the text sensitive positions of the text data to be desensitized based on the multimodal features through a multimodal large model; performing desensitization processing on the image sensitive positions of the image data to be desensitized based on the image desensitization strategy corresponding to the image data to be desensitized to obtain desensitized image data; performing desensitization processing on the audio sensitive positions of the audio data to be desensitized based on the audio desensitization strategy corresponding to the audio data to be desensitized to obtain desensitized audio data; and performing desensitization processing on the text sensitive positions of the text data to be desensitized based on the text desensitization strategy corresponding to the text data to be desensitized to obtain desensitized text data.

[0190] Exemplarily, if the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, the target desensitization strategy includes an image desensitization strategy, an audio desensitization strategy, and a text desensitization strategy; when the determination module 41 determines the target desensitization strategy corresponding to the target data to be desensitized, it is specifically used for: querying the configured desensitization rule library through the multimodal attributes corresponding to the video to be desensitized to obtain the image desensitization strategy corresponding to the multimodal attributes; wherein, the multimodal attributes include at least one of image attributes, audio attributes, and text attributes, and the desensitization rule library includes the corresponding relationship between the multimodal attributes and the image desensitization strategy; the image desensitization strategy includes an image desensitization network model.

[0191] Determining an audio desensitization strategy based on the audio attributes corresponding to the video to be desensitized, the audio desensitization strategy includes an audio desensitization network model and a vocoder, the audio desensitization network model is used to convert the audio data to be desensitized into acoustic features, and the vocoder is used to convert the acoustic features into audio waveforms;

[0192] Determining the configured text desensitization strategy, the text desensitization strategy includes a text encryption algorithm or a text replacement algorithm, the text encryption algorithm is used to perform an encryption operation on the text data to be desensitized, and the text replacement algorithm is used to perform a replacement operation or a blurring operation on the text data to be desensitized.

[0193] Exemplarily, when the processing module 42 adjusts the target desensitization policy to obtain an adjusted desensitization policy, if it is necessary to adjust the image desensitization policy, the network parameters in the image desensitization network model are adjusted; if it is necessary to adjust the audio desensitization policy, the network parameters in the audio desensitization network model are adjusted; if it is necessary to adjust the text desensitization policy, the algorithm type adopted by the text encryption algorithm or the algorithm type adopted by the text replacement algorithm is adjusted.

[0194] Based on the same application concept as the above method, an electronic device is proposed in an embodiment of the present application. Refer to Figure 5 As shown, the electronic device includes: a processor 51 and a machine-readable storage medium 52. The machine-readable storage medium 52 stores machine-executable instructions that can be executed by the processor 51; the processor 51 is used to execute the machine-executable instructions to implement the multi-modal large model-driven data desensitization and video data protection method disclosed in the above examples of the present application.

[0195] Based on the same application concept as the above method, an embodiment of the present application further provides a machine-readable storage medium. A number of computer instructions are stored on the machine-readable storage medium. When the computer instructions are executed by a processor, the multi-modal large model-driven data desensitization and video data protection method disclosed in the above examples of the present application can be implemented.

[0196] Among them, the above machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.

[0197] Based on the same application concept as the above method, an embodiment of the present application further provides a computer program product. The computer program product may include a computer program. When the computer program is executed by a processor, the multi-modal large model-driven data desensitization and video data protection method disclosed in the above examples of the present application is implemented.

[0198] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0199] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A multimodal large model driven data desensitization and video data protection method, characterized in that: The method comprises: Determine the target data to be desensitized based on the acquired videos to be desensitized; Determine a target desensitization strategy corresponding to the target data to be desensitized, and perform desensitization processing on the target data to be desensitized by using a multimodal large model based on the target desensitization strategy to obtain desensitized data; Determine a desensitization effect parameter based on the target data to be desensitized and the desensitized data; If the desensitization effect parameter satisfies the desensitization quality condition, a desensitized video corresponding to the video to be desensitized is generated based on the desensitized data, and the desensitized video is sent; If the desensitization effect parameter does not meet the desensitization quality condition, the target desensitization strategy is adjusted to obtain an adjusted desensitization strategy, the adjusted desensitization strategy is used as the target desensitization strategy, and the operation of performing desensitization processing on the target data to be desensitized based on the target desensitization strategy to obtain desensitized data is returned; Wherein, the target data to be desensitized includes at least two of image data to be desensitized, audio data to be desensitized, and text data to be desensitized; if the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, the desensitized data includes desensitized image data, desensitized audio data, and desensitized text data; the desensitization effect parameter is determined based on the target data to be desensitized and the desensitized data, including: Determine a feature matching degree based on an image matching degree between the desensitized image data and the image data to be desensitized, an audio matching degree between the desensitized audio data and the audio data to be desensitized, and a text matching degree between the desensitized text data and the text data to be desensitized; Determining image quality based on the desensitized image data and the image data to be desensitized; Obtaining a desensitization effect parameter, wherein the desensitization effect parameter includes the feature matching degree and the image quality; if the feature matching degree is greater than a first threshold, and the image quality is greater than a second threshold, the desensitization effect parameter has met the desensitization quality condition; if the feature matching degree is not greater than the first threshold, and / or the image quality is not greater than the second threshold, the desensitization effect parameter does not meet the desensitization quality condition; Wherein, determining the image quality based on the desensitized image data and the image data to be desensitized includes: A first similarity value is determined based on a target brightness weight and a brightness parameter value, a second similarity value is determined based on a target contrast weight and a contrast parameter value, and a third similarity value is determined based on a target structure weight and a structure parameter value; wherein the brightness parameter value is determined based on a mean pixel value of the desensitized image data and a mean pixel value of the image data to be desensitized, the contrast parameter value is determined based on a pixel value variance of the desensitized image data and a pixel value variance of the image data to be desensitized, and the structure parameter value is determined based on a pixel value variance of the desensitized image data, a pixel value variance of the image data to be desensitized, and a covariance of pixel values ​​of the desensitized image data and the image data to be desensitized; determining image quality based on the first similarity value, the second similarity value, and the third similarity value; The process of determining the target brightness weight includes: determining an initial brightness weight corresponding to the frame rate of the video to be desensitized; if the frame rate is less than a first value, the initial brightness weight is a first brightness weight value; if the frame rate is greater than a second value, the initial brightness weight is a second brightness weight value; if the frame rate is not less than the first value and not greater than the second value, the initial brightness weight is between the second brightness weight value and the first brightness weight value, and the initial brightness weight is inversely proportional to the frame rate; determining the target brightness weight based on the initial brightness weight; The determining the target brightness weight based on the initial brightness weight includes: adjusting the initial brightness weight based on the degree of motion of the video to be desensitized, and determining the adjusted brightness weight as the target brightness weight; wherein the degree of motion is determined based on the motion vector between adjacent frames in the video to be desensitized; wherein, if the degree of motion is less than a first motion threshold, the initial brightness weight is increased, and if the degree of motion is greater than a second motion threshold, the initial brightness weight remains unchanged.

2. The method according to claim 1, characterized in that: The process of determining the target contrast weight includes: determining an initial contrast weight corresponding to the frame rate; wherein, if the frame rate is less than a third value, the initial contrast weight is a first contrast weight value; if the frame rate is greater than a fourth value, the initial contrast weight is a second contrast weight value; if the frame rate is not less than the third value and not greater than the fourth value, the initial contrast weight is between the second contrast weight value and the first contrast weight value, and the initial contrast weight is inversely proportional to the frame rate; determining the target contrast weight based on the initial contrast weight; The process of determining the target structure weight includes: determining an initial structure weight corresponding to the frame rate; wherein, if the frame rate is less than a fifth value, the initial structure weight is a first structure weight value; if the frame rate is greater than a sixth value, the initial structure weight is a second structure weight value; if the frame rate is not less than the fifth value and not greater than the sixth value, the initial structure weight is between the first structure weight value and the second structure weight value, and the initial structure weight is proportional to the frame rate; and determining the target structure weight based on the initial structure weight.

3. The method according to claim 2, characterized in that The determining the target contrast weight based on the initial contrast weight comprises: The initial contrast weight is adjusted based on the motion degree, and the adjusted contrast weight is determined as the target contrast weight; wherein if the motion degree is less than a first motion threshold, the initial contrast weight remains unchanged; if the motion degree is greater than a second motion threshold, the initial contrast weight is increased; The step of determining the target structure weight based on the initial structure weight comprises: The initial structure weight is adjusted based on the movement degree, and the adjusted structure weight is determined as the target structure weight; wherein, if the movement degree is less than a first movement threshold, the initial structure weight remains unchanged; if the movement degree is greater than a second movement threshold, the initial structure weight is increased; Among them, if the degree of motion is not less than the first motion threshold and not greater than the second motion threshold, the initial brightness weight, the initial contrast weight and the initial structure weight are adjusted so that the difference between the adjusted brightness weight, the adjusted contrast weight and the adjusted structure weight is not greater than the threshold.

4. The method according to claim 1, characterized in that: If the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, desensitizing the target data to be desensitized by using a multimodal large model based on the target desensitization strategy to obtain desensitized data, including: Extracting features from the desensitized image data through a multimodal large model to obtain image features, extracting features from the desensitized audio data through a multimodal large model to obtain audio features, extracting features from the desensitized text data through a multimodal large model to obtain text features, and fusing the image features, the audio features, and the text features through a multimodal large model to obtain multimodal features; Determine, by means of a multimodal large model, an image sensitive position of the image data to be desensitized, an audio sensitive position of the audio data to be desensitized, and a text sensitive position of the text data to be desensitized based on the multimodal features; Based on the image desensitization strategy corresponding to the image data to be desensitized, the image sensitive positions of the image data to be desensitized are desensitized to obtain desensitized image data; based on the audio desensitization strategy corresponding to the audio data to be desensitized, the audio sensitive positions of the audio data to be desensitized are desensitized to obtain desensitized audio data; based on the text desensitization strategy corresponding to the text data to be desensitized, the text sensitive positions of the text data to be desensitized are desensitized to obtain desensitized text data.

5. The method according to claim 1, characterized in that If the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, the target desensitization strategy includes an image desensitization strategy, an audio desensitization strategy, and a text desensitization strategy; The determining of the target desensitization strategy corresponding to the target data to be desensitized includes: The configured desensitization rule base is queried through the multimodal attribute corresponding to the video to be desensitized, and the image desensitization strategy corresponding to the multimodal attribute is obtained; wherein the multimodal attribute includes at least one of an image attribute, an audio attribute, and a text attribute, and the desensitization rule base includes a correspondence between the multimodal attribute and the image desensitization strategy; wherein the image desensitization strategy includes an image desensitization network model; Determining an audio desensitization strategy based on the audio attributes corresponding to the video to be desensitized, the audio desensitization strategy comprising an audio desensitization network model and a vocoder, the audio desensitization network model is used to convert the audio data to be desensitized into acoustic features, and the vocoder is used to convert the acoustic features into an audio waveform; Determine a configured text desensitization strategy, wherein the text desensitization strategy includes a text encryption algorithm or a text replacement algorithm, wherein the text encryption algorithm is used to perform an encryption operation on the desensitized text data, and the text replacement algorithm is used to perform a replacement operation or a fuzzification operation on the desensitized text data; Among them, when the target desensitization strategy is adjusted to obtain the adjusted desensitization strategy, if the image desensitization strategy needs to be adjusted, the network parameters in the image desensitization network model are adjusted; if the audio desensitization strategy needs to be adjusted, the network parameters in the audio desensitization network model are adjusted; if the text desensitization strategy needs to be adjusted, the algorithm type adopted by the text encryption algorithm is adjusted, or the algorithm type adopted by the text replacement algorithm is adjusted.

6. A multi-modal large model driven data desensitization and video data protection device, characterized in that: The device comprises: A determination module, used to determine the target data to be desensitized based on the acquired video to be desensitized; determine the target desensitization strategy corresponding to the target data to be desensitized; a processing module, used to perform desensitization processing on the target data to be desensitized through a multimodal large model based on the target desensitization strategy to obtain desensitized data; The determination module is further used to determine a desensitization effect parameter based on the target data to be desensitized and the desensitized data, wherein the desensitization effect parameter is used to reflect the desensitization quality of the desensitized data; A generating module, configured to generate a desensitized video corresponding to the video to be desensitized based on the desensitized data if the desensitization effect parameter meets the desensitization quality condition, and send the desensitized video; The processing module is used for adjusting the target desensitization strategy to obtain an adjusted desensitization strategy if the desensitization effect parameter does not meet the desensitization quality condition, taking the adjusted desensitization strategy as the target desensitization strategy, and performing desensitization processing on the target data to be desensitized based on the target desensitization strategy to obtain desensitized data; Among them, the target data to be desensitized includes at least two of the image data to be desensitized, the audio data to be desensitized and the text data to be desensitized; if the target data to be desensitized includes the image data to be desensitized, the audio data to be desensitized and the text data to be desensitized, the desensitized data includes the desensitized image data, the desensitized audio data and the desensitized text data; when the determination module determines the desensitization effect parameters based on the target data to be desensitized and the desensitized data, it is specifically used to: based on the image matching degree between the desensitized image data and the image data to be desensitized, the image matching degree between the desensitized audio data and the audio data to be desensitized The audio matching degree of the desensitized text data and the text matching degree between the desensitized text data and the text data to be desensitized are used to determine the feature matching degree; the image quality is determined based on the desensitized image data and the image data to be desensitized; a desensitization effect parameter is obtained, and the desensitization effect parameter includes the feature matching degree and the image quality; if the feature matching degree is greater than a first threshold value and the image quality is greater than a second threshold value, the desensitization effect parameter has met the desensitization quality condition; if the feature matching degree is not greater than the first threshold value and / or the image quality is not greater than the second threshold value, the desensitization effect parameter does not meet the desensitization quality condition; Wherein, when the determination module determines the image quality based on the desensitized image data and the image data to be desensitized, it is specifically used to: determine a first similarity value based on a target brightness weight and a brightness parameter value, determine a second similarity value based on a target contrast weight and a contrast parameter value, and determine a third similarity value based on a target structure weight and a structure parameter value; the brightness parameter value is determined based on the mean pixel value of the desensitized image data and the mean pixel value of the image data to be desensitized, the contrast parameter value is determined based on the pixel value variance of the desensitized image data and the pixel value variance of the image data to be desensitized, and the structure parameter value is determined based on the pixel value variance of the desensitized image data, the pixel value variance of the image data to be desensitized, and the covariance of the pixel values ​​of the desensitized image data and the image data to be desensitized; the image quality is determined based on the first similarity value, the second similarity value, and the third similarity value; Wherein, when the determination module determines the target brightness weight, it is specifically used to: determine the initial brightness weight corresponding to the frame rate of the video to be desensitized; determine the target brightness weight based on the initial brightness weight; wherein, if the frame rate is less than a first value, the initial brightness weight is a first brightness weight value; if the frame rate is greater than a second value, the initial brightness weight is a second brightness weight value; if the frame rate is not less than the first value and not greater than the second value, the initial brightness weight is between the second brightness weight value and the first brightness weight value, and the initial brightness weight is inversely proportional to the frame rate; Among them, when the determination module determines the target brightness weight based on the initial brightness weight, it is specifically used to: adjust the initial brightness weight based on the degree of motion of the video to be desensitized, and determine the adjusted brightness weight as the target brightness weight; wherein, the degree of motion is determined based on the motion vector between adjacent frames in the video to be desensitized; wherein, if the degree of motion is less than a first motion threshold, the initial brightness weight is increased, and if the degree of motion is greater than a second motion threshold, the initial brightness weight remains unchanged.

7. The device according to claim 6, characterized in that When determining the target contrast weight, the determination module is specifically used to: determine an initial contrast weight corresponding to the frame rate; determine the target contrast weight based on the initial contrast weight; wherein, if the frame rate is less than a third value, the initial contrast weight is a first contrast weight value; if the frame rate is greater than a fourth value, the initial contrast weight is a second contrast weight value; if the frame rate is not less than the third value and not greater than the fourth value, the initial contrast weight is between the second contrast weight value and the first contrast weight value, and the initial contrast weight is inversely proportional to the frame rate; When determining the target structure weight, the determination module is specifically used to: determine the initial structure weight corresponding to the frame rate; determine the target structure weight based on the initial structure weight; wherein, if the frame rate is less than a fifth value, the initial structure weight is a first structure weight value; if the frame rate is greater than a sixth value, the initial structure weight is a second structure weight value; if the frame rate is not less than the fifth value and not greater than the sixth value, the initial structure weight is between the first structure weight value and the second structure weight value, and the initial structure weight is proportional to the frame rate.

8. The device according to claim 7, characterized in that When determining the target contrast weight based on the initial contrast weight, the determination module is specifically used to: adjust the initial contrast weight based on the motion degree, and determine the adjusted contrast weight as the target contrast weight; if the motion degree is less than a first motion threshold, the initial contrast weight remains unchanged; if the motion degree is greater than a second motion threshold, the initial contrast weight is increased; When the determination module determines the target structure weight based on the initial structure weight, it is specifically used to: adjust the initial structure weight based on the movement degree, and determine the adjusted structure weight as the target structure weight; wherein, if the movement degree is less than a first movement threshold, the initial structure weight remains unchanged; if the movement degree is greater than a second movement threshold, the initial structure weight is increased; Among them, if the degree of motion is not less than the first motion threshold and not greater than the second motion threshold, the initial brightness weight, the initial contrast weight and the initial structure weight are adjusted so that the difference between the adjusted brightness weight, the adjusted contrast weight and the adjusted structure weight is not greater than the threshold.

9. The device according to claim 6, characterized in that If the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, the processing module performs desensitization processing on the target data to be desensitized through a multimodal large model based on the target desensitization strategy to obtain desensitized data, and is specifically used for: extracting features of the image data to be desensitized through a multimodal large model to obtain image features, extracting features of the audio data to be desensitized through a multimodal large model to obtain audio features, extracting features of the text data to be desensitized through a multimodal large model to obtain text features, and performing feature fusion of the image features, the audio features, and the text features through a multimodal large model to obtain multimodal features; Determine, by means of a multimodal large model, an image sensitive position of the image data to be desensitized, an audio sensitive position of the audio data to be desensitized, and a text sensitive position of the text data to be desensitized based on the multimodal features; Performing desensitization processing on the image sensitive position of the image data to be desensitized based on the image desensitization strategy corresponding to the image data to be desensitized to obtain desensitized image data; Based on the audio desensitization strategy corresponding to the audio data to be desensitized, the audio sensitive position of the audio data to be desensitized is desensitized to obtain desensitized audio data; Based on the text desensitization strategy corresponding to the text data to be desensitized, the text sensitive positions of the text data to be desensitized are desensitized to obtain desensitized text data.

10. The device according to claim 6, characterized in that If the target data to be desensitized includes image data to be desensitized, audio data to be desensitized, and text data to be desensitized, the target desensitization strategy includes an image desensitization strategy, an audio desensitization strategy, and a text desensitization strategy; The determination module determines the target desensitization strategy corresponding to the target data to be desensitized by querying the configured desensitization rule library through the multimodal attributes corresponding to the video to be desensitized, and obtaining the image desensitization strategy corresponding to the multimodal attributes; wherein the multimodal attributes include at least one of image attributes, audio attributes and text attributes, and the desensitization rule library includes the correspondence between the multimodal attributes and the image desensitization strategy; the image desensitization strategy includes an image desensitization network model; Determining an audio desensitization strategy based on the audio attributes corresponding to the video to be desensitized, the audio desensitization strategy comprising an audio desensitization network model and a vocoder, the audio desensitization network model is used to convert the audio data to be desensitized into acoustic features, and the vocoder is used to convert the acoustic features into an audio waveform; Determine a configured text desensitization strategy, wherein the text desensitization strategy includes a text encryption algorithm or a text replacement algorithm, wherein the text encryption algorithm is used to perform an encryption operation on the desensitized text data, and the text replacement algorithm is used to perform a replacement operation or a fuzzification operation on the desensitized text data; Among them, when the processing module adjusts the target desensitization strategy to obtain the adjusted desensitization strategy, if the image desensitization strategy needs to be adjusted, the network parameters in the image desensitization network model are adjusted; if the audio desensitization strategy needs to be adjusted, the network parameters in the audio desensitization network model are adjusted; if the text desensitization strategy needs to be adjusted, the algorithm type adopted by the text encryption algorithm is adjusted, or the algorithm type adopted by the text replacement algorithm is adjusted.

11. An electronic device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions executable by the processor; The processor is used to execute machine executable instructions to implement the method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Fidelity measurement of digital images

    CN103003841A

  • Privacy parameter optimization method in multi-modal data fusion training

    CN115310122A