Data security risk management method and device based on multi-modal large model

Through the data security risk governance method based on multimodal large models, the security risks in the data flow are identified and dealt with, and the problem that existing technology is difficult to effectively identify and deal with data security risks is solved, and efficient security governance of data assets is achieved.

CN120185931AActive Publication Date: 2025-06-20HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510648759.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

With the abundance of data assets and the transmission and use of data in multiple scenarios, data security risks have increased, and it is difficult for the prior art to effectively identify and deal with these risks.

Method used

The data security risk governance method based on multimodal large models is adopted. By generating multiple data groups, determining data flows, matching risk data flows and strategies, using the big model for risk assessment and response measures generation, and deploying them to the data source for security processing.

Benefits of technology

It realizes accurate identification and response to data flow security risks, reduces the security risks of data assets, and improves the accuracy and efficiency of data security governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120185931A_ABST
    Figure CN120185931A_ABST
Patent Text Reader

Abstract

The invention provides a data security risk management method and device based on a multi-modal large model. The method comprises the following steps: selecting a first type of data streams with security risks from all data streams; selecting target multi-modal sensitive data from the multi-modal sensitive data of the first type of data streams, and inputting the target multi-modal sensitive data to the first multi-modal large model to obtain risk assessment information; selecting a target data stream of which the risk assessment information is unsafe from all the first type of data streams, and inputting multi-modal sensitive data of the target data stream and a plurality of security countermeasures to a second multi-modal large model to obtain a target security countermeasure corresponding to the target data stream; and deploying the target security response measure to a data source corresponding to the target data stream, so that the data source performs security processing on the data based on the target security response measure. According to the scheme, the security risk of the data assets is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security technology, and in particular, to a data security risk governance method and device based on a multimodal large model. Background Art

[0002] With the rapid development of the digital economy, the data carried by the Internet is becoming increasingly rich, generating more and more data assets. A data asset refers to data resources that are owned or controlled by an individual or an enterprise, can bring economic benefits, and are recorded in a physical or electronic manner. A data asset is a data set in the cyber space that has data ownership rights (such as exploration rights, usage rights, ownership rights, etc.), is valuable, measurable, and readable.

[0003] Due to the opaque data sharing rules and inadequate data security protection measures in some scenarios, more concealed data security risks have emerged. In addition, with the increase in scenarios such as data collection, storage, transmission, and use across industries, departments, regions, systems, and businesses, while data creates greater value, data security risks such as tampering and leakage of important data and privacy information are also increasing, which may pose a serious threat to the security of data assets and cannot guarantee the security of data assets. Summary of the Invention

[0004] This application provides a data security risk governance method based on a multimodal large model, including: Generating multiple data groups based on the obtained multiple multimodal sensitive data. For each data group, the data group includes at least one multimodal sensitive data generated for the same security event; Determining the data flow corresponding to each data group; wherein, if the data group includes N multimodal sensitive data, the data flow includes N data sources for generating the N multimodal sensitive data; Selecting the first type of data flow with security risks from all data flows; wherein, for each data flow, if the data flow matches any risk data flow in the knowledge base, the data flow is the first type of data flow, and the knowledge base includes multiple configured risk data flows; if the multimodal sensitive data corresponding to the data flow matches the configured risk policy, the data flow is the first type of data flow; For each first type of data flow, selecting target multimodal sensitive data from the multimodal sensitive data corresponding to the first type of data flow, and inputting the target multimodal sensitive data into the first multimodal large model to obtain risk assessment information for the first type of data flow, where the risk assessment information is secure or insecure; Select the target data stream with insecure risk assessment information from all the first - type data streams, input the multi - modal sensitive data corresponding to the target data stream and multiple security response measures into the second multi - modal large model, and obtain the target security response measure corresponding to the target data stream; Deploy the target security response measure to the data source corresponding to the target data stream, so that the data source performs security processing on the data based on the target security response measure.

[0005] This application provides a data security risk governance device based on a multi - modal large model, including: A determination module, configured to generate multiple data groups based on the obtained multiple multi - modal sensitive data. For each data group, the data group includes at least one multi - modal sensitive data generated for the same security event; determine the data stream corresponding to each data group; wherein, if the data group includes N multi - modal sensitive data, the data stream includes N data sources for generating the N multi - modal sensitive data; A division module, configured to select the first - type data streams with security risks from all data streams; wherein, for each data stream, if the data stream matches any risk data stream in the knowledge base, the data stream is a first - type data stream, and the knowledge base includes multiple configured risk data streams; if the multi - modal sensitive data corresponding to the data stream matches the configured risk policy, the data stream is a first - type data stream; An acquisition module, configured to, for each first - type data stream, select target multi - modal sensitive data from the multi - modal sensitive data corresponding to the first - type data stream, input the target multi - modal sensitive data into the first multi - modal large model, and obtain the risk assessment information of the first - type data stream; wherein, the risk assessment information includes security information, and the security information indicates that the first - type data stream is secure or insecure; The acquisition module is configured to select the target data stream with insecure risk assessment information from all the first - type data streams, input the multi - modal sensitive data corresponding to the target data stream and multiple security response measures into the second multi - modal large model, and obtain the target security response measure corresponding to the target data stream; A sending module, configured to deploy the target security response measure to the data source corresponding to the target data stream, so that the data source performs security processing on the data based on the target security response measure.

[0006] This application provides an electronic device, including: a processor and a machine - readable storage medium, where the machine - readable storage medium stores machine - executable instructions that can be executed by the processor; the processor is configured to execute the machine - executable instructions to implement the data security risk governance method based on the multi - modal large model.

[0007] The present application provides a computer program product, including a computer program, which, when executed by a processor, implements a data security risk governance method based on a multimodal large model.

[0008] The present application provides a machine-readable storage medium storing machine-executable instructions capable of being executed by a processor; wherein, the processor is configured to execute the machine-executable instructions to implement the above-mentioned data security risk governance method based on a multimodal large model.

[0009] As can be seen from the above technical solutions, in the embodiments of the present application, the data flow (not data, not a single data source) is used as the detection object to determine whether there is a security risk in the data flow. If so, the target security response measure corresponding to the data flow is determined, and the target security response measure is deployed to the data source corresponding to the data flow, so that the data source performs security processing on the data based on the target security response measure, thereby achieving accurate identification of security risks and reducing the security risks of data assets. Since security incidents generate sensitive data in multiple data sources, by constructing a data flow composed of multiple data sources and using the data flow as the detection object instead of a single sensitive data or a single data source as the detection object, security risks can be identified more accurately.

[0010] During the detection process of security risks, it is possible to analyze whether there is a security risk in the data flow based on a knowledge base, risk policies, and a multimodal large model, further improving the accuracy of security risk assessment. For the data security protection scenario of Internet of Things enterprises, a data security risk governance method based on a multimodal large model is provided. By using a multimodal large model to implement security risk identification, security risk assessment, and security risk response, it is possible to perform multi-dimensional and intelligent full-life cycle governance of data security risks, thereby reducing the security risks of enterprise data assets, reducing the probability of occurrence of data asset security risks, and realizing the application value of data assets. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a schematic flowchart of a data security risk governance method based on a multimodal large model; Figure 2 is a schematic flowchart of a data security risk governance method based on a multimodal large model; Figure 3 is a schematic flowchart of the security risk identification process in the present application; Figure 4 is a schematic flowchart of the security risk assessment process in the present application; Figure 5 is a schematic flowchart of the security risk response process in the present application; Figure 6 is a schematic flowchart of the security risk detection and effect evaluation process in the present application; Figure 7 It is a schematic structural diagram of a data security risk governance device based on a multimodal large model; Figure 8 It is a hardware structure diagram of an electronic device in an implementation manner of the present application. Specific implementation manners

[0012] In an embodiment of the present application, a data security risk governance method based on a multimodal large model is proposed. Refer to Figure 1 As shown, it is a schematic flow diagram of the method. The method may include: Step 101: Generate a plurality of data groups based on the obtained plurality of multimodal sensitive data. For each data group, the data group includes at least one multimodal sensitive data generated for the same security event.

[0013] Step 102: Determine the data flow corresponding to each data group; wherein, if the data group includes N multimodal sensitive data, the data flow includes N data sources for generating the N multimodal sensitive data.

[0014] Step 103: Select the first type of data flow with security risks from all data flows. Among them, for each data flow, if the data flow matches any risk data flow in the knowledge base, the data flow is the first type of data flow, and the knowledge base includes a plurality of configured risk data flows; if the multimodal sensitive data corresponding to the data flow matches the configured risk policy, the data flow is the first type of data flow.

[0015] Step 104: For each first type of data flow, select target multimodal sensitive data from the multimodal sensitive data corresponding to the first type of data flow, and input the target multimodal sensitive data into the first multimodal large model to obtain risk assessment information of the first type of data flow. Among them, the risk assessment information may include security information, and the security information indicates that the first type of data flow is safe or unsafe.

[0016] Step 105: Select a target data flow whose risk assessment information is unsafe from all the first type of data flows, and input the multimodal sensitive data corresponding to the target data flow and a plurality of security countermeasures into the second multimodal large model to obtain the target security countermeasure corresponding to the target data flow.

[0017] Step 106: Deploy the target security countermeasure to the data source (such as multiple data sources) corresponding to the target data flow, so that the data source performs security processing on the data based on the target security countermeasure.

[0018] Exemplarily, selecting target multimodal sensitive data from the multimodal sensitive data corresponding to the first type of data stream may include, but is not limited to: sorting all the first type of data streams, and determining the reference quantity of each first type of data stream based on the sorting result; wherein, the earlier the first type of data stream is sorted, the larger the reference quantity of the first type of data stream. For each first type of data stream, based on the reference quantity of the first type of data stream, select multimodal sensitive data that matches the reference quantity from the multimodal sensitive data corresponding to the first type of data stream as the target multimodal sensitive data corresponding to the first type of data stream.

[0019] Exemplarily, after selecting the first type of data stream with security risks from all data streams, for each first type of data stream, obtain the portrait parameters corresponding to the first type of data stream. Among them, if the multimodal sensitive data corresponding to the first type of data stream matches the risk policy, the portrait parameters may include, but are not limited to, the text description parameters of the risk policy and the text description parameters of the multimodal sensitive data corresponding to the risk policy. If the first type of data stream matches the risk data stream in the knowledge base, the portrait parameters may include the text description parameters of each multimodal sensitive data corresponding to the first type of data stream.

[0020] Inputting the target multimodal sensitive data into the first multimodal large model to obtain the risk assessment information of the first type of data stream may include: inputting the portrait parameters and the target multimodal sensitive data corresponding to the first type of data stream into the first multimodal large model to obtain the risk assessment information of the first type of data stream.

[0021] Exemplarily, before generating multiple data groups based on the obtained multiple multimodal sensitive data, multiple multimodal raw data may also be obtained. The multiple multimodal raw data may include, but is not limited to, at least one of enterprise financial data, park security data, system operation and maintenance data, product R & D data, production data, human resources data, and employee behavior data; wherein, for each multimodal raw data, the multimodal raw data is text data, or image data, or audio data, or video data, or alarm data.

[0022] Select sensitive data from all multimodal raw data as multimodal sensitive data; wherein, for each multimodal raw data, input the multimodal raw data into the third multimodal large model to obtain a detection result, and the detection result indicates whether the multimodal raw data is sensitive data or not.

[0023] Among them, if the multi-modal raw data is text data, the text detection model in the third multi-modal large model is used to process the multi-modal raw data to obtain a detection result; if the multi-modal raw data is image data, the image detection model in the third multi-modal large model is used to process the multi-modal raw data to obtain a detection result; if the multi-modal raw data is audio data, the audio detection model in the third multi-modal large model is used to process the multi-modal raw data to obtain a detection result; if the multi-modal raw data is video data, the video detection model in the third multi-modal large model is used to process the multi-modal raw data to obtain a detection result; if the multi-modal raw data is alarm data, the alarm detection model in the third multi-modal large model is used to process the multi-modal raw data to obtain a detection result.

[0024] Exemplarily, the training process of the first multi-modal large model may include, but is not limited to: obtaining security risk assessment training data and training the first multi-modal large model based on the security risk assessment training data; among them, the security risk assessment training data includes internal data and external data, the internal data is the data running within the enterprise private network, the external data is the data running within the external network, and the external data includes security events and defect information; or, when receiving a data acquisition request from the model party training server, performing identity authentication and permission authentication on the model party training server; if the identity authentication is successful and the permission authentication is successful, encrypting the internal data to obtain encrypted internal data, and sending the encrypted internal data to the model party training server, so that the model party training server trains the first multi-modal large model based on the encrypted internal data.

[0025] Exemplarily, the second multimodal large model may include a pre-trained model, a position encoding layer, a multi-layer perceptron, and a large language model. Inputting the multimodal sensitive data corresponding to the target data stream and multiple security countermeasures into the second multimodal large model to obtain the target security countermeasure corresponding to the target data stream may include, but is not limited to: based on the pre-trained model, processing the multimodal sensitive data corresponding to the target data stream to obtain an embedding matrix of the target data stream, determining the encoded representation features corresponding to the embedding matrix, and performing an addition operation on the encoded representation features of all target data streams to obtain a comprehensive encoded representation feature; based on the pre-trained model, for each security countermeasure, mapping each input word corresponding to the security countermeasure to obtain an embedding vector of the security countermeasure; based on the position encoding layer, performing a convolution operation on the comprehensive encoded representation feature to obtain a first intermediate feature, performing a convolution operation on the embedding vector of each security countermeasure to obtain a second intermediate feature, performing a fusion operation on the first intermediate feature and the second intermediate feature of each security countermeasure to obtain a third intermediate feature, and performing a pooling operation on the third intermediate feature to obtain a fourth intermediate feature; based on the multi-layer perceptron, projecting the fourth intermediate feature into the same dimensional space as the word embedding of the large language model to obtain a target feature that can be processed by the large language model; based on the large language model, processing the target feature to obtain the target security countermeasure corresponding to the target data stream; wherein, the target security countermeasure may be selected by the large language model from multiple security countermeasures, or may be generated by the large language model.

[0026] In a possible implementation manner, the risk assessment information may further include, but is not limited to, a risk category and a risk level. After inputting the multimodal sensitive data corresponding to the target data stream and multiple security countermeasures into the second multimodal large model to obtain the target security countermeasure corresponding to the target data stream, it is also possible to query a measure information table based on the risk category and risk level corresponding to the target data stream to obtain multiple candidate security countermeasures corresponding to the risk category and the risk level; wherein, the measure information table may include the mapping relationship between the risk category, the risk level, and multiple candidate security countermeasures.

[0027] If the multiple candidate security countermeasures include the target security countermeasure, determine whether the target security baseline of the target security countermeasure meets a predefined security baseline, where the target security baseline represents the degree of security protection of the target security countermeasure for the data; if so, the operation of deploying the target security countermeasure to the data source corresponding to the target data stream may be performed; if not, the target security countermeasure may also be adjusted to obtain an adjusted security countermeasure, and the security baseline of the adjusted security countermeasure meets the predefined security baseline, and update the adjusted security countermeasure to the target security countermeasure, and then perform the operation of deploying the target security countermeasure to the data source corresponding to the target data stream.

[0028] If multiple candidate security response measures do not include the target security response measure, a candidate security policy can be used to generate a security response measure as the target security response measure, and an operation of deploying the target security response measure to the data source corresponding to the target data stream is performed. Among them, the candidate security policy can include, but is not limited to, at least one of a privacy computing security policy, a post-quantum encryption security policy, a data masking security policy, a data watermarking security policy, a content recognition and control security policy, a network blocking security policy, a boundary control security policy, and a terminal protection security policy, and there is no limitation on this candidate security policy.

[0029] Exemplarily, after deploying the target security response measure to the data source corresponding to the target data stream so that the data source performs security processing on the data based on the target security response measure, a first data set and a second data set corresponding to the target data stream can also be obtained. The first data set includes the original data generated by each data source corresponding to the target data stream, and the second data set includes the security data generated by each data source corresponding to the target data stream. The security data generated by the data source is obtained after the data source performs security processing on the original data based on the target security response measure. The first data set is input to the third multi-modal large model to obtain a detection result, and the detection result indicates whether the original data in the first data set is sensitive data or not; a third data set corresponding to the target data stream is obtained, and the third data set includes the sensitive data in the first data set. Based on the third data set, a first type of data stream with a security risk is determined, and risk assessment information for the first type of data stream is obtained based on the first multi-modal large model. The risk assessment information includes security information, a risk category, and a risk level; a fourth data set corresponding to the target data stream is obtained, and the fourth data set includes the sensitive data corresponding to the first type of data stream and the risk assessment information. The fuzzy evaluation scores of the first data set, the second data set, the third data set, and the fourth data set are determined.

[0030] If the fuzzy evaluation score of the first data set is less than the first threshold, the collection method of the multi-modal original data is adjusted, and the adjusted collection method is used to obtain the multi-modal original data; if the fuzzy evaluation score of the second data set is less than the second threshold, the second multi-modal large model is adjusted, and the adjusted second multi-modal large model is used to determine the target security response measure; if the fuzzy evaluation score of the third data set is less than the third threshold, the third multi-modal large model is adjusted, and the adjusted third multi-modal large model is used to identify sensitive data; if the fuzzy evaluation score of the fourth data set is less than the fourth threshold, the first multi-modal large model is adjusted, and the adjusted first multi-modal large model is used to determine the risk assessment information.

[0031] Exemplarily, the input data of the third multimodal large model, the output data of the third multimodal large model, and the model information of the third multimodal large model can also be recorded on the blockchain. The multimodal sensitive data corresponding to each first type of data stream, the output data of the first multimodal large model, and the model information of the first multimodal large model are recorded on the blockchain. The input data of the second multimodal large model, the output data of the second multimodal large model, and the model information of the second multimodal large model are recorded on the blockchain.

[0032] Before generating multiple data groups based on the obtained multiple multimodal sensitive data, it is also possible to determine whether the obtained multiple multimodal sensitive data is the same as the output data of the third multimodal large model recorded on the blockchain. If so, perform the operation of generating multiple data groups based on the obtained multiple multimodal sensitive data. If not, output an exception warning message; wherein, the model information includes at least one of model type, model version, model user, model usage time, model usage frequency, and model resource consumption.

[0033] As can be seen from the above technical solutions, in the embodiments of the present application, the data stream (not data, not a single data source) is used as the detection object to determine whether there is a security risk in the data stream. If so, determine the target security response measure corresponding to the data stream, and deploy the target security response measure to the data source corresponding to the data stream, so that the data source performs security processing on the data based on the target security response measure, thereby achieving accurate identification of security risks and reducing the security risks of data assets. Since security incidents will generate sensitive data in multiple data sources, by constructing a data stream composed of multiple data sources and using the data stream as the detection object instead of a single sensitive data or a single data source as the detection object, security risks can be identified more accurately.

[0034] During the detection process of security risks, it is possible to analyze whether there is a security risk in the data stream based on the knowledge base, risk strategy, and multimodal large model, further improving the accuracy of security risk assessment. For the data security protection scenario of Internet of Things enterprises, a data security risk governance method based on multimodal large models is provided. By using multimodal large models to achieve security risk identification, security risk assessment, and security risk response, it is possible to perform multi-dimensional and intelligent full-life-cycle governance of data security risks, thereby reducing the security risks of enterprise data assets, reducing the probability of occurrence of data asset security risks, and giving play to the application value of data assets.

[0035] The above technical solutions of the embodiments of the present application are described below in combination with specific application scenarios.

[0036] An embodiment of this application proposes a data security risk governance method based on a multimodal large model, which can reduce the probability of data security risks occurring while meeting relevant requirements and give play to the application value of data assets. A multimodal large model refers to a model trained by jointly combining multimodal information such as text, images, videos, and audio. The multimodal large model can understand, fuse, and process various types of information and perform more complex and intelligent tasks. The technologies of the multimodal large model include data preprocessing, feature extraction and representation, modality fusion, models and algorithms, etc. Data security risk governance refers to the full-life cycle management and control of data security risks through links such as risk identification, risk assessment, risk response, risk detection, and effect evaluation for more persistent and concealed data security risks.

[0037] An embodiment of this application proposes a data security risk governance method based on a multimodal large model, which can be applied to electronic devices such as personal computers, servers, routers, switches, etc., and there is no limit to the type of this electronic device. See Figure 2 As shown, it is a schematic flowchart of this method, and this method may include security risk identification, security risk assessment, security risk response, security risk detection, and effect evaluation.

[0038] In the process of security risk identification, it involves the collection of the current situation of enterprise data security governance, the collection of asset information based on the multimodal large model, the security risk identification based on the multimodal large model, etc. The security risk identification process is used to initially identify the data streams with security risks (which can be recorded as the first type of data stream).

[0039] In the process of security risk assessment, it involves the sorting of the first type of data stream, the security risk assessment based on the first multimodal large model, the training of the first multimodal large model, etc. The security risk assessment process is used to finally confirm the data stream with security risks (this data stream can be recorded as the target data stream).

[0040] In the process of security risk response, it involves the generation of security response measures based on the second multimodal large model, the verification of security response measures, the deployment of target security response measures, and the data security protection of the data source based on the target security response measures, etc., which are used to confirm the target security response measures corresponding to the target data stream.

[0041] In the process of security risk detection and effect evaluation, it involves the effect evaluation of data security risk governance, blockchain traceability, etc. The security risk detection and effect evaluation process is used to evaluate the effect of data security risk governance.

[0042] For the above process, the processes of security risk identification, security risk assessment, security risk response, security risk detection, and effect evaluation will be described below in combination with specific application scenarios.

[0043] First, regarding security risk identification. During the security risk identification process, all data streams can be divided into a first type of data stream with security risks and a second type of data stream without security risks. Refer to Figure 3 As shown, it is a flowchart of the security risk identification process. The security risk identification process may include: Step 301: Obtain multiple multimodal raw data.

[0044] Exemplarily, key enterprise data assets can be identified and collected. For the sake of convenience in distinction, key enterprise data assets are referred to as multimodal raw data. Multimodal raw data can be text data, or image data, or audio data, or video data, or alarm data, and there is no limitation thereto. Since key enterprise data assets can be various types of data (such as text, image, audio, etc.), in this embodiment, key enterprise data assets are referred to as multimodal raw data. For multimodal raw data such as text data, image data, audio data, video data, alarm data, etc., they can be structured multimodal data, semi-structured multimodal data, or unstructured multimodal data.

[0045] Key enterprise data assets can be collected through manual review. For example, organize the person in charge of each department to conduct a periodic inventory of department data assets to obtain key enterprise data assets. For example, the information security interface person registers and archives the data assets to obtain key enterprise data assets.

[0046] Key enterprise data assets can be collected through the automatic scanning of collection tools. For example, a DLP (Data Leakage Prevention) system can be deployed on the terminal device, and key enterprise data assets can be automatically scanned through the DLP system. For example, an EDR system (Endpoint Detection and Response) system can be deployed on the access-side device, and key enterprise data assets can be automatically scanned through the EDR system. A traffic analysis system can be deployed on the network boundary device, and key enterprise data assets can be automatically scanned through the traffic analysis system. Of course, the DLP system, the EDR system, and the traffic analysis system are only several examples of collection tools, and there is no limitation to such collection tools.

[0047] The above are only examples of collecting key enterprise data assets. A large number of key enterprise data assets form a data asset list, that is, the data asset list includes a large number of key enterprise data assets. The key enterprise data assets in the data asset list are referred to as multimodal raw data, so that multiple multimodal raw data can be obtained.

[0048] Exemplarily, the multiple multimodal raw data includes but is not limited to at least one of enterprise financial data, park security data, system operation and maintenance data, product R & D data, production data, human resources data, and employee behavior data. For example, enterprise financial data can be multimodal raw data generated by an enterprise financial data source (such as text data, image data, audio data, video data, alarm data, etc.). The enterprise financial data source is a device used for enterprise finance (such as a personal computer, etc.). The data related to enterprise finance is called enterprise financial data, and the data source that generates enterprise financial data is called the enterprise financial data source. Park security data can be multimodal raw data generated by a park security data source. The park security data source is a device used for park security. The data related to park security is called park security data. System operation and maintenance data can be multimodal raw data generated by a system operation and maintenance data source. The system operation and maintenance data source is a device used for system operation and maintenance. The data related to system operation and maintenance is called system operation and maintenance data. Product R & D data can be multimodal raw data generated by a product R & D data source. The product R & D data source is a device used for product R & D. The data related to product R & D is called product R & D data. Production data can be multimodal raw data generated by a production data source. The production data source is a device used for production (i.e., production staff). The data related to production is called production data. Human resources data can be multimodal raw data generated by a human resources data source. The human resources data source is a device used for human resources (i.e., human resources staff). The data related to human resources is called production data. Employee behavior data can be multimodal raw data generated by an employee data source. The employee data source is a device used by employees. The data related to employee behavior is called employee behavior data.

[0049] Enterprise financial data, park security data, system operation and maintenance data, product R & D data, production data, human resources data, and employee behavior data are only examples of multimodal raw data and are not limited thereto.

[0050] Step 302: Select sensitive data from all multimodal raw data as multimodal sensitive data.

[0051] For example, for each multimodal raw data, determine whether the multimodal raw data is sensitive data. If the multimodal raw data is sensitive data, then use the multimodal raw data as multimodal sensitive data. If the multimodal raw data is not sensitive data, then do not use the multimodal raw data as multimodal sensitive data. Based on this, multiple multimodal sensitive data can be obtained, that is, data with sensitive information.

[0052] In a possible implementation, a third multimodal large model can be pre-trained. The third multimodal large model is used to identify whether multimodal raw data is sensitive data, and the training method of the third multimodal large model is similar to that of the first multimodal large model. The training method is described in the subsequent process.

[0053] For each piece of multimodal raw data, the multimodal raw data can be input into the third multimodal large model, and the third multimodal large model processes the multimodal raw data to obtain a detection result of the multimodal raw data. The detection result indicates whether the multimodal raw data is sensitive data or not. Based on this, it can be determined whether the multimodal raw data is multimodal sensitive data.

[0054] Since the third multimodal large model supports the detection of text data, image data, audio data, video data, and alarm data, the third multimodal large model can include a text detection model (i.e., a text detection network), an image detection model, an audio detection model, a video detection model, and an alarm detection model.

[0055] On this basis, if the multimodal raw data is text data, the text detection model in the third multimodal large model can be used to process the multimodal raw data to obtain a detection result. There is no limitation on the processing method of this text detection model, as long as a detection result indicating whether the multimodal raw data is sensitive data or not can be obtained. If the multimodal raw data is image data, the image detection model in the third multimodal large model can be used to process the multimodal raw data to obtain a detection result. If the multimodal raw data is audio data, the audio detection model in the third multimodal large model can be used to process the multimodal raw data to obtain a detection result. If the multimodal raw data is video data, the video detection model in the third multimodal large model can be used to process the multimodal raw data to obtain a detection result. If the multimodal raw data is alarm data, the alarm detection model in the third multimodal large model can be used to process the multimodal raw data to obtain a detection result.

[0056] Exemplarily, steps 301 and 302 can be an asset information collection process based on a multimodal large model, which can collect multiple multimodal raw data and identify multiple multimodal sensitive data based on the multimodal large model. For example, sensitive information recognition and classification and grading processing are performed through the multimodal large model. In the process of sensitive information recognition, for text, images, audio, video, alarm signals, etc., different natural language processing models are used for recognition matching and integration respectively. In the classification and grading processing, through the multimodal large model, the enterprise's key data assets are classified according to dimensions such as enterprise financial data, park security data, system operation and maintenance data, product R & D data, production data, human resources data, and employee behavior data. In this way, for each multimodal sensitive data, the category of the multimodal sensitive data can also be obtained.

[0057] In the classification and grading processing, the enterprise's key data assets can also be graded according to weight ratios such as importance, privacy, and security, such as core data, important data, general data, externally public data, etc. In this way, for each multimodal sensitive data, the level of the multimodal sensitive data can also be obtained.

[0058] For example, a classification and grading processing model can be pre-trained, and the multimodal sensitive data is classified through the classification and grading processing model to obtain the category of the multimodal sensitive data. The multimodal sensitive data is graded through the classification and grading processing model to obtain the level of the multimodal sensitive data.

[0059] For example, for the third multimodal large model and the classification and grading processing model, adjustments can be made according to actual needs, such as model adjustments triggered by industry technology updates and enterprise business adjustments. The third multimodal large model and the classification and grading processing model can be periodically updated with training data and optimized.

[0060] Step 303: Generate multiple data groups based on multiple multimodal sensitive data. For each data group, the data group can include at least one multimodal sensitive data generated for the same security event.

[0061] Exemplarily, when a certain security event occurs, the security event may generate multimodal sensitive data in multiple data sources. For example, 5 multimodal sensitive data are generated in the enterprise financial data source, 1 multimodal sensitive data is generated in the park security data source, and 3 multimodal sensitive data are generated in the system operation and maintenance data source. The multimodal sensitive data generated by these data sources has an association relationship, such as the same or close time, the same data operator, the same event identifier, etc. Based on the association relationship between the multimodal sensitive data, the multimodal sensitive data generated for the same security event can be found, and these multimodal sensitive data can be formed into a data group, and the data group includes multiple multimodal sensitive data generated for the same security event.

[0062] Obviously, based on a large amount of multi-modal sensitive data, when there are multiple security events, multiple data groups corresponding to the multiple security events can be obtained, that is, each security event corresponds to a data group.

[0063] Exemplarily, a security large model can be pre-trained, and the security large model is used to identify multi-modal sensitive data generated for the same security event. Based on this, multiple multi-modal sensitive data can be input into the security large model, and the security large model processes the multiple multi-modal sensitive data, mines and analyzes the correlation relationships of these multi-modal sensitive data, finds the multi-modal sensitive data generated by the same security event, and outputs multiple data groups, and each data group includes the multi-modal sensitive data generated by the same security event.

[0064] For example, the security event can be an internal threat event, such as an employee's misoperation event, an employee's malicious leakage event, an employee's tampering event, etc. The security event can be an external attack event, such as a hacker intrusion event, a virus propagation event, etc. The security event can be a technical defect information event, such as an unauthorized defect information event, a cross-domain data leakage event, an open source compliance event, etc. The above are only several examples of security events, and there is no limitation thereto.

[0065] For these security events, when training the security large model, the security large model can learn the logical relationships and influence associations among various risk factors, and there is no limitation to the training process of the security large model. On this basis, the security large model can identify multi-modal sensitive data generated for the same security event.

[0066] For example, the security large model is based on an expert library, combines the enterprise's proprietary business and private domain knowledge, and uses the direct preference optimization technology for special training and optimization to learn the logical relationships and influence associations among various risk factors. For example, a special security large model based on the video Internet of Things field can be adopted.

[0067] In summary, by inputting multiple multi-modal sensitive data into the security large model, multiple data groups can be obtained, and each data group includes at least one multi-modal sensitive data generated for the same security event.

[0068] Step 304, determine the data stream corresponding to each data group; wherein, if the data group includes N multi-modal sensitive data, the data stream includes N data sources for generating the N multi-modal sensitive data.

[0069] Exemplarily, for each data group (taking one data group as an example hereinafter), if the data group includes N multimodal sensitive data, the category of each multimodal sensitive data can be determined, such as categories like enterprise financial data, park security data, system operation and maintenance data, etc. Based on the category of the multimodal sensitive data, the data source of the multimodal sensitive data can be determined. For example, enterprise financial data corresponds to the enterprise financial data source, park security data corresponds to the park security data source, and so on. After obtaining the data sources corresponding to the N multimodal sensitive data, the N data sources corresponding to the N multimodal sensitive data can be used as a data stream. For example, assume that the data group includes 9 multimodal sensitive data, such as 5 enterprise financial data, 1 park security data, and 3 system operation and maintenance data. Then, the data stream corresponding to this data group includes 9 data sources, such as enterprise financial data source, enterprise financial data source, enterprise financial data source, enterprise financial data source, enterprise financial data source, park security data source, system operation and maintenance data source, system operation and maintenance data source, system operation and maintenance data source.

[0070] In a possible implementation manner, when the security large model outputs a data group, the multimodal sensitive data within the data group has an order relationship. For example, when the security large model mines and analyzes the association relationship of the multimodal sensitive data, the order relationship of these multimodal sensitive data can also be determined. For example, the multimodal sensitive data can be sorted according to time sequence, or sorted according to the data flow situation (such as when multimodal sensitive data A causes multimodal sensitive data B to be generated, multimodal sensitive data A is located in front of multimodal sensitive data B). Within the data stream corresponding to this data group, based on the order relationship of the N multimodal sensitive data, the order relationship of the N data sources is determined.

[0071] For example, if multimodal sensitive data A is located in front of multimodal sensitive data B, then the data source corresponding to multimodal sensitive data A is located in front of the data source corresponding to multimodal sensitive data B.

[0072] In summary, when determining the data stream, the N data sources within the data stream have an order relationship.

[0073] Step 305: For each data stream (taking the processing process of one data stream as an example hereinafter), determine whether the data stream matches any risk data stream in the knowledge base. If so, step 306 can be executed; if not, that is, the data stream does not match all the risk data streams in the knowledge base, step 307 can be executed.

[0074] Step 306: Determine the data stream as the first type of data stream with security risks.

[0075] Step 307: Determine whether the multimodal sensitive data corresponding to the data stream matches the configured risk policies. If so, step 306 can be executed. If not, that is, the multimodal sensitive data (multiple multimodal sensitive data) corresponding to the data stream does not match all risk policies, then step 308 can be executed.

[0076] Step 308: Determine the data stream as a second type of data stream without security risks.

[0077] Exemplarily, a knowledge base can be pre-configured. The knowledge base is a means for matching security threats to data streams. The knowledge base includes multiple risk data streams, that is, data streams with security risks. For example, if it is known that the risk data stream X1 (data source A - data source A - data source B - data source C - data source C) has a security risk, then the risk data stream X1 is configured in the knowledge base. If it is known that the risk data stream X2 (data source A - data source B - data source C) has a security risk, then the risk data stream X2 is configured in the knowledge base, and so on.

[0078] In step 305, it can be determined whether the data stream matches the risk data stream. For example, if the data stream is exactly the same as the risk data stream, then the data stream matches the risk data stream. If some data sources of the data stream are exactly the same as the risk data stream, then the data stream matches the risk data stream. For example, if the data stream is data source A - data source B - data source C, then the data stream is exactly the same as the risk data stream X2, and the data stream matches the risk data stream X2. If the data stream is data source A - data source B - data source C - data source D, then some data sources (data source A - data source B - data source C) of the data stream are exactly the same as the risk data stream X2, and the data stream matches the risk data stream X2. If the data stream is data source A - data source B - data source B - data source C, then the data stream does not match the risk data stream X2.

[0079] Exemplarily, multiple risk policies can be pre-configured. The risk policy is a means for matching security threats to multimodal sensitive data. For example, if it is known that there is a security risk when the operation frequency is greater than 10 times, then the risk policy Y1 can be configured, and the risk policy Y1 is that the operation frequency is greater than 10 times. If it is known that there is a security risk when the operation time is 0:00 - 5:00, then the risk policy Y2 can be configured, and the risk policy Y2 is that the operation time is 0:00 - 5:00. If it is known that there is a security risk when the operation permission (or operation role) is a visitor, then the risk policy Y3 can be configured, and the risk policy Y3 is that the operation permission (or operation role) is a visitor, and so on.

[0080] In step 307, the data stream corresponds to multiple multimodal sensitive data (i.e., multiple multimodal sensitive data within the data group), and each multimodal sensitive data is traversed in sequence. For the currently traversed multimodal sensitive data, it is determined in sequence whether the multimodal sensitive data matches each risk policy. If the multimodal sensitive data matches any risk policy, the data stream can be determined as a first type of data stream. If the multimodal sensitive data does not match all risk policies, it is necessary to continue traversing the next multimodal sensitive data, and so on, until the last multimodal sensitive data is traversed. If the last multimodal sensitive data does not match all risk policies, the data stream can be determined as a second type of data stream.

[0081] For example, when determining whether the multimodal sensitive data matches the risk policy, the statistical features corresponding to the multimodal sensitive data can be determined. For risk policy Y1, the statistical feature can be the operation frequency corresponding to the multimodal sensitive data. For risk policy Y2, the statistical feature can be the operation time corresponding to the multimodal sensitive data. For risk policy Y3, the statistical feature can be the operation permission (or operation role) corresponding to the multimodal sensitive data. In this embodiment, the acquisition method of this statistical feature is not limited. Based on the statistical features corresponding to the multimodal sensitive data, it can be determined whether the multimodal sensitive data matches the risk policy. For example, if the statistical feature indicates that the operation frequency corresponding to the multimodal sensitive data is greater than 10 times, it is determined that the multimodal sensitive data matches risk policy Y1. If the statistical feature indicates that the operation frequency corresponding to the multimodal sensitive data is not greater than 10 times, it is determined that the multimodal sensitive data does not match risk policy Y1.

[0082] Exemplarily, during the collection of the current situation of enterprise data security governance, the causes, consequences, and handling methods of network and data security incidents of Internet of Things enterprises over the years can be collected and analyzed. The management systems and regulatory systems in the field of data security risk governance of the enterprise can be collected. Special questionnaires can be designed to conduct a survey on the current situation of the implementation of security risk governance for groups such as enterprise managers, network security specialists, security interface persons of each business department, IT personnel, R & D personnel of each business department, and factory production specialists, and form a questionnaire. In this way, a large amount of data can be collected, and based on these data, the knowledge base and risk policies can be updated.

[0083] Exemplarily, for each first type of data stream with security risks, the portrait parameters corresponding to the first type of data stream, that is, data security portrait parameters, can also be obtained. The portrait parameters are used to reflect the characteristics of the first type of data stream, such as basic characteristics and preference characteristics, etc. The portrait parameters are not limited in this regard.

[0084] If the multimodal sensitive data corresponding to the first type of data stream matches the risk policy, the portrait parameters include the text description parameters of the risk policy (i.e., the risk policy is described in text), and the text description parameters of the multimodal sensitive data corresponding to the risk policy (i.e., the multimodal sensitive data is described in text, which can be a description of the statistical characteristics of the multimodal sensitive data), and the text description parameters of the multimodal sensitive data correspond to the text description parameters of the risk policy. For example, if the risk policy is that the operation frequency is greater than 10 times, and the multimodal sensitive data 1 matches the risk policy, and the multimodal sensitive data 1 indicates that the operation frequency is 15 times, then the text description parameter of the risk policy is the text information of "the operation frequency is greater than 10 times", and the text description parameter of the multimodal sensitive data 1 is the text information of "the operation frequency is 15 times". Of course, the above are only examples of text description parameters, and there is no limitation on this.

[0085] If the first type of data stream matches the risk data stream in the knowledge base, the portrait parameters include the text description parameters of each multimodal sensitive data corresponding to the first type of data stream. For example, the text description parameter can represent the classification and grading information of the multimodal sensitive data, or other attributes of the multimodal sensitive data. For example, the text description parameter can be the category and level of the multimodal sensitive data. The category can be enterprise financial data, park security data, system operation and maintenance data, etc., and the level can be core data, important data, general data, externally public data, etc. The above are only examples of text description parameters, and there is no limitation on this.

[0086] Exemplarily, steps 303 - 308 can be a security risk identification process based on a multimodal large model, which can include steps such as data preprocessing, construction of data stream logical relationships, security threat matching, construction of data security portraits, etc., and finally identify the first type of data stream that may face security threats.

[0087] For the data preprocessing step, multiple data groups can be generated based on multiple multimodal sensitive data. For the step of constructing the data stream logical relationship, the data stream corresponding to each data group can be determined. If there are the same data sources in the data stream, the data stream can be called an east - west data stream. If there are no same data sources in the data stream, the data stream can be called a north - south data stream. For the security threat matching step, it can be determined whether the data stream matches any risk data stream in the knowledge base, and it can be determined whether the multimodal sensitive data corresponding to the data stream matches the risk policy. For the step of constructing the data security portrait, the data stream can be determined as the first type of data stream or the second type of data stream, and the subsequent process is for the first type of data stream.

[0088] So far, the security risk identification process is completed, and the first type of data stream with potential security risks is initially identified.

[0089] Second, for security risk assessment. During the security risk assessment process, for multiple first-class data streams with security risks, the risk assessment information of each first-class data stream can be determined. Refer to Figure 4 As shown, it is a flowchart of the security risk assessment process. The security risk assessment process may include: Step 401: Sort all first-class data streams.

[0090] During the security risk identification process, multiple first-class data streams can be obtained. For each first-class data stream, the first-class data stream corresponds to multiple multimodal sensitive data. Feature extraction can be performed on these multimodal sensitive data to obtain the feature matrix of the first-class data stream. There is no limitation on this feature extraction method.

[0091] Based on the feature matrix of each first-class data stream, all first-class data streams can be sorted using the target sorting algorithm. For example, the target sorting algorithm may include, but is not limited to, the TOPSIS (Technique for Order Preference by Similarity to Ideal Solution) evaluation algorithm, the AHP (Analytic Hierarchy Process) evaluation algorithm, the fuzzy comprehensive evaluation algorithm, etc. There is no limitation on this target sorting algorithm. The TOPSIS evaluation algorithm will be used as an example for subsequent description.

[0092] For example, the feature matrices of all first-class data streams can be used as the input of the TOPSIS evaluation algorithm. Through the TOPSIS evaluation algorithm, the distances between each first-class data stream and the ideal solution and the anti-ideal solution, and the closeness degrees of each first-class data stream to the ideal solution can be obtained. Sorting is performed according to the magnitudes of the closeness degrees to the ideal solution, which is used as the basis for evaluating the advantages and disadvantages of each first-class data stream, thereby sorting all first-class data streams.

[0093] Based on the feature matrix of each first - type data stream, during the processing of the TOPSIS evaluation algorithm, the following steps may be involved: data pre - processing (during data pre - processing, risk factor dimensions can be determined, including but not limited to data leakage, data corruption, system defect information, network attacks, compliance risks, internal threats, etc., and the collected data can also be standardized), identification of weights (during the identification of weights, a quantitative score can be given according to the degree of impact on the enterprise, such as using a scoring standard of 1 - 5 points, with the degree of impact increasing from low to high), construction of a weighted decision matrix (using the standardized data and comprehensive weights to construct a weighted decision matrix), determination of the positive ideal solution and the negative ideal solution (selecting the optimal value and the worst value among various risk indicators), calculation of distances (calculating the Euclidean distances from each evaluation object (i.e., the first - type data stream) to the positive ideal solution and the negative ideal solution respectively to measure its gap from the ideal state), calculation of relative closeness (calculating the relative closeness of each evaluation object based on its distances to the positive ideal solution and the negative ideal solution), sorting and determination of priorities (sorting all evaluation objects according to the relative closeness to determine the data security risk points that need to be prioritized for attention and processing), etc. Finally, all first - type data streams are sorted. The above is an example of sorting all first - type data streams using the TOPSIS evaluation algorithm, and the sorting method in this embodiment is not limited thereto.

[0094] Step 402: Determine the reference quantity of each first - type data stream based on the sorting result.

[0095] Exemplarily, after sorting all first - type data streams, the more forward the sorting of a first - type data stream, the greater its importance. Therefore, when the sorting of a first - type data stream is more forward, the reference quantity of this first - type data stream is larger, the more reference data there is for risk assessment of this first - type data stream, and the more accurate the risk assessment result for this first - type data stream. The more backward the sorting of a first - type data stream, the smaller its importance. Therefore, when the sorting of a first - type data stream is more backward, the reference quantity of this first - type data stream is smaller, the less reference data there is for risk assessment of this first - type data stream, thereby reducing the amount of data processed by the first multi - modal large model and saving computing resources, that is, reducing the resources occupied by the first - type data streams with low importance.

[0096] Step 403: For each first - type data stream, based on the reference quantity of this first - type data stream, select multi - modal sensitive data that matches the reference quantity from the multi - modal sensitive data corresponding to this first - type data stream, and use the selected multi - modal sensitive data as the target multi - modal sensitive data corresponding to this first - type data stream.

[0097] For example, the reference quantity of the first type of data stream a1 ranked first is 10, and 10 multimodal sensitive data are selected from the multimodal sensitive data corresponding to the first type of data stream a1 as the target multimodal sensitive data. The reference quantity of the first type of data stream a2 ranked second is 9, and 9 multimodal sensitive data are selected from the multimodal sensitive data corresponding to the first type of data stream a2 as the target multimodal sensitive data, and so on.

[0098] Step 404: For each first type of data stream, input the target multimodal sensitive data of this first type of data stream into the first multimodal large model to obtain the risk assessment information of this first type of data stream.

[0099] Exemplarily, the first multimodal large model can be pre-trained. The first multimodal large model is used to identify the risk assessment information of the first type of data stream and conduct a detailed security risk assessment on the first type of data stream. Based on this, after inputting the target multimodal sensitive data of the first type of data stream into the first multimodal large model, the first multimodal large model can process the target multimodal sensitive data to obtain the risk assessment information of this first type of data stream. For example, the risk assessment information may include security information, and this security information may indicate that the first type of data stream is safe or unsafe. If the security information indicates that the first type of data stream is unsafe, the risk assessment information may further include a risk category and a risk level. For example, the risk category indicates what type of security risk exists in the first type of data stream, such as data leakage risk category, data corruption risk category, system defect information risk category, network attack risk category, compliance risk category, internal threat risk category, etc. For example, the risk level indicates what degree of security risk exists in the first type of data stream, such as low risk level, medium risk level, high risk level, etc.

[0100] Exemplarily, in the model function of the first multimodal large model, the tasks of multimodal data understanding and risk assessment generation can be decoupled to reduce task impacts and conflicts and improve the model operation efficiency.

[0101] For example, for different first types of data streams, different quantities of multimodal sensitive data can be input into the first multimodal large model instead of inputting the same quantity of multimodal sensitive data into the first multimodal large model, so as to decouple the tasks and improve the model operation efficiency of the first multimodal large model.

[0102] In summary, after inputting the target multimodal sensitive data of each first type of data stream into the first multimodal large model, the risk assessment information of each first type of data stream can be obtained. Subsequently, select the first type of data stream whose risk assessment information is unsafe from all the first type of data streams as the target data stream. The target data stream is the data stream finally confirmed to have a security risk and needs to respond to the security risk of the target data stream.

[0103] In a possible implementation, for each first type of data stream, the portrait parameters corresponding to the first type of data stream and the target multimodal sensitive data can be input into the first multimodal large model to obtain the risk assessment information of the first type of data stream. For example, for the input data of the first multimodal large model, in addition to the target multimodal sensitive data, portrait parameters can also be included. In this way, the first multimodal large model can process based on the target multimodal sensitive data and portrait parameters to obtain the risk assessment information.

[0104] For example, the first multimodal large model extracts the first feature from the target multimodal sensitive data, extracts the second feature from the portrait parameters, and fuses the first feature and the second feature to obtain the fused feature. The first multimodal large model determines the risk assessment information based on the fused feature.

[0105] For example, all portrait parameters and a reference number of target multimodal sensitive data can be input into the first multimodal large model. Or, all portrait parameters and K target multimodal sensitive data can be input into the first multimodal large model, where K can be the difference between the reference number and the number of portrait parameters. Or, if the number of portrait parameters is greater than the reference number, a reference number of portrait parameters can be input into the first multimodal large model, and no target multimodal sensitive data is input into the first multimodal large model anymore.

[0106] In a possible implementation, for the training process of the first multimodal large model, it may include: Method 1: Train the first multimodal large model through a server within the enterprise private network.

[0107] First, obtain the security risk assessment training data, which can include internal data and external data. The internal data can be the data running within the enterprise private network, such as historical security events and real-time user data, etc. The external data can be the data running within the external network, and the external data can include security events and defect information (such as 0day defect information, etc.). For the external data, data preprocessing can also be performed on the external data. During the data preprocessing process, illegal data and redundant data can be removed to prevent the behavior of illegal data interfering with and attacking the first multimodal large model.

[0108] Then, the first multimodal large model can be trained based on the security risk assessment training data. There is no limitation on the training process of this first multimodal large model, as long as it can optimize the first multimodal large model.

[0109] Method 2: Train the first multimodal large model through the model provider's training server within the external network.

[0110] First, when the model - side training server needs to train the first multi - modal large model, it can receive a data acquisition request sent by the model - side training server. When receiving this data acquisition request, identity authentication and permission authentication can be performed on the model - side training server. For example, identity authentication is used to verify whether the identity of the model - side training server is legal, and permission authentication is used to verify whether the model - side training server has the permission to train the first multi - modal large model. There is no limitation on this identity authentication and permission authentication process.

[0111] Then, if the identity authentication is successful and the permission authentication is successful, the internal data (i.e., the data running within the enterprise private network) is encrypted to obtain ciphertext internal data, and the ciphertext internal data is sent to the model - side training server. After receiving the ciphertext internal data, the model - side training server trains the first multi - modal large model based on the ciphertext internal data. There is no limitation on the training process of this first multi - modal large model. By transmitting the ciphertext internal data and performing model training on the basis of the ciphertext internal data, the leakage of enterprise privacy data is avoided.

[0112] Then, the trained first multi - modal large model can be obtained from the model - side training server.

[0113] In summary, it can be seen that in the training strategy of the first multi - modal large model, for links such as model construction, data training, and model use, the first multi - modal large model is optimized according to the actual situation of the enterprise, supporting the operation of the model within the private network, and realizing that knowledge does not leave the network and privacy data does not leave the domain.

[0114] So far, the security risk assessment process is completed, and the target data flow with security risks is finally confirmed.

[0115] Third, regarding the security risk response process. In the security risk response process, for multiple target data flows with security risks, the target security response measures for each target data flow can be determined. See Figure 5 As shown, it is a flow diagram of the security risk response process. The security risk response process can include: Step 501: Input the multi - modal sensitive data corresponding to the target data flow and multiple security response measures into the second multi - modal large model to obtain the target security response measures corresponding to each target data flow.

[0116] Exemplarily, the second multimodal large model can be pre-trained. The second multimodal large model is used to identify the target security response measures corresponding to the target data stream, that is, to obtain the security risk response method using the second multimodal large model. The training method of the second multimodal large model is similar to that of the first multimodal large model. Based on this, after inputting the multimodal sensitive data and multiple security response measures corresponding to the target data stream into the second multimodal large model, the second multimodal large model can process the multimodal sensitive data and security response measures to obtain the target security response measures corresponding to the target data stream.

[0117] For example, the target security response measure indicates what method to use for secure data processing. For example, using privacy computing method for secure data processing, using post-quantum encryption method for secure data processing, using data masking method for secure data processing, using data watermarking method for secure data processing, using content recognition and control method for secure data processing, using network blocking method for secure data processing, using boundary control method for secure data processing, using terminal protection method for secure data processing, etc. There is no limitation on this target security response measure.

[0118] For example, when inputting multiple security response measures into the second multimodal large model, each security response measure can be text, such as the text content being "using post-quantum encryption method for secure data processing". When the second multimodal large model outputs the target security response measure, the target security response measure can be text.

[0119] In a possible implementation manner, the second multimodal large model can include a pre-trained model, a position encoding layer, a multi-layer perceptron (MLP), and a large language model (LLM). Here is just an example of the second multimodal large model, and there is no limitation on this.

[0120] For each target data stream, input the multimodal sensitive data corresponding to the target data stream into the pre-trained model. The pre-trained model processes the multimodal sensitive data to obtain the embedding matrix corresponding to the target data stream, and determines the encoded representation feature corresponding to the embedding matrix. For example, first extract the features of the multimodal sensitive data, then process the extracted features to obtain the embedding matrix, and then convert the embedding matrix into an encoded representation (that is, encode the embedding matrix) to obtain the encoded representation feature of the target data stream.

[0121] After obtaining the encoded representation features of each target data stream, an addition operation can be performed on the encoded representation features of all target data streams to obtain a comprehensive encoded representation feature (that is, the feature after the addition operation).

[0122] For each security response measure (i.e., the input text), the security response measure can be input into a pre-trained model, and the pre-trained model maps each input word corresponding to the security response measure (such as splitting the security response measure into multiple input words and mapping each split input word) to obtain an embedding vector of the security response measure, and the embedding vector is a more rich semantic representation of the token.

[0123] After obtaining the comprehensive encoding representation feature, input the comprehensive encoding representation feature into a position encoding layer, and perform a convolution operation on the comprehensive encoding representation feature through the position encoding layer to obtain a first intermediate feature. For example, the position encoding layer includes a convolutional layer, and the convolutional layer uses a 5*5 (or other size) matrix as the convolution kernel, with a stride of 1 (or other stride) and a padding of 0 (or other padding), and performs a convolution operation on the comprehensive encoding representation feature through the convolutional layer (i.e., element-wise multiplication and then addition) to obtain the first intermediate feature (i.e., the feature after convolution).

[0124] For each security response measure, after obtaining the embedding vector of the security response measure, input the embedding vector into a position encoding layer, and perform a convolution operation on the embedding vector through the position encoding layer (such as the convolutional layer in the position encoding layer) to obtain a second intermediate feature (i.e., the feature after the convolution operation).

[0125] After obtaining the first intermediate feature and the second intermediate feature of each security response measure, perform a fusion operation on the first intermediate feature and all the second intermediate features through the position encoding layer to obtain a third intermediate feature. For example, the fusion operation can be a feature concatenation operation, a weighted operation, an addition operation, a multiplication operation, a convolution operation, a self-attention-based fusion operation, etc., and the way of this fusion operation is not limited.

[0126] After obtaining the third intermediate feature, a pooling operation can be performed on the third intermediate feature to obtain a fourth intermediate feature. By performing the pooling operation, overfitting of the second multi-modal large model can be prevented.

[0127] After obtaining the fourth intermediate feature, the fourth intermediate feature can be input into a projection-based connector. Taking the projection-based connector as a multi-layer perceptron as an example, after the multi-layer perceptron obtains the fourth intermediate feature, it projects the fourth intermediate feature into the same dimensional space as the word embedding of the large language model to obtain a target feature that can be processed by the large language model, that is, the target feature can be processed by the large language model together with the text tokens.

[0128] For example, assume that the dimension of the word embedding supported by the large language model is dimension A. Then, the multi-layer perceptron is used to generate a feature of dimension A. Based on this, the multi-layer perceptron converts the fourth intermediate feature into a target feature of dimension A, and the target feature of dimension A and the text tokens of dimension A are processed by the large language model together.

[0129] After obtaining the target features, the target features can be input into a large language model, and the large language model processes the target features to obtain the target security response measures corresponding to each target data stream. For example, the large language model can adopt a language model with an expert mixture architecture, or a language model with other network structures. For example, the target features and text tokens (i.e., prompts) can be input into the large language model together, and the large language model performs tasks such as language understanding, generation, or reasoning based on the target features and text tokens to obtain the target security response measures corresponding to each target data stream.

[0130] For each target data stream, the target security response measure corresponding to the target data stream can be selected by the large language model from multiple security response measures (i.e., the input data of the second multi-modal large model). Alternatively, the target security response measure corresponding to the target data stream can be generated by the large language model, that is, the large language model itself generates the target security response measure, and no restrictions are imposed on this generation method. In summary, for each target data stream, the target security response measure corresponding to the target data stream can be obtained.

[0131] Step 502: For each target data stream, query the measure information table (such as the mapping relationship between risk categories, risk levels, and multiple candidate security response measures) based on the risk category and risk level corresponding to the target data stream, and obtain multiple candidate security response measures corresponding to the risk category and the risk level.

[0132] Exemplarily, the measure information table can be pre-configured, and the measure information table is used to record available candidate security response measures. For example, as shown in Table 1, it is an example of the measure information table, and the measure information table includes the mapping relationship between risk categories, risk levels, and multiple candidate security response measures.

[0133] Table 1

[0134] As can be seen from Table 1, for the target data stream with risk category 1 and risk level 1, candidate security response measures P11, candidate security response measures P12,... can be adopted for security governance, and so on.

[0135] For example, during the collection of the current situation of enterprise data security governance, a large amount of data can be collected (see the above description), and based on this data, the measure information table can be updated. Of course, other methods can also be used to update the measure information table, and no restrictions are imposed on the source and content of the measure information table.

[0136] Exemplarily, for each target data stream, a table 1 can be queried based on the risk category and risk level corresponding to the target data stream to obtain multiple candidate security response measures corresponding to the target data stream.

[0137] Step 503: For each target data stream, determine whether the multiple candidate security response measures corresponding to the target data stream include the target security response measure corresponding to the target data stream.

[0138] If not, step 504 can be executed; if so, step 505 can be executed.

[0139] Step 504: Use a candidate security policy to generate a security response measure as the target security response measure corresponding to the target data stream, and based on the currently generated target security response measure, execute step 507.

[0140] Step 505: Determine whether the target security baseline of the target security response measure meets a predefined security baseline, where the target security baseline can represent the degree of security protection of the target security response measure for data.

[0141] If so, step 507 can be executed. If not, step 506 can be executed.

[0142] Step 506: Adjust the target security response measure to obtain an adjusted security response measure, and the security baseline of the adjusted security response measure meets the predefined security baseline. Update the adjusted security response measure as the target security response measure, and based on the target security response measure, execute step 507.

[0143] Step 507: Deploy the target security response measure corresponding to the target data stream to each data source corresponding to the target data stream, so that the data source processes the data securely based on the target security response measure.

[0144] Exemplarily, considering that there is a certain probability of error in the target security response measure output by the second multi-modal large model, therefore, after obtaining the target security response measure output by the second multi-modal large model, instead of directly deploying the target security response measure to the data source, a security verification is performed on the target security response measure. If the multiple candidate security response measures include the target security response measure, it indicates that the security verification is successful. If the multiple candidate security response measures do not include the target security response measure, it indicates that the security verification fails, and a candidate security policy needs to be used to generate the target security response measure and deploy the target security response measure to the data source, so as to ensure that the target security response measure is effective for data security governance.

[0145] Exemplarily, a backup security network can be additionally deployed. After the second multi-modal large model outputs the target security response measures, the backup security network performs a security verification on the target security response measures. If the security verification fails, the backup security network generates the target security response measures using the candidate security policies.

[0146] For example, the backup security network is trained specifically based on technology and management. In terms of technology, the minimum security requirements for information security, network security, and data security are specified, and then the technical solutions are optimized according to different business scenarios. In terms of management, the industry standards for information security, network security, and data security are specified to correct the accuracy of the given security response measures. The backup security network supports using candidate security policies (i.e., security technologies) such as privacy computing, post-quantum encryption, data masking, data watermarking, content recognition and control, network blocking, boundary control, and endpoint protection to generate the target security response measures.

[0147] In this embodiment, the structure of the backup security network is not limited as long as the backup security network can generate the target security response measures using the candidate security policies, and the generation method is not limited.

[0148] Exemplarily, for step 504, when generating the target security response measures using the candidate security policies, the candidate security policies can include but are not limited to at least one of the privacy computing security policy, post-quantum encryption security policy, data masking security policy, data watermarking security policy, content recognition and control security policy, network blocking security policy, boundary control security policy, and endpoint protection security policy. For example, the target security response measures indicate how to perform security processing on the data. For example, if the candidate security policy is the privacy computing security policy, the target security response measures indicate performing security processing on the data using the privacy computing method; if the candidate security policy is the post-quantum encryption security policy, the target security response measures indicate performing security processing on the data using the post-quantum encryption security policy, and so on.

[0149] Exemplarily, if the security verification of the target security response measures is successful, the protection effect of the target security response measures can also be verified. If the target security baseline of the target security response measures meets the predefined security baseline, it indicates that the protection effect verification is successful, and the target security response measures can be deployed to the data source. If the target security baseline does not meet the predefined security baseline, it indicates that the protection effect verification fails, and the target security response measures can be adjusted to obtain the adjusted security response measures, and the security baseline of the adjusted security response measures meets the predefined security baseline. In this way, the adjusted security response measures can be updated as the target security response measures and deployed to the data source.

[0150] For example, a backup security network can be additionally deployed to verify the protection effect of the target security response measure. If the verification of the protection effect of the target security response measure fails, the backup security network adjusts the target security response measure to obtain an adjusted security response measure.

[0151] For example, a predefined security baseline can be preconfigured. For instance, the predefined security baseline indicates that the key length of encryption algorithm A is 1024 bits. Based on this, if the target security baseline of the target security response measure indicates the use of encryption algorithm A and the key length is 256 bits, then the target security baseline does not meet the predefined security baseline, and the target security response measure needs to be adjusted. For example, the security baseline of the adjusted security response measure indicates the use of encryption algorithm A and the key length is 1024 bits. If the target security baseline of the target security response measure indicates the use of encryption algorithm A and the key length is 1024 bits, then the target security baseline meets the predefined security baseline.

[0152] For example, if the target security baseline does not meet the predefined security baseline, this target security response measure can be deleted from multiple security response measures. When multiple security response measures are input to the second multi-modal large model again, this target security response measure is not included, automatically eliminating obsolete rules and releasing storage resources.

[0153] In the above process, the security baseline can represent the degree of security protection of the security response measure for data. Obviously, the higher the security baseline, the higher the degree of security protection. For example, the security protection degree of a key length of 1024 (i.e., the security baseline) is higher than that of a key length of 256.

[0154] Exemplarily, for step 507, the target security response measure can be deployed to each data source corresponding to the target data stream. In this way, for each data source, the data of this data source can be securely processed based on the target security response measure. For example, if the target security response measure indicates using a post-quantum encryption security policy to securely process data, the data source performs post-quantum encryption on the data.

[0155] So far, the security risk response process is completed. Finally, the target security response measure corresponding to the target data stream is confirmed and deployed to each data source corresponding to the target data stream.

[0156] Fourth, regarding the security risk detection and effect evaluation process. In the security risk detection and effect evaluation process, the effect of data security risk governance can be evaluated and blockchain traceability can be performed. See Figure 6 As shown, it is a flowchart of the security risk detection and effect evaluation process. This process can include: Step 601: Obtain a first data set and a second data set corresponding to the target data stream. The first data set may include the raw data generated by each data source corresponding to the target data stream, and the second data set may include the security data generated by each data source corresponding to the target data stream. The security data generated by a data source is obtained after the data source performs security processing on the raw data based on the target security response measures.

[0157] Exemplarily, for each data source corresponding to the target data stream, the data source may generate raw data (i.e., multimodal raw data), and the raw data generated by the data source may be added to the first data set.

[0158] After deploying the target security response measures to the data source, the data source performs security processing (such as data desensitization, etc.) on the raw data based on the target security response measures to obtain the data after security processing, denoted as security data, and the security data generated by the data source may be added to the second data set.

[0159] Step 602: Input the first data set into the third multimodal large model to obtain a detection result, which indicates whether the raw data in the first data set is sensitive data or not. Obtain a third data set corresponding to the target data stream, and the third data set may include the sensitive data in the first data set.

[0160] Referring to step 302, the third multimodal large model can be used to select sensitive data from all multimodal raw data as multimodal sensitive data. Similarly, the third multimodal large model can be used to select sensitive data from all the raw data in the first data set, and these sensitive data can form the third data set.

[0161] Step 603: Determine the first type of data stream with security risks based on the third data set, and obtain the risk assessment information of the first type of data stream based on the first multimodal large model. The risk assessment information may include security information, risk category, and risk degree. Obtain a fourth data set corresponding to the target data stream, and the fourth data set may include the sensitive data and risk assessment information corresponding to the first type of data stream.

[0162] For example, the third data set may include multiple sensitive data. Referring to steps 303 - 308, the first type of data stream with security risks can be determined based on the third data set. Referring to steps 401 - 404, the risk assessment information of each first type of data stream can be obtained based on the first multimodal large model. In this way, for each first type of data stream, the first type of data stream corresponds to multiple sensitive data and risk assessment information, and the fourth data set includes the sensitive data and risk assessment information corresponding to each first type of data stream.

[0163] Step 604: Determine the fuzzy evaluation scores of the first data set, the second data set, the third data set, and the fourth data set.

[0164] Exemplarily, the first data set may include multiple pieces of original data. Feature extraction can be performed on these original data to obtain the first feature matrix of the first data set. This embodiment does not limit the feature extraction method. The second data set may include multiple pieces of security data. Feature extraction can be performed on these security data to obtain the second feature matrix of the second data set. The third data set may include multiple pieces of sensitive data. Feature extraction can be performed on these sensitive data to obtain the third feature matrix of the third data set. The fourth data set may include the sensitive data corresponding to the first type of data stream and risk assessment information. Feature extraction can be performed on these sensitive data and risk assessment information to obtain the fourth feature matrix of the fourth data set.

[0165] Based on the first feature matrix, the second feature matrix, the third feature matrix, and the fourth feature matrix, the fuzzy evaluation scores of the first data set, the second data set, the third data set, and the fourth data set can be determined using the target evaluation algorithm. For example, the target evaluation algorithm may include, but is not limited to, the TOPSIS evaluation algorithm, the AHP evaluation algorithm, the fuzzy comprehensive evaluation algorithm, etc. This embodiment does not limit the target evaluation algorithm. The fuzzy comprehensive evaluation algorithm will be taken as an example hereinafter.

[0166] For example, the first feature matrix, the second feature matrix, the third feature matrix, and the fourth feature matrix can be used as the input of the fuzzy comprehensive evaluation algorithm. The fuzzy comprehensive evaluation algorithm combines multiple fuzzy factors to obtain an evaluation result, and then outputs the fuzzy evaluation scores of the first data set, the second data set, the third data set, and the fourth data set.

[0167] Based on the first feature matrix, the second feature matrix, the third feature matrix, and the fourth feature matrix, during the processing of the fuzzy comprehensive evaluation algorithm, the following aspects may be involved: 1) Determine the evaluation factor set and hierarchically construct indicators: First-level indicators: Each stage of data security risk governance. Second-level indicators: The specific security effects under each stage, where the stage refers to the above 4 data sets. 2) Set the comment set (V), for example, V = {excellent, good, average, poor, very poor}, corresponding to the score intervals. 3) Determine the weight set (A), using the analytic hierarchy process or the expert scoring method. For example, the weight of stage 1 is 0.2, the weight of stage 2 is 0.15, etc. Through the consistency test, ensure that the weight distribution is reasonable (CR < 0.1). 4) Construct the membership degree matrix (R). For single-factor evaluation, for numerical indicators: Calculate the membership degree through the membership function. For qualitative indicators: Score according to the comment set and count the distribution ratio. 5) Use the weighted average model to calculate the comprehensive membership degree. Calculate layer by layer, first synthesize the second-level indicators of each stage, and then synthesize the first-level indicators. 6) Defuzzification and result analysis, using the weighted average method or the maximum membership degree principle, and finally output a report to obtain the scores of each stage, that is, obtain the fuzzy evaluation scores of the first data set, the second data set, the third data set, and the fourth data set.

[0168] The above is an example of obtaining the fuzzy evaluation score using the fuzzy comprehensive evaluation algorithm, and there is no limitation on this.

[0169] In summary, it can be seen that by establishing an intelligent data security risk detection mechanism and using situation awareness technology to predict, identify, isolate, report, and dispose of data security incidents, it is possible to collect the actual project data (i.e., the above security data) after the enterprise's data security risk governance for effect evaluation.

[0170] In a possible implementation manner, after obtaining the fuzzy evaluation scores of each data set, an adjustment object can be determined based on the fuzzy evaluation scores of each data set, and then the adjustment object can be optimized.

[0171] Exemplarily, if the fuzzy evaluation score of the first data set is less than the first threshold (which can be configured according to experience), then the collection method of the multi-modal raw data (i.e., the adjustment object) can be adjusted, and the adjusted collection method is used to obtain the multi-modal raw data. For example, in step 301, the multi-modal raw data can be collected through method A and method B. Based on the adjusted collection method, the multi-modal raw data can be collected through method A, method B, and method C, or through method A and method C. There is no limitation on the adjustment process of this collection method. If the fuzzy evaluation score of the first data set is not less than the first threshold, then the collection method of the multi-modal raw data does not need to be adjusted.

[0172] Exemplarily, if the fuzzy evaluation score of the second data set is less than the second threshold (which can be configured according to experience), then the second multi-modal large model (i.e., the object to be adjusted) can be adjusted (i.e., optimize the performance of the second multi-modal large model, and the optimization process of the second multi-modal large model is not limited), and the adjusted second multi-modal large model is used to determine the target security response measure. If the fuzzy evaluation score of the second data set is not less than the second threshold, then the second multi-modal large model may not be adjusted.

[0173] Exemplarily, if the fuzzy evaluation score of the third data set is less than the third threshold (which can be configured according to experience), then the third multi-modal large model (i.e., the object to be adjusted) can be adjusted (i.e., optimize the performance of the third multi-modal large model), and the adjusted third multi-modal large model is used to identify sensitive data. If the fuzzy evaluation score of the third data set is not less than the third threshold, then the third multi-modal large model is not adjusted.

[0174] Exemplarily, if the fuzzy evaluation score of the fourth data set is less than the fourth threshold (which can be configured according to experience), then the first multi-modal large model (i.e., the object to be adjusted) is adjusted (i.e., optimize the performance of the first multi-modal large model), and the adjusted first multi-modal large model is used to determine the risk assessment information. If the fuzzy evaluation score of the fourth data set is not less than the fourth threshold, then the first multi-modal large model is not adjusted.

[0175] In a possible implementation manner, blockchain traceability can also be supported. For example, data inputs such as data security risk identification, assessment, response, detection, and effectiveness evaluation, model information (type, version, user, usage time, usage frequency, resource consumption), data outputs, and behavior processing are all recorded on the blockchain. In addition, the model operator uploads the key information and hash value to the blockchain, and the risk detector downloads the key information from the blockchain to evaluate the effectiveness of the entire process of data security risk governance, so as to continuously optimize the enabling effect of the multi-modal large model on data security risk governance.

[0176] For the blockchain traceability process, the input data of the third multi-modal large model, the output data of the third multi-modal large model, and the model information of the third multi-modal large model can be recorded on the blockchain. The multi-modal sensitive data corresponding to each first type of data stream, the output data of the first multi-modal large model, and the model information of the first multi-modal large model can be recorded on the blockchain. The input data of the second multi-modal large model, the output data of the second multi-modal large model, and the model information of the second multi-modal large model can be recorded on the blockchain. Of course, the above are just a few examples, and there is no limit to the data that needs to be recorded on the blockchain.

[0177] For the model information of the first multi-modal large model, the model information of the second multi-modal large model, and the model information of the third multi-modal large model, the model information may include, but is not limited to, at least one of the model type, model version, model user, model usage time, model usage frequency, and model resource consumption.

[0178] Exemplarily, before performing each operation, the input data obtained from this operation can also be compared with the data recorded on the blockchain. If the two are consistent, this operation can be executed; if the two are inconsistent, this operation is not executed, and an exception warning message is output. For example, for step 303, after obtaining multiple multi-modal sensitive data, it is determined whether the obtained multiple multi-modal sensitive data is the same as the output data of the third multi-modal large model recorded on the blockchain. If so, step 303, the operation of generating multiple data groups based on the multiple multi-modal sensitive data, is executed; if not, an exception warning message is output.

[0179] As can be seen from the above technical solutions, in the embodiments of the present application, for the data security protection scenario of Internet of Things enterprises, a multi-modal large model is used to perform data training and governance empowerment in stages such as security risk identification, security risk assessment, security risk response, security risk detection, and governance effect evaluation. Technologies such as threat modeling, privacy computing, post-quantum cryptography, and situation awareness are added, as well as joint governance by various departments of the enterprise, to conduct multi-dimensional and intelligent full-life-cycle governance of data security risks, reducing the data security risks of the enterprise. A multi-dimensional and multi-level data security risk governance system is constructed. By introducing a multi-modal large model and combining the enterprise's proprietary business and private domain knowledge for special training and optimization, accurate identification and quantitative assessment of risks are achieved. A data security risk response strategy based on business scenarios and data characteristics is proposed, and a multi-modal large model is used to restore the threat attack path and dynamically orchestrate and optimize data security risk response measures. A guaranteed security network is introduced to output accurate security risk response results. The whole process of risk management is recorded on the blockchain. Before processing each stage, the result record of the previous stage is first verified to ensure data integrity.

[0180] Based on the same application concept as the above method, in the embodiments of the present application, a data security risk governance device based on a multi-modal large model is proposed. Refer to Figure 7 As shown, it is a schematic structural diagram of the data security risk governance device based on the multi-modal large model. The device may include: A determination module 71, configured to generate multiple data groups based on the obtained multiple multi-modal sensitive data. For each data group, the data group includes at least one multi-modal sensitive data generated for the same security event; determine the data stream corresponding to each data group; wherein, if the data group includes N multi-modal sensitive data, the data stream includes N data sources for generating the N multi-modal sensitive data; A partitioning module 72, configured to select a first type of data stream with security risks from all data streams. For each data stream, if the data stream matches any of the risk data streams in the knowledge base, where the knowledge base includes multiple configured risk data streams, then the data stream is a first type of data stream; or if the multimodal sensitive data corresponding to the data stream matches the configured risk policy, then the data stream is a first type of data stream. An acquisition module 73, configured to, for each first type of data stream, select target multimodal sensitive data from the multimodal sensitive data corresponding to the first type of data stream, and input the target multimodal sensitive data into a first multimodal large model to obtain risk assessment information for the first type of data stream. The risk assessment information includes security information, where the security information indicates whether the first type of data stream is secure or insecure. The acquisition module 73 is configured to select a target data stream whose risk assessment information is insecure from all first type of data streams, and input the multimodal sensitive data corresponding to the target data stream and multiple security countermeasures into a second multimodal large model to obtain the target security countermeasure corresponding to the target data stream. A sending module 74, configured to deploy the target security countermeasure to the data source corresponding to the target data stream, so that the data source performs security processing on the data based on the target security countermeasure.

[0181] Exemplarily, when the acquisition module 73 selects target multimodal sensitive data from the multimodal sensitive data corresponding to each first type of data stream, it specifically is configured to: sort all first type of data streams, and determine the reference quantity of each first type of data stream based on the sorting result. When a first type of data stream is ranked higher, the reference quantity of the first type of data stream is larger. Based on the reference quantity of the first type of data stream, select the multimodal sensitive data that matches the reference quantity from the multimodal sensitive data corresponding to the first type of data stream as the target multimodal sensitive data.

[0182] Exemplarily, the obtaining module 73 is further configured to, after selecting the first type of data stream from all data streams, for each first type of data stream, obtain the portrait parameters corresponding to the first type of data stream; wherein, if the multimodal sensitive data corresponding to the first type of data stream matches the risk policy, the portrait parameters include the text description parameters of the risk policy and the text description parameters of the multimodal sensitive data corresponding to the risk policy; if the first type of data stream matches the risk data stream in the knowledge base, the portrait parameters include the text description parameters of each multimodal sensitive data corresponding to the first type of data stream; when the obtaining module 73 inputs the target multimodal sensitive data into the first multimodal large model to obtain the risk assessment information of the first type of data stream, it is specifically configured to: input the portrait parameters corresponding to the first type of data stream and the target multimodal sensitive data into the first multimodal large model to obtain the risk assessment information of the first type of data stream.

[0183] Exemplarily, the determining module 71 is further configured to obtain a plurality of multimodal raw data, where the plurality of multimodal raw data includes at least one of enterprise financial data, park security data, system operation and maintenance data, product R & D data, production data, human resources data, and employee behavior data; wherein, for each multimodal raw data, the multimodal raw data is text data, or image data, or audio data, or video data, or alarm data; select sensitive data from all multimodal raw data as multimodal sensitive data; wherein, for each multimodal raw data, input the multimodal raw data into the third multimodal large model to obtain a detection result, and the detection result indicates that the multimodal raw data is sensitive data or not sensitive data; wherein, if the multimodal raw data is text data, the text detection model in the third multimodal large model is used to process the multimodal raw data to obtain the detection result; if the multimodal raw data is image data, the image detection model in the third multimodal large model is used to process the multimodal raw data to obtain the detection result; if the multimodal raw data is audio data, the audio detection model in the third multimodal large model is used to process the multimodal raw data to obtain the detection result; if the multimodal raw data is video data, the video detection model in the third multimodal large model is used to process the multimodal raw data to obtain the detection result; if the multimodal raw data is alarm data, the alarm detection model in the third multimodal large model is used to process the multimodal raw data to obtain the detection result.

[0184] Exemplarily, the obtaining module 73 is further configured to train a first multimodal large model; when the obtaining module 73 trains the first multimodal large model, it is specifically configured to: obtain security risk assessment training data, and train the first multimodal large model based on the security risk assessment training data; wherein, the security risk assessment training data includes internal data and external data, the internal data is data running within the enterprise private network, and the external data is data running within the external network, and the external data includes security events and defect information; or, when receiving a data acquisition request from the model party training server, perform identity authentication and permission authentication on the model party training server; if the identity authentication is successful and the permission authentication is successful, encrypt the internal data to obtain ciphertext internal data, and send the ciphertext internal data to the model party training server, so that the model party training server trains the first multimodal large model based on the ciphertext internal data.

[0185] Exemplarily, the second multimodal large model includes a pre-trained model, a position encoding layer, a multi-layer perceptron, and a large language model; when the obtaining module 73 inputs the multimodal sensitive data corresponding to the target data stream and multiple security response measures into the second multimodal large model to obtain the target security response measure corresponding to the target data stream, it is specifically configured to: based on the pre-trained model, process the multimodal sensitive data corresponding to the target data stream to obtain the embedding matrix of the target data stream, determine the encoded representation feature corresponding to the embedding matrix, and perform an addition operation on the encoded representation features of all target data streams to obtain a comprehensive encoded representation feature; based on the pre-trained model, for each security response measure, map each input word corresponding to the security response measure to obtain the embedding vector of the security response measure; based on the position encoding layer, perform a convolution operation on the comprehensive encoded representation feature to obtain a first intermediate feature, perform a convolution operation on the embedding vector of each security response measure to obtain a second intermediate feature, perform a fusion operation on the first intermediate feature and the second intermediate feature of each security response measure to obtain a third intermediate feature, and perform a pooling operation on the third intermediate feature to obtain a fourth intermediate feature; based on the multi-layer perceptron, project the fourth intermediate feature into the same dimensional space as the word embedding of the large language model to obtain a target feature that can be processed by the large language model; based on the large language model, process the target feature to obtain the target security response measure corresponding to the target data stream; wherein, the target security response measure is selected by the large language model from the multiple security response measures, or is generated by the large language model.

[0186] Exemplarily, the risk assessment information further includes a risk category and a risk level. The obtaining module 73 is further configured to query a measure information table based on the risk category and risk level corresponding to the target data stream, so as to obtain a plurality of candidate security response measures corresponding to the risk category and the risk level; wherein, the measure information table includes a mapping relationship between the risk category, the risk level and the plurality of candidate security response measures. If the plurality of candidate security response measures include the target security response measure, it is determined whether the target security baseline of the target security response measure meets a predefined security baseline, where the target security baseline represents the degree of security protection of the target security response measure for the data. If so, the sending module 74 deploys the target security response measure to the data source corresponding to the target data stream. If not, the target security response measure is adjusted so that the security baseline of the adjusted security response measure meets the predefined security baseline, and the adjusted security response measure is updated as the target security response measure, and the sending module 74 deploys the target security response measure to the data source corresponding to the target data stream. If the plurality of candidate security response measures do not include the target security response measure, a security response measure is generated as the target security response measure by using a candidate security policy, and the sending module 74 deploys the target security response measure to the data source corresponding to the target data stream.

[0187] Wherein, the candidate security policies include at least one of a privacy computing security policy, a post-quantum encryption security policy, a data masking security policy, a data watermarking security policy, a content recognition and control security policy, a network blocking security policy, a boundary control security policy, and a terminal protection security policy.

[0188] Exemplarily, the obtaining module 73 is further configured to obtain a first data set and a second data set corresponding to the target data stream. The first data set includes the raw data generated by each data source corresponding to the target data stream, and the second data set includes the security data generated by each data source corresponding to the target data stream. The security data generated by a data source is obtained after the data source performs security processing on the raw data based on the target security response measure. Input the first data set into the third multi-modal large model to obtain a detection result, where the detection result indicates whether the raw data in the first data set is sensitive data or not. Obtain a third data set corresponding to the target data stream, where the third data set includes the sensitive data in the first data set. Determine the first type of data stream with a security risk based on the third data set, and obtain the risk assessment information of the first type of data stream based on the first multi-modal large model. The risk assessment information includes security information, risk category, and risk degree. Obtain a fourth data set corresponding to the target data stream, where the fourth data set includes the sensitive data corresponding to the first type of data stream and the risk assessment information. Determine the fuzzy evaluation scores of the first data set, the second data set, the third data set, and the fourth data set. If the fuzzy evaluation score of the first data set is less than the first threshold, adjust the collection method of the multi-modal raw data, and the adjusted collection method is used to obtain the multi-modal raw data. If the fuzzy evaluation score of the second data set is less than the second threshold, adjust the second multi-modal large model, and the adjusted second multi-modal large model is used to determine the target security response measure. If the fuzzy evaluation score of the third data set is less than the third threshold, adjust the third multi-modal large model, and the adjusted third multi-modal large model is used to identify sensitive data. If the fuzzy evaluation score of the fourth data set is less than the fourth threshold, adjust the first multi-modal large model, and the adjusted first multi-modal large model is used to determine the risk assessment information.

[0189] Based on the same application concept as the above method, an electronic device is proposed in an embodiment of the present application. Refer to Figure 8 As shown, it includes: a processor 81 and a machine-readable storage medium 82. The machine-readable storage medium 82 stores machine-executable instructions that can be executed by the processor 81. The processor 81 is configured to execute the machine-executable instructions to implement the data security risk governance method based on the multi-modal large model in the above example.

[0190] Based on the same application concept as the above method, an embodiment of the present application further provides a machine-readable storage medium. A number of computer instructions are stored on the machine-readable storage medium. When the computer instructions are executed by a processor, the data security risk governance method based on the multi-modal large model in the above example can be implemented.

[0191] Among them, the above-mentioned machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, and so on. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.

[0192] Based on the same application concept as the above method, an embodiment of the present application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, can implement the data security risk governance method based on a multi-modal large model in the above example.

[0193] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0194] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A data security risk management method based on a multimodal large model, characterized in that: include: Generate multiple data groups based on the acquired multiple multimodal sensitive data, and for each data group, the data group includes at least one multimodal sensitive data generated for the same security event; Determine a data stream corresponding to each data group; wherein, if the data group includes N multimodal sensitive data, the data stream includes N data sources for generating the N multimodal sensitive data; Selecting a first-category data flow with security risks from all data flows; wherein, for each data flow, if the data flow matches any risk data flow in the knowledge base, the data flow is a first-category data flow, and the knowledge base includes a plurality of configured risk data flows; if the multimodal sensitive data corresponding to the data flow matches the configured risk policy, the data flow is a first-category data flow; For each first-category data stream, select target multimodal sensitive data from the multimodal sensitive data corresponding to the first-category data stream, input the target multimodal sensitive data into the first multimodal large model, and obtain risk assessment information of the first-category data stream, wherein the risk assessment information is safe or unsafe; Selecting a target data stream whose risk assessment information is unsafe from all first-category data streams, inputting multimodal sensitive data and multiple security countermeasures corresponding to the target data stream into a second multimodal large model, and obtaining a target security countermeasure corresponding to the target data stream; The target security countermeasure is deployed to a data source corresponding to the target data stream, so that the data source performs security processing on the data based on the target security countermeasure.

2. The method according to claim 1, characterized in that The selecting target multimodal sensitive data from the multimodal sensitive data corresponding to the first type of data stream includes: Sorting all first-category data flows, and determining a reference number of each first-category data flow based on the sorting result; wherein, the higher the first-category data flow is sorted, the greater the reference number of the first-category data flow; Based on a reference quantity of the first category data stream, multimodal sensitive data matching the reference quantity is selected from the multimodal sensitive data corresponding to the first category data stream as the target multimodal sensitive data.

3. The method according to claim 1, characterized in that After selecting the first type of data stream with security risks from all data streams, the method further includes: for each first type of data stream, obtaining a portrait parameter corresponding to the first type of data stream; wherein, if the multimodal sensitive data corresponding to the first type of data stream matches the risk strategy, the portrait parameter includes a text description parameter of the risk strategy and a text description parameter of the multimodal sensitive data corresponding to the risk strategy; if the first type of data stream matches the risk data stream in the knowledge base, the portrait parameter includes a text description parameter of each multimodal sensitive data corresponding to the first type of data stream; The target multimodal sensitive data is input into the first multimodal large model to obtain risk assessment information of the first type of data stream, including: the portrait parameters corresponding to the first type of data stream and the target multimodal sensitive data are input into the first multimodal large model to obtain risk assessment information of the first type of data stream.

4. The method according to claim 1, characterized in that: Before generating a plurality of data groups based on the acquired plurality of multimodal sensitive data, the method further includes: Acquire multiple multimodal raw data, wherein the multiple multimodal raw data include at least one of enterprise financial data, park security data, system operation and maintenance data, product development data, production data, human data, and employee behavior data; wherein, for each multimodal raw data, the multimodal raw data is text data, or image data, or audio data, or video data, or alarm data; Selecting sensitive data from all multimodal original data as multimodal sensitive data; wherein, for each multimodal original data, the multimodal original data is input into the third multimodal large model to obtain a detection result, wherein the detection result indicates whether the multimodal original data is sensitive data or not sensitive data; Wherein, if the multimodal original data is text data, the multimodal original data is processed by the text detection model in the third multimodal large model to obtain the detection result; If the multimodal raw data is image data, the multimodal raw data is processed by the image detection model in the third multimodal large model to obtain the detection result; If the multimodal original data is audio data, the multimodal original data is processed by the audio detection model in the third multimodal large model to obtain the detection result; If the multimodal original data is video data, the multimodal original data is processed by the video detection model in the third multimodal large model to obtain the detection result; If the multimodal original data is alarm data, the multimodal original data is processed by the alarm detection model in the third multimodal large model to obtain the detection result.

5. The method according to claim 1, characterized in that: The training process of the first multimodal large model specifically includes: Acquire security risk assessment training data, and train the first multimodal large model based on the security risk assessment training data; wherein the security risk assessment training data includes internal data and external data, the internal data is data running in the enterprise private network, the external data is data running in the external network, and the external data includes security events and defect information; or, Upon receiving a data acquisition request from the model training server, identity authentication and authority authentication are performed on the model training server; if the identity authentication is successful and the authority authentication is successful, the internal data is encrypted to obtain ciphertext internal data, and the ciphertext internal data is sent to the model training server, so that the model training server trains the first multimodal large model based on the ciphertext internal data.

6. The method according to claim 1, characterized in that The second multimodal large model includes a pre-trained model, a position encoding layer, a multi-layer perceptron and a large language model; The step of inputting the multimodal sensitive data and multiple security countermeasures corresponding to the target data stream into the second multimodal large model to obtain the target security countermeasures corresponding to the target data stream includes: Based on the pre-trained model, the multimodal sensitive data corresponding to the target data stream is processed to obtain an embedding matrix of the target data stream, a coding representation feature corresponding to the embedding matrix is ​​determined, and the coding representation features of all target data streams are added to obtain a comprehensive coding representation feature; Based on the pre-trained model, for each safety countermeasure, each input word corresponding to the safety countermeasure is mapped to obtain an embedding vector of the safety countermeasure; Based on the position encoding layer, a convolution operation is performed on the comprehensive coding representation feature to obtain a first intermediate feature, a convolution operation is performed on the embedding vector of each safety countermeasure to obtain a second intermediate feature, a fusion operation is performed on the first intermediate feature and the second intermediate feature of each safety countermeasure to obtain a third intermediate feature, and a pooling operation is performed on the third intermediate feature to obtain a fourth intermediate feature; Based on the multilayer perceptron, projecting the fourth intermediate feature to the same dimensional space as the word embedding of the large language model to obtain a target feature that can be processed by the large language model; Based on the large language model, the target features are processed to obtain target security countermeasures corresponding to the target data stream; wherein the target security countermeasures are selected by the large language model from the multiple security countermeasures, or generated by the large language model.

7. The method according to claim 1, characterized in that The risk assessment information also includes risk category and risk degree. After inputting the multimodal sensitive data and multiple security countermeasures corresponding to the target data stream into the second multimodal large model and obtaining the target security countermeasures corresponding to the target data stream, the method further includes: Based on the risk category and risk degree corresponding to the target data flow, a measure information table is queried to obtain a plurality of candidate security countermeasures corresponding to the risk category and the risk degree; wherein the measure information table includes a mapping relationship between the risk category, the risk degree and the plurality of candidate security countermeasures; If the multiple candidate security countermeasures include the target security countermeasure, determine whether the target security baseline of the target security countermeasure meets the predefined security baseline, and the target security baseline represents the degree of security protection of the data by the target security countermeasure; if so, execute the operation of deploying the target security countermeasure to the data source corresponding to the target data stream; if not, adjust the target security countermeasure, the security baseline of the adjusted security countermeasure meets the predefined security baseline, and update the adjusted security countermeasure to the target security countermeasure, and execute the operation of deploying the target security countermeasure to the data source corresponding to the target data stream; If the multiple candidate security countermeasures do not include the target security countermeasure, a security countermeasure is generated using the candidate security policy as the target security countermeasure, and an operation of deploying the target security countermeasure to the data source corresponding to the target data flow is performed; Among them, the candidate security policies include at least one of privacy computing security policy, post-quantum encryption security policy, data desensitization security policy, data watermark security policy, content identification and management security policy, network blocking security policy, border control security policy, and terminal protection security policy.

8. The method according to claim 4, characterized in that After deploying the target security countermeasure to the data source corresponding to the target data flow so that the data source performs security processing on the data based on the target security countermeasure, the method further includes: Acquire a first data set and a second data set corresponding to the target data stream, wherein the first data set includes original data generated by each data source corresponding to the target data stream, and the second data set includes security data generated by each data source corresponding to the target data stream, and the security data generated by the data source is obtained after the data source performs security processing on the original data based on the target security countermeasure; Inputting the first data set into a third multimodal large model to obtain a detection result, wherein the detection result indicates whether the original data in the first data set is sensitive data or not sensitive data; obtaining a third data set corresponding to the target data stream, wherein the third data set includes the sensitive data in the first data set; Determine a first type of data flow with security risks based on the third data set, and obtain risk assessment information of the first type of data flow based on the first multimodal large model, wherein the risk assessment information includes security information, risk category, and risk degree; obtain a fourth data set corresponding to the target data flow, wherein the fourth data set includes sensitive data corresponding to the first type of data flow and the risk assessment information; Determining a fuzzy evaluation score of the first data set, a fuzzy evaluation score of the second data set, a fuzzy evaluation score of the third data set, and a fuzzy evaluation score of the fourth data set; If the fuzzy evaluation score of the first data set is less than a first threshold, adjusting the collection method of the multimodal original data, and the adjusted collection method is used to obtain the multimodal original data; If the fuzzy evaluation score of the second data set is less than a second threshold, the second multimodal large model is adjusted, and the adjusted second multimodal large model is used to determine the target security countermeasures; If the fuzzy evaluation score of the third data set is less than a third threshold, the third multimodal large model is adjusted, and the adjusted third multimodal large model is used to identify sensitive data; If the fuzzy evaluation score of the fourth data set is less than a fourth threshold, the first multimodal large model is adjusted, and the adjusted first multimodal large model is used to determine the risk assessment information.

9. The method according to claim 4, characterized in that The method further comprises: Recording the input data of the third multimodal large model, the output data of the third multimodal large model, and the model information of the third multimodal large model in the blockchain; Recording the multimodal sensitive data corresponding to each first-category data stream, the output data of the first multimodal large model, and the model information of the first multimodal large model in the blockchain; Recording the input data of the second multimodal large model, the output data of the second multimodal large model, and the model information of the second multimodal large model in a blockchain; Before generating multiple data groups based on the multiple multimodal sensitive data that have been obtained, the method further includes: determining whether the multiple multimodal sensitive data that have been obtained are the same as the output data of the third multimodal large model recorded in the blockchain, and if so, performing an operation of generating multiple data groups based on the multiple multimodal sensitive data that have been obtained, and if not, outputting abnormal alarm information; The model information includes at least one of model type, model version, model user, model usage time, model usage frequency, and model resource consumption.

10. A data security risk management device based on a multimodal large model, characterized in that: include: A determination module, configured to generate a plurality of data groups based on the plurality of multimodal sensitive data that have been acquired, wherein each data group includes at least one multimodal sensitive data generated for a same security event; Determine a data stream corresponding to each data group; wherein, if the data group includes N multimodal sensitive data, the data stream includes N data sources for generating the N multimodal sensitive data; A partitioning module, configured to select a first-category data flow with security risks from all data flows; wherein, for each data flow, if the data flow matches any risk data flow in a knowledge base, the data flow is a first-category data flow, and the knowledge base includes a plurality of configured risk data flows; if the multimodal sensitive data corresponding to the data flow matches a configured risk strategy, the data flow is a first-category data flow; an acquisition module, configured to select, for each first-category data flow, target multimodal sensitive data from the multimodal sensitive data corresponding to the first-category data flow, input the target multimodal sensitive data into the first multimodal large model, and obtain risk assessment information of the first-category data flow; wherein the risk assessment information includes security information, and the security information indicates whether the first-category data flow is secure or unsecure; The acquisition module is used to select a target data stream whose risk assessment information is unsafe from all first-category data streams, input the multimodal sensitive data and multiple security countermeasures corresponding to the target data stream into the second multimodal large model, and obtain the target security countermeasures corresponding to the target data stream; The sending module is used to deploy the target security countermeasure to the data source corresponding to the target data flow, so that the data source performs security processing on the data based on the target security countermeasure.

Citation Information

Patent Citations

  • Business data flow security risk analysis method and system, storage medium and terminal

    CN116506217A

  • Security risk dynamic assessment system and method based on multi-source heterogeneous data analysis

    CN118898397A

  • Data fusion correlation analysis method, electronic device, readable medium and program product

    CN119066615A

  • Security event management method and device

    CN119830272A

  • Systems and methods for cybersecurity risk assessment

    US20180146004A1