Data security risk monitoring method and device, equipment and storage medium

By combining the RPA engine with the security risk identification model, automated data collection and multimodal analysis solve the problems of manual dependence and unstructured data analysis in existing technologies, and achieve efficient and accurate data security monitoring.

CN120639346APending Publication Date: 2025-09-12SHENZHEN YUEHUA EXPRESS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510684130.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing data security monitoring solutions rely on manually written code scripts, which are difficult to adapt to new channel iterations, cannot parse unstructured data, and have a high false alarm rate of early warning rules.

Method used

The RPA engine is used to automate data collection, combined with the early warning rule engine and security risk identification model for data screening and analysis, and the pre-trained language model of the Transformer architecture is used to fuse image and text detection models to achieve multimodal data analysis.

Benefits of technology

It realizes the intelligent closed loop of data collection and risk identification, reduces labor costs, improves data monitoring efficiency and risk identification accuracy, and reduces false alarm rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639346A_ABST
    Figure CN120639346A_ABST
Patent Text Reader

Abstract

The invention discloses a data security risk monitoring method and device, electronic equipment and a storage medium, and the method comprises the steps: executing an RPA engine, and obtaining a suspicious data set of a target network source according to a preset task arrangement process; sequentially screening and analyzing the suspicious data set by adopting a constructed early warning rule engine and a security risk identification model, and identifying risk data; and processing the risk data. By using the method disclosed by the invention, the whole-process intelligent closed loop from data acquisition to risk disposal can be realized, the labor cost of data acquisition is saved, and the data monitoring efficiency and the risk identification accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data security prevention and control technology, and in particular to a data security risk monitoring method, device, equipment and storage medium. Background Art

[0002] In the digital age, many businesses require extensive expansion through internet channels. However, with the proliferation of cyberattack methods and the continuous advancement of hacker technology, the risk of data breaches facing businesses has increased dramatically. A data breach not only causes direct financial losses but also severely damages a company's reputation and customer trust. Therefore, a comprehensive internet channel data analysis and early warning mechanism is essential to ensure the stable development of businesses in the internet environment.

[0003] However, existing data security monitoring solutions have shortcomings. The main ones are that the capture of monitoring data relies on manually written code scripts, and continuous operation and iteration to adapt to new channels requires a large amount of R&D costs; the analysis of monitoring data only supports the processing of structured data and cannot parse unstructured data such as images, such as screenshots containing sensitive information; and the warning rules generated based on human experience have a high false alarm rate. Summary of the Invention

[0004] The present invention provides a data security risk monitoring method, device, equipment and storage medium to solve the problems of existing data security monitoring solutions, such as strong dependence on manual monitoring data collection, lack of multimodal support for data analysis, and low accuracy in data risk identification.

[0005] In order to solve the above technical problems, in a first aspect, the present invention provides a data security risk monitoring method, comprising:

[0006] Execute the RPA engine and obtain suspicious data sets from the target network source according to the preset task orchestration process;

[0007] Using the established early warning rule engine and security risk identification model to screen and analyze the suspicious data sets in turn to identify risky data;

[0008] The risk data is processed.

[0009] Optionally, the step of obtaining a suspicious data set of a target network source according to a preset task arrangement process includes:

[0010] Accessing a target web page according to the pre-marked address of the target network source;

[0011] Performing interactive operations according to the interactive operations pre-marked in the interactive area of ​​the target webpage, and collecting webpage content fed back from the interactive operations on the target webpage;

[0012] Data is extracted from the webpage content to obtain a suspicious data set.

[0013] Optionally, the early warning rule engine and security risk identification model constructed are used to screen and analyze the suspicious data set in sequence to identify risk data, including:

[0014] Using an early warning rule engine, the detection rules formed by the alarm feature information are used to filter out the explicit risk content in the suspicious data set;

[0015] The trained security risk identification model is used to analyze the filtered suspicious data set and output risk data including security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information.

[0016] Optionally, the security risk identification model uses a pre-trained language model based on the Transformer architecture as the main model, and integrates a general large model for image detection, a general large model for text detection, and a multimodal general large model; the security risk identification model adopts a multi-task learning framework, including a shared encoder, and multiple task-specific heads including a security risk classification head, a sentiment analysis head, an entity recognition head, and an image analysis head for performing specific task outputs.

[0017] Optionally, the trained security risk identification model is used to analyze the filtered suspicious data set input, and output risk data including security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information, including:

[0018] The filtered suspicious data set is input into the trained security risk identification model, and common semantic features are extracted by the shared encoder; the common semantic features are respectively input into the security risk classification head, sentiment analysis head, entity recognition head and image analysis head, and the corresponding output includes risk data of security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information.

[0019] Optionally, the security risk identification model constructs the following loss function:

[0020] Loss = λ1Loss_Classification+λ2Loss_Sentiment+λ3Loss_Entity+λ4R

[0021] In the formula, Loss represents the total loss function, Loss_classification is the loss function of the security risk classification task, and λ1 is its weight; Loss_sentiment is the loss function of the sentiment analysis task, and λ2 is its weight; Loss_entity is the loss function of the entity localization task, and λ3 is its weight; R represents the regularization term, and λ4 is its weight.

[0022] Optionally, the risk data is processed, including generating a security alert and notifying the corresponding person in charge within a preset time according to the risk level.

[0023] In a second aspect, the present invention provides a data security risk monitoring device, comprising a suspicious data acquisition unit, a data analysis and processing unit, and a risk handling unit;

[0024] The suspicious data collection unit is used to execute the RPA engine and obtain the suspicious data set of the target network source according to the preset task arrangement process;

[0025] The data analysis and processing unit is used to use the constructed early warning rule engine and security risk identification model to screen and analyze the suspicious data set in sequence to identify risky data;

[0026] The risk handling unit is used to handle the risk data.

[0027] In a third aspect, the present invention provides a data security risk monitoring device, comprising a memory and a processor,

[0028] Wherein, the memory is used to store computer programs;

[0029] The processor is used to read the program in the memory and execute the steps of a data security risk monitoring method provided in the first aspect above.

[0030] In a fourth aspect, the present invention provides a computer-readable storage medium having a readable computer program stored thereon, which, when executed by a processor, implements the steps of a data security risk monitoring method provided in the first aspect above.

[0031] Compared with the prior art, the data security risk monitoring method, device, equipment, and storage medium provided by the present invention have the following beneficial effects:

[0032] The RPA engine executes, acquiring suspicious data sets from the target network source according to a pre-set task orchestration process. The pre-built early warning rule engine and security risk identification model are then used to sequentially screen and analyze these suspicious data sets, identifying and processing risky data. This technical solution, combined with the RPA engine that collects data according to a pre-set task coding process and a dual data analysis mechanism constructed from the early warning rule engine and security risk identification model, identifies and processes risky data in the collected suspicious data sets. This creates an intelligent, closed-loop process from data collection, analysis and identification, to risk management, saving labor costs for data collection and improving data monitoring efficiency and risk identification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only part of the embodiments of the present invention, rather than all the embodiments. For ordinary technicians in this field, without paying any creative work, other drawings obtained based on these drawings are all within the scope of protection of this application.

[0034] Figure 1 is a flow chart of a data security risk monitoring method provided by an embodiment of the present invention;

[0035] Figure 2 This is a diagram of the overall architecture of the system implementation of the data security risk monitoring solution provided by an embodiment of the present invention;

[0036] Figure 3 This is a schematic diagram of the RPA engine user interface for a defined task process, provided by an embodiment of the present invention;

[0037] Figure 4 This is a schematic diagram of a process for obtaining a suspicious data set from a target network source according to a preset task arrangement process provided by an embodiment of the present invention;

[0038] Figure 5 This is a flow chart of using a warning rule engine and a security risk identification model to sequentially screen and analyze the suspicious data set to identify risk data, as provided by an embodiment of the present invention;

[0039] Figure 6 This is a data security risk monitoring device provided by an embodiment of the present invention;

[0040] Figure 7 Schematic diagram of the structure of a data security risk monitoring device provided by an embodiment of the present invention;

[0041] Figure 8 It is a schematic diagram of the structure of a computer-readable storage medium provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0043] In order to make the description of the present disclosure more detailed and complete, the following is an illustrative description of the implementation methods and specific examples of the present invention; however, this is not the only form of implementing or using the specific embodiments of the present invention. The implementation methods cover the features of multiple specific embodiments and the method steps and their sequence for constructing and operating these specific embodiments. However, other specific embodiments can also be used to achieve the same or equal functions and step sequences. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0044] Example 1

[0045] like Figure 2 The figure shows the overall architecture diagram of the system implementation of the data security risk monitoring solution provided by the embodiment of the present invention. The overall architecture of the system includes a data collection layer, a data analysis layer and a risk disposal layer. Among them, in the data collection layer, the RPA engine process orchestration center arranges processes according to preset tasks, and can collect suspicious data sets from multiple target network sources including search engines, Internet forum social platforms, dark webs, wool sharing, and e-commerce platforms; the data analysis layer, through big data calculation and analysis, adopts a hierarchical analysis mechanism that combines an early warning rule engine and a security risk identification model to identify risks of the suspicious data sets; the risk disposal layer issues data security risk warnings for identified risks, and can take measures to close the loop of events and conduct risk situation statistics.

[0046] like Figure 1 FIG. 1 is a flow chart of a data security risk monitoring method provided by an embodiment of the present invention, including:

[0047] Step S101: Execute the RPA engine to obtain a suspicious data set from the target network source according to the preset task orchestration process;

[0048] The RPA (Robotic Process Automation) engine is an open source robotic process automation platform that allows users to freely build and customize RPA processes according to their needs and scenarios. The RPA engine provides a visual graphical interface to facilitate users to quickly build automated processes. In a specific embodiment, Figure 3As shown, it is a schematic diagram of the user operation page of the RPA engine for which a certain task process has been formulated according to an embodiment of the present invention. The RPA engine can have various built-in toolboxes, including web automation tools, AI tools, data processing tools, external calls, and process management. According to the preset task process formulated by the user in advance, combined with the toolbox, the RPA engine is executed to complete the acquisition of the suspicious data set of the target network source. When setting the preset process, according to the need to access the final target web page of the target network source, an interactive operation process needs to be formulated in the middle, so as to realize the collection of the suspicious data set of the target web page.

[0049] The suspicious data set is obtained by accessing the target web page. Specifically, according to the web page address marked by the user in advance, the Chrome kernel browser is automatically called to open the web page, and then the interactive operations marked by the user in the interactive area of ​​the web page by clicking the mouse, inputting, etc. are automatically executed, such as automatically filling in the content required by the user, performing click operations, etc., which may include filling in user name, password, verification code and other information, clicking login operations, clicking search operations, etc.

[0050] As an optional implementation, Figure 4 As shown, the process of obtaining a suspicious data set of a target network source according to a preset task arrangement process includes:

[0051] Step S1011, accessing a target web page according to the pre-marked address of the target network source;

[0052] For example, the web page opened corresponding to the target network source address is often not the target web page that needs to collect suspicious data. The RPA engine first opens the web page according to the target network source address pre-marked by the user when formulating the task scheduling process; then, according to the interactive operations marked by the user in the interactive area of ​​the web page through input, mouse click, etc. when formulating the task scheduling process, the RPA engine automatically executes the interactive operations, such as automatically filling in the content, executing click operations, etc. Figure 3As shown, the process for accessing the target web page in this task orchestration process involves opening the initial web page of the target network source based on the web address, then performing login operations, including entering the account and password, recognizing an image / text verification code, and clicking to confirm. During execution, the RPA engine uses the web automation tools in its built-in toolbox to perform this process. Specifically, it performs the following functions: opening the web page, automatically launching the Chrome browser based on the URL pre-marked by the user, opening the web page, and also supporting closing the currently open browser page or all pages; filling in the input field, automatically filling in the user's pre-marked content based on the user's mouse click on the target input box element on the opened web page; and clicking a button, automatically simulating a mouse click on the target button element based on the user's mouse click on the opened web page. Furthermore, the RPA engine has built-in verification code recognition capabilities, including automatic recognition and verification of sliding puzzle verification codes, universal numeric and English verification codes, and image click verification codes.

[0053] Step S1012, performing interactive operations according to the interactive operations pre-marked in the interactive area of ​​the target webpage, and collecting webpage content fed back from the interactive operations on the target webpage;

[0054] For example, after opening the target web page, the interactive operation pre-marked by the user in the interactive area of ​​the target web page when formulating the task arrangement process is automatically executed, and the web page content information fed back by the target web page is collected. The above-mentioned interactive operation can be an input operation, a click operation, etc. The collected web page content information may include the web page content, web page title and web page screenshots. Figure 3 As shown, the data collection process in the task orchestration process is to enter keywords in the search input box of the target web page, click the search button, and collect the feedback web page content. During execution, the RPA engine automatically fills in the keyword content pre-marked by the user according to the input box marked and located by the user when formulating the task orchestration process, and then quickly locates the target button element mark by clicking the mouse when the user formulates the task coding process, automatically simulates the mouse click operation on the area, and the target web page feedback web page content according to the keyword. At this time, the RPA engine collects the page content, web page title and web page screenshot information of the web page. Specifically, it automatically extracts the corresponding page text by parsing the current browser page source code to obtain the web page content; it automatically extracts the corresponding <title> title to obtain the webpage title; export images by parsing the current browser page source code to obtain webpage screenshots.< / title>

[0055] Step S1013: extract data from the webpage content to obtain a suspicious data set.

[0056] For example, data extraction is performed on web content, including web pages, web titles, and web screenshots, to obtain a suspicious data set. Specifically, the RPA engine automatically extracts the corresponding data based on the elements marked and located on the web page when the user sets the task orchestration process. For web screenshots, the RPA engine automatically recognizes and extracts simple character information from the web screenshots using its built-in OCR text recognition tool. The extracted data is saved as a two-dimensional data list variable. When saving, the corresponding data variable names for fields such as the web page title, web content, and web screenshots must be specified. The extracted suspicious data set is then used as a data monitoring target and reported for risk analysis.

[0057] It should be noted that, in some embodiments, steps S1011 to S1013 may be repeatedly executed, with the extracted web page content data as the target object, and target web page access, data collection and extraction are performed cyclically. Figure 3 As shown, the RAP engine has a built-in process management tool. During the loop process, loop traversal, next loop, loop exit, conditional judgment and termination judgment can be executed as needed. Among them, loop traversal is used to traverse and repeat in sequence according to the process nodes of target web page access, data collection and extraction; next loop is used to ignore the current loop and continue to the next loop directly when looping through the process node, which is similar to the statement logic of continue; exit loop is used to loop through the process node and dynamically control the loop to end early, which is similar to the statement logic of break; conditional judgment is used to control the execution logic in process arrangement, and supports the combination logic of multiple condition lists. The conditional relationship involves all satisfaction and satisfaction of any item, which is similar to the statement logic of all and any; termination process is used to control the overall workflow, mainly to dynamically control the early termination of the entire task process, which is similar to the statement logic of return.

[0058] The embodiment of the present invention uses the RPA engine to orchestrate processes according to preset tasks, automatically access target web pages and perform interactive operations, and collect and extract the feedback web page content, thereby realizing the automatic acquisition of suspicious data sets, reducing the complexity of operation and maintenance, achieving zero-code rapid docking, and greatly improving labor efficiency.

[0059] As an optional implementation method, a task execution schedule is formulated that includes several completed task orchestration processes. When polling detects that there is a task execution plan to be executed that meets the task conditions, the RPA engine is executed. For example, after completing the formulation of a task orchestration process, a corresponding task execution plan can be created, and each task execution plan can be summarized to generate a task execution schedule. Each task execution plan in the task execution schedule can be set to perform periodic suspicious data set collection at a fixed time point or interval every day. By periodically polling the task execution schedule, when it is detected that there is a task execution plan to be executed at the current moment, the RPA engine is executed to collect suspicious data sets from the target network source according to the task orchestration process corresponding to the task execution plan.

[0060] The embodiment of the present invention improves the automation and efficiency of data collection by setting the execution time for the task execution plan of each task scheduling process and performing polling to trigger the execution of RPA, so that it collects suspicious data sets according to the corresponding task scheduling process.

[0061] In step S102 , the constructed early warning rule engine and security risk identification model are used to sequentially screen and analyze the suspicious data set to identify risky data.

[0062] Once suspicious data sets collected by the RPA engine are reported, they are analyzed and identified for risk using a layered analysis mechanism that combines an early warning rules engine with a security risk identification model. The early warning rules engine is first used to perform a rough screening analysis of the suspicious data sets. After this screening, the security risk identification model then performs a precise analysis of the suspicious data sets to identify risky data.

[0063] As an optional implementation, Figure 5 As shown, the early warning rule engine and security risk identification model constructed are used to screen and analyze the suspicious data set in turn to identify risk data, including:

[0064] Step S1021: Using an early warning rule engine, and using detection rules formed by the alarm feature information, filter out explicit risk content in the suspicious data set;

[0065] Explicitly risky content refers to content that is clearly low-risk and can generally be understood as non-risky. For example, the early warning rule engine uses the established rules for identifying and judging manual operational risk alerts to establish a blacklist / whitelist rule database. This database combines regular expressions and keyword matching to quickly filter out non-risky content.

[0066] In step S1022, the trained security risk identification model is used to analyze the input filtered suspicious data set, and output risk data including security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information.

[0067] After being quickly filtered by the early warning rule engine, the suspicious data set is input into the trained security risk identification model for further precise analysis and identification. After being processed by the security risk identification model, the output includes risk data such as security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information. For example, security risk category can include label categories such as "data leakage", "system vulnerability", and "sensitive information leakage"; negative / neutral / positive sentiment probability refers to the probability that the sentiment polarity expressed in the information is negative, neutral, and positive respectively; key entity location refers to the identification and location of key information objects, such as "for private data" and "company internal documents"; sensitive image information refers to image content containing sensitive information.

[0068] The embodiment of the present invention first quickly filters out obviously low-risk content through an early warning rule engine, and then uses a security risk identification model to analyze complex feature information such as complex semantics and image features of the filtered suspicious data set, and outputs risk data, which greatly improves the efficiency and accuracy of data risk identification.

[0069] As an optional implementation, the security risk identification model uses a pre-trained language model based on the Transformer architecture as the main model, and integrates a general large model for image detection, a general large model for text detection, and a multimodal general large model; the security risk identification model adopts a multi-task learning framework, including a shared encoder, and multiple task-specific heads including a security risk classification head, a sentiment analysis head, an entity recognition head, and an image analysis head for performing specific task outputs.

[0070] Exemplarily, the main model of the security risk identification model can adopt the RoBERTa or BERT large model. In the fused general large model, the image detection general large model can adopt the DiffusionDet model, the text detection general large model can adopt the GPT-4 or BERT-Large model, and the multimodal general large model can adopt the CLIP model. The security risk identification model adopts a multi-task learning framework that uses a shared encoder and core, integrates the Transformer's multi-head attention mechanism and lightweight convolutional modules, and realizes representation learning of multi-source data such as multi-text and images through cross-modal feature alignment. Based on the shared encoder, multiple specialized task branches are expanded, including a security risk classification head that can identify data risk types, a sentiment analysis head that can mine sentiment tendencies in text, an entity recognition head that can accurately extract key information such as product model names, file names, and company names, and an image analysis head that can understand image content and detect targets.

[0071] As an optional implementation, the trained security risk identification model is used to analyze the filtered suspicious data set and output risk data including security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information, including:

[0072] The filtered suspicious data set is input into the trained security risk identification model, and common semantic features are extracted by the shared encoder; the common semantic features are respectively input into the security risk classification head, sentiment analysis head, entity recognition head and image analysis head, and the corresponding output includes risk data of security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information.

[0073] Exemplarily, the shared encoder may use a RoBER encoder to extract common semantic features from the suspicious data set input into the security risk identification model. It should be noted that before the suspicious data set is input into the trained security risk identification model, data cleaning and data format processing may be performed first, wherein data cleaning includes removing noise data such as advertisements and duplicate content, and data format processing includes extracting text information from key areas of image data and converting unstructured text data into structured feature vectors. The common semantic features extracted by the shared encoder are respectively input into the security risk classification head, sentiment analysis head, entity recognition head and image analysis head, and the corresponding outputs include risk data of security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information.

[0074] The security risk identification model constructed in the embodiment of the present invention is based on multimodal intelligent analysis and deep semantic understanding capabilities, as well as the self-learning ability of large models. It not only realizes the analysis and detection capabilities of unstructured data, but also breaks through the static threshold limitations of traditional rule engines. It can realize dynamic analysis of complex attack chains, adapt to the risks of new attack methods, and greatly reduce the false alarm rate compared with static rule detection.

[0075] Exemplarily, during the training of the security risk identification model, in the preparation stage, data can be collected by combining public corpora and historical enterprise data, and after cleaning and filtering the collected data, data annotation and enhancement work can be performed. Among them, data annotation is performed for specific tasks of the model, including security risk category annotation, sentiment tendency annotation, key entity annotation and sensitive icon annotation. During data enhancement processing, back translation and synonym replacement processing can be used for small sample data, the data set size can be expanded, and encryption and decryption simulation training can be performed on some data to improve the model's recognition ability of steganography. The training stage includes pre-training and fine-tuning. Among them, during fine-tuning, task fine-tuning can be performed for security risk classification tasks and sentiment analysis tasks.

[0076] As an optional implementation, the security risk identification model constructs the following loss function:

[0077] Loss = λ1Loss_Classification+λ2Loss_Sentiment+λ3Loss_Entity+λ4R

[0078] In the formula, Loss represents the total loss function, Loss_Classification is the loss function for the security risk classification task, and λ1 is its weight; Loss_Sentiment is the loss function for the sentiment analysis task, and λ2 is its weight; Loss_Entity is the loss function for the entity localization task, and λ3 is its weight; R represents the regularization term, and λ4 is its weight. The weights are assigned based on task priority, with higher-priority tasks receiving greater weight.

[0079] In an embodiment of the present invention, when designing the loss function of the constructed security risk identification model, each task-specific head realizes parameter sharing and collaborative optimization through a weight distribution mechanism, which effectively improves the efficiency and accuracy of multi-task processing while reducing model redundancy.

[0080] Step S103: Process the risk data.

[0081] After analysis by the early warning rule engine and security risk identification model, the identified risk data is processed.

[0082] As an optional implementation, the risk data is processed, including generating a security alert and notifying the corresponding person in charge within a preset time according to the risk level.

[0083] For example, security alerts can be generated based on actual circumstances, including high-risk warnings and medium- and low-risk warnings, and dual notifications can be made via instant messaging tools and phone calls. High-risk warnings can be notified to the data security team leader within 30 minutes, while medium- and low-risk warnings can be sent to department heads within an hour.

[0084] In addition to security alerts, closed-loop risk event handling can be implemented for specific risk scenarios. In some implementations, this includes real-time interception, including immediate account freezing and interception of any abnormal behavior; optimizing business processes and activity rule design, such as increasing the complexity of verification codes and controlling resource allocation; recovering losses and collaborating with payment platforms to freeze suspicious transactions. Afterward, a data event accumulation mechanism can be established, including a negative case library and risk control strategy library, as well as monthly risk analysis reports to guide business optimization.

[0085] When processing risk data, the embodiment of the present invention distinguishes the risk data levels and issues warnings, thereby improving the efficiency of risk data processing.

[0086] The embodiment of the present invention combines an RPA engine that collects data according to a preset task coding process, and a dual data analysis mechanism constructed by an early warning rule engine and a security risk identification model to identify and handle risk data of the collected suspicious data sets, realizing an intelligent closed loop of the entire process from data collection, analysis and identification to risk disposal, saving the labor cost of data collection, and improving data monitoring efficiency and the accuracy of risk identification.

[0087] Example 2

[0088] Based on the above data security risk monitoring method, an embodiment of the present invention provides a data security risk monitoring device, such as Figure 6 As shown, the data security risk monitoring device includes a suspicious data collection unit 61 , a data analysis and processing unit 62 and a risk handling unit 63 .

[0089] The suspicious data collection unit 61 is used to execute the RPA engine and obtain suspicious data sets from the target network source according to the preset task orchestration process;

[0090] The data analysis and processing unit 62 is used to use the established early warning rule engine and security risk identification model to screen and analyze the suspicious data set in sequence to identify risky data;

[0091] The risk handling unit 63 is used to handle the risk data.

[0092] For other details about how the modules in the above-mentioned data security risk monitoring device implement the above-mentioned technical solution, please refer to the description of the data security risk monitoring method provided in the above-mentioned invention embodiment, which will not be repeated here.

[0093] Based on the above data security risk monitoring method, an embodiment of the present invention further provides a data security risk monitoring device, the structural diagram of which is as follows: Figure 7 As shown, the data security risk monitoring device 7 includes a processor 71 and a memory 72 coupled to the processor 71. The memory 72 stores a computer program, which, when executed by the processor 71, causes the processor 71 to perform the steps of the data security risk monitoring method in the above embodiment.

[0094] For other details about how the processor 71 in the above-mentioned data security risk monitoring device implements the above-mentioned technical solution, please refer to the description of the data security risk monitoring method provided in the above-mentioned invention embodiment, which will not be repeated here.

[0095] Among them, the processor 71 can also be called a CPU (Central Processing Unit), and the processor 71 may be an integrated circuit chip with signal processing capabilities; the processor 71 can also be a general-purpose processor, DSP (Digital Signal Process), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, among which the general-purpose processor can be a microprocessor or the processor 171 can also be any conventional processor, etc.

[0096] The embodiment of the present invention further provides a computer-readable storage medium, the structural diagram of which is as follows: Figure 8 As shown, the storage medium 8 stores a readable computer program 81; wherein, the computer program 81 can be stored in the above-mentioned storage medium in the form of a software product, including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in various embodiments of the present invention. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a magnetic disk or optical disk, ROM (Read-Only Memory), RAM (Random Access Memory), and other media that can store program code, or terminal devices such as computers, servers, mobile phones, and tablets.

[0097] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0098] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.

[0099] In addition, the functional modules in the various embodiments of the present application may be integrated into a single processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may be stored in a computer-readable storage medium.

[0100] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0101] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a server, or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0102] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used in the present application to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

[0103] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0104] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1process or processes and / or

[0105] or box Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0106] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0108] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A data security risk monitoring method, characterized in that: include: Execute the RPA engine and obtain suspicious data sets from the target network source according to the preset task orchestration process; Using the established early warning rule engine and security risk identification model to screen and analyze the suspicious data sets in turn to identify risky data; The risk data is processed.

2. The method according to claim 1, wherein: The process of obtaining a suspicious data set of a target network source according to a preset task arrangement process includes: Accessing a target web page according to the pre-marked address of the target network source; Performing interactive operations according to the interactive operations pre-marked in the interactive area of ​​the target webpage, and collecting webpage content fed back from the interactive operations on the target webpage; Data is extracted from the webpage content to obtain a suspicious data set.

3. The method according to claim 1, wherein: The early warning rule engine and security risk identification model constructed are used to screen and analyze the suspicious data sets in turn to identify risk data, including: Using an early warning rule engine, the detection rules formed by the alarm feature information are used to filter out the explicit risk content in the suspicious data set; The trained security risk identification model is used to analyze the filtered suspicious data set and output risk data including security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information.

4. The method according to claim 3, characterized in that The security risk identification model uses a pre-trained language model based on the Transformer architecture as the main model, and integrates a general large model for image detection, a general large model for text detection, and a general large model for multimodal detection. The security risk identification model adopts a multi-task learning framework, including a shared encoder and multiple task-specific heads that perform specific task outputs, including a security risk classification head, a sentiment analysis head, an entity recognition head, and an image analysis head.

5. The method according to claim 4, characterized in that The trained security risk identification model is used to analyze the filtered suspicious data set and output risk data including security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information, including: The filtered suspicious data set is input into the trained security risk identification model, and common semantic features are extracted by the shared encoder; the common semantic features are respectively input into the security risk classification head, sentiment analysis head, entity recognition head and image analysis head, and the corresponding output includes risk data of security risk category, negative / neutral / positive sentiment probability, key entity location, and sensitive image information.

6. The method according to claim 4, characterized in that: The security risk identification model constructs the following loss function: Loss = λ1Loss_Classification+λ2Loss_Sentiment+λ3Loss_Entity+λ4R In the formula, Loss represents the total loss function, Loss_classification is the loss function of the security risk classification task, and λ1 is its weight; Loss_sentiment is the loss function of the sentiment analysis task, and λ2 is its weight; Loss_entity is the loss function of the entity localization task, and λ3 is its weight; R represents the regularization term, and λ4 is its weight.

7. The method according to claim 1, wherein: Processing the risk data includes generating a security alert and notifying the corresponding person in charge within a preset time according to the risk level.

8. A data security risk monitoring device, characterized in that: It includes suspicious data collection unit, data analysis and processing unit and risk disposal unit; The suspicious data collection unit is used to execute the RPA engine and obtain the suspicious data set of the target network source according to the preset task arrangement process; The data analysis and processing unit is used to use the constructed early warning rule engine and security risk identification model to screen and analyze the suspicious data set in sequence to identify risky data; The risk handling unit is used to handle the risk data.

9. A data security risk monitoring device, comprising a memory and a processor, wherein: The memory is used to store computer programs; The processor is configured to read the computer program in the memory and execute the steps of any one of the data security risk monitoring methods according to claims 1 to 7.

10. A computer-readable storage medium having a readable computer program stored thereon, wherein when the program is executed by a processor, the steps of any one of the data security risk monitoring methods according to claims 1 to 7 are implemented.