Data leakage risk identification method and device and computer storage medium
The acquisition of web page content by combining screen capture and multimodal large model solves the problem of low data leakage risk identification accuracy in the prior art, realizes higher-precision data leakage risk identification, and enhances data security.
Patent Information
- Application Number
- CN202510355533.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the accuracy of obtaining web page data through crawler scripts and analyzing and identifying data leakage risks is low, and it is difficult to fully and accurately identify potential data leakage risks of complex web pages.
Image or video files are obtained by performing screen capture operations on the front-end interface, and these files are processed using a multimodal large model to obtain web page content, and match them with pre-acquisitioned web page content data to identify data leakage risks.
It improves the accuracy and comprehensiveness of data breach risk identification, and can more accurately judge whether the attacker can obtain web page content, thereby enhancing data security.
Smart Images

Figure CN120257320A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data security technology, and in particular to a method, device and computer storage medium for identifying data leakage risks. Background Art
[0002] With the popularization of Internet application services, key business data of enterprises (such as the conversion rules between points and commodity prices in the points redemption mall provided by enterprises for users, historical data and trend analysis data displayed in the data service platform provided by enterprises, etc.) have become one of the core assets of enterprises. These key business data carry the core logic of enterprise operations and are an important part of the core competitiveness of enterprises.
[0003] However, precisely because these key business data have significant economic value, they often become targets of malicious attacks. Once attacked maliciously, it will not only increase the burden on the business operation server, but also may lead to large-scale data leakage. On the one hand, such data leakage will damage the intellectual property rights of the enterprise and cause economic losses to the enterprise. On the other hand, it will pose a serious threat to the privacy of users, thereby destroying the trust relationship between users and enterprises.
[0004] Therefore, strict data security testing needs to be implemented before the application service is released in order to accurately identify potential data leakage risks and take effective response strategies in a timely manner. Summary of the invention
[0005] The inventor of the present disclosure has found that in the related art, the data of a web page is obtained by running a crawler script, and the obtained web page data is analyzed to identify the data leakage risk that may exist in the web page. In this way, the accuracy of identifying the data leakage risk is low.
[0006] In order to solve the above problems, the embodiments of the present disclosure provide a method for identifying data leakage risks and a corresponding identification device, which can improve the accuracy of identifying data leakage risks.
[0007] According to some embodiments of the present disclosure, a method for identifying data leakage risks is provided, comprising: performing a screen capture operation on a front-end interface to obtain a first file containing the content of a web page displayed in the front-end interface, the first file comprising at least one of an image file or a video file; processing the first file using a first multimodal large model to obtain first data, the first data being used to characterize content data of the web page obtained by an attacker through an attacking behavior; matching the first data with pre-acquired content data of the web page to identify the data leakage risk of the web page.
[0008] In some embodiments, performing a screen capture operation on the front-end interface includes: selecting to perform a screenshot operation or a screen recording operation in the screen capture operation on the front-end interface according to the content update frequency of the web page.
[0009] In some embodiments, the selecting to perform a screenshot operation or a screen recording operation in the screen capture operation on the front-end interface according to the content update frequency of the web page includes: in response to the content update frequency being greater than a first threshold, selecting to perform a screen recording operation on the front-end interface to obtain a video file containing the content of the web page as the first file; and / or in response to the content update frequency being less than or equal to the first threshold, selecting to perform a screenshot operation on the front-end interface to obtain an image file containing the content of the web page as the first file.
[0010] In some embodiments, the pre-obtained content data of the web page has a specified format, and the processing the first file by using a first multi-modal large model to obtain first data includes: inputting the first file and the prompt text corresponding to the first file into the first multi-modal large model to obtain the content data of the first file output by the first multi-modal large model, where the prompt text is used to instruct the first multi-modal large model to output the content data of the first file in the specified format; and obtaining the first data based on the content data of the first file.
[0011] In some embodiments, the obtaining the first data based on the content data of the first file includes: removing the hallucination data generated by the first multi-modal large model from the content data of the first file to obtain the first data.
[0012] In some embodiments, the first file includes multiple files, and the inputting the first file and the prompt text corresponding to the first file into the first multi-modal large model to obtain the content data of the first file output by the first multi-modal large model includes: inputting each file in the multiple files and the prompt text corresponding to each file into the first multi-modal large model in parallel to obtain the content data of each file output by the first multi-modal large model; storing the content data of each file in the same specified file.
[0013] In some embodiments, the prompt text corresponding to each file is the same.
[0014] In some embodiments, the matching process of the first data and the content data of the pre-acquired web page to identify the data leakage risk of the web page includes: in response to the result of the matching process that the matching rate of the first data and the content data of the pre-acquired web page is less than or equal to a second threshold, correcting the data in the first data that does not match the content data of the pre-acquired web page to obtain second data; performing a matching process on the second data and the content data of the pre-acquired web page to identify the data leakage risk of the web page.
[0015] In some embodiments, the first data includes a plurality of first sub-data that do not match the content data of the pre-acquired web page and second sub-data that match the content data of the pre-acquired web page. Each first sub-data in the plurality of first sub-data includes a target element that does not match the content data of the pre-acquired web page, and the target element in each first sub-data is the same. The correcting the data in the first data that does not match the content data of the pre-acquired web page to obtain second data includes: correcting the target element in each first sub-data to obtain a plurality of corrected first sub-data; determining the plurality of corrected first sub-data and the second sub-data as the second data.
[0016] In some embodiments, the correcting the data in the first data that does not match the content data of the pre-acquired web page to obtain second data includes at least one of the following methods: performing the screen capture operation on the front-end interface again to obtain a second file containing the content of the web page, and using the first multi-modal large model to process the second file to obtain the second data; preprocessing the first file, and using the first multi-modal large model to process the preprocessed first file to obtain the second data; or using a second multi-modal large model different from the first multi-modal large model to process the first file to obtain the second data, where the number of parameters of the second multi-modal large model is greater than the number of parameters of the first multi-modal large model.
[0017] In some embodiments, the matching process of the first data and the actual content data of the pre-acquired web page to identify the data leakage risk of the web page includes: in response to the result of the matching process that the matching rate of the first data and the content data of the pre-acquired web page is greater than a second threshold, outputting a test report on the data leakage risk of the web page.
[0018] According to some other embodiments of the present disclosure, there is provided an apparatus for identifying data leakage risks, including: a screen capture module configured to perform a screen capture operation on a front-end interface to obtain a first file containing the content of a web page displayed in the front-end interface, where the first file includes at least one of an image file or a video file; a processing module configured to process the first file using a first multi-modal large model to obtain first data of the web page, where the first data is used to characterize the content data of the web page obtained by an attacker through an attack behavior; and an identification module configured to perform a matching process on the first data and the pre-obtained content data of the web page to identify the data leakage risk of the web page.
[0019] According to some other embodiments of the present disclosure, there is provided an apparatus for identifying data leakage risks, including: a memory; and a processor coupled to the memory, where the processor is configured to execute the method for identifying data leakage risks in any of the above embodiments based on instructions stored in the memory device.
[0020] According to some other embodiments of the present disclosure, there is provided a computer-readable storage medium having computer instructions stored thereon, and when the instructions are executed by a processor, the method for identifying data leakage risks in any of the above embodiments is implemented.
[0021] According to some other embodiments of the present disclosure, there is also provided a computer program product including instructions, and when the instructions are executed by a processor, the processor is caused to execute the method for identifying data leakage risks according to any of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0023] Referring to the drawings, the present disclosure can be more clearly understood from the following detailed description, where:
[0024] Figure 1 A flowchart showing some embodiments of the method for identifying data leakage risks of the present disclosure;
[0025] Figure 2 A schematic diagram showing some embodiments of outputting the content data of the first file in CSV format of the present disclosure;
[0026] Figure 3 A flowchart showing some other embodiments of the method for identifying data leakage risks of the present disclosure;
[0027] Figure 4 A block diagram showing some embodiments of the apparatus for identifying data leakage risks of the present disclosure;
[0028] Figure 5 Block diagram showing other embodiments of the apparatus for identifying data leakage risks of the present disclosure;
[0029] Figure 6 Block diagram showing still other embodiments of the apparatus for identifying data leakage risks of the present disclosure. Detailed implementation manners
[0030] Various exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.
[0031] Meanwhile, it should be understood that the following description of at least one exemplary embodiment is merely illustrative and in no way restrictive of the present disclosure and its application or use.
[0032] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.
[0033] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as limitations. Thus, other examples of the exemplary embodiments may have different values.
[0034] The user data (including but not limited to user device data, user personal data, etc.) and web page data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all data and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.
[0035] It should be noted that: similar reference numerals and letters denote similar items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0036] The inventors of the present disclosure have found through research that in the related art, by running a crawler script, the Hyper Text Markup Language (HTML) information of a web page can be obtained, and then the web page content can be extracted by parsing the HTML information, and the potential data leakage risks can be identified by analyzing the extracted web page content. However, it is difficult to accurately identify the data leakage risks of web pages with relatively high complexity in this way.
[0037] For example, a web page may contain various complex elements in different forms, such as watermark text, mixed layout of text and pictures, and dynamically loaded content. In the related art, in the way of only relying on parsing the HTML information of the web page to extract the web page content, on the one hand, since the HTML information often cannot fully reflect the non-text content in the web page (such as the content in the picture), the extracted web page content is not comprehensive enough (for example, the information in the picture is missing from the extracted web page content); on the other hand, even irrelevant content (such as watermark text) will be presented in the HTML information, resulting in a large amount of interference information in the extracted web page content (for example, the extracted web page content is mixed with watermark text), thus making the information extracted from the web page inaccurate.
[0038] That is to say, in the way of the related art, because the extracted web page content is neither comprehensive nor accurate, the accuracy of identifying the data leakage risk of the web page relying on this content is relatively low.
[0039] In view of this, the present disclosure proposes a method for identifying data leakage risk, which can obtain a first file that fully reflects the web page content through screen capture, and then use a multi-modal large model to process the first file to obtain the web page content, so that the obtained web page content is more comprehensive and accurate, and thus identify the data leakage risk of the web page based on the obtained web page content, which can effectively improve the accuracy of identifying the data leakage risk of the web page.
[0040] Figure 1 The flowchart showing some embodiments of the method for identifying data leakage risk of the present disclosure.
[0041] As Figure 1 shown, for example, the identification of data leakage risk may include step 110 to step 130. In some embodiments, the method for identifying data leakage risk is executed by an identification device for data leakage risk.
[0042] In step 110, a screen capture operation is performed on the front-end interface to obtain a first file including the content of the web page displayed in the front-end interface.
[0043] Here, the first file may include at least one of an image file and a video file.
[0044] In some embodiments, the web page accessed through a small program, the web page accessed through a client, and the web page accessed through a browser may be displayed in the front-end interface.
[0045] It can be understood here that the front-end interface is a visual interface for users to interact with application service programs, and can provide an interface for users to interact with application service programs. As a carrier for presenting the front-end of an application service program, a web page provides an access entry to the application service program for users. Through the front-end interface, the web page can display the functions of the application service program to users. In the present disclosure, the web page can include pages of various types of application services and can be presented to users in the front-end interface in various forms. The present disclosure does not limit the source of the web page and the display form of the web page in the front-end interface.
[0046] In some embodiments, the screen capture operation may include at least one of a screenshot operation and a screen recording operation. For example, by performing a screenshot operation on the front-end interface, an image file containing the web page content displayed on the front-end interface can be obtained as the first file. For example, by performing a screen recording operation on the front-end interface, a video file containing the web page content displayed on the front-end interface can be obtained as the first file.
[0047] In some embodiments, a screen capture operation may be performed on a target area in the front-end interface. For example, the target area may be a partial area or the entire area in the front-end interface.
[0048] In some embodiments, the target area may be a partial area in the front-end interface that displays the main content of the web page. For example, other areas in the front-end interface except for the navigation bar and the sidebar may be used as the target area, and a screen capture operation is performed on this target area to obtain the first file. In this way, the interference information contained in the first file obtained through the screen capture operation can be reduced, which helps to improve the accuracy of subsequent processing of the first file.
[0049] In step 120, the first multi-modal large model is used to process the first file to obtain first data.
[0050] Here, the first data is used to represent the content data of the web page obtained by the attacker through an attack behavior.
[0051] It should be understood that a multi-modal large model (for example, a multi-modal large language model (Multimodal Large Language Model, MLLM) or a vision language model (Vision Language Model, VLM)) is an artificial intelligence model that can process and understand various types (modalities) of input data (such as text, images, videos, etc.). The multi-modal large model can fuse data of different modalities to achieve cross-modal understanding and reasoning.
[0052] In some embodiments, different multimodal large models can be used to process different types of first files to improve the accuracy of first file processing. For example, for image files in the first file, the PP-DocBee model focusing on document image understanding can be used to process them. For example, for video files in the first file, the Qwen2-VL model focusing on Vision-Language (VL) tasks can be used to process them.
[0053] In some embodiments, the multimodal large model can run on a Graphics Processing Unit (GPU). After the multimodal large model finishes processing the first file, the memory of the GPU (also known as "video memory") can be released to make reasonable use of the GPU resources, which helps improve system performance.
[0054] In some embodiments, the first file can be input into the multimodal large model to obtain the content data of the first file output by the multimodal large model, and the first data of the web page can be obtained based on the content data of the first file. This will be further described later.
[0055] In step 130, the first data is matched with the pre-acquired content data of the web page to identify the data leakage risk of the web page.
[0056] In some embodiments, the content data of the web page can be pre-acquired from the developer of the web page.
[0057] It should be understood that the pre-acquired content data of the web page can include the web page content provided by the developer of the web page, while the first data includes the data extracted by combining screen capture and a multimodal large model. Since the first data can represent the data that an attacker may obtain through an attack behavior. Based on the result of matching the two, it can be used to evaluate the possibility that an attacker obtains the web page content through an attack behavior, and this possibility can reflect the data leakage risk of the web page.
[0058] In some embodiments, based on the result of the matching process between the first data and the pre-acquired content data of the web page, a risk assessment of data leakage of the web page can be performed. For example, a corresponding scoring system can be used to quantify the data leakage risk of the web page to obtain the result of the risk assessment, where different scores represent different risk levels.
[0059] In some embodiments, according to the result of the risk assessment of data leakage of the web page, it is determined whether the web page passes the data security test. For example, if the data leakage risk of the web page is high, it is determined that the web page fails the data security test. If the data leakage risk of the web page is low, it is determined that the web page has passed the data security test.
[0060] In some embodiments, it is possible to determine whether to allow the output of a test report on the data leakage risk of a web page based on the results of a data security test (i.e., the web page fails the data security test or the web page passes the data security test).
[0061] For example, when the result of the data security test is that the web page fails the data security test, it is allowed to output a test report on the data leakage risk of the web page to prompt that the web page has a high data leakage risk. When the result of the data security test is that the web page passes the data security test, it is not allowed to output a test report on the data leakage risk of the web page and other processing methods will be adopted (which will be further described later).
[0062] For example, the risk points with data leakage risks in the web page can be listed in the test report.
[0063] In some embodiments, the results of the data security test can be provided to developers so that the developers can implement targeted data protection measures for the data leakage risks existing in the web page, thereby improving the data security of the application service corresponding to the web page.
[0064] In some embodiments, after the developers implement data protection measures for the risk points with data leakage risks in the web page, the data security test method provided by the present disclosure can be executed again to confirm the effectiveness of the implemented data protection measures.
[0065] In the above embodiments, through the screen capture operation, any content in the web page can be completely recorded in the first file, that is, the obtained first file can completely cover the content presented in various forms in the web page. On this basis, by virtue of the processing ability of the multi-modal large model for different modal data, more comprehensive and accurate web page content (i.e., the first data) is extracted from the first file. Since the first data can comprehensively and accurately reflect the content data of the web page that the attacker can obtain through the attack behavior, matching the first data with the pre-obtained content data of the web page can effectively determine whether the attacker can obtain the web page content through the attack behavior, thereby effectively improving the accuracy of identifying the data leakage risk of the web page.
[0066] In other words, one of the core ideas of the data security testing method proposed in the present disclosure is to obtain more comprehensive and accurate web page content from the perspective of the attacker by combining screen capture and multimodal large models, so as to effectively expand the content scope covered by data security testing and raise the threshold for web pages to pass data security testing, that is, to make the passing standard of data security testing more stringent. As a result, the accuracy of identifying data leakage risks of web pages is improved, so as to provide more accurate guidance for developers, so that developers can take more stringent data protection measures in a targeted manner to enhance the data security of application services.
[0067] Next, in combination with some embodiments, the relevant implementation of the data security testing method proposed in the present disclosure is further explained by way of example.
[0068] The following first describes the related implementation of obtaining the first file in step 110 in combination with some embodiments.
[0069] In some embodiments, according to the content update frequency of the webpage, a screen capture operation or a screen recording operation is selected to be performed on the front-end interface. For example, it can be determined whether the content update frequency of the webpage is greater than a first threshold, and then according to the result of the determination, a screen capture operation or a screen recording operation is performed on the front-end interface. In other words, it can be dynamically selected whether to obtain the corresponding first file by screen capture or screen recording according to the content update frequency of the webpage.
[0070] It should be noted that the first file obtained by the screenshot operation is mainly analyzed in the form of a single picture in the subsequent model processing process. This method consumes relatively little computing resources, but the lack of contextual information based on inter-frame relationships may make the multimodal large model not accurately understand the content of the first file, which may affect the accuracy of identifying data leakage risks. In contrast, the first file obtained by the screen recording operation can provide contextual information based on inter-frame relationships, so that the multimodal large model can understand the content of the first file more accurately, thereby improving the accuracy of identifying data leakage risks. However, this method consumes relatively much computing resources.
[0071] Therefore, through this flexible selection mechanism, a better balance can be achieved between reducing the consumption of computing resources and improving the identification accuracy of data leakage risks, so as to reasonably control the cost of data security testing on the basis of improving the identification accuracy of data leakage risks.
[0072] In some embodiments, in response to a content update frequency of a web page being greater than a first threshold, a screen recording operation is performed on the front-end interface to obtain a video file containing the content of the web page as a first file.
[0073] It should be noted that the screen recording operation can be understood as a continuous screenshot process. Compared with the screenshot operation, it can capture content at a higher frequency, thus forming a smooth video stream. That is to say, the screen recording operation can not only capture the static elements in the web page, but also completely record the dynamic change process of the elements in the web page.
[0074] For example, in the case where there is dynamic content updated at a high frequency in the web page, using the screenshot operation may capture incomplete information at the moment of content change, resulting in a large deviation in the web page content recorded in the obtained first file. This first file with a large deviation will further affect the processing process of the multimodal large model, resulting in a large deviation in the obtained first data, thus having an adverse impact on the recognition accuracy of the data leakage risk.
[0075] Therefore, if the content update frequency of the web page is relatively high, the screen recording operation can be adopted to obtain the video file corresponding to the web page as the first file. In this way, by recording the complete dynamic update process of the web page content, not only can the complete information at the moment of content change be captured, but also comprehensive analysis can be carried out based on the context information of the inter-frame relationship in the subsequent model processing process, making the obtained first data more accurate, thus effectively improving the recognition accuracy of the data leakage risk of the web page.
[0076] In some embodiments, in response to the content update frequency of the web page being less than or equal to the first threshold, a screenshot operation is performed on the front-end interface to obtain an image file containing the content of the web page as the first file.
[0077] As mentioned above, the screenshot operation consumes relatively less computing resources. When the content update frequency of the web page is low, the screenshot operation is sufficient to capture the changed content in the web page in a timely manner, so that all the content in the web page can be accurately recorded in the first file. However, in this case, if the screen recording operation is still used to obtain the video file corresponding to the web page as the first file, it may result in a large amount of redundant information in the obtained first file, which not only increases the consumption of computing resources but also reduces the subsequent model processing efficiency.
[0078] Therefore, if the content update frequency of the web page is low, the screenshot operation can be adopted to obtain the image file corresponding to the web page as the first file. In this way, the repeated recording of the same content in the web page can be effectively reduced, avoiding unnecessary information accumulation. Thus, on the basis of improving the recognition accuracy of the data leakage risk, the consumption of computing resources is reduced, and the efficiency of data security testing is improved. That is, data security testing can be carried out efficiently and accurately at a relatively low cost.
[0079] In some embodiments, the web pages displayed in the front-end interface may include multiple regions with different content update frequencies. For each of the multiple regions, an operation of taking a screenshot or a screen recording in the screen capture operation may be selected according to the content update frequency of the region to obtain a first file containing the content of the region. For example, for each of the multiple regions, it may be determined whether the content update frequency of the region is greater than a first threshold, and based on the determination result, an operation of taking a screenshot or a screen recording in the screen capture operation may be performed on the region to obtain a first file corresponding to the region. That is, the first file may include a first file corresponding to each of the multiple regions.
[0080] For example, a certain web page may include region A with a content update frequency less than the first threshold and region B with a content update frequency greater than the first threshold. A screenshot operation is performed on region A to obtain an image file corresponding to region A. A screen recording operation is performed on region B to obtain a video file corresponding to region B. That is, the first file corresponding to the web page may include the image file corresponding to region A and the video file corresponding to region B.
[0081] In this way, for a more complex web page with multiple regions having different content update frequencies, it is possible to flexibly adopt an operation of taking a screenshot or a screen recording in the screen capture operation for each of the multiple regions, so as to improve the overall accuracy of the web page content included in the obtained first file. Thus, the recognition accuracy of the data leakage risk is improved.
[0082] Next, some embodiments will be further combined to exemplarily illustrate the related implementation of obtaining the first data in step 120.
[0083] In some embodiments, the first file and the prompt text corresponding to the first file may be input into the first multimodal large model to obtain the content data of the first file output by the first multimodal large model. Then, based on the content data of the first file, the first data is obtained.
[0084] In some embodiments, the content data of the pre-acquired web page has a specified format. The prompt text may be used to instruct the first multimodal large model to output the content data of the first file in the specified format.
[0085] For example, the specified format may include the Comma-Separated Values (CSV) format. A file output in the CSV format (also known as a "CSV file") can store data in plain text form. Each row represents a piece of data, and each piece of data can be composed of multiple fields, with the fields separated by commas. In this way, the extracted content data can be made concise and clear, which helps to improve the efficiency of data processing, and thus improve the efficiency of subsequent data security testing.
[0086] For example, the first file is a video file obtained through a screen recording operation, and the corresponding prompt text can be "Output the product names and corresponding prices in the video file in CSV format."
[0087] For example, the first file is an image file obtained through a screenshot operation, and the corresponding prompt text can be "Output the product names and corresponding prices in the image file in CSV format."
[0088] Figure 2 Schematic diagrams showing some embodiments of the present disclosure for outputting the content data of the first file in CSV format.
[0089] As Figure 2 shown, Figure 2 Schematically shown are 5 pieces of content data including product names and corresponding prices in the first file output in CSV format.
[0090] In the above embodiments, by setting the prompt text for indicating output in a specified format, the content data in the first file output by the multi-modal large model can have a unified structure and format with the content data of the pre-acquired web page. In this way, unnecessary information output is avoided, which helps to improve the efficiency and accuracy of subsequent matching processing, thereby improving the efficiency and reliability of data security testing.
[0091] In some embodiments, the first data obtained based on the content data of the first file and the content data of the pre-acquired web page also have the same specified format for efficient and accurate matching processing. For example, the first data and the content data of the web page can uniformly adopt the CSV format, so that precise comparison at the field level can be performed during the matching processing.
[0092] In this way, through the unified format, the same parts and different parts in the first data and the content data of the web page can be quickly found, improving the efficiency of the matching processing. At the same time, based on this efficient data matching method, potential data leakage risks in the web page can be timely discovered, so as to take corresponding data protection measures in a timely manner, which helps to improve the data security of the application service corresponding to the web page.
[0093] In some embodiments, the first file may include multiple files.
[0094] For example, the first file may contain multiple files of a single type. For example, the multiple files may include multiple video files or multiple image files.
[0095] For example, the first file may contain multiple files of different types. For example, the multiple files may include at least one video file and at least one image file.
[0096] It can be understood here that each of the multiple files corresponds to the same web page.
[0097] Continuing with the above example, the first file corresponding to a certain web page includes an image file corresponding to area A and a video file corresponding to area B. That is, the image file corresponding to area A and the video file corresponding to area B both correspond to this same web page.
[0098] In the case where the first file includes multiple files, each of the multiple files and the prompt text corresponding to each file can be input into the first multi-modal large model in parallel to obtain the content data of each file output by the first multi-modal large model. Then, the content data of each file is stored in the same specified file.
[0099] For example, each of the multiple files and the corresponding prompt text can be used as a pair of input data and input into the first multi-modal large model in parallel, so that the first multi-modal large model can process the multiple files corresponding to the same web page in parallel, thereby efficiently obtaining the content data of each file. After that, the content data of each file can be integrated and stored in the same specified file.
[0100] In this case, for example, the number of files that the first multi-modal model processes in batches can be set. The first file and the corresponding prompt text that do not exceed this number of files are input into the first multi-modal model for batch processing.
[0101] In the above embodiments, since each of the multiple files corresponds to the same web page, using the multi-modal large model to process them in parallel helps to enhance the consistency and coherence of content analysis, improves the accuracy of the obtained content data, and thus improves the recognition accuracy of data leakage risks. And, storing the content data of each obtained file in a specified file for integration can effectively simplify the data management in the subsequent data security testing process, thereby making the data security testing more efficient.
[0102] In some embodiments, the prompt text corresponding to each of the multiple files is the same. For example, the prompt text corresponding to each file can be "Output the product name and the corresponding price in the file in CSV format".
[0103] In this way, the same prompt text can make multiple files follow a unified output standard during the processing of the multi-modal large model, which helps to further enhance the consistency and coherence of content analysis, and thus helps to improve the recognition accuracy of data leakage risks. And, using the same prompt text can avoid the situation where the multi-modal large model needs to be adjusted or trained multiple times due to different prompt texts, thereby reducing the complexity of model processing and improving the efficiency of data security testing.
[0104] In some embodiments, hallucinated data generated by the multimodal large model can be removed from the content data of the first file output by the first multimodal large model to obtain the first data.
[0105] For example, for an image file obtained through a screenshot operation, the number of content data output in CSV format corresponding to it is fixed. Therefore, a quantity threshold for the content data output in CSV format can be set, such as 100 entries. If the number of data in the CSV file output by the first multimodal large model exceeds this quantity threshold, it indicates that there is hallucinated data therein. Subsequently, the hallucinated data can be removed from the CSV file to obtain the first data.
[0106] It should be noted that the "Hallucinated Data" generated by the multimodal large model refers to the incorrect or irrelevant content output by the model after processing multimodal input data (such as images, videos, and texts). For example, during the process of the multimodal large model processing the first file, it will reason and summarize the content in the first file. In the subsequent process of generating output data, it may ignore the content in the input first file and only rely on the information obtained from these summaries, thereby generating hallucinated data with logical errors or irrelevant to the content in the input first file.
[0107] Therefore, after obtaining the content data of the first file output by the multimodal large model, the accuracy of the first data can be improved by removing the hallucinated data therein, thereby enhancing the precision of identifying the data leakage risk of the web page.
[0108] It can be understood here that in the case where the first file includes multiple files, the hallucinated data generated by the multimodal large model can be removed from the content data of each file output by the multimodal large model to obtain the first data. Thereby improving the accuracy of the first data.
[0109] The following will exemplarily illustrate the related implementation of identifying the data leakage risk of the web page by matching the first data with the pre-obtained content data of the web page in step 130 in combination with some embodiments.
[0110] In some embodiments, the first data can be matched with the pre-obtained content data of the web page. According to the result of the matching process, the data leakage risk of the web page can be identified.
[0111] In some embodiments, the content data of the web page provided by the developer can be used as the pre-obtained content data of the web page for the matching process.
[0112] It should be understood that this matching process is essentially a more rigorous data security test. From the perspective of an attacker, it combines screen capture and a multimodal large model to obtain more comprehensive and accurate web page content for a comprehensive assessment of the data security of the web page. This helps to promptly detect potential data leakage risks and provides precise guidance for strengthening subsequent data protection measures, thereby improving the data security of the application service corresponding to the web page.
[0113] In some embodiments, in response to the matching rate of the first data and the content data of the pre-acquired web page being greater than the second threshold, a test report on the data leakage risk of the web page can be output.
[0114] In some embodiments, the matching rate can be characterized by the number of data that the first data successfully matches with the content data of the pre-acquired web page, or can be characterized by the proportion of the number of data that the first data successfully matches with the content data of the web page to the total amount of data. For example, if the first data is consistent with the content data of the pre-acquired web page, it indicates that the two match successfully.
[0115] As mentioned above, the first data represents the data that the attacking party can obtain through an attack behavior. Therefore, if the matching rate of the first data and the content data of the web page is high, it means that the attacking party can obtain more web page content through the attack behavior, that is, the web page fails the data security test. This means that the current data protection measures are insufficient and the web page has a high data leakage risk. In this case, a test report can be output to indicate the data leakage risk of the web page.
[0116] In some embodiments, in response to the matching rate of the first data and the content data of the pre-acquired web page being less than or equal to the second threshold, the data in the first data that does not match the content data of the pre-acquired web page can be corrected to obtain the second data, and the second data and the content data of the pre-acquired web page are subjected to a matching process to identify the data leakage risk of the web page. For example, the first data can include three data, M, P, and Q, where the P data does not match the content data of the pre-acquired web page, then the P data can be corrected. And the M data, Q data, and the corrected P data are used as the second data.
[0117] For example, if the matching result between the second data and the content data of the pre-acquired web page shows that the matching rate of the second data and the content data of the pre-acquired web page is greater than the second threshold, it indicates that the data leakage risk of the web page is relatively high. In this case, a test report can be output to prompt the existence of data leakage risk on the web page. On the contrary, if the matching result between the second data and the content data of the pre-acquired web page shows that the matching rate of the second data and the content data of the pre-acquired web page is less than or equal to the second threshold, it indicates that the data leakage risk of the web page is relatively low. In this case, a test report may not be output, but a prompt message such as the web page has passed the data security test can be output.
[0118] It can be understood here that if the matching rate of the first data and the content data of the pre-acquired web page is relatively low, it means that the attacker can obtain less web page content, that is, the web page has passed the data security test. In this case, to further confirm the data leakage risk of the web page, additional correction measures can be taken for the first data to make the obtained second data more comprehensive and accurate. Then, based on the second data, the data leakage risk of the web page is identified again.
[0119] In this way, the one-sidedness of judging the data leakage risk of the web page based on the result of a single test is avoided, making the passing standard of the data security test more stringent, thereby improving the accuracy of identifying the data leakage risk of the web page.
[0120] The following will exemplarily illustrate, in combination with some embodiments, the correction measures that can be taken to make the second data more comprehensive and accurate.
[0121] In some embodiments, the first data may include a plurality of first sub-data that do not match the content data of the pre-acquired web page and second sub-data that match the content data of the pre-acquired web page. Each first sub-data among the plurality of first sub-data includes a target element that does not match the content data of the pre-acquired web page. The target element in each first sub-data is the same.
[0122] That is to say, the factor causing each first sub-data among the plurality of first sub-data not to match the content data of the pre-acquired web page is the same. For example, the punctuation marks used in each first sub-data are all English punctuation marks, while the punctuation marks used in the content data of the pre-acquired web page are all Chinese punctuation marks. Therefore, the target element in each first sub-data that does not match the content data of the web page is the same, that is, all are English punctuation marks.
[0123] In these embodiments, the target element in each second data can be corrected to obtain the corrected plurality of first sub-data. And the corrected plurality of first sub-data and the second sub-data are determined as the second data.
[0124] Continuing with the above example, for instance, the English punctuation marks in each of the multiple first sub - data can be batch - replaced with Chinese punctuation marks to obtain the corrected multiple first sub - data.
[0125] In this way, factors in the first data that commonly cause data mismatch can be batch - corrected, enabling the resulting second data to more accurately reflect the web content extracted from the first file. Thus, a more reliable data basis is provided for subsequent matching processing, improving the accuracy of identifying the data leakage risk of the web page.
[0126] In some embodiments, the first file can be pre - processed, and then the pre - processed first file can be processed using a first multi - modal large model to obtain the second data.
[0127] For example, the pre - processing can include one or more operations such as cropping, rotating, adjusting brightness, adjusting contrast, and adjusting color of the first file.
[0128] In some embodiments, the pre - processed first file can be processed using a first multi - modal large model to obtain the content data of the pre - processed first file. Then, based on the content data of the pre - processed first file, the second data is obtained.
[0129] For example, the hallucination data generated by the first multi - modal large model can be removed from the content data of the pre - processed first file to obtain the second data. For the relevant description of hallucination data, reference can be made to the description of the relevant embodiments above and will not be elaborated here.
[0130] In this way, considering that the quality of the first file obtained by screenshot or screen recording operations may be low, which affects the accuracy of the first data. Therefore, before processing the first file using the multi - modal large model, a pre - processing operation can be performed to improve the quality of the first file (such as improving the clarity of the first file, etc.) to improve the accuracy of subsequent model processing, thus providing a more reliable data basis for subsequent matching processing and improving the accuracy of identifying the data leakage risk of the web page.
[0131] In some embodiments, a second multi - modal large model different from the first multi - modal large model can be used to process the first file to obtain the second data, where the number of parameters of the second multi - modal large model is greater than that of the first multi - modal large model. For example, the first multi - modal large model and the second multi - modal large model can be of the same type, such as both being multi - modal large prediction models.
[0132] It should be understood that if the number of parameters of the second multimodal large model is greater than that of the first multimodal large model, it indicates that the second multimodal large model has stronger representation ability and better generalization ability, and thus can achieve higher processing accuracy.
[0133] In this way, the first file is processed using the second multimodal large model with higher accuracy to obtain more accurate second data, thereby providing a more reliable data basis for subsequent matching processing and improving the accuracy of identifying the data leakage risk of the web page.
[0134] In some embodiments, the screen capture operation can be performed again on the front-end interface to re-obtain the second file. Then, the first multimodal large model is used to process the re-obtained second file to obtain the second data. For example, the screen capture operation performed again can be the same as or different from the previous screen capture operation.
[0135] In this way, considering that the web page content recorded in the first file may be incomplete due to improper execution of the screen capture operation (such as text truncation), the screen capture operation can be optimized at the source to re-obtain another more accurate second file. In this way, the first data can be made more accurate to provide a more reliable data basis for subsequent matching processing, thereby improving the accuracy of identifying the data leakage risk of the web page.
[0136] Here, it can be understood that the correction measures mentioned in the above embodiments can be used in combination to obtain more accurate second data. For example, the second multimodal large model can be used to process the preprocessed first file to obtain the second data.
[0137] It can also be understood that in the present disclosure, the correction measures mentioned in the above embodiments can be executed multiple times, and then the matching processing is repeated multiple times to improve the accuracy of identifying the data leakage risk of the web page.
[0138] That is to say, even if the result of the previous data security test shows that the web page has temporarily passed the data security test, the correction measures can still be continued, and then the next matching processing is performed. The correction measures executed each time can be different.
[0139] For example, the specified number of times that the web page passes the data security test can be set as the test end condition. Only when the number of times the web page passes the data security test reaches the specified number, the data security test is ended, and the result of the last test is used as the result of the data security test of the web page, thereby improving the accuracy of identifying the data leakage risk of the web page.
[0140] Figure 3Flowchart showing other embodiments of the method for identifying data leakage risks of the present disclosure.
[0141] As Figure 3 shown, for example, the method for identifying data leakage risks may include steps 310 to 380. This method can be executed as Figure 1 a specific implementation of the method in
[0142] In step 310, a test environment is set up.
[0143] In some embodiments, a test environment can be set up based on the user's operating system and network conditions, so that in this test environment, the user can normally access various types of application services.
[0144] Thus, since the test environment is set up based on the user's operating system and network conditions, this test environment can accurately reflect the actual operation behavior of the user during the process of accessing application services. Such a setting makes the result of data security testing closer to the actual situation, thereby improving the identification accuracy of potential data leakage risks in the web pages of application services.
[0145] In step 320, user behavior simulation.
[0146] In some embodiments, an automated testing tool can be used to simulate the user's operations on the web page, including but not limited to operations such as browsing, searching, and clicking. For example, the automated testing tool can include headless browsers, desktop application automated testing tools (such as PyAutoGUI based on Python), etc. Among them, PyAutoGUI can more accurately simulate the user's operation behavior by directly controlling the keyboard and mouse, and supports various operating systems (such as Windows, macOS, and Linux).
[0147] In step 330, it is determined whether the content update frequency of the web page is greater than a first threshold. For example, according to whether the content update frequency of the web page is greater than the first threshold, it is determined whether to perform the screen recording operation in the screen capture operation.
[0148] In some embodiments, a web page monitoring tool (such as Visualping, ChangeTower, etc.) can be used to monitor the content changes of the web page to determine the content update frequency of the web page.
[0149] In response to the determination result that the content update frequency of the web page is greater than the first threshold, steps 340 and 350 can be executed.
[0150] In response to the determination result that the content update frequency of the web page is less than or equal to the first threshold, steps 360 and 370 can be executed.
[0151] In step 340, a screen recording operation is performed to obtain a video file (also known as a "screen recording file") as the first file corresponding to the web page.
[0152] In some embodiments, a screen recording operation (also known as an "automatic screen recording operation") can be performed on a target area (also known as a "screen designated area") in the front-end interface.
[0153] For example, the screen recording operation can be performed in combination with the PyAutoGUI and OpenCV tools. For example, first capture a screenshot of the target area through PyAutoGUI, and then use OpenCV to convert the captured screenshot of the target area into a video file and specify the file name of the output video file. Then, save the video file in a specified location (such as a specified folder).
[0154] In step 350, the video file obtained in step 340 is input into a multimodal large model for processing.
[0155] For example, the video file obtained in step 340 can be input into the Qwen2-VL model for processing.
[0156] For example, the folder path (also known as the "directory path") of the video file and the prompt text corresponding to the video file can be input into the Qwen2-VL model to obtain the content data (i.e., the model output result) of the video file output by the model.
[0157] For example, the prompt text corresponding to the video file can be used to instruct the Qwen2-VL model to output the content data of the video file in a specified format (such as CSV format).
[0158] In some embodiments, the video file obtained in step 340 can include multiple files. For example, each of these multiple files and the prompt text corresponding to each file can be input into the Qwen2-VL model in parallel for batch processing. Then, the content data of each file obtained from the batch processing is stored in the same specified file.
[0159] In step 360, a screenshot operation is performed to obtain an image file as the first file corresponding to the web page.
[0160] In some embodiments, a screenshot operation (also known as an "automatic screenshot operation") can be performed on a target area in the front-end interface. For example, the PyAutoGUI tool can be used to perform the screenshot operation to capture a screenshot of the target area as an image file.
[0161] In step 370, the image file obtained in step 360 is input into a multimodal large model for processing.
[0162] For example, the image file obtained in step 360 can be input into the PP-DocBee model for processing to obtain the first data.
[0163] For example, the folder path of the image file and the prompt text corresponding to the image file can be input into the PP-DocBee model to obtain the content data of the image file output by the model, and then the first data of the web page can be obtained based on the content data of the image file.
[0164] For example, the prompt text corresponding to the image file can be used to instruct the PP-DocBee model to output the content data of the image file in a specified format (such as CSV format).
[0165] In some embodiments, the image file obtained in step 360 may include multiple files. In this case, each of these multiple files and the prompt text corresponding to each file can be input into the PP-DocBee model in parallel for batch processing. Then, the content data of each file obtained from the batch processing is stored in the same specified file.
[0166] In step 380, the output result of the multi-modal large model is obtained. For example, the output result of the multi-modal large model includes the content data of the first file output by the multi-modal large model.
[0167] In step 390, based on the output result of the multi-modal large model, a risk assessment of data leakage on the web page is performed.
[0168] In some embodiments, the first data of the web page can be obtained based on the output result of the multi-modal large model. Then, based on the first data and the pre-obtained content data of the web page, a risk assessment of data leakage on the web page is performed.
[0169] In step 3100, according to the result of the risk assessment, it is determined whether the web page passes the data security test. For example, if the data leakage risk of the web page is high, it is determined that the web page fails the data security test. If the data leakage risk of the web page is low, it is determined that the web page has passed the data security test.
[0170] In some embodiments, the first data can be matched with the pre-obtained content data of the web page to obtain the result of the matching process. Based on the result of the matching process, a risk assessment of data leakage on the web page is performed to obtain the result of the risk assessment.
[0171] For example, in response to the result of the matching process being that the matching rate of the first data and the pre-obtained content data of the web page is greater than the second threshold, that is, the result of the risk assessment is that the data leakage risk of the web page is high, step 3110 is executed.
[0172] In response to the result of the matching process that the matching rate between the first data and the content data of the pre-acquired web page is less than or equal to the second threshold, that is, the result of the risk assessment is that the data leakage risk of the web page is relatively low, step 3120 is executed.
[0173] Here, it can be understood that if the matching rate between the first data and the content data of the pre-acquired web page is greater than the second threshold, it means that more data that may be correctly parsed by the attacker exists, and the data leakage risk of the web page is relatively high. If the matching rate between the first data and the content data of the pre-acquired web page is less than or equal to the second threshold, it means that less data that may be correctly parsed by the attacker exists, and the data leakage risk of the web page is relatively low.
[0174] In step 3110, it is determined that the web page fails the data security test.
[0175] For example, in step 3110, a test report can also be output to prompt the existence of a data leakage risk for the web page.
[0176] In step 3120, it is determined whether to continue to execute the matching process. That is, it is determined whether it is necessary to iteratively optimize the data security test.
[0177] In some embodiments, in response to the result of the matching process that the matching rate between the first data and the content data of the pre-acquired web page is less than or equal to the second threshold, it can be first manually checked by the tester, and then it is determined whether to continue to execute the matching process according to the result of the manual check.
[0178] For example, if the result of the manual check shows that the factors causing each of the multiple first sub-data in the first data not to match the content data of the pre-acquired web page are the same, it can be determined to continue to execute the matching process. And, before continuing to execute the matching process, the target elements in each first sub-data that do not match the content data of the pre-acquired web page are corrected to obtain the second data. Then, the matching process is continued based on the second data.
[0179] This method improves the accuracy of the data used to execute the matching process by batch-correcting the same factors that cause the first data not to match the content data of the web page, thereby improving the recognition accuracy of subsequent data leakage risks.
[0180] For example, if the result of the manual check shows that the quality of obtaining the first file is low, it can be determined to perform a secondary test. And, before continuing to execute the matching process, the first file can be preprocessed, and the preprocessed first file is processed using a multimodal large model to obtain the second data. Then, the matching process is continued based on the second data.
[0181] This method improves the quality of the first file by performing preprocessing operations, thereby enhancing the accuracy of the data used for performing matching processing, and thus improving the recognition accuracy of subsequent data leakage risks.
[0182] For example, if the result of manual inspection shows that the web content recorded in the first file is incomplete (such as problems like text being truncated), it can be determined to continue with the matching processing. And, before continuing with the matching processing, perform the screen capture operation again to re-obtain the second file. Process the re-obtained second file using a multimodal large model to obtain second data. Then continue with the matching processing based on the second data.
[0183] This method solves the problem of inaccurate recording of web content that may be caused by improper initial screen capture operations by re-obtaining a second new first file. That is, optimize the screen capture operation from the source to improve the accuracy of the data used for performing matching processing, and thus improve the recognition accuracy of subsequent data leakage risks.
[0184] Figure 3 Schematically shows this way of optimizing the screen capture operation from the source. For example, as Figure 3 shown, in the case of determining to continue with the matching processing (i.e., iterative optimization is required), it is possible to return to step 320, re-perform the user behavior simulation, and then re-perform the corresponding screen capture operation. And in the case of determining not to continue with the matching processing (i.e., iterative optimization is not required), perform step 3130.
[0185] In step 3130, it is determined that the web page has passed the data security test.
[0186] In the above embodiments, by building a test environment that can accurately reflect the operation behavior of users during the access to application services, and expanding the content scope covered by the data security test from text-form data to include non-text-form data such as pictures and videos, the coverage scope of the data security test becomes more comprehensive, thereby improving the accuracy of identifying the data leakage risk of the web page.
[0187] In addition, by providing an iterative optimization scheme, the process of data security testing can be continuously improved, making the results of the data security test more reliable.
[0188] It should be understood that the relevant steps and descriptions in the above Figures 1 to 2 shown embodiments also equally apply to Figure 3 the data security testing method shown, and the relevant explanations can be referred to the above Figures 1 to 2 shown embodiments, which will not be elaborated here.
[0189] Figure 4A block diagram showing some embodiments of the apparatus for identifying data leakage risks of the present disclosure.
[0190] like Figure 4 As shown, for example, a device 400 for identifying data leakage risks includes a screen capturing module 401 , a processing module 402 , and an identifying module 403 .
[0191] The screen capture module 401 may be configured to perform a screen capture operation on the front-end interface to obtain a first file containing the content of the web page displayed in the front-end interface, wherein the first file includes at least one of an image file and a video file.
[0192] The processing module 402 may be configured to process the first file using the first multimodal large model to obtain first data of the web page. Here, the first data is used to represent content data of the web page that the attacker can obtain through the attack behavior.
[0193] The identification module 403 may be configured to match the first data with pre-acquired content data of the web page to identify data leakage risks of the web page.
[0194] In some embodiments, the screen capture module 401 may be configured to select a screenshot operation or a screen recording operation among the screen capture operations performed on the front-end interface according to the content update frequency of the web page.
[0195] In some embodiments, the screen capture module 401 can be configured to perform at least one of the following operations: in response to the content update frequency being greater than a first threshold, selecting to perform a screen recording operation on the front-end interface to obtain a video file containing the content of the web page as a first file; and in response to the content update frequency being less than or equal to the first threshold, selecting to perform a screenshot operation on the front-end interface to obtain an image file containing the content of the web page as a first file.
[0196] In some embodiments, the content data of the pre-acquired webpage has a specified format. The processing module 402 can be configured to input the first file and the prompt word text corresponding to the first file into the first multimodal large model to obtain the content data of the first file output by the first multimodal large model, and the prompt word text is used to instruct the first multimodal large model to output the content data of the first file according to the specified format; based on the content data of the first file, the first data is obtained.
[0197] In some embodiments, the processing module 402 may be configured to remove the hallucination data generated by the first multimodal large model from the content data of the first file to obtain the first data.
[0198] In some embodiments, the first file includes multiple files. The processing module 402 may be configured to input each file in the multiple files and the corresponding prompt text of each file into the first multi-modal large model in parallel to obtain the content data of each file output by the first multi-modal large model; and store the content data of each file in the same specified file.
[0199] In some embodiments, the prompt text corresponding to each file is the same.
[0200] In some embodiments, the recognition module 403 may be configured to, in response to the matching rate between the first data and the content data of the pre-acquired web page being less than or equal to a second threshold in the result of the matching process, correct the data in the first data that does not match the content data of the pre-acquired web page to obtain second data; and perform a matching process on the second data and the content data of the pre-acquired web page to identify the data leakage risk of the web page.
[0201] In some embodiments, the first data includes multiple first sub-data that do not match the content data of the pre-acquired web page and second sub-data that match the content data of the pre-acquired web page. Each first sub-data in the multiple first sub-data includes a target element that does not match the content data of the pre-acquired web page, and the target elements in each first sub-data are the same. The recognition module 403 may be configured to correct the target elements in each first sub-data to obtain the corrected multiple first sub-data; and determine the corrected multiple first sub-data and the second sub-data as the second data.
[0202] In some embodiments, the recognition module 403 may be configured to perform at least one of the following operations: perform a screen capture operation on the front-end interface again to obtain a second file containing the content of the web page, and process the second file using the first multi-modal large model to obtain second data; preprocess the first file, and process the preprocessed first file using the first multi-modal large model to obtain second data; or process the first file using a second multi-modal large model different from the first multi-modal large model to obtain second data, where the number of parameters of the second multi-modal large model is greater than the number of parameters of the first multi-modal large model.
[0203] In some embodiments, the recognition module 403 may be configured to, in response to the matching rate between the first data and the content data of the pre-acquired web page being greater than the second threshold in the result of the matching process, output a test report on the data leakage risk of the web page.
[0204] Figure 5 The block diagram showing other embodiments of the data leakage risk recognition device of the present disclosure.
[0205] As Figure 5As shown, the identification device 500 for data leakage risk in this embodiment includes: a memory 501 and a processor 502 coupled to the memory 501. The processor 502 is configured to execute the method for identifying data leakage risk in any one of the embodiments of the present disclosure based on the instructions stored in the memory 501.
[0206] Among them, the memory 501 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, a database, and other programs.
[0207] Figure 6 The block diagram showing still some other embodiments of the identification device for data leakage risk of the present disclosure.
[0208] As Figure 6 As shown, the identification device 600 for data leakage risk in this embodiment includes: a memory 601 and a processor 602 coupled to the memory 601. The processor 602 is configured to execute the method for identifying data leakage risk in any one of the foregoing embodiments based on the instructions stored in the memory 601.
[0209] The memory 601 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.
[0210] The electronic device 600 may further include an input / output interface 603, a network interface 604, a storage interface 605, etc. These interfaces 603, 604, 605 and the memory 601 and the processor 602 may be connected through a bus 606, for example. Among them, the input / output interface 603 provides a connection interface for input / output devices such as a display, a mouse, a keyboard, a touch screen, a microphone, a speaker, etc. The network interface 604 provides a connection interface for various networking devices. The storage interface 605 provides a connection interface for external storage devices such as an SD card and a USB flash drive.
[0211] The embodiments of the present disclosure also provide a computer-readable storage medium, including computer program instructions, and when the computer program instructions are executed by a processor, the method for identifying data leakage risk in any one of the above embodiments is implemented.
[0212] The embodiments of the present disclosure also provide a computer program product, including a computer program, and when the computer program is executed by a processor, the method for identifying data leakage risk in any one of the above embodiments is implemented.
[0213] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0214] So far, the technical solutions for data leakage risk identification according to the present disclosure have been described in detail. In order to avoid obscuring the concept of the present disclosure, some details well known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.
[0215] The methods and systems of the present disclosure can be implemented in many ways. For example, the methods and systems of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present disclosure are not limited to the specific order described above unless otherwise specifically stated. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0216] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for identifying data leakage risks, comprising: Performing a screen capture operation on the front-end interface to obtain a first file containing the content of the web page displayed in the front-end interface, where the first file includes at least one of an image file or a video file; Processing the first file using a first multi-modal large model to obtain first data, where the first data is used to characterize the content data of the web page obtained by the attacker through an attack behavior; Performing a matching process on the first data and the pre-obtained content data of the web page to identify the data leakage risk of the web page.
2. The recognition method according to claim 1, wherein The performing a screen capture operation on the front-end interface includes: Selecting to perform a screen capture operation or a screen recording operation in the screen capture operation on the front-end interface according to the content update frequency of the web page.
3. The recognition method according to claim 2, wherein, The selecting to perform a screen capture operation or a screen recording operation in the screen capture operation on the front-end interface according to the content update frequency of the web page includes: In response to the content update frequency being greater than a first threshold, selecting to perform a screen recording operation on the front-end interface to obtain a video file containing the content of the web page as the first file; and / or In response to the content update frequency being less than or equal to the first threshold, selecting to perform a screen capture operation on the front-end interface to obtain an image file containing the content of the web page as the first file.
4. The recognition method according to claim 1, wherein The pre-obtained content data of the web page has a specified format. The processing the first file using a first multi-modal large model to obtain first data includes: Inputting the first file and the prompt text corresponding to the first file into the first multi-modal large model to obtain the content data of the first file output by the first multi-modal large model, where the prompt text is used to instruct the first multi-modal large model to output the content data of the first file in the specified format; Obtaining the first data based on the content data of the first file.
5. The recognition method according to claim 4, wherein, The obtaining the first data based on the content data of the first file includes: Removing the hallucination data generated by the first multi-modal large model from the content data of the first file to obtain the first data.
6. The recognition method according to claim 4, wherein, The first file includes multiple files. The inputting the first file and the prompt text corresponding to the first file into the first multi-modal large model to obtain the content data of the first file output by the first multi-modal large model includes: Inputting each file in the multiple files and the prompt text corresponding to each file into the first multi-modal large model in parallel to obtain the content data of each file output by the first multi-modal large model; Storing the content data of each file in the same specified file.
7. The recognition method according to claim 6, wherein, The prompt text corresponding to each file is the same.
8. The recognition method according to any one of claims 1-7, wherein, The performing a matching process on the first data and the pre-obtained content data of the web page to identify the data leakage risk of the web page includes: In response to the result of the matching process that the matching rate between the first data and the content data of the pre-acquired web page is less than or equal to a second threshold, correct the data in the first data that does not match the content data of the pre-acquired web page to obtain second data; Perform a matching process on the second data and the content data of the pre-acquired web page to identify the data leakage risk of the web page.
9. The recognition method according to claim 8, wherein, The first data includes a plurality of first sub-data that do not match the content data of the pre-acquired web page and second sub-data that match the content data of the pre-acquired web page. Each first sub-data in the plurality of first sub-data includes a target element that does not match the content data of the pre-acquired web page, and the target elements in each first sub-data are the same. The correcting the data in the first data that does not match the content data of the pre-acquired web page to obtain second data includes: Correct the target element in each first sub-data to obtain a plurality of corrected first sub-data; Determine the plurality of corrected first sub-data and the second sub-data as the second data.
10. The recognition method according to claim 8, wherein, The correcting the data in the first data that does not match the content data of the pre-acquired web page to obtain second data includes at least one of the following methods: Perform the screen capture operation on the front-end interface again to obtain a second file containing the content of the web page, and use the first multi-modal large model to process the second file to obtain the second data; Perform preprocessing on the first file, and use the first multi-modal large model to process the preprocessed first file to obtain the second data; Or Use a second multi-modal large model different from the first multi-modal large model to process the first file to obtain the second data, and the number of parameters of the second multi-modal large model is greater than the number of parameters of the first multi-modal large model.
11. The recognition method according to any one of claims 1-7, wherein, The performing a matching process on the first data and the actual content data of the pre-acquired web page to identify the data leakage risk of the web page includes: In response to the result of the matching process that the matching rate between the first data and the content data of the pre-acquired web page is greater than the second threshold, output a test report on the data leakage risk of the web page.
12. An apparatus for identifying data leakage risk, comprising: A screen capture module configured to perform a screen capture operation on a front-end interface to obtain a first file containing the content of a web page displayed in the front-end interface, where the first file includes at least one of an image file or a video file; A processing module configured to use a first multi-modal large model to process the first file to obtain first data of the web page, where the first data is used to represent the content data of the web page obtained by an attacker through an attack behavior; An identification module configured to perform a matching process on the first data and the content data of the pre-acquired web page to identify the data leakage risk of the web page.
13. An apparatus for identifying data leakage risks, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the method for identifying data leakage risks according to any one of claims 1-11 based on instructions stored in the memory.
14. A computer-readable storage medium having computer instructions stored thereon, which when executed by a processor implement the method for identifying data leakage risks according to any one of claims 1-11.
15. A computer program product comprising instructions that, when executed by a processor, cause the processor to execute the method for identifying data leakage risks according to any one of claims 1-11.