Agent-based sensitive data detection method and device, and agent

By combining the decision-making and execution layers of the intelligent agent with a large artificial intelligence model, the problems of accuracy in sensitive data detection and efficiency in parsing unknown data formats are solved, achieving efficient and accurate sensitive data detection and parsing.

CN121029704BActive Publication Date: 2026-06-19BEIJING CHANGDIWANFANG TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CHANGDIWANFANG TECH CO LTD
Filing Date
2025-08-22
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

In the process of data collection, transmission, storage and delivery, existing technologies are unable to effectively detect and identify sensitive data, cannot meet the increasingly stringent legal and regulatory requirements, and lack sufficient efficiency and accuracy in parsing unknown data formats.

Method used

The decision-making layer of the intelligent agent uses a large artificial intelligence model to determine the parsing strategy, the execution layer performs file parsing, and sensitive data is detected through the execution layer. Combined with the feedback and learning layer, a standard parser is generated, which improves the ability to parse unknown data formats and the accuracy of sensitive data detection.

Benefits of technology

It improves the comprehensiveness and accuracy of sensitive data detection, expands the coverage of analysis scenarios, and enhances the efficiency and accuracy of the analysis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029704B_ABST
    Figure CN121029704B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, intelligent agent, electronic device, storage medium, and computer program product for sensitive data detection based on intelligent agents, specifically involving intelligent agents, large-scale artificial intelligence models, and sensitive data detection. It can be applied to sensitive data detection scenarios. The specific implementation scheme is as follows: Through the decision-making layer of the intelligent agent, using a large-scale artificial intelligence model, a parsing strategy for the file to be detected is determined based on the cognitive state of the data format of the file to be detected; through the decision-making layer, the file to be detected is parsed according to the parsing strategy to obtain the target parsed data; through the execution layer of the intelligent agent, whether the target parsed data contains sensitive data is detected, thus obtaining the sensitive data detection result. This disclosure improves the coverage of the parsing operation across parsing scenarios, as well as the accuracy and efficiency of the parsing process, thereby improving the comprehensiveness and accuracy of the sensitive data detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the fields of intelligent agents, large-scale artificial intelligence models, and sensitive data detection, and particularly to a sensitive data detection method, device, intelligent agent, electronic device, storage medium, and computer program product based on intelligent agents, which can be applied to sensitive data detection scenarios. Background Technology

[0002] In fields such as natural resources, intelligent driving, and industrial manufacturing, enterprises often handle sensitive information during data collection, transmission, storage, and delivery. With the continuous improvement of data compliance laws and regulations such as the Data Security Law, the Cybersecurity Law, the Notice of the Ministry of Natural Resources on Promoting the Development of Intelligent Connected Vehicles and Maintaining the Security of Surveying and Mapping Geographic Information, and the Several Provisions on the Management of Vehicle Data Security, relevant enterprises are facing increasingly stringent regulatory requirements in data compliance management. Summary of the Invention

[0003] This disclosure provides a method, apparatus, agent, electronic device, storage medium, and computer program product for sensitive data detection based on an intelligent agent.

[0004] According to the first aspect, a sensitive data detection method based on intelligent agents is provided, comprising: through the decision layer of the intelligent agent, using a large artificial intelligence model, determining the parsing strategy of the file to be detected based on the cognitive state of the data format of the file to be detected; through the decision layer, parsing the file to be detected according to the parsing strategy to obtain target parsed data; and through the execution layer of the intelligent agent, detecting whether the target parsed data includes sensitive data to obtain sensitive data detection results.

[0005] According to the second aspect, a sensitive data detection device based on an intelligent agent is provided, comprising: a strategy determination unit configured to determine a parsing strategy for the file to be detected based on the cognitive state of the data format of the file to be detected by utilizing an artificial intelligence big data model through the decision layer of the intelligent agent; a data parsing unit configured to parse the file to be detected according to the parsing strategy through the decision layer to obtain target parsed data; and a data detection unit configured to detect whether the target parsed data contains sensitive data through the execution layer of the intelligent agent to obtain a sensitive data detection result.

[0006] According to the third aspect, an intelligent agent for detecting sensitive data is provided, comprising: a perception layer for identifying the file to be detected; a decision layer for using a large artificial intelligence model to determine the parsing strategy of the file to be detected based on the cognitive state of the data format of the file to be detected, and parsing the file to be detected according to the parsing strategy to obtain the target parsed data; and an execution layer for detecting whether the target parsed data includes sensitive data to obtain the sensitive data detection result.

[0007] According to a fourth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect.

[0008] According to a fifth aspect, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the method described in any implementation of the first aspect.

[0009] According to a sixth aspect, a computer program product is provided, comprising: a computer program that, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0010] According to the technology disclosed herein, a sensitive data detection method and apparatus based on intelligent agents are provided. By utilizing a large artificial intelligence model, the parsing strategy of the file to be detected is determined based on the cognitive state of the data format of the file to be detected, and the file to be detected is parsed according to the parsing strategy to obtain the target parsed data. This improves the coverage of the parsing operation on the parsing scenario, as well as the accuracy and efficiency of the parsing process, thereby improving the comprehensiveness and accuracy of the sensitive data detection results.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This is an exemplary system architecture diagram that can be applied to an embodiment of this disclosure;

[0014] Figure 2 This is a flowchart of an embodiment of the agent-based sensitive data detection method according to the present disclosure;

[0015] Figure 3 This is a schematic diagram illustrating an application scenario of the agent-based sensitive data detection method according to this embodiment;

[0016] Figure 4 This is a flowchart of yet another embodiment of the agent-based sensitive data detection method according to the present disclosure;

[0017] Figure 5This is a structural diagram of an embodiment of the agent-based sensitive data detection device according to the present disclosure;

[0018] Figure 6 This is a structural diagram of an embodiment of an intelligent agent for detecting sensitive data according to the present disclosure;

[0019] Figure 7 This is a schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present disclosure. Detailed Implementation

[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0022] Figure 1 An exemplary architecture 100 is shown that can be applied to the agent-based sensitive data detection method and apparatus disclosed herein.

[0023] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The communication connections between terminal devices 101, 102, and 103 form a network topology. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0024] Terminal devices 101, 102, and 103 can be hardware or software that supports network connectivity for data interaction and processing. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connectivity, information acquisition, interaction, display, and processing functions, including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as, for example, multiple software programs or software modules to provide distributed services, or as a single software program or software module. No specific limitations are imposed here.

[0025] Server 105 can be a server providing various services, such as a background processing server that receives files to be detected sent by the target object through terminal devices 101, 102, and 103, and detects whether the files contain sensitive data based on an intelligent agent. Optionally, the server can feed back the sensitive data detection results to the terminal devices. As an example, server 105 can be a cloud server.

[0026] It should be noted that a server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules (such as software programs or software modules used to provide distributed services), or as a single software program or software module. No specific limitations are made here.

[0027] It should also be noted that the sensitive data detection method based on intelligent agents provided in the embodiments of this disclosure is generally executed by a server, but the possibility of it being executed by a terminal device, or by the server and the terminal device cooperating with each other, is not excluded. Accordingly, the various parts (e.g., various units) included in the sensitive data detection device based on intelligent agents can be all set in the server, all set in the terminal device, or set in the server and the terminal device respectively.

[0028] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. When the electronic device on which the agent-based sensitive data detection method runs does not need to transmit data with other electronic devices, the system architecture may only include the electronic device (e.g., a terminal device or server) on which the agent-based sensitive data detection method runs.

[0029] Please refer to Figure 2 , Figure 2 A flowchart illustrating a sensitive data detection method based on an intelligent agent, provided in this embodiment of the disclosure. Flowchart 200 includes the following steps:

[0030] Step 201: Through the decision-making layer of the intelligent agent, the large artificial intelligence model is used to determine the parsing strategy of the file to be detected based on the cognitive state of the data format of the file to be detected.

[0031] In this embodiment, the executing entity of the sensitive data detection method based on the intelligent agent (e.g., Figure 1The server in the system can obtain the file to be detected from a remote location or from a local location via a wired or wireless network connection. Then, through the decision-making layer of the intelligent agent, it uses a large artificial intelligence model to determine the parsing strategy of the file to be detected based on its cognitive state of the data format of the file.

[0032] The file to be detected is the file for which sensitive data detection is to be performed. It is sent by the target object to the aforementioned executing entity. The target object can be, for example, a person, other smart devices, or an AI assistant. Sensitive data refers to information that, if accessed, used, disclosed, modified, or destroyed without authorization, may cause damage to the legitimate rights and interests of individuals, organizations, enterprises, or nations (such as privacy breaches, property losses, reputational damage, security risks, etc.). This type of data typically possesses characteristics of privacy, confidentiality, or sensitivity, and its protection must comply with relevant laws and regulations (such as the Personal Information Protection Law, the Data Security Law, etc.) and industry standards. Sensitive data can cover multiple dimensions, for example:

[0033] Sensitive personal data: ID card number, bank account number, biometric information (fingerprint, facial data), medical records, whereabouts, etc.

[0034] Sensitive corporate data: trade secrets (such as core technology solutions, customer lists), financial statements, undisclosed strategic plans, etc.

[0035] Sensitive national data: classified information, critical infrastructure operation data, defense-related data, etc.

[0036] The aforementioned implementing entities are equipped with intelligent agents for detecting sensitive data. The decision-making layer within these agents can invoke a large-scale artificial intelligence model to determine the parsing strategy for the file to be detected based on the model's understanding of the data format of the file to be detected.

[0037] Large-scale artificial intelligence models (hereinafter referred to as "large models") refer to a class of artificial intelligence models with a large number of parameters constructed by artificial neural networks. The main categories include: large language models, large visual models, multimodal large models, and basic science large models. Taking the large language model as an example in this embodiment, it is a large-scale language model built based on deep learning technology, mainly used for processing natural language processing tasks. Through training on large-scale data, it learns language patterns and structures, enabling it to generate natural language text or understand natural language input. This embodiment specifically adopts a multimodal large language model, which typically includes the following modules:

[0038] Input module: Receives the file to be detected from user input. The file to be detected can be at least one modality of multimodal data such as text, image, video, and voice.

[0039] Preprocessing module: preprocesses the input files to be tested.

[0040] The encoding module encodes the preprocessed multimodal data into vector form so that the model can understand and process it. Common encoding methods include word embeddings and encoders in the Transformer architecture. Word embeddings include, for example, Word2Vec (Words to Vector) and GloVe (Global Vectors for Word Representation).

[0041] Model module: The core component, typically based on a deep learning architecture (such as Transformer), responsible for processing encoded data vectors and performing language understanding and generation. The model learns complex patterns and semantic relationships of language through multi-layered neural network structures.

[0042] Decoding module: Decodes the model's output vectors into natural language text, images, videos, and speech, generating responses or processing results for user input. Decoding methods can include greedy decoding, beam search, etc.

[0043] Output module: Outputs the decoded text in a user-readable format.

[0044] Data format includes encoding rules (such as character encoding, compression algorithms, encryption methods, etc.) and data structure (such as field distribution, hierarchical relationships, delimiters, etc.). The AI ​​large-scale model's understanding of the data format of the file to be tested can be either known or unknown. In the known state, the AI ​​large-scale model's knowledge base contains the data format followed by the file to be tested, and there are clear and reusable parsing strategies (parsing rules), eliminating the need for additional reasoning or guessing about the file. Examples of files to be tested in the known state include:

[0045] 1. Common standard formats, such as JSON (JavaScript Object Notation), XML (Extensible Markup Language), PNG (Portable Network Graphics), H.264, etc., have publicly available encoding rules and structural definitions, and their parsing strategies, for example, are to perform targeted parsing according to the standard parsing process corresponding to the data format.

[0046] 2. Standardized proprietary protocols within the industry, such as the DICOM (Digital Imaging and Communications in Medicine) format in the medical field (if the system already understands the meaning of its metadata fields).

[0047] The core of its parsing process is to follow the "structured parsing + verification" process of the format specification, as follows: First, read the binary stream of the file to be tested and check whether it conforms to the basic specifications of the target format (such as JSON needs to verify bracket matching and quotation mark closure; XML needs to check the legality of tag nesting; PNG needs to verify the file header magic number "89 50 4E 47" and CRC check). If it does not conform, a format error is returned.

[0048] Then, the data is broken down according to the syntax rules of the format. For JSON parsing libraries, the text stream is converted into structured data such as dictionaries and lists in memory according to syntax such as "key-value pairs", "arrays", and "nested objects" (e.g., Python's json.loads()); for XML parsers, the DOM (Document Object Model) tree or SAX (Simple API for XML) event stream is built according to the structure such as "tag hierarchy", "attributes", and "text nodes", mapping the parent-child relationship between tags (e.g., Java's DocumentBuilder); for PNG decoders, the metadata and payload of each block are split and parsed according to the block structure of "file header → IHDR (Image Header) block (image size / format) → IDAT (Image Data) block (compressed pixel data) → IEND (Image End) block".

[0049] Then, the parsed structured data is converted into a form that the application can directly use. For JSON / XML data, specific field values ​​are extracted (e.g., from...). <user> <id> 123< / id> < / user> Extracting the id as 123, processing data types (such as converting JSON numbers and booleans to their corresponding programming language types); for PNG data, decompressing the compressed data of the IDAT block (using the Deflate algorithm), and combining it with the bit depth and color type in IHDR to restore it to an RGB / A pixel matrix for image rendering.

[0050] Finally, it outputs the parsed structured data (such as memory objects and pixel arrays), while handling any exceptions that may occur during parsing (such as syntax errors in JSON, unclosed tags in XML, and block validation failures in PNG), and returning specific error information to assist in debugging.

[0051] In the unknown state, the knowledge base of the large artificial intelligence model does not include the data format followed by the file to be tested, and there is no ready-made parsing strategy. It is necessary to dynamically explore data format information such as data structure and encoding logic.

[0052] Files to be detected in an unknown state include, for example:

[0053] 1. There are standardized proprietary protocols in the industry, but their data format knowledge is not included in the knowledge base, or due to the flexibility of the data contained within them, it is not clear which encoding formats are included, such as MF4 (Measurement Data Format 4) for custom channels.

[0054] 2. Completely custom proprietary protocols, such as sensor data formats developed internally by the enterprise (without publicly disclosed structure definitions) or special files that are encrypted or obfuscated (with opaque encoding rules).

[0055] 3. Mixed format files, such as a single file containing multiple unlabeled sub-protocols (e.g., a binary stream containing both text and image data with unknown encodings, without clear delimiters).

[0056] For files with unknown data formats to be detected, the AI ​​model formulates a parsing strategy based on the file characteristics. For example, the file is first divided into fixed-size segments, and the most frequent byte sequence in each segment is extracted as a candidate delimiter. The file is then split using the delimiter. If sub-blocks of similar length are obtained, the byte distribution within each sub-block is analyzed. If a certain type of sub-block contains consecutive 0x00s, it is inferred to be a padding field and removed, and the remaining part is analyzed. For segments without obvious delimiters, a sliding window (e.g., 100 bytes) is used to scan, recording special values ​​(e.g., 0xFF, 0x0A) appearing within the window, associating them with known format markers (e.g., image frame start), and attempting to parse them using the corresponding format parsing logic. If the parsing attempt fails, the window size is adjusted and the scan is re-engaged. The data dimension is inferred by combining the temporal relationship between sub-blocks (e.g., increasing values), and the parsing strategy is gradually optimized.

[0057] In some optional implementations of this embodiment, the execution entity can perform step 201 as follows: in response to the fact that the cognitive state of the artificial intelligence big model of the file to be detected is unknown, the artificial intelligence big model is used to determine the parsing strategy according to the file type of the file to be detected.

[0058] In response to the fact that the AI ​​model's understanding of the file to be detected is unknown, the AI ​​model is used to determine the data storage method of the file to be detected based on its file type; and based on the data storage method, a parsing strategy is determined.

[0059] Taking MF4 files as an example, their data storage method involves organizing and storing data according to channels. Each channel represents a physical quantity or ECU (Electronic Control Unit) signal, such as engine speed, vehicle speed, and temperature. Based on the channel storage method, the parsing strategy can be determined to be to parse channel by channel as the parsing unit.

[0060] Taking PCAP (Packet Capture Library) files as an example, they employ a sequential storage method based on data packets. Based on this data packet storage method, the parsing strategy can be determined as parsing data packets one by one.

[0061] In this implementation, for files whose data format is unknown, the corresponding parsing strategy is determined based on the file type, which improves the accuracy of the parsing strategy and its compatibility with the files to be detected, and helps to improve the speed and accuracy of file parsing.

[0062] Step 202: The decision-making layer parses the file to be detected according to the parsing strategy to obtain the target parsing data.

[0063] In this embodiment, the aforementioned execution entity can parse the file to be detected according to the parsing strategy through the decision layer in the intelligent agent to obtain the target parsing data.

[0064] For a file to be tested whose data format is known, the file is parsed according to the "structured parsing + verification" process that follows the format specifications.

[0065] For files whose data format is unknown, the parsing strategy specified by the large-scale artificial intelligence model is used. For example, the file is first divided into fixed-size segments, and the most frequent byte sequence in each segment is extracted as a candidate delimiter. The file is split using the delimiter. If sub-blocks of similar length are obtained, the byte distribution within the sub-blocks is statistically analyzed. If a certain type of sub-block contains consecutive 0x00s, it is inferred to be a padding field and removed, and the remaining part is analyzed. For segments without obvious delimiters, a sliding window (e.g., 100 bytes) is used to scan, and special values ​​appearing in the window (e.g., 0xFF, 0x0A) are recorded. These values ​​are associated with known format markers (e.g., the start of an image frame), and parsing is attempted using the corresponding format parsing logic. If the parsing attempt fails, the window size is adjusted and the scan is repeated. The data dimension is inferred by combining the temporal relationship between sub-blocks (e.g., increasing values), and the parsing strategy is gradually optimized.

[0066] In some optional implementations of this embodiment, the execution entity can perform step 202 as follows:

[0067] The first step is to determine the parsing units in the file to be tested based on the parsing strategy.

[0068] A parsing unit refers to a basic unit in the file parsing process that possesses independent metadata (such as type, format, and range), can be located and processed independently, and does not depend on other units for complete parsing. It generally has clear boundaries (such as data start / end positions and identifier fields), contains independent format information (such as encoding method and data structure), and can be parsed as an independent step (without waiting for the parsing results of other units).

[0069] Taking MF4 files as an example, its parsing strategy is to parse each channel as a parsing unit, thus determining that the parsing unit in the file to be detected is the channel.

[0070] Taking the PCAP file as an example, its parsing strategy is to parse data packets one by one, thus determining that the parsing unit in the file to be tested is the data packet.

[0071] The second step is to determine the encoding format of the parsing unit based on its data characteristics, and then parse the parsing unit according to the encoding format to obtain the unit parsing data.

[0072] For each parsing unit, the encoding format of the parsing unit is determined according to the data characteristics of the parsing unit, and the parsing unit is parsed according to the encoding format to obtain the unit parsing data.

[0073] The file to be tested is generally stored in binary. First, the binary features of the parsing unit are extracted, including the first 16 bytes of the file header sequence, the overall byte entropy value (to determine whether it is compressed or encrypted), and the repeating byte pattern. The file header is compared with a known format signature library. If an approximate format is matched (e.g., 6 bytes overlap with the first 8 bytes of a PNG file header), the corresponding format parsing tool is called for tentative parsing. If a logical structure (e.g., image size, pixel values) is parsed, the same strategy is used.

[0074] The third step is to combine the unit parsing data of each parsing unit to obtain the target parsing data.

[0075] For each parsing unit, the aforementioned execution entity may successfully parse the parsing unit and obtain the unit parsing data of the parsing unit, or the aforementioned execution entity may fail to parse the parsing unit and not obtain the unit parsing data of the parsing unit.

[0076] For each successfully parsed unit, the target parsed data is obtained by combining the unit parsed data of each parsed unit.

[0077] This implementation provides a method for parsing the file to be detected based on a parsing strategy. It attempts to parse each parsing unit according to the parsing strategy, which further improves the coverage of the parsing operation on the parsing scenario, as well as the accuracy and efficiency of the parsing process.

[0078] In some optional implementations of this embodiment, the execution entity can perform the second step as follows:

[0079] The following parsing operations are performed iteratively until the preset termination condition is met:

[0080] First, the encoding format of the parsing unit is determined based on its data characteristics. Then, the parsing unit is parsed according to the encoding format to obtain the parsed data. Finally, in response to the parsed data not being presented as the target data modality, the next parsing operation is performed.

[0081] As an example, for each parsing unit in the file to be detected, firstly, the encoding format of the parsing unit is predicted based on its data characteristics; then, based on the encoding format, the corresponding parser is determined, and the parsing unit is parsed to obtain the unit parsed data; next, it is determined whether the unit parsed data presents the target data modality, which is generally human-readable data, such as text, images, etc. If the unit parsed data is in a human-unreadable format, such as garbled text, it indicates that the parsing process for the parsing unit has failed. Finally, the next parsing operation is performed, that is, the encoding format of the parsing unit is predicted again based on its data characteristics.

[0082] The preset termination condition is, for example, successful parsing of a parsed unit, that is, the parsed data of the unit is presented as the target data modality, or the possible encoding formats of the parsed unit predicted based on the data characteristics have all been tried, but the parsed unit is still not successfully parsed.

[0083] In this implementation, the parsing data of the unit obtained by the parsing operation when the preset termination condition is met can be used as the parsing data of the corresponding parsing unit.

[0084] This implementation provides a method for attempting to parse parsing units in the file to be detected based on a loop, which helps to improve the success rate of the parsing process and thus helps to improve the comprehensiveness of the sensitive data detection results.

[0085] In some optional implementations of this embodiment, the execution entity can perform the second step as follows: using an artificial intelligence big data model, determine the encoding format of the parsing unit based on the file characteristics of the parsing unit, and parse the parsing unit according to the encoding format to obtain the unit parsing data.

[0086] Specifically, using a large artificial intelligence model, the following parsing operations are iteratively executed until a preset termination condition is met: First, the encoding format of the parsing unit is determined based on the data characteristics of the parsing unit; then, the parsing unit is parsed according to the encoding format to obtain the parsed data of the unit; then, in response to the fact that the parsed data of the unit does not present the target data modality, the next parsing operation is executed.

[0087] In this implementation, the large AI model acts as a determiner of the encoding format, a multimodal parser of the parsing unit, and a determiner of whether parsing is successful. It iteratively executes the parsing process of the parsing unit. With the help of the powerful logical analysis and data processing capabilities of the large AI model, it helps to further improve the success rate and efficiency of the parsing operation.

[0088] In some optional implementations of this embodiment, in response to the fact that the cognitive state of the artificial intelligence big model of the file to be detected is known, the above-mentioned execution subject can perform the above step 202 in the following way: use the standard parser corresponding to the data format to parse the file to be detected and obtain the target parsing data.

[0089] First, a parser is used to parse the file according to its data format. Then, a standard parser is used to parse the file to obtain the target parsing data.

[0090] For example, for JSON files, a JSON parsing library is used to convert the text stream into structured data such as dictionaries and lists in memory, based on syntax such as "key-value pairs," "arrays," and "nested objects" (e.g., Python's json.loads()). For XML files, an XML parser is used to build a DOM tree or SAX event stream based on structures such as "tag hierarchy," "attributes," and "text nodes," mapping the parent-child relationships between tags (e.g., Java's DocumentBuilder). For PNG files, a PNG decoder is used to split and parse the metadata and payload of each block according to the block structure of "file header → IHDR block (image size / format) → IDAT block (compressed pixel data) → IEND block."

[0091] In this implementation, for a file to be detected whose data format is known, the decoding strategy is to use the standard parser corresponding to the data format to parse the file to be detected, which helps to further improve the parsing efficiency and parsing success rate of the file to be detected.

[0092] Step 203: Through the execution layer of the intelligent agent, detect whether the target parsed data includes sensitive data, and obtain the sensitive data detection result.

[0093] In this embodiment, the aforementioned execution entity can detect whether the target parsed data includes sensitive data through the execution layer of the intelligent agent, and obtain the sensitive data detection result.

[0094] As an example, the target parsed data can be compared with sensitive data in the sensitive database. If the similarity between the target parsed data and the sensitive data is higher than a preset similarity threshold, it indicates that the target parsed data includes sensitive data; otherwise, it indicates that the target parsed data does not include sensitive data, and the final sensitive data detection result is obtained.

[0095] In some optional implementations of this embodiment, the execution entity can perform step 203 as follows:

[0096] The first step is to determine the target toolchain used to detect whether the target parsing data contains sensitive data, based on the target data modality to which the target parsing data belongs.

[0097] The target toolchain includes at least one tool, which, when executed sequentially, can perform the task of detecting sensitive data. For each data modality, a corresponding target toolchain can be preset, so that after determining the target data modality to which the target parsed data belongs, the target toolchain used to detect whether the target parsed data contains sensitive data can be determined according to the preset correspondence.

[0098] For example, for the text modality, the corresponding target toolchain includes text preprocessing tools and text scanners. The text preprocessing tools are used to perform preprocessing on the text data, such as word segmentation, while the text scanners are used to determine whether the preprocessed text data contains sensitive data.

[0099] For example, for image modalities, the corresponding target toolchain includes an OCR (Optical Character Recognition) engine, text preprocessing tools, a text scanner, and an image processing engine. The OCR engine is used to recognize text in image data, while the image processing engine is used to understand the image content.

[0100] The second step involves using tools from the target toolchain to detect whether the target parsed data includes sensitive data, thus obtaining the sensitive data detection results.

[0101] Each tool in the target toolchain is called sequentially, and the data output by the previous tool is used as input for data processing. Finally, the sensitive data detection task is completed, and the sensitive data detection result is obtained.

[0102] In some implementations, the tools in the toolchain can be based on a large AI model. For example, different tools can be implemented using preset prompts corresponding to their respective functions, all based on the large AI model. For instance, the preset prompt might be "As an OCR engine, please recognize the text in the image." Inputting this preset prompt into the large AI model instructs it to run as an OCR engine.

[0103] In this implementation, a target toolchain is determined based on the target data modality to which the target parsed data belongs, in order to perform the sensitive data detection task of the parsed data in a targeted manner, which helps to improve the accuracy and efficiency of sensitive data detection results.

[0104] In some optional implementations of this embodiment, the target parsing data includes data from multiple target data modalities. For example, the target parsing data includes image data and text data.

[0105] In this implementation, the aforementioned execution entity can perform the second step described above to determine the sensitive data detection result in the following manner:

[0106] First, for the data of multiple target data modalities associated in the target parsing data, the tools in the target toolchain corresponding to the target data modal are used to detect whether the data of the target data modal includes sensitive data, and the initial detection results are obtained.

[0107] In this implementation, data from different modalities with the same or adjacent timestamps in the file to be detected can be used as associated data; data from different modalities that have semantic relationships can be used as associated data, for example, image content and text content in image modal data can be used as associated data.

[0108] For data from multiple related target data modalities, tools from the target toolchain corresponding to the target data modal are used to detect whether the data of the target data modal includes sensitive data, and initial detection results and result confidence are obtained.

[0109] Then, the initial detection results of each of the multiple target data modalities are cross-validated to obtain the sensitive data detection results corresponding to each of the multiple target data modalities.

[0110] For the initial detection results of various target data modalities, cross-validation is performed based on their confidence levels. For example, if the text modal data contains sensitive data with a high confidence level, and its associated image data also contains sensitive data with a high confidence level, then the detection results for sensitive data corresponding to each of the various target data modalities are determined to contain sensitive data.

[0111] In some implementations, to further improve the accuracy of cross-validation results, the initial detection results of multiple target data modalities are cross-validated according to the type of sensitive data. For example, sensitive data included in image data and text data represent sensitive entities such as the same location, so as to obtain the detection results of sensitive data corresponding to multiple target data modalities, thereby ensuring the correlation of sensitive data in related different modalities.

[0112] In this implementation, cross-validation based on detection results from different modal data improves the accuracy of the final sensitive data detection results.

[0113] In some implementations, in response to the sensitive data detection results indicating the presence of sensitive data in the data to be detected, a data anonymization method is determined based on the data type of the sensitive data; the sensitive data is then processed according to the data anonymization method to obtain the anonymized data. For example, for video data, the data anonymization method is to replace the video frame containing the sensitive data with a black frame; for text data, sensitive words and other sensitive data parts are deleted.

[0114] See also Figure 3 , Figure 3 This is a schematic diagram 300 illustrating an application scenario of the agent-based sensitive data detection method according to this embodiment. First, an agent 302 is deployed in the server 301. The agent includes a perception layer 3021, a decision layer 3022, and an execution layer 3023. The target object 303 sends the file to be detected to the server through the corresponding terminal device 304. The server, through the agent's decision layer, utilizes a large-scale artificial intelligence model to determine the parsing strategy for the file to be detected based on its understanding of the file's data format. The decision layer then parses the file according to the parsing strategy to obtain the target parsed data. Finally, the agent's execution layer detects whether the target parsed data contains sensitive data, thus obtaining the sensitive data detection result.

[0115] This embodiment provides a sensitive data detection method based on an intelligent agent. By utilizing a large artificial intelligence model, the method determines the parsing strategy for the file to be detected based on the cognitive state of the data format of the file to be detected, and parses the file to be detected according to the parsing strategy to obtain the target parsed data. This improves the coverage of the parsing operation on the parsing scenario, as well as the accuracy and efficiency of the parsing process, thereby improving the comprehensiveness and accuracy of the sensitive data detection results.

[0116] In some optional implementations of this embodiment, the execution entity may also perform the following operations: through the feedback and learning layer in the intelligent agent, standardize the processing of the file to be detected for unknown data formats, and generate a standard parser corresponding to the data format of the file to be detected.

[0117] As an example, firstly, after completing the iterative parsing of the unknown data format file, the agent automatically extracts all valid operation trajectories during this parsing process, including successful encoding format predictions (such as encoding method, field separator), called toolchains (such as specific decoders, feature extractors), parameter configurations (such as byte order, block size), and manual feedback correction records, forming the original parsing rule set.

[0118] Then, the original rule set is structurally abstracted, separating "encoding layer rules" (such as compression algorithms and character sets) and "structural layer rules" (such as field length and hierarchical nesting relationships). Through feature tagging tools, core features such as file header signatures and key delimiters are bound to the rules to form a preliminary format description template.

[0119] Then, the format description template is validated using multiple unknown files of the same type within the same domain (those not involved in the initial parsing). If the parsing success rate exceeds a preset threshold, the core rules of the template are retained. If there are failed cases, the rules that are not covered (such as special field variants) are located by comparing differences and added to the template to form a candidate parser rule base.

[0120] Then, a standardized parser code framework is generated based on the candidate rule base, encapsulating input and output interfaces (such as receiving binary streams and outputting structured dictionaries), integrating tool calling modules (such as automatically associating with the corresponding decoders), and embedding error handling logic (such as a fallback strategy when field parsing fails) to ensure that the standard parser can be directly called by the agent.

[0121] Finally, the newly generated standard parser is incorporated into the system knowledge base, and dynamic optimization trigger conditions are set. For example, when multiple new errors occur in parsing similar files, the parsing rules of the new cases are automatically extracted and merged with the existing parser rules to achieve continuous iteration of the parser and gradually form a standardized parsing capability covering a class of unknown formats.

[0122] The feedback and learning layer is a functional layer independent of the perception, decision-making, and execution layers, but it relies on the outputs of these three layers to achieve its core function, forming a closed-loop collaborative relationship. The feedback and learning layer does not directly participate in data perception (perception layer function), strategy formulation (decision layer function), or operation execution (execution layer function). Instead, it independently undertakes the function of "generating a standard parser based on the normalization of collected data," possessing its own rule base and learning module. However, the operation of the feedback and learning layer depends on the perception, decision-making, and execution layers. The data it collects includes the data to be processed determined by the perception layer, historical parsing process data from the decision-making layer, and operational process data from the execution layer (such as execution process data based on the toolchain). Through the aforementioned normalization process, the feedback and learning layer extracts the standard parser, which is then fed back to the decision-making layer to guide the iteration of subsequent parsing strategies, ultimately achieving the autonomous evolution of the entire system's parsing capabilities.

[0123] In this implementation, based on the standardized processing of files to be detected with unknown data formats, a standard parser corresponding to the data format of the file to be detected is generated. This gradually enriches the understanding of the data formats and the types of standard parsers, which helps to improve the efficiency and accuracy of the agent in the subsequent parsing process.

[0124] In some optional implementations of this embodiment, the execution entity may also perform the following operations:

[0125] First, for the remaining data in the file to be tested that failed to be parsed, the remaining data is parsed according to the received parsing operation to obtain auxiliary parsing data.

[0126] The parsing operation can be a target object, such as an experienced technician, to assist in parsing the remaining data and obtain auxiliary parsing data.

[0127] Then, based on the target data modality to which the auxiliary parsing data belongs, an auxiliary toolchain is determined for detecting whether the auxiliary parsing data includes sensitive data.

[0128] Finally, tools from the auxiliary toolchain are used to detect whether the auxiliary parsing data includes sensitive data.

[0129] In this implementation, you can refer to the relevant examples of the target toolchain mentioned above.

[0130] Furthermore, for data from multiple target data modalities associated with the auxiliary parsing data, tools from the auxiliary toolchain corresponding to the target data modal are used to detect whether the data of the target data modal includes sensitive data, and initial detection results are obtained. Cross-validation is then performed on the initial detection results of each of the multiple target data modalities to obtain the sensitive data detection results corresponding to each of the multiple target data modalities.

[0131] In this implementation, the remaining data in the file to be detected that failed to be parsed is parsed based on the received parsing operation, which improves the comprehensiveness and flexibility of the parsing process.

[0132] In some optional implementations of this embodiment, the execution entity may also perform the following operations: through the feedback and learning layer in the agent, standardize the processing of the remaining data and generate a standard parser corresponding to the data format of the remaining data.

[0133] As an example, firstly, based on the manual review results and parsing logs during the processing of the remaining data, the agent automatically extracts all valid operation trajectories during this parsing process, including successful encoding format predictions (such as encoding method, field separator), called toolchains (such as specific decoders, feature extractors), parameter configurations (such as byte order, block size), and manual feedback correction records, forming the original parsing rule set.

[0134] Then, the original rule set is structurally abstracted, separating "encoding layer rules" (such as compression algorithms and character sets) and "structural layer rules" (such as field length and hierarchical nesting relationships). Through feature tagging tools, core features such as file header signatures and key delimiters are bound to the rules to form a preliminary format description template.

[0135] Then, the template is validated using multiple unknown files of the same type within the same domain (those not involved in the initial parsing). If the parsing success rate exceeds a preset threshold, the core rules of the template are retained. If there are failed cases, the rules that are not covered (such as special field variants) are located by comparing differences and added to the template to form a candidate parser rule base.

[0136] Then, a standardized parser code framework is generated based on the candidate rule base, encapsulating input and output interfaces (such as receiving binary streams and outputting structured dictionaries), integrating tool calling modules (such as automatically associating with the corresponding decoders), and embedding error handling logic (such as a fallback strategy when field parsing fails) to ensure that the standard parser can be directly called by the agent.

[0137] Finally, the newly generated standard parser is incorporated into the system knowledge base, and dynamic optimization trigger conditions are set. For example, when multiple new errors occur in parsing similar files, the parsing rules of the new cases are automatically extracted and merged with the existing parser rules to achieve continuous iteration of the parser and gradually form a standardized parsing capability covering a class of unknown formats.

[0138] In this implementation, based on the standardized processing of remaining data for unknown data formats, a standard parser corresponding to the data format of the file to be detected is generated. This gradually enriches the understanding of the data format and the standard parser, which helps to improve the efficiency and accuracy of the agent in the subsequent parsing process.

[0139] Continue to refer to Figure 4 This illustrates an illustrative flow 400 of yet another embodiment of the agent-based sensitive data detection method according to the present disclosure. Flow 400 includes the following steps:

[0140] Step 401: In response to the fact that the AI ​​model's understanding of the data format of the file to be detected is unknown, the AI ​​model is used to determine the parsing strategy for the file to be detected based on its file type.

[0141] Step 402: Determine the parsing units in the file to be detected according to the parsing strategy.

[0142] Step 403: Using a large-scale artificial intelligence model, iteratively execute the following parsing operations until the preset termination condition is met:

[0143] Step 4031: Determine the encoding format of the parsing unit based on its data characteristics;

[0144] Step 4032: Parse the parsing unit according to the encoding format to obtain the unit parsing data;

[0145] Step 4033: In response to the fact that the parsed data of the cell is not presented as the target data modality, perform the next parsing operation.

[0146] Step 404: Combine the unit parsing data of each parsing unit to obtain the target parsing data.

[0147] Step 405: Determine the target toolchain for detecting whether the target parsing data includes sensitive data, based on the target data modality to which the target parsing data belongs.

[0148] Step 406: Use tools from the target toolchain to detect whether the target parsed data includes sensitive data, and obtain the sensitive data detection results.

[0149] Step 407: Standardize the processing procedure for files to be tested with unknown data formats, and generate a standard parser corresponding to the data format of the files to be tested.

[0150] The process 400 of the agent-based sensitive data detection method in this embodiment specifically illustrates the parsing process of the file to be detected, the detection process of sensitive data, and the generation process of the standard parser. This improves the coverage of the parsing operation on the parsing scenario, as well as the accuracy and efficiency of the parsing process, thereby improving the comprehensiveness and accuracy of the sensitive data detection results.

[0151] Continue to refer to Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a sensitive data detection device based on an intelligent agent. This system embodiment is similar to... Figure 2 Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.

[0152] like Figure 5 As shown, the sensitive data detection device 500 based on intelligent agents includes: a strategy determination unit 501, configured to determine the parsing strategy of the file to be detected by using an artificial intelligence big data model through the decision layer of the intelligent agent, based on the cognitive state of the data format of the file to be detected; a data parsing unit 502, configured to parse the file to be detected according to the parsing strategy through the decision layer to obtain target parsed data; and a data detection unit 503, configured to detect whether the target parsed data contains sensitive data through the execution layer of the intelligent agent to obtain a sensitive data detection result.

[0153] In some optional implementations of this embodiment, the strategy determination unit 501 is further configured to: in response to the fact that the cognitive state of the artificial intelligence big model of the file to be detected is unknown, use the artificial intelligence big model to determine the parsing strategy according to the file type of the file to be detected.

[0154] In some optional implementations of this embodiment, the data parsing unit 502 is further configured to: determine the parsing units in the file to be detected according to the parsing strategy; determine the encoding format of the parsing units according to the data characteristics of the parsing units, and parse the parsing units according to the encoding format to obtain unit parsing data; and combine the unit parsing data of each parsing unit to obtain target parsing data.

[0155] In some optional implementations of this embodiment, the data parsing unit 502 is further configured to iteratively execute the following parsing operations until a preset termination condition is met: determine the encoding format of the parsing unit based on the data characteristics of the parsing unit; parse the parsing unit according to the encoding format to obtain the parsed data of the unit; and execute the next parsing operation in response to the fact that the parsed data of the unit does not present as the target data modality.

[0156] In some optional implementations of this embodiment, the data parsing unit 502 is further configured to: use an artificial intelligence big data model to determine the encoding format of the parsing unit based on the file characteristics of the parsing unit, and parse the parsing unit according to the encoding format to obtain the parsed data of the unit.

[0157] In some optional implementations of this embodiment, the data parsing unit 502 is further configured to: in response to the cognitive state being known, use a standard parser corresponding to the data format to parse the file to be detected and obtain the target parsing data.

[0158] In some optional implementations of this embodiment, the data detection unit 503 is further configured to: determine a target toolchain for detecting whether the target parsing data includes sensitive data based on the target data modality to which the target parsing data belongs; use the tools in the target toolchain to detect whether the target parsing data includes sensitive data, and obtain a sensitive data detection result.

[0159] In some optional implementations of this embodiment, the target parsing data includes data of multiple target data modalities, and the data detection unit 503 is further configured to: for the data of multiple target data modalities associated in the target parsing data, use the tools in the target toolchain corresponding to the target data modal to detect whether the data of the target data modal includes sensitive data, and obtain an initial detection result; perform cross-validation on the initial detection results of each of the multiple target data modalities to obtain the sensitive data detection results corresponding to each of the multiple target data modalities.

[0160] In some optional implementations of this embodiment, the above apparatus further includes: a feedback learning unit (not shown in the figure), configured to: standardize the processing of a file to be detected with an unknown data format through the feedback and learning layer in the agent, and generate a standard parser corresponding to the data format of the file to be detected.

[0161] In some optional implementations of this embodiment, the above apparatus further includes: an auxiliary parsing unit (not shown in the figure), configured to parse the remaining data in the file to be detected that failed to be parsed, according to the received parsing operation, to obtain auxiliary parsing data; and the data detection unit 503 is further configured to: determine an auxiliary toolchain for detecting whether the auxiliary parsing data includes sensitive data according to the target data modality to which the auxiliary parsing data belongs; and use the tools in the auxiliary toolchain to detect whether the auxiliary parsing data includes sensitive data.

[0162] In some optional implementations of this embodiment, the feedback learning unit (not shown in the figure) is further configured to: standardize the processing of the remaining data through the feedback and learning layer in the agent, and generate a standard parser corresponding to the data format of the remaining data.

[0163] This embodiment provides a sensitive data detection device based on an intelligent agent. By utilizing a large artificial intelligence model, the device determines the parsing strategy for the file to be detected based on the cognitive state of the data format of the file to be detected, and parses the file to be detected according to the parsing strategy to obtain the target parsed data. This improves the coverage of the parsing operation on the parsing scenario, as well as the accuracy and efficiency of the parsing process, thereby improving the comprehensiveness and accuracy of the sensitive data detection results.

[0164] Continue to refer to Figure 6 This disclosure provides an embodiment of an intelligent agent for detecting sensitive data, and the system embodiment is similar to... Figure 2 Corresponding to the method embodiments shown, this intelligent agent can be specifically applied to various electronic devices.

[0165] like Figure 6 As shown, an intelligent agent for detecting sensitive data includes: a perception layer 601, used to determine the file to be detected; a decision layer 602, used to use an artificial intelligence big data model to determine the parsing strategy of the file to be detected based on the cognitive state of the data format of the file to be detected, and to parse the file to be detected according to the parsing strategy to obtain the target parsed data; and an execution layer 603, used to detect whether the target parsed data includes sensitive data and to obtain the sensitive data detection result.

[0166] In some optional implementations of this embodiment, the decision layer 602 is further used to: in response to the fact that the cognitive state of the artificial intelligence big model of the file to be detected is unknown, to use the artificial intelligence big model to determine the parsing strategy according to the file type of the file to be detected.

[0167] In some optional implementations of this embodiment, the decision layer 602 is further used to: determine the parsing units in the file to be detected according to the parsing strategy; determine the encoding format of the parsing units according to the data characteristics of the parsing units, and parse the parsing units according to the encoding format to obtain unit parsing data; and combine the unit parsing data of each parsing unit to obtain target parsing data.

[0168] In some optional implementations of this embodiment, the decision layer 602 is further used to: iteratively execute the following parsing operations until a preset termination condition is met: determine the encoding format of the parsing unit based on the data characteristics of the parsing unit; parse the parsing unit according to the encoding format to obtain the unit parsing data; and execute the next parsing operation in response to the unit parsing data not being presented as the target data modality.

[0169] In some optional implementations of this embodiment, the decision layer 602 is further used to: utilize a large artificial intelligence model to determine the encoding format of the parsing unit based on the file characteristics of the parsing unit, and parse the parsing unit according to the encoding format to obtain the unit parsing data.

[0170] In some optional implementations of this embodiment, the decision layer 602 is further used to: in response to the known cognitive state, use a standard parser corresponding to the data format to parse the file to be detected and obtain the target parsing data.

[0171] In some optional implementations of this embodiment, the execution layer 603 is further used to: determine a target toolchain for detecting whether the target parsed data includes sensitive data based on the target data modality to which the target parsed data belongs; use the tools in the target toolchain to detect whether the target parsed data includes sensitive data, and obtain a sensitive data detection result.

[0172] In some optional implementations of this embodiment, the target parsing data includes data of multiple target data modalities, and the execution layer 603 is further used to: for the data of multiple target data modalities associated in the target parsing data, use the tools in the target toolchain corresponding to the target data modal to detect whether the data of the target data modal includes sensitive data, and obtain an initial detection result; perform cross-validation on the initial detection results of each of the multiple target data modalities to obtain the sensitive data detection results corresponding to each of the multiple target data modalities.

[0173] In some optional implementations of this embodiment, the intelligent agent further includes a feedback learning layer (not shown in the figure), which is used to: standardize the processing of the file to be detected for unknown data formats through the feedback and learning layer in the intelligent agent, and generate a standard parser corresponding to the data format of the file to be detected.

[0174] In some optional implementations of this embodiment, the decision layer 602 is further configured to: an auxiliary parsing unit (not shown in the figure) is configured to parse the remaining data in the file to be detected that failed to be parsed, according to the received parsing operation, to obtain auxiliary parsing data; and the execution layer 603 is further configured to: determine an auxiliary toolchain for detecting whether the auxiliary parsing data includes sensitive data according to the target data modality to which the auxiliary parsing data belongs; and use the tools in the auxiliary toolchain to detect whether the auxiliary parsing data includes sensitive data.

[0175] In some optional implementations of this embodiment, the feedback learning layer (not shown in the figure) is used to: standardize the processing of the remaining data through the feedback and learning layer in the agent, and generate a standard parser corresponding to the data format of the remaining data.

[0176] In this embodiment, an intelligent agent for detecting sensitive data is provided. Utilizing a large artificial intelligence model, the agent determines the parsing strategy for the file to be detected based on its cognitive state of the data format of the file to be detected, and parses the file to be detected according to the parsing strategy to obtain the target parsed data. This improves the coverage of the parsing operation for the parsing scenario, as well as the accuracy and efficiency of the parsing process, thereby improving the comprehensiveness and accuracy of the sensitive data detection results.

[0177] According to embodiments of this disclosure, this disclosure also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the agent-based sensitive data detection method described in any of the above embodiments.

[0178] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to implement the agent-based sensitive data detection method described in any of the above embodiments when executed.

[0179] This disclosure provides a computer program product that, when executed by a processor, can implement the agent-based sensitive data detection method described in any of the above embodiments.

[0180] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0181] like Figure 7As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded into random access memory (RAM) 703 from storage unit 708. The RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0182] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0183] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the agent-based sensitive data detection method. For example, in some embodiments, the agent-based sensitive data detection method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the agent-based sensitive data detection method described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform an agent-based sensitive data detection method by any other suitable means (e.g., by means of firmware).

[0184] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0185] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to the processor or controller of a general-purpose computer, special-purpose computer, or other programmable agent-based sensitive data detection device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0186] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0187] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0188] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0189] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service system to address the management difficulties and weak business scalability inherent in traditional physical hosts and Virtual Private Servers (VPS) services; they can also be servers for distributed systems or servers incorporating blockchain technology.

[0190] According to the technical solution of the embodiments of this disclosure, a sensitive data detection method and apparatus based on intelligent agents are provided. By utilizing a large artificial intelligence model, the parsing strategy of the file to be detected is determined based on the cognitive state of the data format of the file to be detected, and the file to be detected is parsed according to the parsing strategy to obtain the target parsed data. This improves the coverage of the parsing operation on the parsing scenario, as well as the accuracy and efficiency of the parsing process, thereby improving the comprehensiveness and accuracy of the sensitive data detection results.

[0191] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.

[0192] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A sensitive data detection method based on intelligent agents, comprising: Through the decision-making layer of the intelligent agent, using a large artificial intelligence model, the parsing strategy of the file to be detected is determined based on the cognitive state of the data format of the file to be detected; The decision layer parses the file to be detected according to the parsing strategy to obtain target parsing data, including: for a file to be detected in an unknown state, determining the parsing units in the file to be detected according to the parsing strategy; determining the encoding format of the parsing unit according to the data characteristics of the parsing unit, and parsing the parsing unit according to the encoding format to obtain unit parsing data; and combining the unit parsing data of each parsing unit to obtain the target parsing data. The execution layer of the intelligent agent detects whether the target parsed data includes sensitive data, and obtains the sensitive data detection result.

2. The method of claim 1, wherein, The method of utilizing a large-scale artificial intelligence model to determine the parsing strategy for the file to be detected based on the understanding of the data format of the file to be detected includes: In response to the fact that the AI ​​model's perception of the file to be detected is unknown, the AI ​​model is used to determine the parsing strategy based on the file type of the file to be detected.

3. The method of claim 1, wherein, The step of determining the encoding format of the parsing unit based on its data characteristics, and parsing the parsing unit according to the encoding format to obtain unit parsing data, includes: The following parsing operations are performed iteratively until the preset termination condition is met: The encoding format of the parsing unit is determined based on the data characteristics of the parsing unit; The parsing unit is parsed according to the encoding format to obtain the parsed data of the unit; If the parsed data from the unit does not present the target data modality, the next parsing operation is performed.

4. The method of claim 1, wherein, The step of determining the encoding format of the parsing unit based on its data characteristics, and parsing the parsing unit according to the encoding format to obtain unit parsing data, includes: Using the aforementioned large-scale artificial intelligence model, the encoding format of the parsing unit is determined based on the file characteristics of the parsing unit, and the parsing unit is parsed according to the encoding format to obtain the parsed data of the unit.

5. The method of claim 1, wherein, In response to the known cognitive state, the step of parsing the file to be detected according to the parsing strategy to obtain target parsing data includes: The file to be detected is parsed using the standard parser corresponding to the data format to obtain the target parsing data.

6. The method of any one of claims 1-5, wherein, The process of detecting whether the target parsed data includes sensitive data and obtaining sensitive data detection results includes: Based on the target data modality to which the target parsed data belongs, determine the target toolchain used to detect whether the target parsed data includes sensitive data; Using the tools in the target toolchain, detect whether the target parsed data includes sensitive data, and obtain the sensitive data detection result.

7. The method of claim 6, wherein, The target parsing data includes data from multiple target data modalities, and The step of using tools from the target toolchain to detect whether the target parsed data includes sensitive data, and obtaining the sensitive data detection result, includes: For the data of multiple target data modalities associated with the target parsing data, the tools in the target toolchain corresponding to the target data modal are used to detect whether the data of the target data modal includes sensitive data, and an initial detection result is obtained; Cross-validation is performed on the initial detection results of each of the various target data modalities to obtain the sensitive data detection results corresponding to each of the various target data modalities.

8. The method of claim 1 or 3, wherein, Also includes: Through the feedback and learning layers in the intelligent agent, the processing procedure for files to be detected with unknown data formats is standardized, and a standard parser corresponding to the data format of the files to be detected is generated.

9. The method of claim 1, wherein, Also includes: For the remaining data in the file to be detected that failed to be parsed, the remaining data is parsed according to the received parsing operation to obtain auxiliary parsing data; Based on the target data modality to which the auxiliary parsing data belongs, determine an auxiliary toolchain for detecting whether the auxiliary parsing data includes sensitive data; The tools in the auxiliary toolchain are used to detect whether the auxiliary parsing data includes sensitive data.

10. The method of claim 9, wherein, Also includes: Through the feedback and learning layers in the intelligent agent, the processing procedure for the remaining data is standardized, and a standard parser corresponding to the data format of the remaining data is generated.

11. A sensitive data detection device based on intelligent agents, comprising: The strategy determination unit is configured to determine the parsing strategy of the file to be detected by using a large artificial intelligence model through the decision layer of the intelligent agent, based on the cognitive state of the data format of the file to be detected. The data parsing unit is configured to parse the file to be detected according to the parsing strategy through the decision layer to obtain target parsing data, including: for the file to be detected in an unknown state, determining the parsing units in the file to be detected according to the parsing strategy; determining the encoding format of the parsing unit according to the data characteristics of the parsing unit, and parsing the parsing unit according to the encoding format to obtain unit parsing data; and combining the unit parsing data of each parsing unit to obtain the target parsing data. The data detection unit is configured to detect whether the target parsed data includes sensitive data through the execution layer of the intelligent agent, and obtain the sensitive data detection result.

12. An intelligent agent for detecting sensitive data, comprising: The perception layer is used to identify the file to be detected; The decision layer utilizes a large-scale artificial intelligence model to determine the parsing strategy for the file to be tested based on its cognitive state regarding the data format of the file to be tested, and parses the file to be tested according to the parsing strategy to obtain target parsed data. This includes: for a file to be tested in an unknown state, determining the parsing units in the file to be tested according to the parsing strategy; determining the encoding format of the parsing unit based on its data characteristics, and parsing the parsing unit according to the encoding format to obtain unit parsing data; and combining the unit parsing data of each parsing unit to obtain the target parsed data. The execution layer is used to detect whether the target parsed data includes sensitive data and obtain the sensitive data detection result.

13. An electronic device, comprising: include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.

14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

15. A computer program product comprising: A computer program that, when executed by a processor, implements the method according to any one of claims 1-10.