A method for privacy-secured identification of data and related devices

CN117633227BActive Publication Date: 2026-08-21JIAXING RES INST ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311648269.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2026-08-21
Estimated Expiration
2043-12-04

AI Technical Summary

Technical Problem

另外,相关技术中所检测的隐私数据类型是预定义好的,输出结果是封闭的,没有预定义外开放类型的识别能力

Benefits of technology

[0034]综上,本申请实施例的数据的隐私安全识别方法包括:获取原始数据信息和文本提示信息,其中,上述文本提示信息是用户基于自定义场景需求仿照最佳实践的文本提示模板进行编辑或修改后生成的;对上述原始数据信息分别基于数据的模态类型进行对应的预处理操作以获取模态数据信息;基于上述模态数据信息和上述文本提示信息进行跨模态交互以文本提示的方式识别提示中的开放对象,以获取第一隐私安全类型识别结果,其中,上述第一隐私安全类型识别结果是基于上述文本提示信息的每个上述文本提示模板引导生成的;根据所有的上述第一隐私安全类型识别结果经过简明化后处理获取隐私安全识别结果。本申请实施例提出的一种数据的隐私安全识别方法,本申请针对解决非结构化数据的识别,基于提示工程的跨模态技术设计的模型在识别范围上能以文本提示信息识别预定义外的结果,理论识别范围是开放的,由文本提示的灵活性,能够做到指代表达理解。该方法支持多种数据类型,包括图片、视频、音频和文本数据,可以用于分析和保护不同数据形式的隐私,提供了广泛的适用性。同时,本方法相对于需要大量精心标注样本的监督模型,对高质量训练样本的需求较少,仅需要在对齐阶段构造上千个高质量样本即可,在其他阶段大量的训练样本都是公开获取的,相比于需要大量精心标注样本的监督模型样本标注成本更低。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117633227B_ABST
    Figure CN117633227B_ABST
Patent Text Reader

Abstract

The application discloses a privacy security identification method of data and related equipment, and relates to the field of data identification. The method comprises the following steps: obtaining original data information and text prompt information, wherein the text prompt information is generated by selecting, editing or modifying a text prompt template based on the best practice according to the user's self-defined scene demand; performing corresponding preprocessing operations on the original data information based on the modal type of the data to obtain modal data information; performing cross-modal interaction based on the modal data information and the text prompt information to identify the open object in the prompt in the form of text prompt, so as to obtain a first privacy security type identification result, wherein the first privacy security type identification result is generated based on the text prompt template of the text prompt information; and obtaining a privacy security identification result after simplifying all the first privacy security type identification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of data identification, and more specifically, this application relates to a method and related equipment for identifying data privacy and security. Background Technology

[0002] Currently, the types of privacy-related data structures processed by technology are still limited to structured data, namely the initial data in specific implementations, interface parameters, interface request fields, database SQL statements, and interface request return values, etc. Generally speaking, structured data is highly organized and neatly formatted data with a predefined data model, and its storage and arrangement in the computer follows certain rules. In addition, the privacy-related data types detected in related technologies are predefined, and the output results are closed, lacking the ability to identify predefined externally exposed types. Summary of the Invention

[0003] The summary section introduces a series of simplified concepts, which will be further explained in detail in the detailed description section. This summary section is not intended to limit the key and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.

[0004] Firstly, this application proposes a method for identifying data privacy and security, the method comprising:

[0005] Obtain raw data information and text prompt information. The text prompt information is generated by the user after editing or modifying the text prompt template based on the best practice according to the user's custom scenario requirements.

[0006] The above raw data information is preprocessed according to the modal type of the data to obtain modal data information;

[0007] Based on the above modal data information and the above text prompt information, cross-modal interaction is performed to identify open objects in the prompts in the form of text prompts, so as to obtain a first privacy and security type identification result, wherein the above first privacy and security type identification result is generated based on each of the above text prompt templates of the above text prompt information;

[0008] The privacy security identification result is obtained by simplifying and post-processing all the above-mentioned first privacy security type identification results.

[0009] In one feasible implementation, the aforementioned raw data information includes at least one modality such as image data information, video data information, audio data information, and text data information, and the aforementioned preprocessing operations include anomaly removal operations and input standardization operations.

[0010] In one feasible implementation, the text prompt template of the above best practice includes a concise privacy and security type template and a privacy and security type referential expression understanding template. Users can edit and modify the text prompt template of the above best practice according to their own usage scenario needs to form the above text prompt information. The concise privacy and security type template is formed by directly splicing the concise type with specific separators. The privacy and security type referential expression understanding template is a template with a specific scenario-specific potential identification object and only identifies objects or behaviors that specifically refer to the state described. The privacy and security type referential expression understanding template is used to help the model understand the referential expression in the text prompt in order to determine the specific object or behavior referred to in the text prompt.

[0011] In one feasible implementation, the above method further includes:

[0012] Based on the above modal data information and predefined supervised models, the identification results of the second privacy security type are obtained;

[0013] The privacy security identification result is obtained by simplifying all the above first privacy security type identification results and the above second privacy security type identification results.

[0014] In one feasible implementation, the above-mentioned cross-modal interaction based on the above-mentioned modal data information and the above-mentioned text prompt information to identify open objects in the prompts in the form of text prompts, in order to obtain a first privacy and security type identification result, includes:

[0015] The modal data is processed using a variant of the Transformer architecture to extract modal features.

[0016] The above text prompt information is segmented and vectorized for feature extraction to obtain text features;

[0017] Open object recognition is performed through cross-modal fusion using a variant of the Transformer architecture based on the aforementioned modal features and text features to obtain the first privacy-preserving type recognition result.

[0018] In one feasible implementation, the open object recognition based on the modal features and text features through a Transformer architecture variant to obtain a first privacy-preserving type recognition result includes:

[0019] Feature interaction is performed on the above modal features and text features to obtain enhanced text features and enhanced modal features;

[0020] Based on the enhanced text features and enhanced modality features described above, a language-guided query operation is performed to obtain cross-modal query features;

[0021] Based on the aforementioned enhanced text features, enhanced modal features, and cross-modal query features, cross-modal decoding operations are performed to obtain the model decoding output;

[0022] Based on the decoding output of the above model, the identification result of the first privacy and security type is obtained.

[0023] In one feasible implementation, the above method further includes:

[0024] The contrastive loss is determined based on the enhanced text features, the output of the decoding model, and the contrastive learning function described above.

[0025] The mode-specific loss is determined based on the model decoding output and the mode-specific loss function described above;

[0026] Based on the aforementioned contrastive loss and modality-specific loss, we train an open set object recognition model and adjust the model parameters to obtain the model's detection capabilities.

[0027] Secondly, this application also proposes a data privacy and security identification device, comprising:

[0028] The first acquisition unit is used to acquire raw data information and text prompt information. The text prompt information is generated by the user after editing or modifying the text prompt template based on the best practice according to the user's custom scenario requirements.

[0029] The second acquisition unit is used to perform corresponding preprocessing operations on the above-mentioned raw data information based on the modal type of the data to obtain modal data information.

[0030] The third acquisition unit is used to perform cross-modal interaction based on the above modal data information and the above text prompt information to identify open objects in the prompts in the form of text prompts, so as to obtain a first privacy and security type identification result, wherein the above first privacy and security type identification result is generated based on each of the above text prompt templates of the above text prompt information;

[0031] The fourth acquisition unit is used to obtain the privacy and security identification result after simplifying and processing all the above-mentioned first privacy and security type identification results.

[0032] Thirdly, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program stored in the memory to implement the steps of the privacy-secure identification method for data as described in any of the first aspects above.

[0033] Fourthly, this application also proposes a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the privacy-secure identification method for data of any of the above claims in the first aspect.

[0034] In summary, the data privacy and security identification method of this application embodiment includes: acquiring original data information and text prompt information, wherein the text prompt information is generated by the user after editing or modifying a text prompt template based on best practices according to custom scenario requirements; performing corresponding preprocessing operations on the original data information based on the modal type of the data to obtain modal data information; performing cross-modal interaction based on the modal data information and the text prompt information to identify open objects in the prompts in the form of text prompts to obtain a first privacy and security type identification result, wherein the first privacy and security type identification result is generated based on each of the text prompt templates of the text prompt information; and obtaining a privacy and security identification result after simplification post-processing based on all the first privacy and security type identification results. This application embodiment proposes a data privacy and security identification method that addresses the identification of unstructured data. The model designed based on cross-modal technology of prompt engineering can identify results beyond the predefined scope using text prompt information. The theoretical identification scope is open, and the flexibility of text prompts enables comprehension of the indicated meaning. This method supports multiple data types, including images, videos, audio, and text data, and can be used to analyze and protect the privacy of different data forms, providing broad applicability. Meanwhile, compared to supervised models that require a large number of carefully labeled samples, this method requires fewer high-quality training samples. It only needs to construct thousands of high-quality samples in the alignment stage, while a large number of training samples in other stages are publicly available. Therefore, the sample labeling cost is lower than that of supervised models that require a large number of carefully labeled samples.

[0035] The privacy and security identification method for data proposed in this application, as well as other advantages, objectives, and features of this application, will be partly apparent from the following description and partly understood by those skilled in the art through research and practice of this application. Attached Figure Description

[0036] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit this specification. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0037] Figure 1 This application provides a schematic flowchart of a data privacy and security identification method.

[0038] Figure 2 A schematic diagram of another data privacy and security identification method provided in this application embodiment;

[0039] Figure 3This is a schematic diagram illustrating the processing principle of an open object recognition model for cross-modal interaction in a data privacy and security identification method provided in an embodiment of this application.

[0040] Figure 4 This is a structural schematic diagram of a data privacy and security identification device provided in an embodiment of this application;

[0041] Figure 5 This is a schematic diagram of the structure of an electronic device for identifying data privacy and security, provided as an embodiment of this application. Detailed Implementation

[0042] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. The technical solutions of the embodiments of this application will now be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0043] Please see Figure 1 This is a schematic flowchart of a data privacy and security identification method provided in an embodiment of this application, which may specifically include:

[0044] S110. Obtain raw data information and text prompt information, wherein the above text prompt information is generated by the user after editing or modifying the text prompt template based on the best practice according to the user's custom scenario requirements;

[0045] For example, the raw data may include image data, video data, audio data, and text data, and the raw data is unstructured. Additionally, user-provided text prompts also need to be obtained. These text prompts can be entered by the user based on prompts from a text prompt template, or they can be generated by the user selecting text options from the provided prompt template. The purpose of the text prompts is to guide the subsequent privacy and security type identification process.

[0046] It should be noted that the method used in this application is implemented by training a conventional neural network model. Before applying the method provided in this application, the model needs to undergo generalized training and secondary fine-tuning. By collecting easily accessible Internet modal data and text description phrases such as Weibo, Instagram, and Facebook, the target model corresponding to the method used in this application is trained to learn the most generalized deep semantic connections. Finally, a small number of carefully labeled high-quality sample pairs are used to fine-tune the model so that the target model can play a guiding role in semantic reach through an open set during training.

[0047] Understandably, this method can be widely applied in industries such as finance, internet, telecommunications, and data security. This unstructured data privacy and security identification method is deployed on a server and invoked via an API, which can retrieve the raw unstructured data.

[0048] S120. Perform corresponding preprocessing operations on the original data information based on the modal type of the data to obtain modal data information;

[0049] For example, preprocessing operations are performed on the collected raw data. These preprocessing operations may include anomaly removal to detect and remove outliers or noise in the data to ensure data quality. Input normalization operations may also be included to standardize different types of data into the same format or range for subsequent processing. The raw data, categorized according to different modalities, may include image data, video data, audio data, and text data. Each modality corresponds to a specific preprocessing operation.

[0050] S130. Based on the above modal data information and the above text prompt information, perform cross-modal interaction to identify open objects in the prompts in the form of text prompts, so as to obtain a first privacy and security type identification result, wherein the above first privacy and security type identification result is generated based on each of the above text prompt templates of the above text prompt information;

[0051] For example, preprocessed modal data and user-provided text prompts are used to perform open object recognition for cross-modal interaction. The text prompts can be of various types, each generated by the user based on a desired scenario template. During the open object recognition process, cross-modal interactions can be performed separately based on each type of text prompt and modal data, resulting in multiple recognition results for different privacy and security categories.

[0052] S140. Obtain the privacy security identification result by simplifying and post-processing all the above-mentioned first privacy security type identification results.

[0053] For example, based on all the first privacy security type identification results, including at least two different types of first privacy security type identification results, these results are combined to generate the final privacy security identification result. The privacy security identification result includes a more in-depth analysis or classification of potential privacy security issues. It may include specific outputs such as concise type, confidence level, and modality.

[0054] In summary, the data privacy and security identification method proposed in this application addresses the identification of unstructured data. Based on cross-modal technology using prompting engineering, the model can identify results beyond the predefined scope using text prompts. The theoretical identification range is open, and the flexibility of text prompts enables comprehension of the intended meaning. This method supports various data types, including images, videos, audio, and text data, and can be used to analyze and protect the privacy of different data formats, providing broad applicability. Furthermore, compared to supervised models that require a large number of carefully labeled samples, this method requires fewer high-quality training samples, only needing to construct thousands of high-quality samples during the alignment stage. In other stages, a large number of training samples are publicly available, resulting in lower sample labeling costs compared to supervised models that require a large number of carefully labeled samples.

[0055] In one feasible implementation, the aforementioned raw data information includes at least one of image data information, video data information, audio data information, and text data information, and the aforementioned preprocessing operations include anomaly removal operations and input standardization operations.

[0056] For example, Figure 2 This is a schematic flowchart of another data privacy and security identification method provided in an embodiment of this application. Figure 2 As shown, the raw data includes image data, video data, audio data, and text data. The interface obtains the raw unstructured data, performs anomaly removal and data normalization, and then enters the corresponding preprocessing pipeline according to the data modality to obtain the standard model input.

[0057] In this embodiment, the data autoloader refers to the initial data used to receive unstructured data and automatically classify it into various modalities according to conventional computer storage formats. This data serves as input for the corresponding cross-modal model, which is then invoked to identify data privacy and security vulnerabilities. The raw data read by the data autoloader enters the preprocessing pipeline corresponding to its respective storage format for anomaly removal and input standardization.

[0058] It should be noted that anomaly removal mainly refers to situations such as data corruption or storage formats that are not mainstream and require special reading methods or conversions. Input standardization, on the other hand, is a data normalization transformation caused by the characteristics of data between modalities, as well as the representation transformation required by cross-modal model inputs. Typical data normalization transformations include statistical transformations, normalization, and standardization. Data representation transformations include converting images from images to multidimensional arrays, audio from Melp plots, and text from positive integer sequences that encode words with numbers.

[0059] In one feasible implementation, the text prompt template of the above best practice includes a concise privacy and security type template and a privacy and security type referential expression understanding template. Users can edit and modify the text prompt template of the above best practice according to their own usage scenario needs to form the above text prompt information. The concise privacy and security type template is formed by directly splicing the concise type with specific separators. The privacy and security type referential expression understanding template is a template with a specific scenario-specific potential identification object and only identifies objects or behaviors that specifically refer to the state described. The privacy and security type referential expression understanding template is used to help the model understand the referential expression in the text prompt in order to determine the specific object or behavior referred to in the text prompt.

[0060] For example, user-defined privacy and security type prompt templates are provided. It should be noted that this application achieves open object recognition by guiding the model to detect corresponding privacy information through deep semantic interaction between text prompt engineering and modal data. Specifically, the difference between this and a supervised model with predefined recognition types is that a predefined supervised model aims to learn a mapping or correspondence between modal data and closed, predefined privacy risk types. In contrast, cross-modal interactive open object recognition technology aims to learn the abstract semantic connection between the prompt text sequence and the modal data. Due to the flexibility of the text sequence, for the same modal data sample, a specific text prompt guides the model to focus on the corresponding semantic information, thereby identifying privacy risk information not predefined. Although the flexibility of text prompt engineering is theoretically arbitrary, this implementation example is only a partial illustration, but the potential of prompt engineering stems from the emergent general intelligence of high-level large language models.

[0061] For custom text prompt templates, one could be an extended concise privacy and security type template. For example, for images, options like "safe deposit box," "smart lock," "postcard," or "house number" could be used. These concise privacy types, not easily handled by predefined models, could be directly used as text prompt input instances by connecting them with separators. For audio, options like "stocks," "psychotherapy," or "banking consultation" could be used, but predefined models might not provide a high return on investment for these categories.

[0062] Text prompt templates can also refer to templates representing the recipient's understanding, such as "person driving," "smart lock entering password," etc., indicating the identification of the recipient. In audio, examples include "banking consultation," "telecom password change." Examples of text prompt templates represent best practices and include, but are not limited to, the two types mentioned above.

[0063] The method proposed in this application, by extending the concise privacy and security type template, allows users to customize and input privacy types that are not easily processed by predefined models. This better adapts to the specific needs of different industries and users, improving flexibility and adaptability. The template for understanding referential expressions helps the model better understand the referential expressions in text prompts. Identifying the specific objects or behaviors referred to in text prompts improves the model's understanding and analysis capabilities, reducing misunderstandings or confusion. Unlike traditional predefined supervised models, this scheme uses cross-modal interactive open object recognition technology, which can flexibly adapt to different prompt text sequences, thereby guiding the model to detect undefined privacy risk information. User-defined privacy types and referential expression understanding templates can provide more specific information, helping the system to more accurately identify and understand privacy and security issues. This reduces false alarms and misjudgments, improving the effectiveness of privacy and security management. Building this scheme on a high-level large language model, it has universality and adaptability, and can be applied to multiple industries and fields to meet the needs of different users and organizations.

[0064] In one feasible implementation, the above method further includes:

[0065] Based on the above modal data information and predefined supervised models, the identification results of the second privacy security type are obtained;

[0066] The privacy security identification result is obtained by simplifying all the above first privacy security type identification results and the above second privacy security type identification results.

[0067] For example, after obtaining the user-defined template, the recognition system first processes the concise prompt template and outputs a privacy and security type that is not predefined. Next, it processes the pointer expression template. Unlike the concise prompt template, which has a special separator, the pointer expression privacy and security type is output as a category from the named entities extracted from the template. The specific type string may differ slightly from that in the concise prompt template, requiring further post-processing to normalize and remove duplicates, thereby generating the first privacy and security type recognition result.

[0068] A second privacy-preserving type identification result is generated based on modal data information using an existing mature predefined supervised model. The final output, the privacy-preserving identification result, is obtained by combining the first and second privacy-preserving type identification results. It should be noted that the predefined supervised model covers almost all accessible AI domain supervision algorithms, including open-source algorithms and those deployed using company-owned resources to label data samples and design and train models. Deep learning supervised models have matured considerably over the past decade and are not the focus of this implementation case, but they are a standard component, covering areas including but not limited to visual, speech, and natural language supervision tasks. This primarily addresses situations where sufficient predefined labeled samples are available for model training.

[0069] The embodiments of this application combine the first privacy and security type identification result obtained by cross-modal interaction using user-defined templates with the second privacy and security type identification result generated by a mature supervised model, thereby improving the accuracy, personalization, and adaptability of privacy and security type identification.

[0070] In one feasible implementation, the above-mentioned cross-modal interaction based on the above-mentioned modal data information and the above-mentioned text prompt information to identify open objects in the prompts in the form of text prompts, in order to obtain a first privacy and security type identification result, includes:

[0071] The modal data is processed using a variant of the Transformer architecture to extract modal features.

[0072] The above text prompt information is segmented and vectorized for feature extraction to obtain text features;

[0073] Open object recognition is performed through cross-modal fusion using a variant of the Transformer architecture based on the aforementioned modal features and text features to obtain the first privacy-preserving type recognition result.

[0074] For example, Transformer is a deep learning model architecture that has achieved significant success in the field of natural language processing. It is a model for sequence-to-sequence tasks, and variants of the Transformer architecture can include swimTramsformer, BERT, etc. In this embodiment, the modal data information is processed using a Transformer architecture variant, such as SwimTransformer, to extract features and obtain modal features. The text prompt information is then segmented and vectorized, such as using BERT feature extraction, to obtain text features. Based on the modal features and the text features, cross-modal fusion open object recognition is performed using a Transformer architecture variant, such as the modal bidirectional module in GLIP, to obtain a first privacy-preserving type recognition result.

[0075] This embodiment includes two inputs: one is the original data information, and the other is the text prompt information generated based on the text prompt template. After acquiring the preprocessed modal data and the text prompt information, the method encodes the modal data (such as images, audio, video, and text) into abstract vector feature sequences using a modal-side encoder within the model. The text prompt information is then encoded into word vector sequences using a text encoder. In other words, feature extraction is performed on the modal data information to obtain modal features, and feature extraction is performed on the aforementioned text prompt information to obtain text features.

[0076] The Transformer self-attention mechanism, or its variants, extracts deep features through several cross-modal information interactions of varying designs. These features are then fed into a modality-related prediction head, which outputs the first privacy-preserving type identification result. After appropriate post-processing, the result typically outputs a specific privacy-preserving category along with its corresponding confidence level and probability. Categories exceeding a certain threshold are classified as having privacy risks.

[0077] In one feasible implementation, the open object recognition based on the modal features and text features through a Transformer architecture variant to obtain a first privacy-preserving type recognition result includes:

[0078] S310. Perform feature interaction on the above modal features and the above text features to obtain enhanced text features and enhanced modal features;

[0079] S320. Perform language-guided query operations based on the above-mentioned enhanced text features and enhanced modal features to obtain cross-modal query features;

[0080] S330. Based on the above-mentioned enhanced text features, enhanced modal features, and cross-modal query features, perform cross-modal decoding operations to obtain the model decoding output;

[0081] S340. Based on the above model decoding output, obtain the above first privacy security type identification result.

[0082] For example, such as Figure 3 The diagram shown illustrates the principle of an open object recognition model for cross-modal interaction in a data privacy and security identification method provided in this application. Input modalities and text are abstracted and represented through modules such as encoders, decoders, and language query guidance. The interaction between text and modal data is further facilitated by a query module that guides the model to focus on features in the modal data related to the semantic representation of the text, ultimately outputting the result.

[0083] Specifically, the modal features and text features mentioned above are subjected to feature interaction through a feature enhancement module to obtain enhanced text features and enhanced modal features. Based on the enhanced text features and enhanced modal features, a language-guided query operation is performed through a covariance-guided query module to obtain cross-modal query features. Based on the enhanced text features, enhanced modal features, and cross-modal query features, a cross-modal decoding operation is performed through a cross-modal decoder to obtain model decoding output. Based on the model decoding output, the first privacy and security type identification result is obtained.

[0084] In one feasible implementation, the above method further includes:

[0085] The contrastive loss is determined based on the enhanced text features, the output of the decoding model, and the contrastive learning function described above.

[0086] The mode-specific loss is determined based on the above model decoding output and the mode-specific loss function;

[0087] Based on the aforementioned contrastive loss and modality-specific loss, we train an open set object recognition model and adjust the model parameters to obtain the model's detection capabilities.

[0088] For example, enhancing text features involves improving or enhancing text feature engineering through modal features to increase the expressiveness of text data. This may include using more advanced natural language processing techniques, text embedding vectorization methods, or other feature engineering techniques to capture more information and semantics from the text.

[0089] The output of the decoding model includes the model's classification prediction of the input data or the generation of text descriptions.

[0090] Contrastive learning is typically used to compare the similarity or difference between a model's output and the true label or other samples. It can be used to train a model to make its output more accurate or discriminative. The contrastive loss function is a measure of similarity and difference; minimizing this loss function optimizes the model.

[0091] Modal loss functions are used to measure modality-specific loss, focusing on the loss of specific types of data (such as images, text, audio, etc.), helping the model to better adapt to different types of data and features, and improving the model's generalization performance.

[0092] This approach aims to improve the classification performance of Transformer models by enhancing text features, contrastive learning, and modal loss functions, enabling them to better handle text and other modal data. By introducing enhanced text features, contrastive learning, and modal loss functions at different levels, it seeks to improve the classification performance of deep learning models on multimodal data.

[0093] It should be noted that the principle of contrastive learning in this application is to use a proxy task or proxy target loss to update the feature representation similarity of the text and modal data pairs, thereby connecting the generalized semantic association between the prompt text and the modal data to realize the text semantic guidance of the identification focus of the modal data.

[0094] In one implementation, Figure 2 A method for identifying data privacy and security is presented, which specifically includes:

[0095] S210: Obtain raw unstructured data through the port, and then perform anomaly removal, data normalization, and standardization processing.

[0096] S220. Input text prompt information through a text prompt template, wherein the text prompt template includes concise object type, referential understanding object type and exploration type.

[0097] S230. According to the data mode, the data enters the corresponding preprocessing pipeline to obtain standard modal data information. The modal processing pipeline includes image modal preprocessing pipeline, video modal preprocessing pipeline, audio modal preprocessing pipeline and text modal preprocessing pipeline, etc.

[0098] S240. Through a predefined supervised model, learn a mapping relationship or correspondence between modal data and closed predefined privacy risk types, and identify the corresponding modal data information to obtain the second privacy security type identification result;

[0099] S250. In the corresponding cross-modal interaction open object recognition model, the modal encoder encodes modal data such as images, audio, video, and text into abstract vector feature sequences. The text prompt information is encoded into word vector sequences by the text encoder. Through the Transformer self-attention mechanism architecture or variant, deep features are extracted through several cross-modal information interactions with different designs and enter the modality-related prediction head to output multiple first privacy and security type recognition results. The first privacy and security type recognition results include predefined concise privacy and security type recognition results guided by text prompt templates, privacy and security type recognition results with specific feature indicators representing the understanding of the object guided by text prompt templates, and exploratory privacy and security recognition results that extract high-level large model domain knowledge guided by text prompt templates.

[0100] S260. By integrating the first privacy security type identification result and the second privacy security type identification result, the final privacy security identification result is obtained.

[0101] Please see Figure 4 One embodiment of the data privacy and security identification device in this application may include:

[0102] The first acquisition unit 21 is used to acquire raw data information and text prompt information, wherein the text prompt information is generated by the user after editing or modifying the text prompt template based on the best practice according to the user's custom scenario requirements;

[0103] The second acquisition unit 22 is used to perform corresponding preprocessing operations on the above-mentioned raw data information based on the modal type of the data to obtain modal data information;

[0104] The third acquisition unit 23 is used to perform cross-modal interaction based on the above modal data information and the above text prompt information to identify open objects in the prompts in the form of text prompts, so as to obtain a first privacy and security type identification result, wherein the above first privacy and security type identification result is generated based on each of the above text prompt templates of the above text prompt information;

[0105] The fourth acquisition unit 24 is used to simplify and acquire privacy and security identification results based on all the above-mentioned first privacy and security type identification results.

[0106] like Figure 5 As shown, this application embodiment also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any of the above-mentioned methods for privacy and security identification of data.

[0107] Since the electronic device described in this embodiment is a device used to implement a data privacy and security identification device in the embodiments of this application, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the method described in the embodiments of this application. Therefore, how the electronic device implements the method in the embodiments of this application will not be described in detail here. Any device used by those skilled in the art to implement the method in the embodiments of this application falls within the scope of protection of this application.

[0108] In practical implementation, when the computer program 311 is executed by the processor, it can achieve the following: Figure 1 Any of the corresponding implementation methods in the embodiments.

[0109] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0110] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0111] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0112] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0113] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0114] This application also provides a computer program product, which includes computer software instructions. When the computer software instructions are executed on a processing device, the processing device performs the data privacy and security identification process in the corresponding embodiment.

[0115] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0117] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0118] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0119] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0120] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0121] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for identifying data privacy and security, characterized in that, include: Obtain raw data information and text prompt information, wherein the text prompt information is generated by the user after editing or modifying a text prompt template based on best practices according to the user's custom scenario requirements; The raw data information is preprocessed according to the modal type of the data to obtain modal data information; Based on the modal data information and the text prompt information, cross-modal interaction is performed to identify open objects in the prompts in the form of text prompts, so as to obtain a first privacy and security type identification result, wherein the first privacy and security type identification result is generated based on each text prompt template of the text prompt information; The privacy security identification result is obtained by simplifying and post-processing all the first privacy security type identification results; The best practice text prompt templates include a concise privacy and security type template and a privacy and security type referential expression understanding template. Users can edit and modify the best practice text prompt templates according to their own usage scenario needs to form the text prompt information. The concise privacy and security type template is formed by directly splicing concise types with specific separators. The privacy and security type referential expression understanding template is a template with a specific scenario-specific potential identification object and only identifies objects or behaviors that specifically refer to the described state. The privacy and security type referential expression understanding template is used to help the model understand the referential expression in the text prompt to determine the specific object or behavior referred to in the text prompt.

2. The data privacy and security identification method according to claim 1, characterized in that, The raw data information includes at least one of image data information, video data information, audio data information, and text data information, and the preprocessing operation includes anomaly removal operation and input standardization operation.

3. The data privacy and security identification method according to claim 1, characterized in that, Also includes: The second privacy and security type identification result is obtained based on the modal data information and the predefined supervision model; The privacy security identification result is obtained by simplifying all the first privacy security type identification results and combining them with the second privacy security type identification results.

4. The data privacy and security identification method according to claim 1, characterized in that, Based on the modal data information and the text prompt information, cross-modal interaction is performed to identify open objects in the prompts via text prompts, in order to obtain a first privacy and security type identification result, including: The modal data information is subjected to feature extraction operation through a variant of the Transformer architecture to obtain modal features; The text prompt information is segmented and vectorized for feature extraction to obtain text features; Open object recognition is performed through cross-modal fusion using a variant of the Transformer architecture based on the modal features and the text features to obtain a first privacy-preserving type recognition result.

5. The data privacy and security identification method according to claim 4, characterized in that, The open object recognition based on the modal features and the text features through a variant of the Transformer architecture, performing cross-modal fusion to obtain a first privacy-preserving type recognition result, includes: The modal features and text features are subjected to feature interaction to obtain enhanced text features and enhanced modal features; Perform language-guided query operations based on the enhanced text features and the enhanced modality features to obtain cross-modal query features; Perform cross-modal decoding operations based on the enhanced text features, enhanced modal features, and cross-modal query features to obtain the decoding model output; The first privacy and security type identification result is obtained based on the output of the decoding model.

6. The data privacy and security identification method according to claim 5, characterized in that, Also includes: The contrastive loss is determined based on the enhanced text features, the decoding model output, and the contrastive learning function. The mode-specific loss is determined based on the output of the decoding model and the mode-specific loss function; The open set object recognition model is trained based on the contrastive loss and the modality-specific loss, and the model parameters are adjusted to obtain the model's detection capability.

7. A data privacy and security identification device, characterized in that, include: The first acquisition unit is used to acquire raw data information and text prompt information, wherein the text prompt information is generated by the user after editing or modifying a text prompt template based on the best practice according to the user's custom scenario requirements; The second acquisition unit is used to perform corresponding preprocessing operations on the original data information based on the modal type of the data to obtain modal data information. The third acquisition unit is used to perform cross-modal interaction based on the modal data information and the text prompt information to identify open objects in the prompts in the form of text prompts, so as to obtain a first privacy and security type identification result, wherein the first privacy and security type identification result is generated based on each text prompt template of the text prompt information; The fourth acquisition unit is used to obtain the privacy and security identification result after simplifying and processing all the first privacy and security type identification results; The best practice text prompt templates include a concise privacy and security type template and a privacy and security type referential expression understanding template. Users can edit and modify the best practice text prompt templates according to their own usage scenario needs to form the text prompt information. The concise privacy and security type template is formed by directly splicing concise types with specific separators. The privacy and security type referential expression understanding template is a template with a specific scenario-specific potential identification object and only identifies objects or behaviors that specifically refer to the described state. The privacy and security type referential expression understanding template is used to help the model understand the referential expression in the text prompt to determine the specific object or behavior referred to in the text prompt.

8. An electronic device, comprising: A memory and a processor, characterized in that the processor, when executing a computer program stored in the memory, implements the steps of the privacy-secure identification method for data as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the data privacy and security identification method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Personal sensitive information knowledge graph query method and device

    CN116127013A

  • System and method for query authorization and response generation using machine learning

    US20210343295A1