Method and device for judging application value of data, electronic equipment and storage medium
By acquiring and analyzing the title and content data of the data and determining its correlation and diversity, the problem of combining list samples in the prior art interfering with the identification of data breach belonging area, achieving more efficient and accurate judgment of data breach belonging area.
Patent Information
- Application Number
- CN202510172432.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art lacks the ability to filter combined list samples in the identification of data breach attribution areas, resulting in low recognition efficiency and poor accuracy.
By acquiring the title data and content data of the data to be applied, information extraction is performed to obtain the first entity set and the second entity set respectively, the degree of correlation and diversity are determined based on these sets, and the application value of the data is judged.
Accurately screening out data that does not have application value improves the efficiency and accuracy of data leakage belonging to areas, reduces noise interference, and improves the identification capabilities of hot spots in data leakage.
Smart Images

Figure CN120257987A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and particularly to a method, apparatus, electronic device, and storage medium for determining the application value of data. Background Art
[0002] Data is the lifeblood of national development and enterprise survival. With the development of new technologies such as cloud computing, big data, and blockchain, data has become the most valuable production factor in the process of digital economy development. Data security is the cornerstone of the digital economy and also the bottom-line guarantee for the development of the digital economy. However, with the rapid development of information technology, the problem of data leakage has also increased. The problem of data leakage has great security risks: user privacy data is used by hacker groups for criminal acts such as telecom fraud, and enterprise sensitive information causes serious losses to enterprises. Therefore, the supervision of data leakage incidents is particularly important, especially the identification of the attribution region of data leakage, which can determine which countries and regions have had data leakage incidents.
[0003] The current related work on attribution region identification all focuses on finding a better named entity recognition method to achieve higher extraction accuracy in the target data set, but lacks the ability to filter out low-application-value data in the above target data set, resulting in low efficiency in the application process of identifying regions. Summary of the Invention
[0004] In view of this, the purpose of the present application is to propose a method, apparatus, electronic device, and storage medium for determining the application value of data to overcome all or part of the deficiencies in the prior art.
[0005] Based on the above purpose, the present application provides a method for determining the application value of data, including: obtaining data to be applied, where the data to be applied includes title data and content data; performing information extraction on the title data to obtain a first entity set, and performing information extraction on the content data to obtain a second entity set; determining the degree of association between the title data and the content data based on the first entity set and the second entity set; in response to determining that the degree of association is less than a predetermined degree of association, determining the degree of diversity of the data to be applied based on the first entity set and the second entity set; in response to determining that the degree of diversity is greater than a predetermined degree, determining that the data to be applied does not have the application value.
[0006] Optionally, the first entity set includes a first regional entity set, and the second entity set includes a second regional entity set; determining the degree of association between the title data and the content data based on the first entity set and the second entity set includes: performing the following processing operations for each second regional entity in the second regional entity set: using a predetermined inclusion relationship, determining whether there is a first regional entity in the first regional entity set that has an inclusion relationship with the second regional entity; in response to determining that there is a first regional entity that has an inclusion relationship with the second regional entity, using the first regional entity to replace the second regional entity in the second regional entity set; based on the first regional entity set and the second regional entity set after the processing operation, determining the degree of association between the title data and the content data.
[0007] Optionally, the first entity set includes a first regional entity set and a first identity entity set, the second entity set includes a second regional entity set and a second identity entity set, and the degree of diversity includes the degree of regional diversity and the degree of identity diversity; determining the degree of diversity of the data to be applied based on the first entity set and the second entity set includes: based on the first regional entity set and the second regional entity set, determining the degree of regional diversity through the following formula: where H1 is the degree of regional diversity, p i is the i-th regional entity in the first regional entity set and the second regional entity set, and s1 is the total number of regional entities in the first regional entity set and the second regional entity set; based on the first identity entity set and the second identity entity set, determining the degree of identity diversity through the following formula: where H2 is the degree of identity diversity, p j is the j-th identity entity in the first identity entity set and the second identity entity set, and s2 is the total number of identity entities in the first identity entity set and the second identity entity set.
[0008] Optionally, the predetermined degree includes a first predetermined degree and a second predetermined degree; responding to determining that the degree of diversity is greater than the predetermined degree and determining that the data to be applied does not have the application value includes: responding to determining that the degree of regional diversity is greater than the first predetermined degree and the degree of identity diversity is greater than the second predetermined degree, determining that the data to be applied does not have the application value.
[0009] Optionally, before performing information extraction on the title data to obtain a first entity set and performing information extraction on the content data to obtain a second entity set, the method includes: performing a data cleaning operation on the data to be applied.
[0010] Optionally, performing information extraction on the title data and performing information extraction on the content data includes: using a pre-trained information extraction model to perform information extraction on the title data and the content data respectively.
[0011] Optionally, the loss function used to train the information extraction model is determined by the following formula: where Loss is the loss function, y is the correct entity label, x is the input data to be applied, s is the structure pattern director, p represents the probability that the structure pattern director s outputs the correct entity label y given the data to be applied x, and θ UIE represents all the neural network parameters to be adjusted of the information extraction model, and D train is the training data of the information extraction model.
[0012] Based on the same inventive concept, the present application also provides a device for judging the application value of data, including: an acquisition module configured to acquire data to be applied, where the data to be applied includes title data and content data; an information extraction module configured to perform information extraction on the title data to obtain a first entity set and perform information extraction on the content data to obtain a second entity set; a first determination module configured to determine the degree of association between the title data and the content data based on the first entity set and the second entity set; a second determination module configured to, in response to determining that the degree of association is less than a predetermined degree of association, determine the diversity degree of the data to be applied based on the first entity set and the second entity set; and a judgment module configured to, in response to determining that the diversity degree is greater than a predetermined degree, judge that the data to be applied does not have the application value.
[0013] Based on the same inventive concept, the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor, where the processor implements the method as described above when executing the computer program.
[0014] Based on the same inventive concept, the present application also provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method as described above.
[0015] As can be seen from the above, the method, apparatus, electronic device, and storage medium for determining the application value of data provided by the present application, the method includes obtaining data to be applied, where the data to be applied includes title data and content data. Extracting information from the title data to obtain a first entity set, and extracting information from the content data to obtain a second entity set, achieving the purpose of reducing the data volume of the title data and the content data. Based on the first entity set and the second entity set, determining the degree of association between the title data and the content data, being able to accurately determine the degree of association between the title data and the content data of the data to be applied, achieving the purpose of preliminarily determining the application value of the data to be applied. In response to determining that the degree of association is less than a predetermined degree of association, based on the first entity set and the second entity set, determining the degree of diversity of the data to be applied, being able to accurately determine the degree of diversity of the title data and the content data of the data to be applied, achieving the purpose of further determining the application value of the data to be applied. In response to determining that the degree of diversity is greater than a predetermined degree, determining that the data to be applied does not have the application value, making a judgment on the application value of the data to be applied, and being able to accurately judge the application value of the data to be applied. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 Flow chart of the method for determining the application value of data in the embodiment of the present application;
[0018] Figure 2 Administrative division entity distribution map of data leakage in the embodiment of the present application;
[0019] Figure 3 Statistical chart of data leakage hot spots when the Combolist sample (combination list sample) is not filtered in the embodiment of the present application;
[0020] Figure 4 Statistical chart of data leakage hot spots after filtering the Combolist sample in the embodiment of the present application;
[0021] Figure 5 Structural diagram of the apparatus for determining the application value of data in the embodiment of the present application;
[0022] Figure 6 Hardware structure diagram of the electronic device in the embodiment of the present application. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0024] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be the general meanings understood by those with ordinary skills in the field to which the present application belongs. The "first", "second" and similar terms used in the embodiments of the present application do not represent any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0025] As described in the background art section, data is the lifeblood of national development and enterprise survival. With the development of new technologies such as cloud computing, big data, and blockchain, data has become the most valuable production factor in the process of digital economy development. Data security is the cornerstone of the digital economy and also the bottom-line guarantee for the development of the digital economy. However, with the rapid development of information technology, the problem of data leakage has also increased. Data leakage refers to the situation where sensitive data is leaked to unauthorized external personnel or systems without the authorization of the subject, resulting in the leakage of privacy, trade secrets or other sensitive information, which mostly occurs in the darker web environment with a higher degree of anonymity. The problem of data leakage has great security risks: user privacy data is used by hacker groups for criminal acts such as telecommunications fraud, and enterprise sensitive information causes serious losses to enterprises. Therefore, the supervision of data leakage incidents is particularly important, especially the identification of the place of origin of data leakage, which can determine which countries and regions have had data leakage incidents.
[0026] The content of the data leakage description posted by hackers on the forum will involve the region to which the data leakage event belongs, sample information, advertising information, etc. Accurately identifying the region to which a data leakage event belongs will significantly improve network supervision capabilities and maintain network security. However, identifying the region to which a data leakage event belongs faces the following challenges: Data leakage posts in some Combolist samples seriously affect the identification of the true region to which the data belongs and the judgment of the overall data leakage hotspots. A Combolist sample refers to a sample in which the region information mentioned in the data leakage post content is inconsistent with the region information described in the title, and can be summarized into two types: First, Combolist samples often contain junk advertising information, and there are many useless addresses in the advertising information, so not all addresses belong to the attributed addresses. For example, a sample of data leakage for sale in country A in the title mentions multiple countries such as country B and country C in its content, causing supervisors to misjudge the country to which this data leakage event occurred as country A, country B, country C, etc. Second, the post content contains personal identity information, and the personal identity information contains addresses in multiple different regions. For example, in a sample event titled "Data of homeowners in city D", the content mentions the native places such as city E and city F. The actual attribution of the above data leakage sample is city D, but due to the example content containing a large amount of other administrative division information, it may be misjudged as belonging to city E, city F, etc. According to statistics, less than 5% of the Combolist samples in the total sample volume contain administrative division entities exceeding 30% of the total number of entities, thus seriously affecting the identification of the region to which the data leakage belongs and the analysis of hotspots. The other region information in such samples will interfere with the determination of the actual attribution. The emergence of Combolist samples seriously affects the accuracy of judging the region to which the data leakage belongs, so an effective filtering method is needed.
[0027] Current related work on identifying the region to which it belongs focuses on finding a better named entity recognition method to achieve higher extraction accuracy under the target data set, but lacks the ability to filter the noise of the above-defined Combolist samples. On the one hand, it results in the region entities identified not necessarily being the regions to which they belong. On the other hand, it leads to low efficiency in the application process of the identified regions.
[0028] In view of this, the embodiments of this application propose a method for judging the application value of data, referring to Figure 1 , including the following steps:
[0029] Step 101, obtain the data to be applied, where the data to be applied includes title data and content data.
[0030] In this step, with the rapid development of information technology, the problem of data leakage has also increased. In order to determine which region's data has been leaked, supervisors need to judge the region to which the data belongs, and then take corresponding measures to recover the losses. The application scenario corresponding to the data to be applied can be a scenario for judging the region to which the leaked data belongs. First, the data to be applied needs to be obtained. The data to be applied includes title data and content data. Exemplarily, the title data of the data to be applied can be "Resort data of Country A", and the content data includes resort data of multiple countries. It should be noted that the data to be applied is obtained by a method that complies with relevant laws and regulations, and supervisors can use the method provided in this application to judge the application value of the data for data supervision.
[0031] Step 102: Extract information from the title data to obtain a first entity set, and extract information from the content data to obtain a second entity set.
[0032] In this step, the title data and the content data include relatively much text data. If the title data and the content data are directly processed for subsequent judgment, the efficiency of the judgment process will be low. Therefore, it is necessary to extract the predetermined entities in the subsequent processing from the title data and the content data, where the predetermined entity is an entity used to determine the region to which the data to be applied belongs in advance. Extract information from the title data to obtain a first entity set, and extract information from the content data to obtain a second entity set, so as to reduce the data volume of the title data and the content data.
[0033] Step 103: Determine the degree of association between the title data and the content data based on the first entity set and the second entity set.
[0034] In this step, there may be a situation where the title data and content data of the data to be applied do not have a corresponding relationship. Exemplarily, the title data of the data to be applied is "Resort data of Country A", while the content data includes data in other aspects of multiple countries. The authenticity of the data to be applied that meets the above situation is relatively low and can be regarded as false data. Even if the region of attribution of the data to be applied is judged, it has no judgment significance and will only affect the efficiency of the application process and the accuracy of the application result. Therefore, the application value of the data to be applied is relatively low, and it is necessary to screen out the data to be applied with relatively low application value. Determine the degree of association between the title data and content data of the data to be applied. The first entity set can reflect the key information used by the title data in subsequent processing, and the second entity set can reflect the key information used by the content data in subsequent processing. Therefore, through the first entity set and the second entity set, the degree of association between the title data and content data of the data to be applied can be accurately determined, achieving the purpose of preliminarily judging the application value of the data to be applied.
[0035] Step 104, in response to determining that the degree of association is less than a predetermined degree of association, based on the first entity set and the second entity set, determine the degree of diversity of the data to be applied.
[0036] In this step, in the case where it has been determined that the degree of association is less than the predetermined degree of association, it indicates that the possibility of the title data and content data of the data to be applied having a corresponding relationship is relatively small. For example, the title data of the data to be applied is "Resort data of Country A", while the content data introduces data in other aspects of multiple countries. The title data and content data of the data to be applied do not have a corresponding relationship, and the application value of the data to be applied may be relatively low. It should be noted that the predetermined degree of association is determined through historical experience. To further accurately determine the application value of the data to be applied, it is also necessary to determine the degree of diversity of the data to be applied. The degree of diversity can accurately reflect the possibility of the title data and content data of the data to be applied having a corresponding relationship. The first entity set can reflect the key information used by the title data in subsequent processing, and the second entity set can reflect the key information used by the content data in subsequent processing. Therefore, by using the first entity set and the second entity set, the degree of diversity of the title data and content data of the data to be applied can be accurately determined, achieving the purpose of further determining the application value of the data to be applied.
[0037] It should also be supplemented that, in response to determining that the degree of association is greater than or equal to a predetermined degree of association, it is determined that the data to be applied has application value. When the degree of association is greater than or equal to the predetermined degree of association, it indicates that there is a corresponding relationship between the title data and the content data of the data to be applied, and both the title data and the content data reflect the situation of the same region. Judging the region of attribution of the data to be applied has judgment significance, and the application value of the data to be applied is relatively high.
[0038] Step 105, in response to determining that the degree of diversity is greater than a predetermined degree, it is determined that the data to be applied does not have the application value.
[0039] In this step, when the degree of diversity is greater than the predetermined degree, the data in the data to be applied has diversity, indicating that there is no corresponding relationship between the title data and the content data of the data to be applied, and the data to be applied does not have application value. It should be noted that the predetermined degree is determined based on historical experience. Even if the region of attribution of the data to be applied is judged, it has no judgment significance and will only affect the efficiency of the application process and the accuracy of the application result. Since the determined degree of diversity is accurate, by comparing the degree of diversity with the predetermined degree, the application value of the data to be applied is judged, and the application value of the data to be applied can be accurately judged. When it is determined that the data to be applied does not have application value, it is determined that the data to be applied does not participate in the predetermined application, thereby improving the efficiency of the application process of the predetermined application. Among them, the predetermined application can be an application for identifying the region of attribution of data. By filtering out the data without application value in the data leakage sample set, the influence of the data without application value on the identification of the region of attribution of the data leakage event is reduced, and the identification efficiency, accuracy, and analysis ability of the overall data leakage hot spot region are improved.
[0040] Through the above solution, the data to be applied is obtained, where the data to be applied includes title data and content data. Information extraction is performed on the title data to obtain a first entity set, and information extraction is performed on the content data to obtain a second entity set, so as to reduce the data volume of the title data and the content data. Based on the first entity set and the second entity set, the correlation degree between the title data and the content data is determined, and the correlation degree between the title data and the content data of the data to be applied can be accurately determined, so as to initially judge the application value of the data to be applied. In response to determining that the correlation degree is less than a predetermined correlation degree, based on the first entity set and the second entity set, the diversity degree of the data to be applied is determined, and the diversity degree of the title data and the content data of the data to be applied can be accurately determined, so as to further determine the application value of the data to be applied. In response to determining that the diversity degree is greater than a predetermined degree, it is determined that the data to be applied does not have the application value, and the application value of the data to be applied is judged, and the application value of the data to be applied can be accurately judged.
[0041] In some embodiments, the first entity set includes a first regional entity set, and the second entity set includes a second regional entity set; the determining the correlation degree between the title data and the content data based on the first entity set and the second entity set includes: performing the following processing operations on each second regional entity in the second regional entity set: using a predetermined inclusion relationship, determining whether there is a first regional entity in the first regional entity set that has an inclusion relationship with the second regional entity; in response to determining that there is a first regional entity that has an inclusion relationship with the second regional entity, using the first regional entity to replace the second regional entity in the second regional entity set; based on the first regional entity set and the second regional entity set after the processing operation, determining the correlation degree between the title data and the content data.
[0042] In this embodiment, it is not only when the regions in the title data are the same as those in the content data that it indicates a corresponding relationship between the title data and the content data. When the region in the title data contains the region in the content data, it can also indicate a corresponding relationship between the title data and the content data. Therefore, a predetermined inclusion relationship is established in advance. The first region entity set includes at least one first region entity, and the second region entity set includes at least one second region entity. For each second region entity, the following processing operations are performed: Using the predetermined inclusion relationship, determine whether there is a first region entity that has an inclusion relationship with this second region entity. If there is a first region entity that has an inclusion relationship with the second region entity, replace this second region entity in the second entity region set with the above-mentioned first region entity. In this way, the difference between the second region entity set and the first region entity set is reduced, making the degree of association between the title data determined based on the first region entity set and the second region entity set after the processing operation and the content data accurate.
[0043] It should be noted that the Jaccard similarity calculation formula can be used to calculate the regional entity consistency between the title data and the content data to determine whether the regional entities mentioned in the title data and the content data are the same. The degree of association between the title data and the content data is between 0 and 1, and the larger the value, the higher the degree of association. If the regions are the same, the degree of association is 1, indicating that the data to be applied has application value and is not a Combolist sample. If the degree of association is not 1, it indicates that there are differences in the administrative division entities between the title data and the content data, and subsequent judgment is required.
[0044] The Jaccard similarity calculation formula is specifically as follows: Among them, (A, B) represents the degree of association corresponding to the first region entity set and the second region entity set, A represents the first region entity in the first region entity set, and B represents the second region entity in the second region entity set.
[0045] It should be noted that the predetermined inclusion relationship is an inclusion relationship between regions established in advance. For example, country C has an inclusion relationship with city D, and city D has an inclusion relationship with sub-region E. The predetermined inclusion relationship can exist in the form of a knowledge graph or in the form of a table.
[0046] In some embodiments, the first entity set includes a first regional entity set and a first identity entity set, the second entity set includes a second regional entity set and a second identity entity set, and the diversity degree includes a regional diversity degree and an identity diversity degree; determining the diversity degree of the data to be applied based on the first entity set and the second entity set includes: determining the regional diversity degree based on the first regional entity set and the second regional entity set through the following formula: where H1 is the regional diversity degree, and p i is the i-th regional entity in the first regional entity set and the second regional entity set, and s1 is the total number of regional entities in the first regional entity set and the second regional entity set; determining the identity diversity degree based on the first identity entity set and the second identity entity set through the following formula: where H2 is the identity diversity degree, and p j is the j-th identity entity in the first identity entity set and the second identity entity set, and s2 is the total number of identity entities in the first identity entity set and the second identity entity set.
[0047] In this embodiment, the diversity degree includes a regional diversity degree and an identity diversity degree. Determining the diversity degree of the data to be applied from multiple perspectives makes the determination of the diversity degree accurate. Both the determination of the regional diversity degree and the identity diversity degree are based on the Shannon-Wiener index. The Shannon-Wiener index is usually used to measure the richness of information and the magnitude of entropy in a text. The higher the index, the richer the information in the text, containing more different contents or words. Therefore, the regional diversity degree reflects the degree to which the content data is richer than the title data in administrative division information; the identity diversity degree reflects the richness of information of personal information identity entities. Numeralizing the regional diversity degree and the identity diversity degree achieves the purpose of accurately calculating the regional diversity degree and the identity diversity degree.
[0048] In some embodiments, the predetermined degree includes a first predetermined degree and a second predetermined degree; responding to determining that the diversity degree is greater than the predetermined degree and determining that the data to be applied does not have the application value includes: responding to determining that the regional diversity degree is greater than the first predetermined degree and the identity diversity degree is greater than the second predetermined degree, and determining that the data to be applied does not have the application value.
[0049] In this embodiment, the diversity degree of the data to be applied is determined from multiple perspectives. Only when the regional diversity degree is greater than the first predetermined degree and the identity diversity degree is greater than the second predetermined degree, can it be accurately stated that there is no corresponding relationship between the title data and the content data of the data to be applied, and thus the purpose of accurately determining that the data to be applied has no application value can be achieved. Exemplarily, the first predetermined degree can be 0.5 and the second predetermined degree can be 1.0.
[0050] In some embodiments, before extracting information from the title data to obtain a first entity set and extracting information from the content data to obtain a second entity set, the method includes: performing a data cleaning operation on the data to be applied.
[0051] In this embodiment, the data to be applied may include data with no practical meaning, and the existence of data with no practical meaning will lead to low judgment efficiency for the data to be applied. Therefore, data cleaning operations are respectively performed on the title data and the content data in the data to be applied. The data cleaning operation includes removing special characters, HTML tags, meaningless stop words, and standardizing the text format in the data to be applied, so as to achieve the purpose of reducing the data volume of the data to be applied.
[0052] In some embodiments, extracting information from the title data and extracting information from the content data include: using a pre-trained information extraction model to extract information from the title data and the content data respectively.
[0053] In this embodiment, UIE (Universal Information Extraction) is a unified framework for general information extraction proposed in ACL 2022, which realizes a unified modeling model for tasks such as entity extraction, relationship extraction, event extraction, and sentiment analysis, and can have good transfer and generalization learning capabilities among different tasks. The information extraction model in this embodiment can be a UIE model fine-tuned with the cross-entropy loss of Teacher Forcing. Teacher Forcing is a method for training sequence generation models, which accelerates convergence by using the true labels as input during training. When fine-tuning the UIE model, the cross-entropy loss function is used to measure the difference between the model prediction and the true label. Specifically, when the model generates an output at each time step, the true label rather than the prediction result of the previous moment is used as the input, thereby reducing error accumulation. By minimizing the cross-entropy loss, the model gradually optimizes the parameters and improves the accuracy and generalization ability of information extraction.
[0054] Since the present application can be applied to the scenario of identifying the region to which the data to be applied belongs, when training the information extraction model, the information extraction model needs to be able to extract region entities. Exemplarily, the region entities include four types of region entities: "country", "province", "city", and "county (district)". Since the present application also needs to determine the degree of identity diversity, when training the information extraction model, the information extraction model needs to be able to extract identity entities. Exemplarily, the identity entities include three types of identity entities: "name", "telephone", and "email". According to the extraction results of the information extraction model, the entity types can be divided into three categories: region entities of title data (title_region set), region entities of content data (content_region set), and identity entities of content data (content_person_set).
[0055] The information extraction model is based on the T5 generation architecture of text to structure, and the information extraction model also has a built-in prompt-based structural pattern instructor. The information extraction model is based on the T5 generation architecture and adopts the text-to-structure generation method to convert input data into structured output. The model has a built-in structural pattern instructor, which guides the model to understand the task requirements through the prompt mechanism to ensure that the generated results conform to the predefined structural pattern. This design improves the accuracy and generalization ability of the model in information extraction tasks, enabling it to better adapt to a variety of application scenarios. The information extraction model is used to extract information from the title data to obtain the first entity set, and the information extraction model is used to extract information from the content data to obtain the second entity set, ensuring the efficiency and accuracy of information extraction.
[0056] In some embodiments, the loss function used to train the information extraction model is determined by the following formula: Where Loss is the loss function, y is the correct entity label, x is the input data to be applied, s is the structural mode director, p represents the probability that the structural mode director s outputs the correct entity label y given the data to be applied x, θ UIE represents all neural network parameters to be adjusted for the information extraction model, D train Training data for the information extraction model.
[0057] In this embodiment, the loss function trains the information extraction model by maximizing the log likelihood of the correct label (i.e., minimizing the negative log likelihood). The role of the structural pattern guide is to provide additional structured information to the information extraction model, helping the information extraction model to better understand the task requirements and generate expected outputs. By training the information extraction model with the formula in this embodiment, the information extraction model can extract information more accurately and adapt to different task scenarios.
[0058] In another embodiment provided by the present application, the present invention collects 7,000 actually occurred data leakage events as a data set to evaluate the effectiveness of the solution of the present invention. The details of the data set are shown in Table 1 and Figure 2 as follows.
[0059] Table 1 Data Leakage Event Data Set
[0060]
[0061] Among them, the Combolist samples account for 4.34%, including 304; the non-Combolist samples account for 95.66%, including 6,696. The Combolist samples are samples containing data to be applied that have no application value, and the non-Combolist samples are samples containing data to be applied that have application value. According to Figure 2 the display, although the Combolist samples only account for 4.34%, the regional entities they contain account for a high proportion, reaching 32.4% of the total number of entities in the data set, and exceeding 50% in the specific "province", "city", and "county" entities, proving that the Combolist samples can cause great noise interference to the identification of the data leakage attribution location. On the premise of accurately judging the combined list samples (Combolist), the method of the present invention successfully recalled 252 data leakage samples among 7,000 data leakage texts (304 are Combolist samples and 6,696 are non-Combolist samples). Among them, the recall rate of the combined list samples reached 82.89%, and the precision rate reached 96.18%.
[0062] The solution of the present invention was experimented on the basis of the foregoing data set. For the identification of Combolist, an accuracy rate of 96.18%, a recall rate of 82.89%, and an F1 score of 89.05% were achieved. The specific identification effects are shown in Table 2.
[0063] Table 2 Experimental Results of Filtering Combined List Samples
[0064]
[0065] Figure 3 and Figure 4 respectively show the statistics of the top 15 hot regions before and after filtering the Combolist samples. It can be seen from Figure 4 this that many administrative division entities such as Shenzhen, Taiwan, and Zhejiang Province that were only mentioned were wrongly included in the scope of the attribution region before filtering the Combolist samples, affecting the judgment of the popular data leakage regions, while the present invention effectively reduces the influence of the Combolist samples. Figure 3 and Figure 4By comparison, it can be seen that the distribution of hotspots has changed significantly before and after filtering the Combolist, verifying the effectiveness and importance of the present invention.
[0066] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.
[0067] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0068] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides an apparatus for determining the application value of data.
[0069] Reference Figure 5 , the apparatus for determining the application value of data includes:
[0070] An acquisition module 10, configured to acquire data to be applied, where the data to be applied includes title data and content data.
[0071] An information extraction module 20, configured to perform information extraction on the title data to obtain a first entity set, and perform information extraction on the content data to obtain a second entity set.
[0072] A first determination module 30, configured to determine the degree of association between the title data and the content data based on the first entity set and the second entity set.
[0073] A second determination module 40, configured to, in response to determining that the degree of association is less than a predetermined degree of association, determine the degree of diversity of the data to be applied based on the first entity set and the second entity set.
[0074] A judgment module 50, configured to, in response to determining that the degree of diversity is greater than a predetermined degree, judge that the data to be applied does not have the application value.
[0075] Through the above device, the data to be applied is obtained, where the data to be applied includes title data and content data. Information extraction is performed on the title data to obtain a first entity set, and information extraction is performed on the content data to obtain a second entity set, so as to reduce the data volume of the title data and the content data. Based on the first entity set and the second entity set, the correlation degree between the title data and the content data is determined, and the correlation degree between the title data and the content data of the data to be applied can be accurately determined, so as to achieve the purpose of preliminarily judging the application value of the data to be applied. In response to determining that the correlation degree is less than a predetermined correlation degree, based on the first entity set and the second entity set, the diversity degree of the data to be applied is determined, and the diversity degree of the title data and the content data of the data to be applied can be accurately determined, so as to achieve the purpose of further determining the application value of the data to be applied. In response to determining that the diversity degree is greater than a predetermined degree, it is determined that the data to be applied does not have the application value, and the application value of the data to be applied is judged, and the application value of the data to be applied can be accurately judged.
[0076] In some embodiments, the first determination module 30 is further configured such that the first entity set includes a first regional entity set, and the second entity set includes a second regional entity set; for each second regional entity in the second regional entity set, the following processing operations are performed: using a predetermined inclusion relationship, determining whether there is a first regional entity in the first regional entity set that has an inclusion relationship with the second regional entity; in response to determining that there is a first regional entity that has an inclusion relationship with the second regional entity, using the first regional entity to replace the second regional entity in the second regional entity set; based on the first regional entity set and the second regional entity set after the processing operation, determining the correlation degree between the title data and the content data.
[0077] In some embodiments, the second determination module 40 is further configured such that the first entity set includes a first regional entity set and a first identity entity set, the second entity set includes a second regional entity set and a second identity entity set, and the diversity degree includes a regional diversity degree and an identity diversity degree; based on the first regional entity set and the second regional entity set, the regional diversity degree is determined by the following formula: where H1 is the regional diversity degree, p i is the i-th regional entity in the first regional entity set and the second regional entity set, and s1 is the total number of regional entities in the first regional entity set and the second regional entity set; based on the first identity entity set and the second identity entity set, the identity diversity degree is determined by the following formula: wherein, H2 is the degree of identity diversity, and p j is the j-th identity entity in the first identity entity set and the second identity entity set, and s2 is the total number of identity entities in the first identity entity set and the second identity entity set.
[0078] In some embodiments, the determination module 50 is further configured such that the predetermined degree includes a first predetermined degree and a second predetermined degree; in response to determining that the regional diversity degree is greater than the first predetermined degree and the identity diversity degree is greater than the second predetermined degree, it is determined that the data to be applied does not have the application value.
[0079] In some embodiments, it further includes a data cleaning module, and the data cleaning module is configured to perform a data cleaning operation on the data to be applied before performing information extraction on the title data to obtain a first entity set and performing information extraction on the content data to obtain a second entity set.
[0080] In some embodiments, the information extraction module 20 is further configured to perform information extraction on the title data and the content data respectively by using a pre-trained information extraction model.
[0081] In some embodiments, the information extraction module 20 is further configured such that the loss function for training the information extraction model is determined by the following formula: where Loss is the loss function, y is the correct entity label, x is the input data to be applied, s is the structure pattern guide, p represents the probability that the structure pattern guide s outputs the correct entity label y given the data to be applied x, and θ UIE represents all neural network parameters to be adjusted of the information extraction model, and D train is the training data of the information extraction model.
[0082] For the convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the present application, the functions of each module can be implemented in one or more software and / or hardware.
[0083] The device in the above embodiments is used to implement the method for determining the application value of the corresponding judgment data in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0084] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method for determining the application value of the data as described in any of the above embodiments.
[0085] Figure 6 shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0086] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0087] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0088] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0089] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0090] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0091] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0092] The electronic device in the above embodiment is used to implement the method for judging the application value of data corresponding to any one of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0093] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions for causing the computer to execute the method for judging the application value of data as described in any of the foregoing embodiments.
[0094] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0095] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the method for judging the application value of data as described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0096] Based on the same concept, corresponding to the method of any of the above embodiments, the present application also provides a computer program product, including computer program instructions, which when running on a computer, cause the computer to execute the method for judging the application value of data as described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0097] It should be noted that the embodiments of the present application can also be further described in the following manner:
[0098] It can be understood that before using the technical solutions of the various embodiments in the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0099] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0100] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0101] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0102] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of brevity.
[0103] In addition, for simplicity of explanation and discussion, and so as not to make the embodiments of the present application difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be considered illustrative rather than restrictive.
[0104] Although the present application has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0105] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the present application. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A method for judging the application value of data, characterized in that, Including: Obtain data to be applied, where the data to be applied includes title data and content data; Perform information extraction on the title data to obtain a first entity set, and perform information extraction on the content data to obtain a second entity set; Based on the first entity set and the second entity set, determine the degree of association between the title data and the content data; In response to determining that the degree of association is less than a predetermined degree of association, based on the first entity set and the second entity set, determine the diversity degree of the data to be applied; In response to determining that the diversity degree is greater than a predetermined degree, determine that the data to be applied does not have the application value.
2. The method according to claim 1, characterized in that, The first entity set includes a first regional entity set, and the second entity set includes a second regional entity set; The determining the degree of association between the title data and the content data based on the first entity set and the second entity set includes: Perform the following processing operations for each second regional entity in the second regional entity set: use a predetermined inclusion relationship to determine whether there is a first regional entity in the first regional entity set that has an inclusion relationship with the second regional entity; In response to determining that there is a first regional entity that has an inclusion relationship with the second regional entity, use the first regional entity to replace the second regional entity in the second regional entity set; Based on the first regional entity set and the second regional entity set after the processing operation, determine the degree of association between the title data and the content data.
3. The method according to claim 1, characterized in that The first entity set includes a first regional entity set and a first identity entity set, the second entity set includes a second regional entity set and a second identity entity set, and the diversity degree includes a regional diversity degree and an identity diversity degree; The determining the diversity degree of the data to be applied based on the first entity set and the second entity set includes: Based on the first regional entity set and the second regional entity set, determine the regional diversity degree through the following formula: Among them, H1 is the degree of regional diversity, and p i is the i-th regional entity in the first regional entity set and the second regional entity set, and s1 is the total number of regional entities in the first regional entity set and the second regional entity set; Based on the first identity entity set and the second identity entity set, determine the identity diversity degree through the following formula: Among them, H2 is the degree of identity diversity, and p j is the j-th identity entity in the first identity entity set and the second identity entity set, and s2 is the total number of identity entities in the first identity entity set and the second identity entity set.
4. The method according to claim 3, characterized in that, The predetermined degree includes a first predetermined degree and a second predetermined degree; The determining that the data to be applied does not have the application value in response to determining that the diversity degree is greater than a predetermined degree includes: In response to determining that the regional diversity degree is greater than the first predetermined degree and the identity diversity degree is greater than the second predetermined degree, determine that the data to be applied does not have the application value.
5. The method according to claim 1, characterized in that, Before performing information extraction on the title data to obtain a first entity set and performing information extraction on the content data to obtain a second entity set, the method includes: Perform a data cleaning operation on the data to be applied.
6. The method according to claim 1, wherein The performing information extraction on the title data and performing information extraction on the content data includes: Use a pre-trained information extraction model to perform information extraction on the title data and the content data respectively.
7. The method according to claim 6, characterized in that, The loss function for training the information extraction model is determined by the following formula: Among them, Loss is the loss function, y is the correct entity label, x is the input data to be applied, s is the structure pattern guidance, p represents the probability that the structure pattern guidance s outputs the correct entity label y given the input data x to be applied, and θ UIE represents all the neural network parameters to be adjusted of the information extraction model, and D train is the training data of the information extraction model.
8. An apparatus for determining the application value of data, characterized in that, Including: An acquisition module, configured to acquire data to be applied, where the data to be applied includes title data and content data; An information extraction module, configured to perform information extraction on the title data to obtain a first entity set, and perform information extraction on the content data to obtain a second entity set; A first determination module, configured to determine the degree of association between the title data and the content data based on the first entity set and the second entity set; A second determination module, configured to, in response to determining that the degree of association is less than a predetermined degree of association, determine the degree of diversity of the data to be applied based on the first entity set and the second entity set; A judgment module, configured to, in response to determining that the degree of diversity is greater than a predetermined degree, judge that the data to be applied does not have the application value.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method described in any one of claims 1 to 7.