Geographic information risk sensitivity assessment method and device for internet social media

By extracting geographical subjects and related information from Internet social media, a multi-dimensional quantitative evaluation model is constructed, which solves the problem of identification and quantification of geographical information security risks in Internet social media, and effectively evaluates and deals with geographical information risks.

CN120336529AActive Publication Date: 2025-07-18CENT SOUTH UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510828494.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In Internet social media, it is difficult to effectively identify and quantify the security risks of geographic information, especially sensitive geospatial information leakage accidents that may be caused by multimodal data, which poses huge security risks.

Method used

Through natural language processing technology, information such as geographical subjects, sensitive application types, sensitive region coverage types and timeliness are extracted from text information on Internet social media, and information on geographic subjects, sensitive application types, and timeliness, and a multi-dimensional quantitative evaluation model is constructed to calculate the risk sensitivity of geographical information, providing a basis for further risk treatment.

Benefits of technology

It has realized the effective identification and quantitative assessment of geographic information risks in Internet social media, and can discover hidden sensitive geographic information and provide scientific quantitative values, providing an important basis for the treatment of geographic information security risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336529A_ABST
    Figure CN120336529A_ABST
Patent Text Reader

Abstract

The invention discloses a geographic information risk sensitivity assessment method and device for Internet social media, and the method comprises the steps: collecting the original published content of the Internet social media including geographic information text description and user comment data, and carrying out the preprocessing; extracting information in the dialogue corpus by utilizing a natural language processing technology, wherein the information comprises a geographic subject, geographic position information, a sensitive region coverage type, a sensitive application type and timeliness; merging geographical location information and sensitive region coverage types belonging to the same geographical subject in the extracted information; for each geographic subject, secret-related quantitative evaluation is carried out from a sensitive-related application type, a sensitive-related region coverage type, a geographic position and timeliness dimension; and calculating a risk sensitivity quantification result of each geographic subject in the dialogue corpus according to the confidential quantitative evaluation value of each dimension and the risk influence weight thereof. According to the method, the geographic information security risk involved in the Internet social media information can be identified and judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of geographic information security, and particularly relates to a method and device for evaluating the risk sensitivity of geographic information of Internet social media. Background Art

[0002] In recent years, with the rapid development of the Internet, especially mobile Internet, Internet media social networking has become an important way for people to communicate and interact. In a situation where everyone is a video recorder and a positioning device, the security of geospatial information is ubiquitous. This is mainly manifested in the ubiquity of the collection and dissemination of geospatial information and the multimodality of geospatial information data entities. In addition to traditional basic surveying and mapping data files, with the explosive emergence of various forms of social media (such as WeChat, Weibo, Douyin, etc.), in the process of information release and interaction of posting and commenting casually, there are deeper risks for the security of geographic information. These ubiquitously spread data, from the perspective of a single individual, the sensitivity of the geographic information contained is not strong, but after the geographic information in a large amount of multimodal data is associated through a certain connection, it may lead to serious accidents of sensitive geospatial information leakage, causing huge potential security risks. Therefore, it is urgent to identify and protect the geographic information risks of social media information on the Internet. Summary of the Invention

[0003] The present invention provides a method and device for evaluating the risk sensitivity of geographic information of Internet social media. By analyzing the text information of Internet social media, relevant geographic entities and address accuracy, application type, classified coverage area type, and timeliness of the information that affect the sensitivity risk of the geographic entity are extracted, so as to calculate the dimensions that affect the sensitivity risk of the geographic entity, and finally obtain the risk sensitivity evaluation result of the geographic entity information, providing an important basis for further geographic information security risk disposal work.

[0004] To achieve the above technical objectives, the present invention adopts the following technical solutions: A method for evaluating the risk sensitivity of geographic information of Internet social media, comprising: Step 1, collect the original published content including the text description of geographic information and its user comment data of Internet social media, record each original published content and all its comments as a piece of dialogue corpus, and preprocess each piece of collected dialogue corpus; Step 2, use natural language processing technology to extract information from the preprocessed dialogue corpus, including: geographic entities, geographic location information of each geographic entity, classified geographic coverage type of each geographic location information, classified application type of each geographic entity, and timeliness of each geographic entity; Step 3, merge the geographic location information and classified geographic coverage type that belong to the same geographic entity in the extracted information; Step 4: For each geographical entity, calculate the classified quantification evaluation value respectively from four dimensions: sensitive application type, sensitive geographical coverage type, geographical location, and timeliness. Step 5: For each geographical entity, calculate the final risk sensitivity quantification result according to the classified quantification evaluation value of each dimension and its risk impact weight.

[0005] Furthermore, the preprocessing of the data collected in Step 1 includes: completing the abbreviations of the geographical location information therein, and correcting the grammar errors and typos therein.

[0006] Furthermore, Step 2 specifically includes: S2-1: For the original published content in the conversation corpus, use the natural language processing engine to identify the geographical entities and their corresponding set of all address location information ; among them, country, province, city name, street name, house number, GPS location and azimuth information all belong to geographical location information; S2-2: For the identified geographical entities , analyze the sensitive application type of the original published content where the geographical entity is located, including three preset sensitive application types, which are respectively recorded as , , ; S2-3: For each geographical location information of the geographical entity , identify the sensitive geographical coverage type where it is located, including important type , non-disclosure type , non-disclosable shape type ; S2-4: For each comment in the conversation corpus, perform Steps S2-1, S2-2, and S2-3 to obtain the geographical entities, the set of geographical location information P of the geographical entity and the sensitive application type , as well as the sensitive geographical coverage type of each geographical location information ; S2-5: Identify the release time of the original published content and each comment, and obtain the timeliness of the corresponding geographical entity information according to the sorting of the release dates , ; where represents the distance from the statistical time.

[0007] Furthermore, Step 3 specifically includes: S3-1: Integrate all the geographical location information in the set of geographical location information of the geographical entity to obtain the geographical entity ​ The only geographical location information ; S3-2. For the geographical entity , merge the sensitive area coverage types of all its geographical location information, that is: ; In the formula: represents the statistical value of the sensitive area coverage type associated with the geographical entity s, which are the statistical values of the important type , the statistical value of the non-disclosed type , and the statistical value of the non-disclosable shape type ; represents the sensitive area coverage type of the i-th geographical location information in the original geographical location information set P of the geographical entity , where the three elements correspond to the important type, the non-disclosed type, and the non-disclosable shape type in sequence; n is the number of geographical location information included in the original geographical location information set P.

[0008] Furthermore, in step 4, for each geographical entity, calculating the classified quantification evaluation value from the sensitive application type is specifically as follows: First, count the sensitive application type risk matrix R: ; Among them: represents the number of events of the three sensitive levels of the preset first sensitive application type, with the sensitive levels from high to low; represents the number of events of the three sensitive levels of the preset second sensitive application type, with the sensitive levels from high to low; represents the number of events of the three sensitive levels of the preset third sensitive application type, with the sensitive levels from high to low; among them, the statistical method for the number of events of each sensitive level of each sensitive application type is: regarding the original published content of the conversation corpus and each comment as a piece of speech, in all speeches related to the geographical entity, if a piece of speech includes the currently counted sensitive application type and belongs to the currently counted sensitive level, the number of events of the currently counted sensitive level of the currently counted sensitive application type is incremented by 1.

[0009] Then, construct the sensitive level setting matrix of the sensitive application type and the sensitive application type weight matrix : , ; Among them: , , respectively represent the sensitivity values set for different sensitive levels, sorted from high to low; , , respectively represent the weight values of the preset first, second, and third sensitivity-related application types; Finally, based on the three matrices R constructed above, , , calculate the quantitative evaluation value : .

[0010] Furthermore, in step 4, for each geographical entity, calculating the classified quantification evaluation value from the sensitivity-related geographical coverage types is specifically as follows: First, count the risk matrix of the sensitivity-related geographical coverage types : ; wherein: represents the statistical values of the sensitivity-related geographical coverage types associated with the geographical entity s, which are the statistical values of the important type , the statistical values of the non-disclosure type , and the statistical values of the non-disclosable shape type ; Then, construct the sensitivity level setting matrix of the sensitivity-related geographical coverage types : ; wherein: represents the sensitivity value of the important type , represents the sensitivity value of the non-disclosure type , the sensitivity value of the non-disclosable shape type ; Finally, based on the above two constructed matrices and , calculate the sensitivity quantification evaluation value of the geographical entity s in the dimension of geographical coverage type: .

[0011] Furthermore, step 5 specifically includes: S5-1. Construct a one-dimensional horizontal matrix V of the multi-dimensional sensitivity values of the geographical entity , including the sensitivity quantification evaluation values in 4 dimensions of classified application types, sensitivity-related geographical coverage types, geographical locations, and timeliness: ; wherein: respectively represent the sensitivity quantification evaluation values in the dimensions of sensitivity-related application types, sensitivity-related geographical coverage types, geographical locations, and timeliness; S5-2. For the geographical entity For each sensitive dimension, set its related risk impact weight value one-dimensional vertical matrix K, that is: ; Where: respectively represent the risk impact weight value of the sensitive application type, the risk impact weight value of the sensitive geographical coverage type, the risk impact weight value of the geographical location, and the risk impact weight value of the timeliness; S5-3. According to the two matrices of S5-1 and S5-2, calculate the final risk-sensitive quantization result of the geographical information entity s : .

[0012] An electronic device includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor implements the above-mentioned method.

[0013] According to the analysis of the text information of Internet social media, the present invention extracts relevant geographical entities and address accuracy, application type, classified coverage area type, timeliness of the geographical entity that affect the sensitive risk of the geographical entity, etc. Through the calculation of the dimensions affecting the sensitive risk of the geographical entity extracted above, the risk sensitivity assessment result of the geographical entity information is finally obtained. The present invention can not only discover sensitive geographical information hidden in the content of Internet social topics, but also evaluate the risk sensitivity of the geographical information in a more reasonable and effective way, give a more scientific and effective quantization value, and provide a basis for further geographical information risk disposal.

[0014] Under the background of the high-speed spread of Internet information including mobile Internet, the present invention can identify and determine the geographical information security risks involved in the process of everyone participating in the production of network information content. It has important practical significance in preventing the leakage of sensitive geographical information. The present invention provides an effective quantization assessment method for the risk sensitivity of geographical information in Internet media information and comments. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flowchart of the method described in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENT

[0016] In order to better understand the technical solution of the present invention, the specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings, so that those skilled in the art can understand the present invention. It should be clear that the described implementations and all other implementations obtained by ordinary technicians in the field without creative work are within the scope of protection of the present invention.

[0017] This embodiment provides a method for assessing the risk sensitivity of geographic information in Internet social media. The following steps are explained in sequence. Figure 1 shown.

[0018] Step 1: Collect the original published content and comment data of the Internet social media including the text description of geographic information. Each original published content and its comment data together constitute a dialogue corpus, and pre-process the collected data.

[0019] S1-1. Targeted completion based on the abbreviations of geographic information. Some published content and user comments contain colloquial abbreviations for certain places or venues, which are not conducive to the analysis of natural language algorithms and may be treated as noise. Therefore, the geographic location information database is used to complete some regional abbreviations in order to more accurately identify and understand the geographic location information.

[0020] S1-2. Correct grammatical errors. Users often make grammatical errors or typos when entering published content and comments, which results in the sentence itself not conforming to normal semantics and cannot be recognized by the natural language system, or is recognized incorrectly. Therefore, it is necessary to automatically improve the sentence through a grammatical correction program in order to more accurately extract and understand the geographic information in the published content and comments.

[0021] Step 2: Use natural language processing technology to extract information from the preprocessed conversation topic corpus, including: geographic entities, geographic location information of each geographic entity, sensitive regional coverage type of each geographic location information, sensitive application type of each geographic entity, and timeliness of each geographic entity.

[0022] S2-1. Use the natural language processing engine to analyze the original published content in the dialogue corpus and identify the geographic subjects in the original published content. And its corresponding set of all address location information The geographic location information may include country, province, city name, street name, house number, GPS location and direction information, etc.

[0023] S2-2. Geographical Subjects , analyze the sensitive application types of the original published content, including the three preset sensitive application types, which are recorded as , , , taking values of 0 or 1. If the corresponding sensitive application type exists in the original published content, the corresponding sensitive type of the geographical entity takes the value of 1; otherwise, it is 0. For example, the sensitive application type of a certain geographical entity is , indicating that the sensitive application types of the geographical entity include the preset first and second sensitive application types.

[0024] S2-3. For each geographical location information , identify the sensitive geographical coverage type where the geographical location is located, including the important type , the non-disclosure type , and the non-disclosable shape type ; , , take values of 0 or 1. If the geographical location has the corresponding sensitive geographical coverage type, the value is 1; otherwise, it is 0. For example, the sensitive geographical type of a certain geographical location information indicates that the sensitive geographical coverage type of the geographical location information is the important type .

[0025] S2-4. For each comment in the dialogue corpus, perform steps S2-1, S2-2, and S2-3 to obtain the geographical entity, the set P of geographical location information of the geographical entity, and the sensitive application type in each comment, as well as the sensitive geographical coverage type of each geographical location information.

[0026] For comments where no geographical entity is identified, by default, classify the identified content in the comment into the geographical entity of the original published content.

[0027] S2-5. Identify the publication time of the original published content and each comment, and obtain the timeliness of the corresponding geographical entity information according to the sorting of the publication dates , . Among them, represents within one month, represents within 6 months, represents within 12 months, represents within 24 months. Information over two years old has relatively low value and such information will not be extracted and will be directly ignored. By default, determine its timeliness based on the most recently commented date. For example, even if the content was published two years ago but was commented within one month, it indicates that the published content is still being concerned about and has timeliness, and the timeliness is .

[0028] Through the above steps, the following will be obtained from the original publication content of a conversation corpus and its comments: a set of geographical entities S = ; for each geographical entity the corresponding geographical location information set , that is ; for each geographical location information the corresponding sensitive area coverage type, that is ; for each geographical entity the corresponding sensitive application type, that is ; for each geographical entity the corresponding timeliness value, that is , .

[0029] Step 3: For each segment of the conversation corpus, merge the geographical location information and the sensitive area coverage type that belong to the same geographical entity in the extracted information.

[0030] S3-1. Merge the geographical location information sets of geographical entities.

[0031] The relationship between the geographical entity and its geographical location information obtained in Step 2 is . For the geographical location information set P of the geographical entity , use the address library matching analysis technology to integrate all geographical location information to obtain the more accurate unique geographical location information of the geographical entity . If these address location information sets can be directly located to GPS coordinate points, then use the coordinate point information as the unique geographical location information of the geographical entity. Finally, obtain the unique corresponding relationship between the geographical entity and the geographical location information.

[0032] S3-2. Merge the sensitive area coverage types of each geographical location of the geographical entity identified in S2-3, that is: ; In the formula: represents the sensitive area coverage type of the i-th geographical location information in the original geographical location information set P of the geographical entity , where the three elements correspond to the important type, non-disclosure type, and non-disclosable shape type in sequence; n is the number of geographical location information included in the original geographical location information set P; represents the statistical value of the sensitive area coverage type associated with the geographical entity s, which are the statistical values of the important type , the statistical value of the non-disclosure type , and the statistical value of the non-disclosable shape type respectively.

[0033] After step 3, each geographical entity corresponds to 1 geographical location information, 1 sensitive area coverage type, 1 sensitive application type, and 1 timeliness metric value, expressed as .

[0034] Step 4: For each geographical entity, calculate the classified quantification evaluation value from the four dimensions of classified application type, sensitive area coverage type, geographical location, and timeliness.

[0035] S4-1. Construct the risk matrix R of the sensitive application type of the geographical entity and calculate the sensitive quantification evaluation value of the dimension of the geographical information application type: First, count the risk matrix R of the sensitive application type: ; where: represents the number of events at the three sensitive levels of the first preset sensitive application type, with the sensitive level from high to low; represents the number of events at the three sensitive levels of the second preset sensitive application type, with the sensitive level from high to low; represents the number of events at the three sensitive levels of the second preset sensitive application type, with the sensitive level from high to low. Among them, the statistical method for the number of events at each sensitive level of each sensitive application type is: regarding the original published content of the conversation corpus and each comment as a piece of speech, among all the speeches related to the geographical entity, if a certain speech includes the currently counted sensitive application type and belongs to the currently counted sensitive level, then the number of events at the currently counted sensitive level of the currently counted sensitive application type is incremented by 1.

[0036] Then, construct the sensitive level setting matrix of the sensitive application type and the weight matrix of the sensitive application type: , ; where: , , respectively represent the sensitivity values of different sensitive level settings, sorted from high to low; , , respectively represent the weight values of the first, second, and third preset sensitive application types. Among them, the sensitivity values of different sensitive level settings can be specifically set according to standards, such as the national standard GB / T 43697-2024 "Data Security Technology - Data Classification and Grading Rules".

[0037] Finally, according to the three matrices R, , , calculate the quantitative evaluation value of the geographical entity s in the dimension of sensitive application types : .

[0038] S4-2. Construct the risk matrix of the sensitive geographical coverage type of the geographical entity , and calculate the sensitive quantitative evaluation value in the dimension of sensitive geographical coverage type : : First, count the risk matrix of the sensitive geographical coverage type : ; Among them: represents the statistical value of the sensitive geographical coverage type associated with the geographical entity s, which are the statistical values of the important type , the statistical value of the non-public type , the statistical value of the non-disclosable shape type . Among them, the sensitive geographical coverage types in this embodiment are classified according to preset requirements for statistics. For example, the sensitive geographical coverage types are counted according to "Research on Policies and Laws of Geographical Information Security in China".

[0039] Then, construct the sensitive level setting matrix of the sensitive geographical coverage type : ; Among them: represents the sensitivity value of the important type , represents the sensitivity value of the non-public type , the sensitivity value of the non-disclosable shape type .

[0040] Finally, according to the above two constructed matrices and , calculate the sensitive quantitative evaluation value of the geographical coverage type dimension of the geographical entity s : .

[0041] S4-3. According to the geographical location information of the geographical entity , determine its accuracy range and assign the classified quantitative evaluation value of the geographical location dimension . In practical applications, the higher the accuracy of the geographical location information , the larger the classified quantitative evaluation value of the geographical location dimension .

[0042] S4-4. Assign a classified quantitative assessment of the timeliness dimension according to the timeliness of the topic of the geographical entity. In actual applications, the older the corpus is from the current time, the lower its timeliness, and the smaller the classified quantitative assessment value of the timeliness dimension. of the topic timeliness In actual applications, the older the corpus is from the current time, the lower its timeliness, and the smaller the classified quantitative assessment value of the timeliness dimension. is.

[0043] Step 5. For each geographical entity, calculate the final risk sensitivity quantification result according to the classified quantitative assessment value of each dimension and its risk impact weight.

[0044] S5-1. According to the calculation results of step S4, construct a one-dimensional horizontal matrix V of multi-dimensional sensitivity values of the geographical entity , including sensitive application types, sensitive geographical coverage types, geographical locations, and timeliness: ; Where: respectively represent the classified quantitative assessment values of the sensitive application type dimension, the sensitive geographical coverage type dimension, the geographical location dimension, and the timeliness dimension; S5-2. For each sensitive dimension of the geographical entity , preset a one-dimensional vertical matrix K of its related risk impact weight values according to empirical values, that is: ; Where: respectively represent the risk impact weight values of the sensitive application type, the risk impact weight values of the sensitive geographical coverage type, the risk impact weight values of the geographical location, and the risk impact weight values of the timeliness; S5-3. According to the two matrices of S5-1 and S5-2, calculate the final risk sensitivity quantification result of the geographical information entity s: .

[0045] In summary, the present invention can not only discover sensitive geographical information hidden in the original published content and its comment content on Internet social media, but also be able to evaluate the risk sensitivity of the geographical information in a more reasonable and effective way, give a more scientific and effective quantification value, and provide a basis for further geographical information risk disposal.

[0046] The above embodiments are the preferred embodiments of the present application. Those of ordinary skill in the art can also make various transformations or improvements on this basis. Without departing from the overall concept of the present application, these transformations or improvements should all fall within the scope of protection required by the present application.

Claims

1. A method for evaluating the geographical information risk sensitivity of an Internet social media, characterized in that, Including: Step 1: Collect the original published content including geographical information text descriptions on Internet social media and its user comment data. Denote each original published content and all its comments as a conversation corpus, and preprocess each collected conversation corpus. Step 2: Use natural language processing techniques to extract information from the preprocessed conversation corpus, including: geographical entities, geographical location information of each geographical entity, sensitive area coverage type of each geographical location information, sensitive application type of each geographical entity, timeliness of each geographical entity. Step 3: Merge the geographical location information and sensitive area coverage type that belong to the same geographical entity in the extracted information. Step 4: For each geographical entity, calculate the classified quantification evaluation value respectively from the four dimensions of sensitive application type, sensitive area coverage type, geographical location, and timeliness. Step 5: For each geographical entity, calculate the final risk sensitivity quantification result according to the classified quantification evaluation values of each dimension and their risk impact weights.

2. The method for evaluating the geographical information risk sensitivity of an Internet social media according to claim 1, wherein The preprocessing of the data collected in Step 1 includes: completing the abbreviations of the geographical location information therein, and correcting the grammar errors and typos therein.

3. The method for evaluating the geographical information risk sensitivity of an Internet social media according to claim 1, wherein Step 2 specifically includes: S2-1. For the original published content in the dialogue corpus, use a natural language processing engine to identify the geographical entities and their corresponding set of all address location information ; among them, country, province, city name, street name, house number, GPS location and azimuth information all belong to geographical location information; S2-2, for the identified geographical entity , analyze the sensitive application types of the original published content where the geographical entity is located, including three preset sensitive application types, which are respectively recorded as , , ; S2-3, for geographical entities For each geographical location information, identify the sensitive area coverage type it belongs to, including important types , non-disclosure types , non-disclosable shape types ; S2-4. For each comment in the dialogue corpus, perform steps S2-1, S2-2, and S2-3 to obtain the geographical entities, the geographical location information set P of the geographical entities, and the sensitive application types , as well as the sensitive regional coverage types of each geographical location information ; S2-5. Identify the publication time of the original content and each comment, and obtain the timeliness of the corresponding geographical entity information according to the sorting of the publication dates , ; Among them represents the distance from the statistical time 4. The method for evaluating the geographical information risk sensitivity of an Internet social media according to claim 1, characterized in that, Step 3 specifically includes: S3-1. Integrate all the geographical location information in the geographical location information set of the geographical entity to obtain the unique geographical location information of the geographical entity ; ;​ S3-2. For the geographical entity , merge the sensitive area coverage types of all its geographical location information, that is: ; In the formula: represents the statistical value of the sensitive area coverage type associated with the geographical entity s, which are the statistical values of the important type , the statistical value of the non-disclosure type , and the statistical value of the non-disclosable shape type respectively; represents the i-th geographical location information in the original geographical location information set P of the geographical entity of the sensitive area coverage type, where the three elements correspond to the important type, the non-disclosure type, and the non-disclosable shape type in sequence; n is the number of geographical location information included in the original geographical location information set P.

5. The method for evaluating the geographical information risk sensitivity of an Internet social media according to claim 1, characterized in that, In Step 4, for each geographical entity, the specific calculation of the classified quantification evaluation value from the sensitive application type is: First, count the risk matrix R of the sensitive application type: ; Wherein: Represents the number of events of three sensitivity levels of a preset first sensitivity-related application type, with the sensitivity levels from high to low; Represents the number of events of three sensitivity levels of a preset second sensitivity-related application type, with the sensitivity levels from high to low; Represents the number of events of three sensitivity levels of a preset third sensitivity-related application type, with the sensitivity levels from high to low; wherein, the statistical method for the number of events of each sensitivity level of each sensitivity-related application type is: regarding the original published content of the conversation corpus and each comment as a piece of speech, among all the speeches related to the geographical entity, if a piece of speech includes the currently counted sensitivity-related application type and belongs to the currently counted sensitivity level, then the number of events of the currently counted sensitivity level of the currently counted sensitivity-related application type is incremented by 1; Then, construct a sensitivity level setting matrix for sensitivity-related application types and a weight matrix for sensitivity-related application types : , ; Wherein: , , respectively represent the sensitivity values set for different sensitivity levels, sorted from high to low; , , respectively represent the weight values of the preset first, second, and third types of sensitivity-related applications; Finally, based on the three matrices R, , constructed above, calculate the quantitative evaluation value of the geographical entity s in the dimension of sensitive application types: 。 6. The method for evaluating the geographical information risk sensitivity of an Internet social media according to claim 1, wherein In Step 4, for each geographical entity, the specific calculation of the classified quantification evaluation value from the sensitive area coverage type is: First, count the risk matrix for the risk area coverage types related to sensitivity : ; Wherein: represents the statistical values of the sensitive area coverage types associated with the geographical entity s, which are the statistical values of the important type the statistical value of the non-disclosure type the statistical value of the non-disclosable shape type the statistical value; Then, construct a sensitivity level setting matrix for the sensitivity-affected area coverage type : ; Wherein: represents the sensitivity value of the important type , represents the sensitivity value of the non-disclosed type , the sensitivity value of the non-disclosable shape type; ​ Finally, based on the above two matrices constructed and , calculate the sensitivity-related quantification evaluation value of the geographical coverage type dimension of the geographical entity s : 。 7. The method for evaluating the geographical information risk sensitivity of an Internet social media according to claim 1, wherein Step 5 specifically includes: S5-1. Construct a geographical entity A one-dimensional horizontal matrix V of multi-dimensional sensitive values, including sensitive quantification evaluation values in 4 dimensions: classified application type, sensitive area coverage type, geographical location, and timeliness ; Wherein: respectively represent the sensitive quantification evaluation values of the sensitive application type dimension, the sensitive geographical coverage type dimension, the geographical location dimension, and the timeliness dimension; S5-2. For geographical entities For each sensitive dimension, set a one-dimensional longitudinal matrix K of its related risk impact weight values, that is: ; Wherein: respectively represent the risk impact weight value of the sensitivity-related application type, the risk impact weight value of the sensitivity-related geographical coverage type, the risk impact weight value of the geographical location, and the risk impact weight value of timeliness; S5-3. Calculate the final risk-sensitive quantification result of the geographical information subject s based on the two matrices of S5-1 and S5-2 : 。 8. An electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the computer program is executed by the processor, the processor is caused to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Internet secret map detection algorithm based on Spark

    CN109446288A

  • Risk assessment method and device of IP address, storage medium and electronic equipment

    CN117692237A

  • Social media event detection method sensitive to space-time factors

    CN118410167A

  • Intelligent distinguishing method for sensitive information of micro-map text content

    CN119782542A

  • Efficient statistical techniques for detecting sensitive data

    US11599667B1