A security-resilient urban theme target text data processing method and device
By performing word operations, syntactic analysis, and thematic causal polarity analysis on the safety and resilience urban planning documents of the target region, this study solves the problem that existing technologies cannot reveal the semantic relationships and causal logic of resilience planning policy texts. It achieves scientific and automated processing of text data and improves the ability to understand and evaluate policy logical relationships.
Patent Information
- Application Number
- CN202511248953.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing technologies cannot effectively reveal the semantic relationships and causal logic in resilience planning policy texts from different regions, and cannot meet the processing requirements of resilience planning texts.
By acquiring the text of target planning documents for the theme of safe and resilient cities in the target region, we perform word operation preprocessing, syntactic analysis, structured polarity result generation, document generation model theme matching, and theme causal polarity analysis to output the logical relationship of the development plan for safe and resilient cities in the target region.
It improved the scientific rigor and automation of the processing of textual data on the theme of regional security resilience cities, revealed the underlying logical relationships in policy texts, and supported the comprehensive evaluation of policy strategies.
Smart Images

Figure CN120805931B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of text information processing, in particular to a safety and resilience city theme target text data processing method and device. BACKGROUND
[0002] Different regions have significant differences in policies for resilience development. For example, some regions focus on engineering resilience, improving the ability of the region to resist disasters and risks by strengthening the anti-seismic and flood control capacity of infrastructure; while other regions pay more attention to social resilience, focusing on housing, employment, emergency, education and medical services, and especially on the disaster response capacity of vulnerable groups, so as to improve the overall comprehensive ability of the region to prevent and reduce disasters.
[0003] From a global or multi-regional perspective, the scientific text semantic mining method and system for resilience planning policy research is not sufficient, and the technical means for revealing the semantic association and causal logic in the text of different regions' resilience planning policy cannot meet the requirements of resilience planning text processing. SUMMARY
[0004] The technical problem to be solved by the present application is to provide a safety and resilience city theme target text data processing method and device. The scientificity and automation level of the safety and resilience city theme target text data processing of the region can be improved.
[0005] To solve the above technical problems, the technical solutions of the present application are as follows:
[0006] A safety and resilience city theme target text data processing method, comprising:
[0007] Obtaining a safety and resilience city theme target planning file text of a target region;
[0008] Performing word operation preprocessing on the target planning file text to obtain a target processing text;
[0009] Performing syntax analysis on the target processing text to obtain a syntax semantic unit;
[0010] Obtaining a structured polarity result according to the syntax semantic unit and the verb polarity;
[0011] Obtaining a document generation model theme according to the probability of word allocation to the theme in the target processing text;
[0012] Matching the structured polarity result with the document generation model theme to obtain a matching theme;
[0013] Performing theme causal polarity analysis on the matching theme to obtain a causal relationship between themes;
[0014] According to the inter-topic causal relationship, output a logical relationship of a target area safety and resilience urban development planning.
[0015] Optionally, the target planning file text is subjected to word operation preprocessing to obtain a target processing text, including:
[0016] The target planning file text is subjected to at least one of word segmentation processing, part-of-speech tagging, and lemmatization to obtain a target processing text.
[0017] Optionally, the target processing text is subjected to syntactic analysis to obtain a syntactic semantic unit, including:
[0018] The text with a parallel structure in the target processing text is subjected to recursive parsing to obtain a first intermediate text.
[0019] The text with a passive structure in the first intermediate text is subjected to passive structure identification to obtain a second intermediate text.
[0020] The second intermediate text is subjected to extraction according to the syntactic roles of subject, predicate, and object to obtain a syntactic semantic unit.
[0021] Optionally, a structured polarity result is obtained according to the syntactic semantic unit and a verb polarity, including:
[0022] The verbs in the syntactic semantic unit are subjected to sentiment polarity analysis to obtain a first verb polarity score.
[0023] The structured polarity result is obtained according to the first verb polarity score and the syntactic semantic unit.
[0024] Optionally, a document generation model topic is obtained according to the probability of a word in the target processing text being assigned to a topic, including:
[0025] The number of words in a document that are assigned to a topic and the number of times a topic word appears in a topic are obtained according to the probability of a word in the target processing text being assigned to a topic.
[0026] A document-topic distribution and a topic-word distribution are obtained according to the number of words in the document that are assigned to a topic and the number of times a topic word appears in a topic.
[0027] A document generation model topic is obtained according to the document-topic distribution and the topic-word distribution.
[0028] Optionally, the structured polarity result is matched with the document generation model topic to obtain a matching topic, including:
[0029] performing similarity calculation on the word item in the structured polarity result and a topic word of the document generation model topic to obtain a target similarity;
[0030] obtaining a similarity score of the word item and the document generation model topic according to the target similarity;
[0031] obtaining a matching topic according to the similarity score.
[0032] Optionally, subject causal polarity analysis is performed on the matching topic to obtain an inter-topic causal relationship, including:
[0033] mapping the first verb polarity score of the structured polarity result to a second verb polarity score of the matching topic;
[0034] obtaining an inter-topic causal relationship according to the second verb polarity score.
[0035] Optionally, the subject text data processing method for the safety and resilience city theme category also includes:
[0036] performing word item extraction processing on the target processing text to obtain a noun category item;
[0037] performing statistical analysis on the noun category item to obtain a standard word frequency;
[0038] performing co-occurrence analysis on the noun category item to obtain a co-occurrence matrix;
[0039] outputting a correlation mode of the target region according to the standard word frequency and the co-occurrence matrix.
[0040] Embodiments of the present application also provide a safety and resilience city theme category target text data processing device, including:
[0041] an acquisition module configured to acquire a safety and resilience city theme category target planning file text of a target region;
[0042] a processing module configured to perform word operation preprocessing on the target planning file text to obtain a target processing text, perform syntax analysis on the target processing text to obtain a syntax semantic unit, obtain a structured polarity result according to the syntax semantic unit and a verb polarity, obtain a document generation model topic according to a probability of word allocation to a topic in the target processing text, perform matching on the structured polarity result and the document generation model topic to obtain a matching topic, perform subject causal polarity analysis on the matching topic to obtain an inter-topic causal relationship, and output a safety and resilience city development planning logical relationship of the target region according to the inter-topic causal relationship.
[0043] Embodiments of the present invention also provide a computing device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the target text data processing method for the theme of secure and resilient cities as described in the present invention.
[0044] The above-described technical solution of the present invention has at least the following technical effects:
[0045] The above-mentioned method for processing target text data on the theme of safe and resilient cities of the present invention involves: acquiring target planning document text on the theme of safe and resilient cities in a target region; performing word operation preprocessing on the target planning document text to obtain target processed text; performing syntactic analysis on the target processed text to obtain syntactic and semantic units; obtaining structured polarity results based on the syntactic and semantic units and verb polarity; obtaining document generation model topics based on the probability of words in the target processed text being assigned to topics; matching the structured polarity results with the document generation model topics to obtain matched topics; performing topic causal polarity analysis on the matched topics to obtain causal relationships between topics; and outputting the logical relationship of the development plan for safe and resilient cities in the target region based on the causal relationships between topics. This method can improve the scientific rigor and automation level of target text data processing on the theme of safe and resilient cities in a region. Attached Figure Description
[0046] Figure 1 This is a flowchart illustrating the method for processing target text data related to the theme of safe and resilient cities according to the present invention.
[0047] Figure 2 This is a schematic diagram of the framework of the method for processing target text data on the theme of safe and resilient cities according to the present invention;
[0048] Figure 3 This is a schematic diagram of the target text data processing device for the theme of safe and resilient cities according to the present invention. Detailed Implementation
[0049] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0050] like Figure 1 As shown, embodiments of the present invention propose a method for processing target text data on the theme of safe and resilient cities, including:
[0051] Step S1, obtaining a target planning file text of at least one target region;
[0052] Step S2, performing word operation preprocessing on the target planning file text to obtain a target processing text;
[0053] Step S3, performing syntax analysis on the target processing text to obtain a syntax semantic unit;
[0054] Step S4, obtaining a structured polarity result according to the syntax semantic unit and a verb polarity;
[0055] Step S5, obtaining a document generation model theme according to a probability of a word in the target processing text being assigned to a theme;
[0056] Step S6, matching the structured polarity result with the document generation model theme to obtain a matching theme;
[0057] Step S7, performing theme causal polarity analysis on the matching theme to obtain an inter-theme causal relationship;
[0058] Step S8, outputting a target region safety resilience urban development planning logical relationship according to the inter-theme causal relationship.
[0059] In the embodiment, the target planning file text of the target region can be a relevant policy file, a planning file, a core issue file in resilience construction, etc. of the target region. For the target planning file text of each target region in the multiple target regions, uniform word operation preprocessing is performed to ensure the accuracy and consistency of subsequent analysis.
[0060] The content in each target planning file text is subjected to theme extraction and subject-predicate-object structure extraction, and the subject and object are associated to the theme through semantic similarity, and then a causal relationship network of theme hierarchy is constructed, potential logical relationships in resilience policies of regions are revealed, structured analysis and causal logic modeling of the target planning file texts of multiple regions are realized, deep understanding and application value of the text content are improved, and this is helpful to understanding the construction logic of resilience region development policies and providing technical support for comprehensive evaluation of region policy strategies.
[0061] In an optional embodiment of the present application, in step S2, the target planning file text is subjected to word operation preprocessing to obtain a target processing text, including:
[0062] Step S21, at least one of the following is performed on the target planning text: word segmentation processing, part-of-speech tagging, and morphological restoration, to obtain a target processing text.
[0063] In this embodiment, first, at least one of converting capital letters in the text into lowercase, deleting numeric characters, deleting punctuation marks and abbreviation protection is performed to obtain an operation-processed text, at least one of word segmentation processing, part-of-speech tagging and lemmatization is performed on the operation-processed text to obtain a target processing text.
[0064] Word segmentation is a process of converting continuous text data into a structured word sequence, that is, the original target planning file text is disassembled into independent word units; the purpose of word segmentation is to split the text into meaningful language units (such as words or tokens) to facilitate subsequent processing (such as part-of-speech tagging, syntax analysis, etc.); in the word segmentation process, through a preset abbreviation dictionary (such as NYC, CPU), regular expressions are used to identify and retain abbreviations to avoid segmentation errors, and abbreviation protection can ensure the integrity of professional terms and lay a foundation for subsequent semantic analysis.
[0065] Then, lemmatization is performed, which is to remove the affixes of a word and extract the stem of the word, and the extracted word is usually a word in the dictionary.
[0066] Again, part-of-speech tagging is a process of assigning a grammatical category to each word in the text, such as nouns, verbs, adjectives, etc. Part-of-speech tagging establishes a structured representation of the text to facilitate computer understanding and processing of natural language. The intermediate text is part-of-speech tagged by combining the en_core_web_sm model in the natural language processing library (spaCy); by analyzing the grammatical structure of the sentence, the subject (nsubj represents the active voice subject, nsubjpass represents the passive voice subject), predicate (the core verb marked with ROOT) and object (dobj is the direct object, and attr is used to represent the attribute object) are accurately identified. For example, in the sentence "The regional management department formulates a new policy", it is clear that "the regional management department" is the nsubj active voice subject, "formulates" is the ROOT verb, and "a new policy" is the dobj direct object. After preprocessing, the target processing text is obtained for subsequent processing.
[0067] In an optional embodiment of the present application, in step S3, the target processing text is subjected to syntax analysis to obtain a syntax semantic unit, comprising:
[0068] In step S31, a subtext with a parallel structure in the target processing text is recursively parsed to obtain a first intermediate text;
[0069] In step S32, a subtext with a passive structure in the first intermediate text is subjected to passive structure identification to obtain a second intermediate text;
[0070] Step S33, the second intermediate text is extracted according to the syntactic role of the subject, predicate and object, and the syntactic semantic unit is obtained.
[0071] In this embodiment, in the face of complex parallel structure texts, a recursive parsing strategy is adopted, and taking "A and B implement C and D" as an example, the method will be disassembled into four groups of syntactic semantic units: (A, implement, C) (A, implement, D) (B, implement, C) (B, implement, D). The recursive parsing strategy breaks through the limitations of traditional syntactic analysis, can effectively deal with the common multi-agent, multi-behavior combination expression in policy texts, and ensures the integrity and accuracy of information extraction; after recursive parsing of the subtext with parallel structure in the target processing text, the first intermediate text is obtained.
[0072] On the basis of the first intermediate text, the passive structure of the subtext in the first intermediate text is identified, and the second intermediate text is obtained; specifically, for the diversified language expression in the text, a corresponding processing mechanism is designed. In the passive structure processing, such as identifying "C is implemented by A", it is automatically converted into active semantic expression (A, implement, C); for negative expression, when detecting words such as "not to implement", the polarity of the verb will be marked as negative. This mechanism can adapt to the complex semantic logic in policy texts and improve the reliability of information extraction.
[0073] Finally, the subject (nsubj), object (dobj), agent (representing the action relationship between the subject and the object), predicate and other syntactic roles of the second intermediate text are identified by using the natural language processing library spaCy, and the subject-predicate-object (SVO) triple is extracted.
[0074] After the dependency syntactic analysis and expression conversion processing, artificial auditing is performed by professional personnel. The accuracy of the semantic triple is checked, and the missing information is supplemented, and finally the (S, V, O) structure in which the subject and the object are both nouns is reserved. For example, (management agency, implements, garbage classification policy), this kind of standardized triple provides a solid data foundation for subsequent causal relationship analysis, policy target disassembly and other deep processing.
[0075] The processing flow disassembles the complex regional policy text into standardized semantic units by accurately capturing the action subject, behavior and object in the text. These atomic units provide reliable semantic support for subsequent causal relationship extraction, policy correlation analysis and other high-level natural language processing tasks.
[0076] In an optional embodiment of the present application, in step S4, a structured polarity result is obtained according to the syntactic semantic unit and the verb polarity, including:
[0077] Step S41, sentiment polarity analysis is performed on the verb in the syntactic semantic unit to obtain a first verb polarity score.
[0078] Step S42, obtaining a structured polarity result according to the first verb polarity score and the syntactic semantic unit.
[0079] In this embodiment, the sentiment polarity analysis is performed on the verbs of the syntactic semantic unit, and a professional verb sentiment dictionary is used, such as the VADER sentiment analysis tool in the NLTK library of the Python language open source tool set, and the verb V in the text is recognized through the part-of-speech tagging tool.
[0080] For each verb, a first verb polarity score polarity(V) is determined, which takes a value range of {-1, 0, 1}, corresponding to negative sentiment, neutral sentiment and positive sentiment respectively. For example, the polarity value of the verb "destroy" is -1, the polarity value of the verb "maintain" is 0, and the polarity value of the verb "push" is 1.
[0081] By matching the verb polarity sentiment dictionary, a structured four-tuple of subject, verb, object and verb polarity is constructed.
[0082] The sentiment polarity of the verb in the syntactic semantic unit is taken as a new element, combined with the syntactic semantic unit, and a structured four-tuple data model (S, V, O, p) is constructed, that is, a structured polarity result, wherein S represents the action subject, V is the sentiment verb, O is the action object, and p represents the polarity(V) verb sentiment polarity. By setting a filtering condition, the neutral four-tuple with polarity = 0 is automatically removed, and the key semantic units with emotional tendency are focused. For example, for the sentence "policy optimization promotes economic growth", it can be extracted as ("policy optimization", "promote", "economic growth", 1), and "project continues to run" will be screened out because the verb "run" has a polarity of 0.
[0083] The structured polarity result can give the emotional direction attribute to the causal relationship network in the text, and by quantifying the sentiment polarity, the positive influence or negative consequences of the policy measures on the regional development can be accurately distinguished. This step provides structured emotional data support for the subsequent value evaluation, risk warning and other depth analysis of the strategy text.
[0084] In an optional embodiment of the present application, a document generation model is used to output the topic probability distribution of each document and the word probability distribution under each topic, so as to realize the topic extraction and analysis of the text. In step S5, the document generation model topic is obtained according to the probability of assigning words in the target processing text to a topic, which can include:
[0085] In step S51, the number of words in the document assigned to the topic and the number of times of occurrence of the topic word in the topic are obtained according to the probability of assigning words in the target processing text to a topic.
[0086] Step S52: Based on the number of words assigned to topics in the document and the number of times topic words appear in the topics, obtain the document-topic distribution and topic-word distribution;
[0087] Step S53: Based on the document-topic distribution and topic-word distribution, obtain the document generation model topic.
[0088] Specifically, in the process of generating topics for the document generation model of the target processed text, based on the probability of words in the target processed text being assigned to topics, the number of words assigned to topics in the document and the frequency of topic words appearing in the topics are obtained, including:
[0089] Define parameters α (a prior parameter for document-topic distribution) and β (a prior parameter for topic-word distribution). Randomly assign words from each document to one of K topics to obtain the initial topic assignments. , where i represents the i-th document and j represents the j-th word in the i-th document;
[0090] Count the number of words in each document m that are assigned to topic k. And the number of times word t appears in each topic k. .
[0091] Calculate global statistics: , The total number of words in document m. , The total number of words for topic k. N represents the total number of words in all documents, M is the total number of documents, and V0 is the size of the vocabulary.
[0092] For each word in each document, remove the topic assignment for the current word and update the statistics.
[0093] The probability of a word being assigned to any topic is calculated using the following formula:
[0094]
[0095] in, This indicates the topic allocation of all words except the current word. This represents the statistics of topic k in document m after removing the current word. This represents the statistic after removing the current word t when it appears in topic k. Indicates in Statistics after removing the current word Indicates in The statistics after removing the current word are displayed. The symbol ∝ represents proportionality, indicating a proportional relationship between two quantities. The word item is represented. Based on the calculated probability, the word item is re-assigned to a topic, and the statistics are updated.
[0096] In the generation process of the target document generation model topic, according to the number of word items assigned to the topic in the document and the number of times the topic word appears in the topic, a document-topic distribution and a topic-word distribution are obtained, comprising:
[0097] Parameter estimation and output, according to the final statistical quantity, estimate the document-topic distribution and the topic-word distribution :
[0098]
[0099]
[0100] For example: set the number of topics num_topics=5, indicating that the regional policy text is divided into 5 potential topics, 5 topics T1-T5 are extracted, and the corresponding topic word set under each topic is obtained, which are {t11, t12,..., t1k},..., {t51, t52,..., t5m}. These topic words intuitively reflect the core semantics of each topic, for example, {t11, t12} under T1 topic may be "traffic planning" and "road construction", which implies that the topic focuses on the regional transportation field.
[0101] In the generation process of the target document generation model topic, according to the document-topic distribution and the topic-word distribution, the document generation model topic is obtained, comprising:
[0102] The document-topic distribution and the topic-word distribution are combined to obtain the document generation model topic.
[0103] In this embodiment, perplexity and coherence score are used to verify the stability of the topic. Perplexity is used to measure the prediction ability of the model on the corpus, and the lower the value, the better the understanding and fitting effect of the model on the text; the coherence score evaluates the closeness between topic words from the perspective of semantic similarity, with a value range of 0-1, and the closer to 1, the stronger the semantic correlation between topic words. In specific implementation, topic words with a coherence score lower than 0.4 are removed, for example, there are words deviating from the core semantics under a certain topic. Through this threshold screening, high-representative topic words can be effectively retained, and each topic can accurately summarize the corresponding text content.
[0104] The original text is abstracted from the vocabulary level to the topic dimension through the topic model, realizing the high-level extraction of the text semantics. This process not only reduces the data dimension, but also provides a higher level of perspective for subsequent analysis of causal logic. For example, through the analysis of the topic dimension, the causal relationship between the "green energy theme" and the "regional carbon emission policy theme" can be explored, providing management agencies with more systematic and forward-looking decision-making basis, thereby realizing the dimensional leap from text data to deep logic.
[0105] In an optional embodiment of the present application, in step S6, the structured polarity result is matched with the document generation model topic to obtain a matching topic, comprising:
[0106] In step S61, the word in the structured polarity result is similarity calculated with the topic word of the document generation model topic to obtain a target similarity;
[0107] In step S62, the similarity score of the word and the document generation model topic is obtained according to the target similarity;
[0108] In step S63, a matching topic is obtained according to the similarity score.
[0109] In this embodiment, step S61 can include:
[0110] Loading the word vector model to calculate the similarity of the word and the topic word:
[0111]
[0112] where w and t are the word and the topic word respectively, · is the inner product, and ||·|| is the L2 norm, is the target similarity.
[0113] For a given word w, by constructing a bidirectional attention mechanism, the feature vector representation of the word in different semantic contexts is considered comprehensively, and the cosine similarity, semantic correlation and co-occurrence frequency of all topic words in the topic Tk are calculated respectively. On this basis, a dynamic weight adjustment factor is introduced, and according to the importance of the topic word in the document, the TF-IDF (Term Frequency-Inverse Document Frequency) value of the word itself and the context semantic coherence, the weighted sum of each similarity index is weighted and summed, and finally the comprehensive weighted topic score is generated.
[0114] TF-IDF (Term Frequency-Inverse Document Frequency) measures the importance of a word in a document by combining term frequency (TF) and inverse document frequency (IDF). Term frequency counts the frequency of a word in a single document, and the higher the frequency, the stronger the relevance of the word to the document theme. Inverse document frequency reflects the universality of a word in the entire corpus by calculating the logarithm of the ratio of the total number of documents in the corpus to the number of documents containing the word, avoiding over-weighting common words (such as "of" and "is"). The final TF-IDF score is equal to the product of TF and IDF, and the higher the score, the more the word represents the core content of the document.
[0115] The integrated weighted topic score can more accurately reflect the semantic association degree of the word item w and the topic Tk, effectively improving the accuracy and robustness of topic recognition and text classification.
[0116] In this embodiment, step S62 can include:
[0117] For a word item w, the similarity score of all topic words t of a topic Tk is calculated:
[0118]
[0119] where freq(t) is the normalized term frequency of the topic word t.
[0120] Two threshold values are introduced to control the matching quality: a similarity threshold (SIMILARITY_THRESHOLD) and a minimum topic score threshold (MIN_TOPIC_SCORE). The similarity threshold is used to filter word items that are semantically close to the topic words, and the minimum topic score threshold is used to determine whether the mapping to the target topic is successful. Both thresholds are systematically adjusted, with a step size of 5 percentage points, and then multiple rounds of performance comparison and verification are performed, and finally the similarity threshold SIMILARITY_THRESHOLD = 25% is determined, filtering Sim(w, t) ≥ 25% of the topic words; the topic score threshold MIN_TOPIC_SCORE = 15%, only keeping score(w, Tk) ≥ 15% of the mapping. By setting the double thresholds, accurate mapping from vocabulary to topic is achieved, avoiding semantic deviation.
[0121] In an optional embodiment of the present application, in step S7, topic causal polarity analysis is performed on the matching topic to obtain the causal relationship between topics, including:
[0122] Step S71 maps the first verb polarity score of the structured polarity result to the second verb polarity score of the matching topic;
[0123] Step S72 obtains the causal relationship between topics according to the second verb polarity score.
[0124] In the embodiment, the structured polarity result (S, V, O, p) is mapped to a subject relationship (Ts, V, To, p), wherein Ts is a subject to which S belongs, and To is a subject to which O belongs. According to the verb polarity, the influence direction of the subject to the object is determined, and the second verb polarity score is added by 1 for a positive relationship (p = 1) or subtracted by 1 for a negative relationship (p = -1), that is, the connection strength strength (Ts, To) is accumulated to realize the quantification of the direction and strength of the causal relationship between the subjects.
[0125] In an optional embodiment of the present application, in step S8, a target area safety resilience urban development planning logical relationship is output according to the causal relationship between the subjects, including:
[0126] In step S81, the causal relationship between the subjects of a preset number of subjects is visualized according to the causal relationship between the subjects and the second verb polarity score, and a target area safety resilience urban development planning logical relationship is output.
[0127] In the embodiment, the NetworkX is used to construct a causal network, the preset number of core subjects of the target processing text are used as nodes, such as subjects T1-T5, then the second verb polarity score is used as the connection strength strength (Ts, To) between the subjects, only the positive relationship (strength > 0) is retained, and the connection strength strength (Ts, To) is marked between the related subjects, so as to output a target area safety resilience urban development planning logical relationship.
[0128] In an optional embodiment of the present application, the above method can further include:
[0129] In step S9, a term extraction process is performed on the target processing text to obtain a noun class item.
[0130] In step S10, a statistical analysis is performed on the noun class item to obtain a standard word frequency.
[0131] In step S11, a co-occurrence analysis is performed on the noun class item to obtain a co-occurrence matrix.
[0132] In step S12, a correlation mode of a target area is output according to the standard word frequency and the co-occurrence matrix.
[0133] In this embodiment, the text preprocessing is divided into words, and the part of speech is obtained after part-of-speech tagging. The words and phrases with the part of speech as nouns are called noun phrases. When extracting the words from the target processing text, the text is first divided into independent words and phrases with no more than four words. The output format is that each word or phrase occupies a separate line, and the sentences are separated by blank lines to keep the text structure clear. Then, the word frequency of each line is calculated, and the lines containing only stop words are deleted. Subsequently, according to the part-of-speech tagging results of spaCy, only the word frequency of nouns and noun phrases is retained, and other parts of speech are filtered. Finally, the word frequency is standardized to convert it into a frequency of every ten thousand words to ensure the data comparability between different regional target processing texts.
[0134] Specifically, for phrases in the text, the Gensim library is used to mine candidate phrases. By setting min_count=5, only continuous word combinations that appear at least 5 times in the text corpus are included in the candidate range, which effectively eliminates meaningless phrases that occur by chance. At the same time, threshold=100 is set as the phrase scoring threshold, and only word combinations with a score exceeding this threshold are recognized as phrases to improve the quality of candidate phrases.
[0135] In the candidate phrase screening link, based on the part-of-speech tagging results, only the phrase combinations composed of nouns (NN represents singular nouns, NNS represents plural nouns), adjectives (JJ), and nouns are retained. For example, when analyzing regional transportation policy texts, phrases like "public transportation system" (NN+NN+NN) and "efficient urban planning" (JJ+NN+NN) are retained. At the same time, to ensure that the extracted phrases have sufficient generality and practicality, the length of the phrase is strictly limited to 1-4 words, and non-core phrase structures such as transitive verbs (e.g., "develop infrastructure", VB+NNS) and adverbials (e.g., "through collaboration", IN+NN) are filtered out to avoid including non-key expressions that describe specific actions or relationships into core concepts.
[0136] Taking a typical regional policy text as an example, "sustainable development" (JJ+NN) as a core concept can accurately reflect the policy theme and is therefore retained. However, phrases like "implement policies" (VB+NNS) that emphasize specific implementation actions are filtered out as they cannot directly reflect the core content of the policy.
[0137] By accurately extracting core concept phrases from policy texts, the interference of irrelevant phrases on subsequent word frequency analysis is effectively reduced. During in-depth analysis such as high-frequency word statistics and topic modeling, focusing on core concepts can significantly improve the accuracy and effectiveness of the analysis results, helping researchers or policymakers more efficiently grasp the core content and key issues of policy texts.
[0138] When counting noun phrases, traverse the preprocessed text (1 word / phrase per line, separated by blank lines between sentences), and count the initial word frequency count(w).
[0139] Filter out pure stopword lines (e.g., lines containing only "the" and "and") using the stopword list (nltk.corpus.stopwords). Based on spaCy part-of-speech tagging, only keep noun word frequencies (NN, NNS, NNP, NNPS) to build a noun word frequency library (fn(w)).
[0140] The standardization formula is:
[0141]
[0142] Where nf(w) represents the standardized word frequency of word w, fn(w) represents the occurrence frequency of word w, and tnw represents the total number of noun words / phrases in the text.
[0143] Standardized word frequency can eliminate the influence of text length differences, making word frequencies of policy texts from different regions horizontally comparable.
[0144] Conduct co-occurrence analysis on noun phrases to obtain a co-occurrence matrix. Based on spaCy tagging, extract words with part-of-speech tags such as nouns (NN, NNS), proper nouns (NNP, NNPS), gerunds (NN gerund), and adjective-noun phrases (JJ+NN) from sentence-level text (separated by blank lines). Form a sentence-level noun list S=[w1,w2,...,wn]. Count all word pairs (wi,wj) (i<j) in sentence S with sentence as the co-occurrence boundary. Use collections.defaultdict(int) to store co-occurrence frequency co_occur(wi,wj). Express the word pair co-occurrence frequency in matrix form to construct the co-occurrence matrix. Then convert the co-occurrence matrix to a symmetric matrix that satisfies co_occur(wi,wj)=co_occur(wj,wi). Store the symmetric matrix in CSV format using pandas, with row / column indices as word items and cell values as co-occurrence frequencies, as shown in Table 1.
[0145] Table 1 Co-occurrence Matrix Example
[0146] Emergency Reserve Risk Emergency 0 36 19 Reserve 36 0 11 Risk 19 11 0
[0147] The co-occurrence matrix can completely capture all concept associations within a sentence, and provide structured data for semantic network analysis.
[0148] The co-occurrence weight is then calculated by the following formula:
[0149]
[0150] wherein, represents the co-occurrence weight, co_occur(wi, wj) represents the co-occurrence frequency, represents the normalized word frequency.
[0151] In an optional embodiment of the present application, the above method can further include:
[0152] Step S13: outputting the fusion analysis result of at least one target region according to the association mode of the target region and the logical relationship of the target region development.
[0153] In this embodiment, in the co-occurrence weight, the 30 word pairs with the highest weight are selected, a chord diagram is drawn using a data visualization tool, the word frequency is represented by the arc length, the chord width represents the co-occurrence weight, the node is used as the visualization carrier of the keyword, the node size intuitively maps the appearance frequency of the keyword in the text corpus, the high-frequency word corresponds to a larger node, which facilitates quick identification of core concepts; the connection between nodes represents the semantic association relationship of the keywords, and the thickness of the edge accurately quantifies the association strength, the thicker the connection line, the higher the frequency of the two keywords appearing in the text and the stronger the semantic coupling degree. At the same time, the node distribution can be optimized by a layout algorithm, and different semantic categories can be distinguished by color coding, and finally a knowledge graph with information density and readability is generated, which helps policy makers intuitively understand the concept network structure and potential logical context in the regional strategy text. According to the association mode of the target region corresponding to the chord diagram and the logical relationship of the target region development, the association mode and the logical relationship of the target region are intuitively displayed and further analyzed, and the fusion analysis result of at least one target region is output.
[0154] As shown in Figure 2 , the present embodiment takes the resilience policy texts of a first region, a second region, a third region and a fourth region as an example, and systematically processes and analyzes the policy documents of each region. Of course, it is not limited to four regions, and specifically includes the following steps:
[0155] Step 1: collect the resilience policy texts of the four regions, i.e. the above-mentioned target planning document texts, to ensure that the text sources are authoritative and the content is complete. For the target planning document text of each region, uniform text preprocessing operations are performed to ensure the accuracy and consistency of subsequent analysis.
[0156] Step 2: For the preprocessed text of the four regions, word frequency statistics are performed, high-frequency keywords are extracted, and the focus and core content of the resilience policy of each region are reflected.
[0157] Step 3: For the preprocessed text of the four regions, a non-window limited noun pair co-occurrence relationship is constructed based on sentences, a co-occurrence matrix of each region is formed, and the top 30 co-occurrence pairs are selected for visualization.
[0158] Step 4: LDA topic extraction and subject-predicate-object structure extraction are performed for the four regions, and the subject and object are associated with the theme through semantic similarity, and a causal relationship network of the theme hierarchy is constructed to reveal the potential logical relationship in the resilience policy of each region.
[0159] Step 5: The analysis results of the four regions are analyzed to clarify the commonalities and differences of the resilience policies of different regions, and to provide data support and theoretical basis for cross-regional resilience policy comparison research.
[0160] The above method can reveal the core issues and policy priorities of these regions in resilience construction by analyzing the common high-frequency words in the resilience planning policy documents of the regions, and can provide important theoretical basis and technical support for the scientific formulation and optimization of regional resilience related policies. By analyzing the top 10 high-frequency keywords, the significant differences and characteristics of the resilience construction methods of each region can be revealed. The structured analysis and causal logic modeling of multi-regional resilience policy text is achieved, which improves the deep understanding and application value of policy text.
[0161] Through co-occurrence analysis of the resilience policy text, the invention constructs a policy keyword co-occurrence network, and systematically reveals the structural association and semantic features in the policy content. It can effectively extract the core elements of the policy and their internal relationships, reflect the logical structure and key orientation of the policy. It helps to improve the depth of analysis and processing efficiency of resilience regional policy text, and provides reliable basis for related policy analysis and decision support.
[0162] Through the topic model extraction of policy text, the subject-predicate-object structure is extracted by dependency syntax analysis, the subject and object are classified into corresponding themes by using semantic similarity, and the causal relationship network with direction and strength between themes is constructed according to the emotional polarity of verbs. The invention reveals the structural logic of the policy document and the internal association between themes. It can effectively extract the deep semantic features of regional resilience development policy and mine the influence path and logic chain between different policy themes.
[0163] Through the application of the invention, it is helpful to understand the construction logic of resilience regional development policy, and to provide technical support for the horizontal comparison and comprehensive evaluation of different regional policy strategies.
[0164] For example, Figure 3As shown, the embodiment of the present application also provides a security and resilience urban theme target text data processing device 30, comprising:
[0165] An acquisition module 31 is configured to acquire a security and resilience urban theme target planning file text of a target region.
[0166] A processing module 32 is configured to perform word operation preprocessing on the target planning file text to obtain a target processing text, perform syntax analysis on the target processing text to obtain a syntax semantic unit, obtain a structured polarity result according to the syntax semantic unit and a verb polarity, obtain a document generation model theme according to a probability of a word in the target processing text being assigned to a theme, match the structured polarity result with the document generation model theme to obtain a matching theme, perform theme causal polarity analysis on the matching theme to obtain an inter-theme causal relationship, and output a logical relationship of a security and resilience urban development planning of the target region according to the inter-theme causal relationship.
[0167] Optionally, the word operation preprocessing on the target planning file text to obtain the target processing text comprises:
[0168] At least one of the following is performed on the target planning file text: word segmentation processing, part-of-speech tagging, and lemmatization, to obtain the target processing text.
[0169] Optionally, the syntax analysis on the target processing text to obtain the syntax semantic unit comprises:
[0170] The target processing text whose syntax structure is a parallel structure is recursively parsed to obtain a first intermediate text.
[0171] The first intermediate text whose syntax structure is a passive structure is identified to obtain a second intermediate text.
[0172] The second intermediate text is extracted according to the syntax roles of subject, predicate, and object to obtain the syntax semantic unit.
[0173] Optionally, the structured polarity result is obtained according to the syntax semantic unit and the verb polarity, comprising:
[0174] The verbs in the syntax semantic unit are subjected to sentiment polarity analysis to obtain a first verb polarity score.
[0175] The structured polarity result is obtained according to the first verb polarity score and the syntax semantic unit.
[0176] Optionally, the document generation model theme is obtained according to the probability of the word in the target processing text being assigned to the theme, comprising:
[0177] According to the probability of the word in the target processing text being assigned to the topic, the number of words in the document being assigned to the topic and the number of times of the topic word appearing in the topic are obtained;
[0178] According to the number of words in the document being assigned to the topic and the number of times of the topic word appearing in the topic, a document-topic distribution and a topic-word distribution are obtained;
[0179] According to the document-topic distribution and the topic-word distribution, a document generation model topic is obtained.
[0180] Optionally, the structured polarity result is matched with the document generation model topic to obtain a matched topic, including:
[0181] A term in the structured polarity result is subjected to similarity calculation with a topic word of the document generation model topic to obtain a target similarity;
[0182] According to the target similarity, a similarity score of the term and the document generation model topic is obtained;
[0183] According to the similarity score, a matched topic is obtained.
[0184] Optionally, the matched topic is subjected to topic causal polarity analysis to obtain an inter-topic causal relationship, including:
[0185] A first verb polarity score of the structured polarity result is mapped to a second verb polarity score of the matched topic;
[0186] According to the second verb polarity score, an inter-topic causal relationship is obtained.
[0187] Optionally, the processing module 32 is further configured to:
[0188] The target processing text is subjected to term extraction processing to obtain a noun class term;
[0189] The noun class term is subjected to statistical analysis to obtain a standard word frequency;
[0190] The noun class term is subjected to co-occurrence analysis to obtain a co-occurrence matrix;
[0191] According to the standard word frequency and the co-occurrence matrix, an association mode of a target region is output.
[0192] It should be noted that all implementation manners in the method embodiments are applicable to the device embodiments, and the same technical effects can be achieved.
[0193] The embodiment of the present application also provides a computing device, comprising: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the safety and resilience urban theme target text data processing method described in the present application. All implementation manners in the above method embodiment are applicable to the embodiment of the computing device, and the same technical effects can also be achieved.
[0194] Those skilled in the art can clearly understand that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software mode depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0195] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0196] In the embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0197] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0198] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0199] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, and various program code storage media.
[0200] In addition, it should be noted that in the device and method of the present application, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombination should be considered as equivalent solutions of the present application. Moreover, the steps of performing the above series of processes can naturally be executed in time sequence according to the order of description, but do not necessarily have to be executed in time sequence, and some steps can be executed in parallel or independently of each other. It can be understood by those skilled in the art that all or any steps or components of the method and device of the present application can be implemented in hardware, firmware, software or a combination thereof in any computing device (including processors, storage media, etc.) or network of computing devices, which can be implemented by those skilled in the art using their basic programming skills after reading the description of the present application.
[0201] Therefore, the object of the present application can also be achieved by running a program or a set of programs on any computing device. The computing device can be a commonly known general-purpose device. Therefore, the object of the present application can also be achieved by providing a program product containing program code for implementing the method or device. That is, such a program product also constitutes the present application, and a storage medium storing such a program product also constitutes the present application. Obviously, the storage medium can be any commonly known storage medium or any storage medium developed in the future. It should be noted that in the device and method of the present application, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombination should be considered as equivalent solutions of the present application. Moreover, the steps of performing the above series of processes can naturally be executed in time sequence according to the order of description, but do not necessarily have to be executed in time sequence. Some steps can be executed in parallel or independently of each other.
[0202] The above is the preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should also be considered within the scope of protection of the present application.
Claims
1. A secure and resilient urban theme-based target text data processing method, characterized in that, The method comprises the following steps: obtaining a target planning document text of a safety and resilience city theme in a target region; performing word operation preprocessing on the target planning document text to obtain a target processing text; performing syntax analysis on the target processing text to obtain a syntax semantic unit; obtaining a structured polarity result according to the syntax semantic unit and a verb polarity; obtaining a document generation model theme according to a probability of a word in the target processing text being assigned to a theme; matching the structured polarity result with the document generation model theme to obtain a matching theme; performing theme causal polarity analysis on the matching theme to obtain a causal relationship between themes; outputting a logical relationship of a safety and resilience city development plan in the target region according to the causal relationship between themes; wherein, obtaining a structured polarity result according to the syntax semantic unit and a verb polarity comprises: performing sentiment polarity analysis on a verb in the syntax semantic unit to obtain a first verb polarity score; obtaining a structured polarity result according to the first verb polarity score and the syntax semantic unit; performing theme causal polarity analysis on the matching theme to obtain a causal relationship between themes comprises: mapping the first verb polarity score of the structured polarity result to a second verb polarity score of the matching theme; obtaining a causal relationship between themes according to the second verb polarity score.
2. The secure resilient urban theme-based target text data processing method according to claim 1, wherein, performing word operation preprocessing on the target planning document text to obtain a target processing text comprises: performing at least one of the following processing on the target planning document text: word segmentation processing, part-of-speech tagging, and morphological restoration, to obtain a target processing text.
3. The secure resilient urban theme-based target text data processing method according to claim 1, wherein, performing syntax analysis on the target processing text to obtain a syntax semantic unit comprises: performing recursive parsing on a text with a parallel structure in the target processing text to obtain a first intermediate text; performing passive structure identification on a text with a passive structure in the first intermediate text to obtain a second intermediate text; extracting a syntax semantic unit according to the syntax roles of subject, predicate, and object in the second intermediate text.
4. The secure-resilient urban-themed target text data processing method of claim 1, wherein, obtaining a document generation model theme according to a probability of a word in the target processing text being assigned to a theme comprises: obtaining the number of words assigned to a theme in a document and the number of times a theme word appears in a theme according to the probability of a word in the target processing text being assigned to a theme; obtaining document-theme distribution and theme-word distribution according to the number of words assigned to a theme in a document and the number of times a theme word appears in a theme; obtaining a document generation model theme according to the document-theme distribution and the theme-word distribution.
5. The secure resilient urban themed object text data processing method of claim 1, wherein, matching the structured polarity result with the document generation model theme to obtain a matching theme comprises: performing similarity calculation on a word in the structured polarity result and a theme word of the document generation model theme to obtain a target similarity; obtaining a similarity score of the word and the document generation model theme according to the target similarity; obtaining a matching theme according to the similarity score.
6. The secure resilient urban theme-based target text data processing method according to any one of claims 1 to 5, characterized in that, The method further comprises the following steps: performing word extraction processing on the target processing text to obtain a noun item; performing statistical analysis on the noun item to obtain a standard word frequency; Performing co-occurrence analysis on the noun class items to obtain a co-occurrence matrix; According to the standard word frequency and the co-occurrence matrix, output the association mode of the target area.
7. A secure and resilient urban theme-based target text data processing device, characterized by, Comprise: An acquisition module is configured to acquire a security and resilience city theme class target planning file text of a target area; A processing module is configured to perform word operation preprocessing on the target planning file text to obtain a target processing text; Performing syntax analysis on the target processing text to obtain a syntax semantic unit; obtaining a structured polarity result according to the syntax semantic unit and the verb polarity; obtaining a document generation model theme according to the probability of assigning a word in the target processing text to a theme; matching the structured polarity result with the document generation model theme to obtain a matching theme; performing theme causal polarity analysis on the matching theme to obtain an inter-theme causal relationship; and outputting a security and resilience city development planning logic relationship of the target area according to the inter-theme causal relationship; According to the syntax semantic unit and the verb polarity, the structured polarity result is obtained, comprising: Performing sentiment polarity analysis on the verbs in the syntax semantic unit to obtain a first verb polarity score; According to the first verb polarity score and the syntax semantic unit, the structured polarity result is obtained; The matching theme is subjected to theme causal polarity analysis to obtain an inter-theme causal relationship, comprising: Mapping the first verb polarity score of the structured polarity result to a second verb polarity score of the matching theme; According to the second verb polarity score, the inter-theme causal relationship is obtained.
8. A computing device, comprising: Comprise: One or more processors; A storage device is configured to store one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the method as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Text semantic analysis method
CN109271626A
Internal and external emotion causal relationship analysis method, system, equipment and medium
CN120562553A