Emergency sensing method including semantic extraction
By constructing a historical lexicon and expanding entity vocabulary through a two-layer model, and combining user association weights and knowledge graphs, the problem of inaccurate vocabulary extraction in existing situational awareness methods is solved, enabling accurate situational analysis and prediction of emergencies.
Patent Information
- Application Number
- CN202511018294.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-07
AI Technical Summary
Existing situational awareness methods fail to fully utilize real-time text data on the network and neglect user relationships, resulting in inaccurate semantic extraction of words and making it impossible to accurately predict the development of emergencies.
A historical lexicon is constructed, and a variant lexicon is generated through a two-layer model. By combining user association weights and knowledge graphs, entity words are expanded and matched, and the knowledge graph is updated to adapt to changes in network vocabulary.
It improves the reliability of vocabulary expansion and the accuracy of semantic extraction, enabling more accurate prediction of the development of emergencies and real-time updates to the knowledge graph.
Smart Images

Figure CN120911604A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to a sudden event perception method containing semantic extraction. BACKGROUND
[0002] In the prior art, the situation perception method adopts the combination of the semantic extraction combining the language model, the knowledge graph based on the entity knowledge, the situation prediction model based on the information propagation model and the network structure, and thus obtains the collection and identification of the data required for situation prediction, the knowledge reasoning and evolution analysis, and finally the identification and prediction of the occurrence and development situation of the sudden event. However, in the prior art, the semantic extraction model of the entity vocabulary only performs semantic extraction based on the existing semantic analysis model, and can only perform separate vocabulary extraction based on the collected text data, which often ignores the approximate words and other forms of variant words related to the vocabulary, or further expands the vocabulary by means of the pre-collected and stored variant word library, but this can only expand the vocabulary based on the external vocabulary relationship, which obviously does not utilize the text data continuously collected by the model, and also ignores the vocabulary use habits of the users in the network. In the case of the wide use of the Internet and mobile network, the evolution of language is rapid and drastic, the network hot words and the special use of various vocabularies not only develop and spread in an explosive manner, but also have obvious circle characteristics. Therefore, the existing situation perception method does not fully consider the characteristics of the evolution of the existing vocabulary, and needs to be further improved. SUMMARY
[0003] The purpose of the application is to provide a sudden event perception method containing semantic extraction, which solves the technical problem that the prior art ignores the real-time text data in the network and fails to perform more accurate and reliable vocabulary semantic extraction based on user association.
[0004] The sudden event perception method containing semantic extraction comprises the following steps:
[0005] S1, a historical vocabulary library about sudden events is constructed, an entity vocabulary library contains a plurality of vocabulary groups, and an association weight between the vocabulary groups and the users is formed by training the historical data;
[0006] S2, an entity variant word expansion model is constructed based on a double-layer model to generate a variant word library, the double-layer model comprises a proper noun expansion layer of a first layer and a historical vocabulary library expansion layer of a second layer;
[0007] S3, text data on the network is collected, the text data is classified and processed based on a text classification rule, and entity words are preliminarily extracted;
[0008] S4, the extracted entity words are input into the variant word expansion model to obtain corresponding variant vocabulary groups, and the identification and extraction of the entity words are completed.
[0009] S5, performing situation analysis based on the knowledge graph to predict the occurrence and development situation of the emergency.
[0010] Preferably, in step S1, the semantic features and frequency features of the corresponding entity words are extracted from the historical data by a pre-trained language feature model, the similarity between the semantic features of different entity words is calculated, and a vocabulary group is formed by a clustering algorithm, the vocabulary group has a specific core word as a clustering center and contains approximate words, the similarity between the approximate words and the core word is greater than a threshold value, in the clustering analysis, the initial core word is an entity word with the highest frequency feature value, and the association weight between the user and the vocabulary group is calculated by using the frequency features of the entity words and the approximate words used by the specific user.
[0011] Preferably, in step S1, the frequency feature calculation formula is as follows:
[0012]
[0013] wherein f(x) is the frequency feature, x represents an entity word, fr i (x) represents the number of occurrences of the entity word x in the i th data source text, t i represents the vocabulary size of the i th data source text, N is the number of data source texts in the time window, and T Δ represents the duration of the time window Δ.
[0014] The calculation formula of the association weight between the user and the vocabulary group is as follows: wherein ω u,Gi represents the association weight between the user u and the vocabulary group G i , G i is the i th vocabulary group, g ki is the k th word in the i th vocabulary group, including the entity word x and its approximate words, f u (x) is the frequency feature of the entity word x used by the user u.
[0015] Preferably, in step S1, when the association weight is greater than a certain threshold value, the user and the corresponding vocabulary group are classified into a strong association relationship, when the number of vocabulary groups strongly associated by a plurality of users is greater than a set threshold value, these users are classified into a user group with approximate words, the vocabulary groups strongly associated by the users in the user group are fused according to the relevance of the core words, the core word with the minimum average distance between the core word and other approximate words is taken as the core word of the new vocabulary group after fusion, and the association weight between the relevant users and the new vocabulary group is recalculated.
[0016] Preferably, in step S2, the proper noun lexicon contains collected proper nouns. When expanding the vocabulary with variant words, the vocabulary group is first expanded at the first level based on the proper noun lexicon, and the similarity between the core word and the proper noun in the vocabulary group is calculated. When the core word and the proper noun are the same, the corresponding variant word of the proper noun in the proper noun lexicon is merged into the corresponding vocabulary group, while the core word remains unchanged. When the proper noun meets the requirements of a similar word to a specific core word, the proper noun and the relevant variant word in the proper noun lexicon are merged into the vocabulary group, and the core word of the vocabulary group is modified to the corresponding proper noun. When the proper noun does not meet the requirements of a similar word to any core word, the proper noun and its relevant variant word are added to the entity lexicon as a new vocabulary group.
[0017] Preferably, in step S2, a second layer of expansion is performed based on the historical lexicon. For a specific word, a word group containing the specific word is queried from the historical lexicon. If it does not exist, the similarity between the specific word and the core words in other word groups is calculated. When the similarity is less than the similarity threshold of the corresponding word group, it means that the corresponding word group can contain the specific word. Then, it is determined whether the corresponding word can be classified into the corresponding variant word group. After the variant word expansion of the two-layer model, the variant word group of the specific word is obtained, and the variant word group further forms a variant lexicon.
[0018] Preferably, in step S2, the sum of the association weights between the specific word and all related similar words and each user is calculated as the score for determining whether the corresponding word can be classified into the corresponding variant word group. When the score is greater than a set threshold, the corresponding word is included in the variant word group of the specific word. The formula for calculating the score is: in, Let w be a set of a specific word and its similar words, and a be a set of words. The words in Let ω be the set of all users u. u,a The association weight between word a and user u.
[0019] Preferably, in step S2, a contrastive learning approach is used to train the variant vocabulary. Entity words and their semantic features are extracted based on historical data. Entity words serve as query instances, and semantic features are used to determine similarity. During each training iteration, positive samples of words with similarity less than a similarity threshold are found based on the query instances, while other words that do not meet the similarity threshold are treated as negative samples. The model is trained by minimizing the contrastive loss function, and the variant vocabulary group is updated during training. The contrastive loss function L... CE as follows:
[0020] Where sim(·) represents calculating similarity, q is the query instance, and p + For positive samples, p kFor the kth negative sample, K is the number of negative samples, and tau represents a temperature scaling factor.
[0021] Preferably, in step S3, a preset text classification rule is used to divide the text in a sentence into classified texts such as entity text, relationship text and condition text according to grammatical rules, the entity text is the subject or object in grammar, the relationship text is the predicate in grammar, and the condition text is the complement, attributive or adverbial in grammar.
[0022] Preferably, in step S4, after obtaining the variant vocabulary set, the variant vocabulary set is matched with the corresponding vocabulary in the knowledge graph, first, the extracted original word is queried in the knowledge graph, when the corresponding node is queried, the corresponding vocabulary in the variant vocabulary set is marked, thereafter, the words in the variant vocabulary set are queried, the query is preferentially based on the marked vocabulary, then, the nodes in the knowledge graph are bidirectionally expanded, the priority of the search direction is determined by the node connection degree, and the search is pruned by the semantic correlation of other entity words in the text data, and finally, the knowledge subgraph content corresponding to the text data is queried; when the matching result is that the corresponding node cannot be queried, the entity words extracted from the related text are used to form new meta-knowledge, and the knowledge graph is updated.
[0023] The application has the advantages that: the application realizes the expansion of vocabulary based on the analysis of the association relationship between users and the vocabulary in the historical vocabulary library, and can further expand the vocabulary group according to the relationship between users with similar vocabulary habits. In this way, the historical data is effectively utilized, and the vocabulary expansion effect is targeted for customers, the association weight can be better used to mine the variant vocabulary library suitable for the whole network in the association relationship between different users and different similar words, and the reliability of vocabulary expansion is improved.
[0024] On the other hand, the method first expands the variant words of the special nouns in the field based on the special noun library, and then further expands the expanded vocabulary based on the historical vocabulary library based on the association relationship between users and vocabulary, realizes the expansion and matching based on special nouns preferentially, and takes into account the expansion mode of the similarity of the variant words in the historical vocabulary library, realizes more accurate expansion effect, improves the accuracy of semantic extraction, and thus can more accurately predict the development trend of the sudden event.
[0025] Meanwhile, the method can also add new knowledge elements and update the knowledge graph when new vocabulary not contained in the knowledge graph appears based on the results of semantic extraction and entity word matching. In this way, the accurate trend analysis of public safety sudden events is realized, and real-time updating can be realized according to newly collected text data. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1A flow chart of a burst event perception method comprising semantic extraction. DETAILED DESCRIPTION
[0027] The specific embodiments of the present application will be further described in detail below with reference to the drawings, by describing the embodiments, to help the skilled in the art have a more complete, accurate and in-depth understanding of the inventive concept and technical solutions of the present application.
[0028] As shown in the drawings, Figure 1 The present application provides a burst event perception method comprising semantic extraction, comprising the following steps:
[0029] S1, constructing a historical vocabulary about the burst event, the entity vocabulary comprising a plurality of vocabulary groups, and forming the association weight between the vocabulary group and the user through historical data training.
[0030] The semantic features and frequency features of the corresponding entity words are extracted from the historical data through the pre-trained language feature model. The frequency feature is calculated by setting the appearance frequency of the entity word in the corresponding data source text within the time window, and the calculation formula is as follows:
[0031]
[0032] Wherein, f(x) is the frequency feature, x represents the entity word, fr i (x) represents the number of times the entity word x appears in the i-th data source text, t i represents the number of words in the i-th data source text, N is the number of data source texts within the time window, T Δ represents the length of the time window Δ.
[0033] By calculating the similarity between the semantic features of different entity words, the clustering algorithm is used to form the vocabulary group, and the vocabulary group takes a specific core word as the clustering center and contains approximate words with a similarity greater than a threshold. The expression is as follows: s.t.sim(s x ,s ci )≥θ s :V i ←V i ∪{x}
[0034] Wherein, s ci represents the semantic feature of the core word of the i-th vocabulary group, s x represents the semantic feature of the entity word x, V i represents the approximate word set of the i-th vocabulary group.
[0035] In the clustering analysis, the initial core word is the entity word with the highest frequency feature value. The association weight between the user and the vocabulary group is calculated by using the frequency feature of the specific user using the entity word and its approximate word, and the calculation formula is: where ω u,Gi denotes the association weight between user u and the glossary group G i , G i is the ith glossary group, g ki is the kth word in the ith glossary group, including the entity word x and its approximate words, f u (x) is the frequency feature of user u using the entity word x. When the association weight is greater than a certain threshold, the user and the corresponding glossary group are classified as a strong association relationship.
[0036] When the number of glossary groups strongly associated with several users is greater than a set threshold, these users are classified as a group of users using approximate words. The glossary groups commonly strongly associated within the user group are merged according to the relevance of the core words, and the core word with the smallest average distance between it and other approximate words is taken as the core word of the new glossary group after merging. The association weight between the relevant users and the new glossary group is recalculated.
[0037] S2, based on the double-layer model, an entity variant word expansion model is constructed to realize the generation of the variant word library. The double-layer model includes a first layer of proper noun expansion layer and a second layer of historical word library expansion layer.
[0038] The proper noun library contains collected proper nouns. When expanding the variant words, first, based on the proper noun library, the glossary group is expanded in the first layer, and the approximation degree between the core word in the glossary group and the proper noun is calculated. When the core word is the same as the proper noun, the corresponding variant word of the proper noun in the proper noun library is merged into the corresponding glossary group, and the core word remains unchanged. When the proper noun meets the approximate word requirement of a specific core word, the proper noun and the related variant words in the proper noun library are merged into the glossary group, and the core word of the glossary group is modified to the corresponding proper noun. When the proper noun does not meet the approximate word requirement of any core word, the proper noun and its related variant words are added to the entity word library as a new glossary group.
[0039] After the first layer expansion, the glossary groups in the historical word library include unchanged glossary groups, newly added new glossary groups, and merged new glossary groups. Then, based on the historical word library, the second layer expansion is performed. For a specific word, the glossary group containing the specific word is queried from the historical word library. If it does not exist, the similarity between the specific word and the core word in other glossary groups is calculated. When the similarity is less than the similarity threshold of the corresponding glossary group, it means that the corresponding glossary group can contain the specific word. Then, the sum of the association weights between the specific word and all related approximate words and each user is calculated as a score for judging whether the corresponding word can belong to the corresponding variant glossary group. When the score is greater than a set threshold, the corresponding word is included in the variant glossary group of the specific word. The calculation formula of the score is: where ω is the set of the specific word w and its approximate words, and a is the set the words in the text, is a set of all users u, ω u,a is an association weight between the word a and the user u.
[0040] After the variant word expansion of the two-layer model, the variant vocabulary set of the specific vocabulary is obtained, and all entity words in the variant vocabulary set are the corresponding variant words of the specific vocabulary, and the variant vocabulary set further forms a variant word library.
[0041] In the training process, the variant word library is trained in a contrast learning manner, the entity words and semantic features are extracted based on historical data, the entity words are query instances, and the semantic features are used to judge similarity. During each round of training iteration, based on the query instance, the vocabulary positive sample with a similarity less than a similarity threshold is found, and other vocabularies that do not meet the similarity threshold requirement are used as negative samples. The model is trained by minimizing the contrast loss function, and the variant vocabulary set is also updated with the training. Wherein, the contrast loss function L CE is as follows:
[0042]
[0043] Wherein, sim(·) represents calculating similarity, q is a query instance, p + is a positive sample, p k is the kth negative sample, K is the number of negative samples, and τ represents a temperature scaling factor.
[0044] S3, collect text data on the network, classify and process the text data based on text classification rules, and preliminarily extract entity words.
[0045] Collect text data on the network, and preset text classification rules. According to the grammar rules, the text in a sentence is divided into classified texts such as entity text, relationship text and condition text. The entity text is the subject or object in the grammar, the relationship text is the predicate in the grammar, and the condition text is the complement, attributive, adverbial or complement in the grammar. For long sentences in this paper, they can be further divided into short sentences, so as to facilitate the division of text types. The classified texts are input into the pre-trained language feature model, and the corresponding entity words and semantic features are extracted.
[0046] The text classification rules are as follows: "The patient with fever admitted by ×× Hospital yesterday, the name, phone number and ID number of the whole family are in this picture. Everyone, please forward it quickly, don't let your own people step on the mine." In the text, "×× Hospital", "fever patient", "name", "phone number", "ID number", "everyone" and "own people" are entity texts, "admit", "in this picture", "forward" and "step on the mine" are relationship texts, and "yesterday", "of the whole family" and "quickly" are condition texts.
[0047] S4, input the extracted entity word into a variant word expansion model to obtain a corresponding variant word set, and complete the recognition and extraction of the entity word.
[0048] After obtaining the variant word set in this step, the variant word set is matched with the corresponding word in the knowledge graph. When the matching result is completely the same or the word similarity is higher than the threshold, the corresponding node is queried, that is, the word corresponding to the knowledge graph is recognized and extracted. When the matching result is that the word similarity is not higher than the threshold, the extracted entity word in the related text forms new meta-knowledge (triplet or quadruplet), and the knowledge graph is updated.
[0049] When matching, the extracted original word is first queried in the knowledge graph, for example, the original word is recognized as an entity name, then the corresponding node in the knowledge graph is queried according to the classification, if the original word cannot be queried, the variant word is queried until the corresponding node is queried. When the corresponding node is queried, the corresponding word in the variant word set is marked, and thereafter the words in the variant word set are queried, the words are queried based on the marked words. Then, the nodes of the knowledge graph are searched in both directions, the priority of the search direction is determined by the connection degree of the nodes, and the search pruning is performed by the semantic correlation of other entity words in the text data, unnecessary query paths are reduced, and finally the knowledge subgraph content corresponding to the text data is queried.
[0050] S5, based on the knowledge graph, the situation analysis is performed to predict the occurrence and development situation of the emergency event.
[0051] The knowledge subgraph content can preliminarily determine whether an emergency event occurs and the type of the emergency event through reasoning, but the current time development state and the subsequent event development situation need to be predicted. This step also needs to combine the information propagation model related to the emergency event and the characteristics of the social network, and use a reasonable existing situation prediction model to predict. Then, the situation prediction model can be trained and updated based on the actual emergency event development situation data.
[0052] The above describes the present application in conjunction with the drawings, and it is obvious that the specific implementation of the present application is not limited by the above method. Various non-essential improvements using the inventive concept and technical solution of the present application, or directly applying the inventive concept and technical solution to other occasions without improvement, are all within the protection scope of the present application.
Claims
1. A crisis awareness method comprising semantic extraction, characterized in that: The method comprises the following steps: S1, constructing a historical library of emergencies, the entity library containing a plurality of vocabulary groups, and forming an association weight between the vocabulary groups and users through historical data training; S2, generating an entity variant vocabulary expansion model based on a double-layer model to realize the generation of a variant vocabulary library, the double-layer model comprising a first layer of a proper noun expansion layer and a second layer of a historical library expansion layer; S3, collecting text data on the network, classifying and processing the text data based on a text classification rule, and preliminarily extracting entity words; S4, inputting the extracted entity words into the variant vocabulary expansion model to obtain corresponding variant vocabulary groups, and completing the recognition and extraction of the entity words; S5, performing situation analysis based on a knowledge graph to predict the occurrence and development situation of the emergency.
2. The method for detecting sudden events including semantic extraction according to claim 1, characterized in that: In step S1, the semantic features and frequency features of the corresponding entity words are extracted from the historical data through a pre-trained language feature model, the approximation degrees between the semantic features of different entity words are calculated, and a clustering algorithm is used to form a vocabulary group, the vocabulary group taking a specific core word as a clustering center and containing an approximate word, the approximation degree between the approximate word and the core word being greater than a threshold value, in the clustering analysis, the initial core word being an entity word with the highest value of the frequency feature, and the association weight between the user and the vocabulary group being calculated using the frequency features of the entity words and the approximate words used by the specific user.
3. The burst perception method with semantic extraction as claimed in claim 2, wherein: In step S1, the frequency feature calculation formula is as follows: wherein f(x) is a frequency feature, x represents an entity word, fr i (x) represents the number of occurrences of the entity word x in the i-th data source text, t i represents the number of words in the i-th data source text, N is the number of data source texts within the time window, T Δ represents the length of the time window Δ; The calculation formula of the association weight between the user and the vocabulary group is: where ω u,Gi represents the association weight between the user u and the vocabulary group G i , G i is the i-th vocabulary group, g ki is the k-th word in the i-th vocabulary group, including the entity word x and its approximate words, f u (x) is the frequency feature of the user u using the entity word x.
4. The method for detecting sudden events including semantic extraction according to claim 3, characterized in that: In step S1, when the association weight is greater than a certain threshold value, the user and the corresponding vocabulary group are classified into a strong association relationship, when the number of vocabulary groups strongly associated by a plurality of users is greater than a set threshold value, the users are classified into a user group using approximate words, the vocabulary groups strongly associated in the user group are fused according to the correlation of the core words, the core word with the smallest average distance between the core word and other approximate words is taken as the core word of the new vocabulary group after fusion, and the association weight between the relevant users and the new vocabulary group is recalculated.
5. The burst perception method with semantic extraction as claimed in claim 1, wherein: In step S2, the proper noun library contains collected proper nouns, when performing variant vocabulary expansion, first, the vocabulary groups are expanded in the first layer based on the proper noun library, and the approximation degrees between the core words in the vocabulary groups and the proper nouns are calculated; when the core word is the same as the proper noun, the corresponding variant word of the proper noun in the proper noun library is fused into the corresponding vocabulary group, and the core word remains unchanged; when the proper noun meets the approximate word requirement of a specific core word, the proper noun and the related variant words in the proper noun library are fused into the vocabulary group, and the core word of the vocabulary group is modified to the corresponding proper noun; when the proper noun does not meet the approximate word requirement of any core word, the proper noun and the related variant words are added to the entity library as a new vocabulary group.
6. The burst perception method with semantic extraction as claimed in claim 5, wherein: In step S2, the second layer expansion is performed based on the historical vocabulary library. For a specific vocabulary, a vocabulary group containing the specific vocabulary is queried from the historical vocabulary library. If the specific vocabulary does not exist, the similarity between the specific vocabulary and the core vocabulary in other vocabulary groups is calculated. When the similarity is less than the similarity threshold of the corresponding vocabulary group, it is indicated that the corresponding vocabulary group can contain the specific vocabulary. Then, it is determined whether the corresponding vocabulary can be attributed to the corresponding variant vocabulary group. After the variant vocabulary expansion of the two layers of models, the variant vocabulary group of the specific vocabulary is obtained. The variant vocabulary group further forms a variant vocabulary library.
7. The burst perception method with semantic extraction as claimed in claim 6, wherein: In step S2, the sum of the association weights between the specific word and all related similar words and each user is calculated as the score for determining whether the corresponding word can be classified into the corresponding variant word group. When the score is greater than a set threshold, the corresponding word is included in the variant word group of the specific word. The formula for calculating the score is: in, Let w be a set of a specific word and its similar words, and a be a set of words. The words in Let ω be the set of all users u. u,a The association weight between word a and user u.
8. The burst perception method with semantic extraction as claimed in claim 7, wherein: In step S2, the variant vocabulary is trained in a contrastive learning manner, entity words and semantic features thereof are extracted based on historical data, the entity words are query instances, and the semantic features are used to determine whether they are similar. During each training iteration, positive samples of words with similarity less than a similarity threshold are found based on the query instances, while other words that do not meet the similarity threshold requirement are used as negative samples. The model is trained by minimizing a contrastive loss function, and the variant vocabulary group is also updated during the training. The contrastive loss function L CE is as follows: where sim(·) denotes the similarity calculation, q is the query instance, p + is the positive sample, p k is the kth negative sample, K is the number of negative samples, and τ denotes the temperature scaling factor.
9. The method as claimed in claim 1 comprising semantic extraction for disaster awareness, wherein: In step S3, a preset text classification rule is provided. According to the text classification rule, the text in a sentence is divided into classified texts such as entity text, relationship text and condition text. The entity text is the subject or object in the syntax, the relationship text is the predicate in the syntax, and the condition text is the complement, attributive or adverbial in the syntax.
10. The burst perception method with semantic extraction as claimed in claim 1 wherein: In step S4, after the variant vocabulary group is obtained, the variant vocabulary group is matched with the corresponding vocabulary in the knowledge graph. First, the extracted original word is queried in the knowledge graph. When the corresponding node is queried, the corresponding vocabulary in the variant vocabulary group is marked. Thereafter, when the words in the variant vocabulary group are queried, the query is preferentially performed based on the marked vocabulary. Then, the node of the knowledge graph is bidirectionally expanded and searched. The priority of the search direction is determined by the node connection degree. The search is pruned by the semantic correlation of other entity words in the text data. Finally, the knowledge subgraph content corresponding to the text data is queried. When the matching result is that the corresponding node cannot be queried, the entity word extracted from the related text forms new meta-knowledge, and the knowledge graph is updated.
Citation Information
Cited By
Scientific and technological data processing method and device
CN121919781A