Earthquake disaster data acquisition method and system based on multi-source data
Through the earthquake disaster data collection method based on multi-source data, ordinary network term groups and keyword groups are screened out, keyword group collections are dynamically updated, and single sentences related to disaster situations are screened out, solving the problem of falsification information collection and vague network terminology, and achieving comprehensive and accurate collection of earthquake disaster information.
Patent Information
- Application Number
- CN202510534786.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the prior art, the collection of earthquake disaster information has the problem of poor accuracy in information screening due to fake information and fuzzy network terms, especially because user descriptions tend to be more life-oriented and colloquial, resulting in insufficient effectiveness of earthquake disaster information collection.
The earthquake disaster data collection method based on multi-source data is adopted. By obtaining disaster information in different earthquake areas, ordinary network terms and keyword groups are selected, keyword group collections are dynamically updated, and single sentences related to disaster situations are selected, and information irrelevant to disaster situations are eliminated to ensure the comprehensiveness and accuracy of the information.
It improves the effectiveness of earthquake disaster data collection, reduces noise interference, ensures the comprehensiveness and accuracy of disaster situation description, adapts to dynamic changes in disaster situations, and achieves real-time and accuracy of information.
Smart Images

Figure CN120046611B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of seismic data processing, and specifically relates to a method and system for collecting seismic disaster data based on multi-source data. Background Art
[0002] Usually, text information related to earthquakes on a certain network platform is obtained, and the severity of disasters in different regions is obtained according to the frequency of occurrence of earthquake-related keywords. However, in actual applications, the information released by a single user is one-sided, and there are forged information or obvious deviations in sentence meanings although the keywords are repeated in the collected text information.
[0003] In the prior art, regular expression keyword matching etc. are used to distinguish effective disaster information. However, since users usually use online terms and other vocabulary in their descriptions, their description words tend to be more life-like and colloquial, and there may be certain forged information, resulting in vague expressions and poor accuracy of information screening, and the effectiveness of seismic disaster information collection is worse. Summary of the Invention
[0004] In order to solve the technical problem that the effectiveness of seismic disaster information collection becomes worse due to the lack of consideration of description vocabulary such as online terms, the purpose of the present invention is to provide a method and system for collecting seismic disaster data based on multi-source data, and the specific technical solutions adopted are as follows:
[0005] The present invention proposes a method for collecting seismic disaster data based on multi-source data, and the method includes:
[0006] Obtain the disaster information in different earthquake regions at each moment, obtain the word groups of each single sentence in the disaster information, and a set of keyword groups;
[0007] For any moment, form suspected online term groups by any two adjacent word groups in each single sentence; according to the word group distribution of each suspected online term group in the single sentence, screen out habitual online term groups; according to the correlation between different word groups in the habitual online term groups and the set of keyword groups, update the set of keyword groups to obtain an updated set of keyword groups;
[0008] According to the word group distribution of the updated set of keyword groups appearing in each single sentence, and the word group correlation between the corresponding single sentence and the updated set of keyword groups, obtain the possibility of disaster-unrelated information for each single sentence, and screen out single sentences related to disasters; for the single sentences related to disasters corresponding to any earthquake region, according to the subject difference between each single sentence and the other single sentences, and the possibility of disaster-unrelated information of the other single sentences, screen out effective disaster description single sentences;
[0009] Obtain the retained single sentences of the disaster situation description for each earthquake area at the real-time moment according to the number of valid single sentences of the disaster situation description in each earthquake area at different moments.
[0010] Furthermore, the method for obtaining the conventional Internet buzzword groups includes:
[0011] For any moment, obtain the conventional Internet buzzword possibility of each suspected Internet buzzword group according to the distribution of each suspected Internet buzzword group in the single sentence.
[0012] If the conventional Internet buzzword possibility of the suspected Internet buzzword group is greater than the preset possibility threshold, use the corresponding suspected Internet buzzword group as the conventional Internet buzzword group.
[0013] Furthermore, the method for obtaining the conventional Internet buzzword possibility includes:
[0014] For any moment, obtain the maximum value of the number of other characters within all the occurrence ranges in the single sentence for each suspected Internet buzzword group, and the ratio between the number of characters in the single sentence where each occurrence is located as the local coefficient of the interval distance for each occurrence; obtain the average value of the local coefficients of the interval distances for all occurrences of each suspected Internet buzzword group as the overall coefficient of the interval distance.
[0015] Obtain the ratio of the occurrence frequency of each suspected Internet buzzword group in all single sentences to the number of single sentences as the first possible coefficient.
[0016] Perform a negative correlation mapping on the overall coefficient of the interval distance, multiply the result of the negative correlation mapping by the first possible coefficient, and perform normalization as the conventional Internet buzzword possibility of each suspected Internet buzzword group.
[0017] Furthermore, the method for obtaining the updated set of keyword groups includes:
[0018] Obtain the cosine similarity of the corresponding word vectors of different phrases between each conventional Internet buzzword group and the set of keyword groups, and perform a negative correlation normalization mapping as the first difference.
[0019] Obtain the average value of all the first differences between each conventional Internet buzzword group and the set of keyword groups as the mean difference.
[0020] If the mean difference of the conventional Internet buzzword group is less than the preset difference threshold, add the corresponding conventional Internet buzzword group to the set of keyword groups to form the updated set of keyword groups.
[0021] Furthermore, the method for obtaining the possibility of disaster situation - irrelevant information includes:
[0022] Obtain the ratio of the number of characters in each single sentence to the total number of characters of all the phrases in the updated set of keyword groups that appear in the corresponding single sentence as the first possible coefficient.
[0023] Obtain multiple triples for each single sentence, and obtain the average relative distance of the word vectors corresponding to different phrases between all triples and the updated set of keyword phrases as the average distance between each single sentence and the updated set of keyword phrases.
[0024] According to the first possible coefficient, the number of occurrences of phrases in the earthquake keyword update library in each single sentence, and the average distance between each single sentence and the updated set of keyword phrases, obtain the possibility of disaster-unrelated information for each single sentence. Both the first possible coefficient and the average distance are positively correlated with the possibility of disaster-unrelated information, and the number of occurrences of phrases is negatively correlated with the possibility of disaster-unrelated information.
[0025] Further, the method for obtaining the disaster-related single sentences includes:
[0026] Normalize the possibility of disaster-unrelated information for all single sentences. If the normalization result corresponding to a single sentence is less than or equal to the preset unrelated threshold, regard the corresponding single sentence as a disaster-related single sentence.
[0027] Further, the method for obtaining the effective disaster description single sentences includes:
[0028] For the disaster-related single sentences corresponding to any earthquake area, according to the subject difference between each single sentence and the rest of the single sentences, and the possibility of disaster-unrelated information of other single sentences, obtain the degree of disaster description difference between each single sentence and other single sentences.
[0029] Normalize the degree of disaster description difference between all single sentences and other single sentences. If there is a single sentence whose normalization result with other single sentences is less than the preset difference threshold, regard the corresponding single sentence as an effective disaster description single sentence.
[0030] Further, the method for obtaining the degree of disaster description difference includes:
[0031] Among all the triples of each single sentence, select the triples containing the phrases in the updated set of keyword phrases as effective triples.
[0032] For the disaster-related single sentences in any earthquake area, obtain the cosine similarity between each effective triple of each single sentence and all the effective triples of each other single sentence, and accumulate the negatively correlated mapped cosine similarities as the local description difference of each effective triple of each single sentence relative to each other single sentence.
[0033] According to the local difference in the description of each single sentence's different valid triples relative to all other single sentences, as well as the possibility of disaster-unrelated information corresponding to other single sentences, the degree of difference in disaster situation description between each single sentence and other single sentences is obtained. The local difference in description is positively correlated with the degree of difference in disaster situation description, and the possibility of disaster-unrelated information is negatively correlated with the degree of difference in disaster situation description.
[0034] Furthermore, the method for obtaining single sentences retaining disaster situation description includes:
[0035] For any earthquake area, obtain the average value of the ratio of the number of single sentences with valid disaster situation description at different times to the number of all single sentences as the effective occupancy coefficient;
[0036] Obtain the product of the number of single sentences with valid disaster situation description at the real-time moment and the effective occupancy coefficient, and round up to obtain the acquisition quantity at the real-time moment;
[0037] Select the number of acquisitions with the smallest possible value of disaster-unrelated information among all single sentences with valid disaster situation description as the single sentences retaining disaster situation description.
[0038] The present invention also provides a seismic disaster situation data acquisition system based on multi-source data, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned seismic disaster situation data acquisition methods based on multi-source data are implemented.
[0039] The present invention has the following beneficial effects:
[0040] For any moment, the present invention forms suspected Internet buzzword groups from any adjacent word groups in each single sentence, capturing possible Internet buzzwords or catchphrases in the sentence; according to the word group distribution of each suspected Internet buzzword group in the single sentence, it screens out conventional Internet buzzword groups, screens out frequently-occurring Internet buzzwords, and improves the accuracy of Internet buzzword recognition; according to the correlation between different word groups in the conventional Internet buzzword groups and the keyword group set, it updates the keyword group set to obtain an updated keyword group set, and through correlation analysis, dynamically updates the keyword group set so that it can reflect the latest Internet buzzword and catchphrase trends; according to the word group distribution of the updated keyword group set that appears in each single sentence, and the word group correlation between the corresponding single sentence and the updated keyword group set, it obtains the possibility of disaster-unrelated information for each single sentence, and screens out single sentences related to the disaster; for the single sentences related to the disaster corresponding to any earthquake area, according to the subject difference between each single sentence and the rest of the single sentences, and the possibility of disaster-unrelated information of other single sentences, it screens out effective disaster description single sentences, eliminates single sentences unrelated to the disaster, and reduces the interference of noise on subsequent analysis; according to the number of effective disaster description single sentences of each earthquake area at different moments, it obtains the disaster description retained single sentences of each earthquake area at the real-time moment, achieving a balance between information redundancy and loss, and ensuring the comprehensiveness and accuracy of disaster descriptions. The present invention improves the effectiveness of earthquake disaster data collection by analyzing the accurate performance of disaster information during earthquakes and screening out effective disaster description single sentences. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a flowchart of a method for collecting earthquake disaster data based on multi-source data provided by an embodiment of the present invention;
[0043] Figure 2 It is a flowchart of a method for obtaining the possibility of conventional Internet buzzwords provided by an embodiment of the present invention;
[0044] Figure 3 It is a flowchart of a method for obtaining the possibility of disaster-unrelated information provided by an embodiment of the present invention;
[0045] Figure 4 It is a flowchart of a method for obtaining the degree of difference in disaster descriptions provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a method and system for collecting earthquake disaster situation data based on multi-source data proposed according to the present invention, including its specific implementation manner, structure, features and effects, as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0048] The following specifically describes the specific solution of a method and system for collecting earthquake disaster situation data based on multi-source data provided by the present invention in conjunction with the accompanying drawings.
[0049] Please refer to Figure 1 , which shows the flowchart of a method for collecting earthquake disaster situation data based on multi-source data provided by an embodiment of the present invention, specifically including:
[0050] Step S1: Obtain the disaster situation information of different earthquake regions at each moment, obtain the phrases of each single sentence in the disaster situation information, and the set of keyword phrases.
[0051] In the embodiment of the present invention, in order to ensure the effectiveness of earthquake disaster situation information, it is necessary to analyze the information text in the earthquake disaster situation; first, in order to reduce the impact of incomplete description of the disaster situation by a single user, web crawler technology is used to crawl the relevant text information of the earthquake disaster situation on multiple network platforms for analysis, obtain the disaster situation information of different earthquake source regions at each moment, obtain the disaster situation information of different earthquake regions at each moment, obtain the phrases of each single sentence in the disaster situation information, and the set of keyword phrases. Among them, the single sentences are divided by punctuation marks.
[0052] It should be noted that in the embodiment of the present invention, each single sentence is segmented by jieba to remove meaningless function words and prepositions, and the corresponding multiple phrases are obtained; the specific jieba algorithm is a technical means well-known to those skilled in the art and will not be elaborated here.
[0053] It should be noted that the keyword group set is a set of earthquake-related phrases. The method for obtaining the keyword group set is to obtain the cosine similarity of the corresponding word vectors between each phrase and the earthquake phrase, and perform normalization to obtain the earthquake relevance of each phrase. The greater the cosine similarity, the more likely each phrase is related to the description of the earthquake phrase. If the earthquake relevance of a phrase is greater than the preset relevance threshold, the corresponding phrase is used as a keyword phrase. It should be noted that in an embodiment of the present invention, the sentences are segmented into word vectors through a word2vec model. In other embodiments of the present invention, the BERT algorithm can also be used to obtain word vectors; the specific word2vec and BERT algorithms are well-known technical means to those skilled in the art and will not be elaborated here.
[0054] It should be noted that in an embodiment of the present invention, the interval between moments is 5 minutes, and the preset relevance threshold is taken as 0.7. In other embodiments of the present invention, the interval between moments and the preset relevance threshold can be specifically set according to specific situations and will not be limited or elaborated here.
[0055] Step S2: For any moment, form suspected Internet slang phrase groups by any two adjacent phrases in each single sentence; screen out habitual Internet slang phrase groups according to the phrase distribution of each suspected Internet slang phrase group in the single sentence; update the keyword group set according to the relevance between the habitual Internet slang and different phrases in the keyword group set to obtain an updated keyword group set.
[0056] Internet slang usually consists of multiple words and has a specific combination pattern in terms of semantics and grammar. By traversing the phrases, the combination characteristics of the phrases can be captured; form suspected Internet slang phrase groups by any two adjacent phrases in each single sentence.
[0057] It should be noted that the method for obtaining suspected Internet slang phrase groups includes: starting from the first phrase in the single sentence, successively incorporating any number of subsequent adjacent phrases to obtain all possible adjacent phrase combinations, that is, suspected Internet slang phrase groups; taking an example, if the phrases of a certain single sentence are ccac, then the suspected Internet slang phrase groups include c, cc, cca, ccac, ca, cac, a, ac.
[0058] Since when Internet users post comments, their descriptive words tend to be more life-like and colloquial, and there may be word order differences in narration, but Internet slang is usually a fixed combination commonly used by most users. Therefore, according to the distribution of each suspected Internet slang phrase group in the single sentence, habitual Internet slang phrase groups are screened out.
[0059] Preferably, in an embodiment of the present invention, the method for obtaining habitual Internet slang phrase groups includes:
[0060] At any moment, according to the distribution of each suspected Internet buzzword group in a single sentence, obtain the likelihood of the habitual Internet buzzword for each suspected Internet buzzword group;
[0061] Preferably, in an embodiment of the present invention, for the method of obtaining the likelihood of the habitual Internet buzzword, please refer to Figure 2 , which shows a flowchart of a method for obtaining the likelihood of the habitual Internet buzzword, including:
[0062] Step S201: At any moment, obtain the maximum value of the number of other characters within the occurrence range of each suspected Internet buzzword group in a single sentence, and the ratio between the number of characters in the single sentence where each occurrence is located, as the local coefficient of the interval distance for each occurrence; obtain the average value of the local coefficients of the interval distance for all occurrences of each suspected Internet buzzword group, as the overall coefficient of the interval distance.
[0063] Since when some users use Internet buzzwords, they will add characters according to their own language habits. By analyzing the maximum value of the number of other characters within the occurrence range of each suspected Internet buzzword group in a single sentence, and the ratio between the number of characters in the single sentence where each occurrence is located, it can be found that the less the number of other characters appears, the more it indicates that it is a habitual Internet buzzword.
[0064] It should be noted that in the embodiments of the present invention, the occurrence range of each suspected Internet buzzword group in a single sentence refers to traversing all single sentences at each moment, and taking the phrase range corresponding to the appearance of the suspected Internet buzzword group in the single sentence as the occurrence range. That is, for the suspected Internet buzzword group cc in the phrase ccac of a certain single sentence, for the first occurrence on this single sentence, the occurrence range is cc, and the number of other characters is 0. For the second occurrence, the occurrence range is cac, and the number of other characters is the number of characters of the other occurrence phrase a.
[0065] Step S202: Obtain the ratio of the occurrence frequency of each suspected Internet buzzword group in all single sentences to the number of single sentences, as the first possible coefficient.
[0066] The occurrence frequency reflects the universality of the phrase in the text. The higher the occurrence frequency of the phrase, the more likely it is to be an Internet buzzword or a buzzword, and the larger the first possible coefficient.
[0067] Step S203: Perform a negative correlation mapping on the overall coefficient of the interval distance, multiply the result of the negative correlation mapping by the first possible coefficient, and perform normalization, as the likelihood of the habitual Internet buzzword for each suspected Internet buzzword group.
[0068] In an embodiment of the present invention, the formula for the likelihood of the habitual Internet buzzword is expressed as:
[0069] ;
[0070] ;
[0071] Wherein, represents the likelihood of the conventional Internet term for the th suspected Internet term group; represents the occurrence frequency of the th suspected Internet term group in all single sentences; represents the ratio of the number of single sentences; represents the overall coefficient of the interval distance at which the th suspected Internet term group appears in a single sentence; represents the maximum number of other characters within the range of all occurrences of the th suspected Internet term group in a single sentence; represents the number of characters in the single sentence where the th suspected Internet term group appears for the th time in a single sentence; represents the number of times the th suspected Internet term group appears in a single sentence; represents the logistic function.
[0072] In the formula for the likelihood of a conventional Internet term, adding 0.01 is to avoid a zero denominator in the formula, making the formula meaningless; the larger the overall coefficient of the interval distance, the more other characters there are within the occurrence range, and the smaller the likelihood of it being a conventional description; represents obtaining the occurrence frequency of each suspected Internet term group in all single sentences and the ratio of the number of single sentences, that is, the first possible coefficient. The larger the first possible coefficient, the more frequently each suspected Internet term group appears in all single sentences, the greater the credibility of the analyzed interval distance. The smaller the overall coefficient of the interval distance, the greater the likelihood of it being a conventional description. The larger the first possible coefficient, the more frequently the suspected Internet term group appears, the greater the credibility of it being a conventional description, and the greater the likelihood of the conventional Internet term.
[0073] If the likelihood of a conventional Internet term for a suspected Internet term group is greater than the preset likelihood threshold, the corresponding suspected Internet term group is regarded as a conventional Internet term group.
[0074] It should be noted that in an embodiment of the present invention, the size of the preset likelihood threshold is 0.7; in other embodiments of the present invention, the size of the preset likelihood threshold can be specifically set according to specific circumstances, which will not be limited and elaborated herein.
[0075] Conventional Internet phrase groups represent Internet phrases or buzzwords and have high reference value; the keyword group set is candidate keywords extracted from disaster information. Through correlation analysis, conventional Internet phrase groups highly relevant to the keyword group can be screened out to expand the relevant content of the keyword group. According to the correlation between different phrases in the conventional Internet phrase group and the keyword group set, the keyword group set is updated to obtain an updated keyword group set.
[0076] Preferably, in an embodiment of the present invention, the method for obtaining the earthquake keyword update library includes:
[0077] Obtain the cosine similarity of the corresponding word vectors of different phrases between each conventional Internet phrase group and the keyword group set, and perform negative correlation normalization mapping as the first difference;
[0078] Obtain the mean value of all the first differences between each conventional Internet phrase group and the keyword group set as the difference mean value;
[0079] If the difference mean value of the conventional Internet phrase group is less than the preset difference threshold, add the corresponding conventional Internet phrase group to the keyword group set to form an updated keyword group set.
[0080] It should be noted that in an embodiment of the present invention, the size of the preset difference threshold is 0.3; in other embodiments of the present invention, the size of the preset difference threshold can be specifically set according to specific situations and will not be limited and elaborated here.
[0081] Step S3: According to the phrase distribution of the updated keyword group set in each single sentence, and the phrase correlation between the corresponding single sentence and the updated keyword group set, obtain the possibility of disaster-unrelated information for each single sentence, and screen out single sentences related to the disaster; for any single sentence related to the disaster corresponding to an earthquake area, according to the subject difference between each single sentence and the remaining single sentences, and the possibility of disaster-unrelated information of other single sentences, screen out effective disaster description single sentences.
[0082] When the main sentence meaning of the keyword is irrelevant to the earthquake disaster, the proportion of earthquake-related keyword groups in their corresponding conventional Internet phrase groups is small, and the relevance between the main components in the single sentence and the earthquake disaster is small. Therefore, according to the phrase distribution of the updated keyword group set in each single sentence, and the phrase correlation between the corresponding single sentence and the updated keyword group set, obtain the possibility of disaster-unrelated information for each single sentence, and screen out single sentences related to the disaster.
[0083] Preferably, in an embodiment of the present invention, for the method of obtaining the possibility of disaster-unrelated information, please refer to Figure 3 , which shows a flowchart of a method for obtaining the possibility of disaster-unrelated information, including:
[0084] Step S301: Obtain the number of characters in each single sentence, and the ratio of the total number of characters of all phrases in the keyword group update set that appear in the corresponding single sentence, as the first possible coefficient.
[0085] The smaller the first possible coefficient, the higher the proportion of the keyword group in the single sentence, indicating that the single sentence has a stronger correlation with the keyword group and is more likely to be a sentence related to the disaster situation.
[0086] Step S302: Obtain multiple triples for each single sentence, and obtain the average value of the relative distances of the word vectors corresponding to different phrases between all triples and the keyword group update set, as the average distance between each single sentence and the keyword group update set.
[0087] Word vectors can capture the semantic information of phrases and reflect the semantic correlation between phrases; relative distance is used to quantify the semantic similarity between triples and keyword groups. The smaller the relative distance, the closer the semantics, and the stronger the semantic correlation between the single sentence and the keyword group.
[0088] It should be noted that in some embodiments of the present invention, the relative distance can be calculated by existing distance calculation methods such as calculating the Euclidean distance or Manhattan distance. The specific means are well-known technical means to those skilled in the art and will not be elaborated here.
[0089] It should be noted that a triple usually consists of a subject, a predicate, and an object. Entities in a single sentence are obtained according to the existing named entity recognition technology, and then the dependency syntactic analysis tool is used to analyze the grammatical structure of the sentence to determine the dependency relationship between phrases. Furthermore, a relationship extraction tool is used to determine the subject, predicate, and object to form a triple. The specific means are well-known technical means to those skilled in the art and will not be elaborated here.
[0090] Step S303: According to the first possible coefficient, the number of times the earthquake keyword update library phrases appear in each single sentence, and the average distance between each single sentence and the keyword group update set, obtain the possibility of disaster-unrelated information for each single sentence. Both the first possible coefficient and the average distance are positively correlated with the possibility of disaster-unrelated information, and the number of times of the phrase is negatively correlated with the possibility of disaster-unrelated information.
[0091] It should be noted that the smaller the first possible coefficient, the higher the proportion of the character count of the keyword group in the single sentence, the stronger the correlation between the single sentence and the keyword group, and the lower the possibility of disaster-unrelated information, showing a positive correlation; the larger the number of times of the phrase, the more likely it is related to the disaster situation information, and the smaller the possibility of disaster-unrelated information, showing a positive correlation; the larger the average distance, the greater the difference in the meanings represented by the phrases, and the greater the possibility of disaster-unrelated information.
[0092] In an embodiment of the present invention, the formula for the possibility of disaster-unrelated information is expressed as:
[0093] ;
[0094] Among them, represents the possibility of disaster-unrelated information of the th single sentence; represents the number of characters of the th single sentence; represents the total number of characters in which all the phrases in the keyword group update set appear in the th single sentence; represents the average value of the relative distances of the word vectors corresponding to different phrases between all the triples in the th single sentence and the keyword group update set, that is, the average distance; represents the number of times the phrases in the earthquake keyword update library appear in the th single sentence.
[0095] In the formula for the possibility of disaster-unrelated information, and add 0.01 to avoid the denominator of the formula being 0 and the formula being meaningless; represents the ratio of the number of characters of the th single sentence to the total number of characters in which all the phrases in the keyword group update set appear in the corresponding single sentence, that is, the first possibility coefficient. The larger the first possibility coefficient, the smaller the total number of characters in which all the phrases in the keyword group update set appear, the greater the possibility of disaster-unrelated information, the greater the number of times of the phrases, the more likely it is to be related to the disaster information, and the smaller the possibility of disaster-unrelated information; the larger the average distance, the greater the difference in the meanings represented between the phrases, and the greater the possibility of disaster-unrelated information.
[0096] Preferably, in an embodiment of the present invention, the method for obtaining disaster-related single sentences includes:
[0097] Normalize the possibility of disaster-unrelated information of all single sentences. If the normalization result corresponding to a single sentence is less than or equal to a preset unrelated threshold, the corresponding single sentence is used as a disaster-related single sentence.
[0098] It should be noted that, in an embodiment of the present invention, the size of the preset unrelated threshold is 0.5; in other embodiments of the present invention, the size of the preset unrelated threshold can be specifically set according to specific situations, and no limitation and elaboration are made here.
[0099] Since there may be some forged or incorrect information in online comments, that is, the described disaster situation may deviate seriously from the objective situation of the disaster situation in the corresponding earthquake area. At the same time, there is a high consistency in the objective descriptions of the disaster situation in the same area by users, while the consistency between false information and the overall text information of the area is poor. By analyzing the theme differences between single sentences and the possibility of information unrelated to the disaster situation, it helps to evaluate effective earthquake-related information. For any single sentence related to the disaster situation corresponding to an earthquake area, according to the subject difference between each single sentence and the rest of the single sentences, as well as the possibility of information unrelated to the disaster situation in other single sentences, effective disaster situation description single sentences are screened out.
[0100] Preferably, in an embodiment of the present invention, the method for obtaining effective disaster situation description single sentences includes:
[0101] For any single sentence related to the disaster situation corresponding to an earthquake area, according to the subject difference between each single sentence and the rest of the single sentences, as well as the possibility of information unrelated to the disaster situation in other single sentences, the degree of difference in disaster situation description between each single sentence and other single sentences is obtained;
[0102] Preferably, in an embodiment of the present invention, for the method for obtaining the degree of difference in disaster situation description, please refer to Figure 4 , which shows a flowchart of a method for obtaining the degree of difference in disaster situation description, including:
[0103] Step S401: Among all the triples of each single sentence, select the triples containing the phrases in the updated set of keyword phrases as effective triples.
[0104] The updated set of keyword phrases contains phrases highly related to the disaster situation. Effective triples can reduce the subsequent calculation amount and focus on effective information.
[0105] Step S402: For any single sentence related to the disaster situation of an earthquake area, obtain the cosine similarity of the corresponding word vectors of different phrases between each effective triple of each single sentence and all the effective triples of each other single sentence, and accumulate after negative correlation mapping of all the cosine similarities, as the local description difference of each effective triple of each single sentence relative to each other single sentence.
[0106] Cosine similarity is used to quantify the semantic similarity between two triples. The higher the similarity, the closer the semantics. After negative correlation mapping, the similarity is converted into a difference. The higher the similarity, the smaller the difference, and the smaller the local description difference.
[0107] Step S403: According to the local descriptive differences of different valid triples of each single sentence relative to all other single sentences, and the possibility of disaster-unrelated information corresponding to other single sentences, obtain the degree of disaster situation description difference between each single sentence and other single sentences. The local descriptive difference is positively correlated with the degree of disaster situation description difference, and the possibility of disaster-unrelated information is negatively correlated with the degree of disaster situation description difference.
[0108] In an embodiment of the present invention, for the single sentences related to the disaster situation in any earthquake area, the formula for the degree of disaster situation description difference is expressed as:
[0109] ;
[0110] where, represents the degree of disaster situation description difference between each single sentence and other single sentences; represents the local descriptive difference of the th valid triple of each single sentence relative to the th other single sentence; represents the possibility of disaster-unrelated information of the th other single sentence; represents the number of the th other single sentence; represents the number of valid triples in each single sentence.
[0111] In the formula for the degree of disaster situation description difference, represents the accumulation of the local descriptive differences of all valid triples of each single sentence relative to the th other single sentence, as the overall descriptive difference of each single sentence for the th other single sentence; represents the ratio of the overall descriptive difference of each single sentence for the th other single sentence to the corresponding possibility of disaster-unrelated information. The larger the ratio, the greater the overall descriptive difference, the greater the local descriptive difference, the smaller the possibility of disaster-unrelated information, and the more it indicates that the th other single sentence is an earthquake situation, and the greater the contribution degree to the authenticity of the overall descriptive difference. The greater the degree of disaster situation description difference between each single sentence and other single sentences, the local descriptive difference is positively correlated with the degree of disaster situation description difference, and the possibility of disaster-unrelated information is negatively correlated with the degree of disaster situation description difference.
[0112] Normalize the degree of disaster situation description difference between all single sentences and other single sentences. If the normalization result between a single sentence and other single sentences is less than the preset difference threshold, the corresponding single sentence is used as a valid disaster situation description single sentence.
[0113] It should be noted that, in an embodiment of the present invention, the size of the preset difference threshold is 0.5; in other embodiments of the present invention, the size of the preset difference threshold can be specifically set according to specific circumstances, and no limitation and elaboration will be made here.
[0114] Step S4: Obtain the disaster situation description retained single sentences of each earthquake area at the real-time moment according to the number of valid disaster situation description single sentences of each earthquake area at different moments.
[0115] Dynamically adjusting the number of retained single sentences can adapt to the dynamic changes of the disaster situation, ensure the timeliness and comprehensiveness of the disaster situation description, and improve the effectiveness of the earthquake information in this area.
[0116] Preferably, in an embodiment of the present invention, the method for obtaining the disaster situation description retained single sentences includes:
[0117] For any earthquake area, obtain the ratio mean of the number of valid disaster situation description single sentences and the number of all single sentences at different moments as the effective proportion coefficient;
[0118] Obtain the product of the number of valid disaster situation description single sentences at the real-time moment and the effective proportion coefficient, and round up to obtain the acquisition quantity at the real-time moment;
[0119] Select the smallest number of acquisition quantities with the smallest possibility value of disaster situation irrelevant information among all valid disaster situation description single sentences as the disaster situation description retained single sentences.
[0120] It should be noted that, in another embodiment of the present invention, after obtaining the disaster situation description retained single sentences of each earthquake area, the information acquisition results of different regions can also be output; obtain the corresponding keyword group set, and obtain the cosine similarity mean of the phrase corresponding word vectors between different disaster situation description retained single sentences and the keyword group set. The greater the cosine similarity mean, the more serious the disaster situation of the corresponding earthquake area, and a deeper color is output for annotation on the visualization platform; it helps to understand the severity of the disaster situation in different earthquake areas.
[0121] In summary, the present invention forms suspected Internet buzzword groups from any adjacent phrases in each single sentence; obtains the updated keyword group set according to the phrase distribution of each suspected Internet buzzword group in the single sentence and the correlation between different phrases in the keyword group set; filters out the valid disaster situation description single sentences according to the phrase distribution of the updated keyword group set that appears in each single sentence and the phrase correlation between the corresponding single sentence and the updated keyword group set; obtains the disaster situation description retained single sentences of each earthquake area at the real-time moment according to the number of valid disaster situation description single sentences of each earthquake area at different moments. The present invention filters out the valid disaster situation description single sentences by analyzing the accurate performance of the disaster situation information during the earthquake, and improves the effectiveness of earthquake disaster situation data acquisition.
[0122] The present invention also provides a seismic disaster situation data acquisition system based on multi-source data, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any one of the seismic disaster situation data acquisition methods based on multi-source data are implemented.
[0123] It should be noted that the above-mentioned sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0124] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized.
Claims
1. A method for collecting seismic disaster situation data based on multi-source data, characterized in that The method includes: Obtaining the disaster situation information of different earthquake regions at each moment, obtaining the phrases of each single sentence in the disaster situation information, and a set of keyword phrases; For any moment, forming suspected Internet buzzword phrases from any adjacent phrases in each single sentence; screening out the habitual Internet buzzword phrases according to the phrase distribution of each suspected Internet buzzword phrase in the single sentence; updating the set of keyword phrases according to the correlation between the habitual Internet buzzword phrases and different phrases in the set of keyword phrases to obtain an updated set of keyword phrases; Obtaining the possibility of disaster situation - unrelated information for each single sentence according to the phrase distribution of the phrases in the updated set of keyword phrases that appear in each single sentence and the phrase correlation between the corresponding single sentence and the updated set of keyword phrases, and screening out the single sentences related to the disaster situation; for the single sentences related to the disaster situation corresponding to any earthquake region, screening out the effective disaster situation description single sentences according to the subject difference between each single sentence and the other single sentences and the possibility of disaster situation - unrelated information of the other single sentences; Obtaining the retained single sentences for the disaster situation description of each earthquake region at the real - time moment according to the number of effective disaster situation description single sentences of each earthquake region at different moments.
2. The method for collecting earthquake disaster situation data based on multi-source data according to claim 1, wherein The method for obtaining the habitual Internet buzzword phrases includes: For any moment, obtaining the habitual Internet buzzword possibility of each suspected Internet buzzword phrase according to its distribution in the single sentence; If the habitual Internet buzzword possibility of a suspected Internet buzzword phrase is greater than a preset possibility threshold, taking the corresponding suspected Internet buzzword phrase as a habitual Internet buzzword phrase.
3. A method for collecting earthquake disaster situation data based on multi-source data according to claim 2, characterized in that The method for obtaining the habitual Internet buzzword possibility includes: For any moment, obtaining the maximum value of the number of other characters within the range of all occurrences of each suspected Internet buzzword phrase in the single sentence, and the ratio between the number of characters of the single sentence where each occurrence is located as the local coefficient of the interval distance for each occurrence; obtaining the average value of the local coefficients of the interval distance for all occurrences of each suspected Internet buzzword phrase as the overall coefficient of the interval distance; Obtaining the ratio of the occurrence frequency of each suspected Internet buzzword phrase in all single sentences to the number of single sentences as the first possible coefficient; Performing a negative - correlation mapping on the overall coefficient of the interval distance, multiplying the result of the negative - correlation mapping by the first possible coefficient, and normalizing it as the habitual Internet buzzword possibility of each suspected Internet buzzword phrase.
4. A method for collecting earthquake disaster situation data based on multi-source data according to claim 1, characterized in that, The method for obtaining the updated set of keyword phrases includes: Obtaining the cosine similarity of the word vectors corresponding to different phrases between each habitual Internet buzzword phrase and the set of keyword phrases, and performing a negative - correlation normalization mapping as the first difference; Obtaining the average value of all the first differences between each habitual Internet buzzword phrase and the set of keyword phrases as the mean difference; If the mean difference of a habitual Internet buzzword phrase is less than a preset difference threshold, adding the corresponding habitual Internet buzzword phrase to the set of keyword phrases to form an updated set of keyword phrases.
5. A method for collecting earthquake disaster situation data based on multi-source data according to claim 1, characterized in that, The method for obtaining the possibility of disaster situation - unrelated information includes: Obtaining the ratio of the number of characters of each single sentence to the total number of characters of all the phrases in the updated set of keyword phrases that appear in the corresponding single sentence as the first possible coefficient; Obtain multiple triples for each single sentence, and obtain the average relative distance of the word vectors corresponding to different phrases between all triples and the updated set of keyword phrases as the average distance between each single sentence and the updated set of keyword phrases. According to the first possible coefficient, the number of occurrences of the phrases in the updated set of keyword phrases in each single sentence, and the average distance between each single sentence and the updated set of keyword phrases, obtain the possibility of disaster-unrelated information for each single sentence. Both the first possible coefficient and the average distance are positively correlated with the possibility of disaster-unrelated information, and the number of occurrences of the phrases is negatively correlated with the possibility of disaster-unrelated information.
6. A method for collecting earthquake disaster data based on multi-source data according to claim 1, characterized in that The method for obtaining the single sentences related to the disaster situation includes: Normalize the possibility of disaster-unrelated information for all single sentences. If the normalized result corresponding to a single sentence is less than or equal to the preset unrelated threshold, regard the corresponding single sentence as a single sentence related to the disaster situation.
7. A method for collecting seismic disaster situation data based on multi-source data according to claim 1, characterized in that, The method for obtaining the single sentences with effective disaster situation descriptions includes: For any single sentence related to the disaster situation corresponding to an earthquake area, obtain the degree of difference in disaster situation description between each single sentence and other single sentences according to the subject difference between each single sentence and the rest of the single sentences, as well as the possibility of disaster-unrelated information of other single sentences. Normalize the degree of difference in disaster situation description between all single sentences and other single sentences. If there is a single sentence whose normalized result with other single sentences is less than the preset difference threshold, regard the corresponding single sentence as a single sentence with an effective disaster situation description.
8. A method for collecting earthquake disaster situation data based on multi-source data according to claim 7, characterized in that, The method for obtaining the degree of difference in disaster situation description includes: Among all the triples of each single sentence, select the triples containing the phrases in the updated set of keyword phrases as effective triples. For any single sentence related to the disaster situation in an earthquake area, obtain the cosine similarity of the word vectors corresponding to different phrases between each effective triple of each single sentence and all the effective triples of each other single sentence, and accumulate the cosine similarities after negative correlation mapping as the local difference in description of each effective triple of each single sentence relative to each other single sentence. According to the local difference in description of different effective triples of each single sentence relative to all other single sentences, and the possibility of disaster-unrelated information of the corresponding other single sentences, obtain the degree of difference in disaster situation description between each single sentence and other single sentences. The local difference in description is positively correlated with the degree of difference in disaster situation description, and the possibility of disaster-unrelated information is negatively correlated with the degree of difference in disaster situation description.
9. A method for collecting seismic disaster situation data based on multi-source data according to claim 1, characterized in that, The method for obtaining the single sentences retained for disaster situation description includes: For any earthquake area, obtain the average value of the ratio of the number of single sentences with effective disaster situation descriptions at different times to the number of all single sentences as the effective proportion coefficient. Obtain the product of the number of single sentences with effective disaster situation descriptions at the real-time moment and the effective proportion coefficient, and round up as the acquisition quantity at the real-time moment. Select the smallest number of acquisition quantities of the possibility values of disaster-unrelated information among all single sentences with effective disaster situation descriptions as the single sentences retained for disaster situation description.
10. A seismic disaster situation data acquisition system based on multi-source data, the system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for collecting earthquake disaster situation data based on multi-source data according to any one of claims 1 to 9.
Citation Information
Patent Citations
Target event marking method and device, storage medium and electronic device
CN110458296A
Public opinion analysis and arrangement system based on AIGC
CN115982473A