A disaster metadata automatic matching method and system based on word2vec model

Through the automatic matching method of disaster metadata based on word2vec model, the problem of low efficiency of disaster metadata matching in the existing technology is solved, efficient and accurate disaster data processing is achieved, and disaster response and risk management capabilities are enhanced.

CN118114060BActive Publication Date: 2025-05-23ZHENGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410142877.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-05-23
Estimated Expiration
2044-02-01

AI Technical Summary

Technical Problem

In the matching process of disaster metadata, the existing technology has problems such as high computing resource requirements, huge model parameters and low efficiency, which is difficult to meet the needs of modern disaster management.

Method used

The automatic matching method of disaster metadata based on word2vec model is adopted, and the automatic matching of disaster metadata is achieved through steps such as metadata acquisition and preprocessing, model training, confidence threshold setting, metadata extraction and word vector conversion.

Benefits of technology

It improves the efficiency and accuracy of disaster data processing, enhances the capabilities of disaster response and risk management, and ensures the accuracy and efficiency of data matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118114060B_ABST
    Figure CN118114060B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of disaster prevention and reduction, and discloses a disaster metadata automatic matching method and system based on a word2vec model, which specifically includes the following steps: natural disaster metadata collection and preprocessing, using the CBOW architecture of the word2vec model, and model training on the LCQMC corpus; based on the training results, determining the mean cosine distance of metadata matching and metadata mismatching as a confidence threshold; metadata extraction, word segmentation, word vector conversion, and calculation of the cosine distance between word vectors. If the calculated cosine distance between word vectors is greater than the confidence threshold, it is determined that the metadata is matched, and the matching result is evaluated to ensure its integrity, consistency and accuracy. The disaster metadata automatic matching method proposed by the present invention makes the disaster metadata processing more efficient, thereby enhancing the timeliness and accuracy of disaster response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of disaster prevention and reduction, and relates to a disaster metadata automatic matching method and system based on a word2vec model. Background Art

[0002] Natural disasters, such as geological disasters such as earthquakes, collapses, landslides, mudslides, and meteorological disasters such as floods, have caused great damage to the ecological environment and human living environment due to their randomness and suddenness, and have caused great losses of life and property to the people. Disaster management, as an important social cause, is of great significance for protecting people's lives and property and promoting social stability and development. Metadata, as the core of key information such as the structure, characteristics and content of natural disaster data, is crucial to maximizing the use of disaster data. It not only constitutes the basis for effective management of disaster information, but also an indispensable basis for scientific decision-making and emergency response. However, a large amount of disaster metadata requires efficient extraction and matching, and traditional manual processing methods can no longer meet the needs of modern disaster management. Therefore, the research on disaster information management combined with deep learning has important theoretical and practical significance. In the existing technology, although models such as BERT, GPT, and RoBERTa have achieved remarkable success in capturing semantic information, they have problems such as high computing resource requirements, large model parameters, and low efficiency, which makes the ability of disaster response and risk management poor. The Word2Vec model has a relatively simple training process and an effective similarity measure. It performs well when processing large-scale corpora and has good generalization ability. At present, this model is often used in text similarity analysis, but there is still a gap in solving the problem of inconsistent data names and list names in metadata matching with the database. At the same time, when applying it, it is also necessary to consider the problem of massive multi-source structured metadata extraction and unstructured metadata extraction, but there is no relevant technical solution in the existing technology. Summary of the invention

[0003] In view of the above problems, the present invention proposes a disaster metadata automatic matching method based on the word2vec model, aiming to improve the efficiency and accuracy of disaster data processing, thereby enhancing the ability of disaster response and risk management.

[0004] In a first aspect, the present invention provides a disaster metadata automatic matching method based on a word2vec model, the method comprising:

[0005] Step 1: Metadata collection and preprocessing: Collect natural disaster metadata, divide the collected metadata into structured metadata and unstructured metadata, convert the unstructured metadata and perform part-of-speech tagging to form a text sequence x;

[0006] Step 2: Model training: Use the CBOW architecture of the word2vec model and train it on the LCQMC corpus;

[0007] Step 3: Set the confidence threshold: Based on the training results, determine the mean cosine distance between metadata matches and metadata mismatches as the confidence threshold;

[0008] Step 4: Metadata extraction: Use data parsing rules to extract structured metadata, and use NLP and UIE to extract unstructured metadata;

[0009] Step 5: Word segmentation: Use the jieba tool to segment the extracted metadata;

[0010] Step 6: Word vector conversion: convert the metadata name into the corresponding word vector;

[0011] Step 7: Calculate the cosine distance between word vectors: If the calculated cosine distance between word vectors is greater than the confidence threshold set in step 3, the metadata is determined to be a match.

[0012] Furthermore, the structured metadata includes data in shp, tif, xlsx / xls, and csv formats; and the unstructured metadata includes data in doc / docx formats.

[0013] Furthermore, the step of converting the unstructured metadata and then performing part-of-speech tagging specifically includes: using GIS software and python software to convert the unstructured metadata into text format data, and using doccano tool to perform part-of-speech tagging.

[0014] Furthermore, the CBOW architecture using the word2vec model is trained on the LCQMC corpus as follows: Given a fixed context window 2c, the objective function of CBOW is:

[0015]

[0016] in,

[0017]

[0018] Among them, v i is the vector representation of the input word, u 0 is the output word w 0 The corresponding vector, context window 2c, takes c words on the left and right of the current word as context.

[0019] Furthermore, the extraction of unstructured metadata using NLP and UIE is specifically as follows: UIE takes a given structural pattern guide s and a text sequence x as input to generate a linearized SEL(y)

[0020]

[0021] Where x=[x 1 , …, x |x| ] is a text sequence, s=[s 1 ,…,s |s| ] is the structural pattern indicator, y=[y 1 ,…,y |y| ] is a SEL sequence that can be converted into record information for metadata extraction.

[0022] Furthermore, the calculation formula of the cosine distance is as follows:

[0023]

[0024] Among them, A and B are two word vectors respectively.

[0025] Furthermore, after step 7, the method further includes: evaluating the matching results to ensure their completeness, consistency and accuracy.

[0026] Furthermore, the evaluating of the matching results to ensure their completeness, consistency and accuracy specifically includes: evaluating the matching results in the data table dimension and the data item dimension.

[0027] In a second aspect, the present invention provides a disaster metadata automatic matching system based on a word2vec model, characterized in that it includes:

[0028] Metadata collection and preprocessing module: collect natural disaster metadata, divide the collected metadata into structured metadata and unstructured metadata, convert the unstructured metadata and perform part-of-speech tagging to form a text sequence x;

[0029] Model training module: The CBOW architecture of the word2vec model is used for training on the LCQMC corpus. Based on the training results, the mean cosine distance between metadata matches and metadata mismatches is determined as the confidence threshold.

[0030] Metadata extraction module: extract structured metadata using data parsing rules, and extract unstructured metadata using NLP and UIE;

[0031] Word segmentation module: Use the jieba tool to segment the extracted metadata;

[0032] Word vector conversion module: converts the name of metadata into the corresponding word vector;

[0033] Judgment module: Calculate the cosine distance between word vectors: If the calculated cosine distance between word vectors is greater than the set confidence threshold, the metadata is determined to be matched;

[0034] Evaluation module: Quality assessment is performed on the extracted and matched metadata tables and metadata items from the three dimensions of consistency, completeness, and accuracy.

[0035] In a third aspect, the present invention provides a storage medium storing program instructions, wherein when the program instructions are executed, the device where the storage medium is located is controlled to execute any one of the above-mentioned disaster metadata automatic matching methods based on the word2vec model.

[0036] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0037] (1) Data parsing rules are used to extract the required information from structured metadata, and natural language processing technology NLP and information extraction framework UIE are used to process unstructured metadata to extract the required information from it; different data are processed in different ways to make data extraction more accurate, thereby improving the accuracy of data matching.

[0038] (2) By using the CBOW architecture for training on the LCQMC corpus, the elements of the well-defined metadata model are matched, and the automatic filling of disaster data metadata is finally achieved, making the model training more efficient.

[0039] (3) Evaluate the matching structure from the three aspects of consistency, completeness and accuracy to ensure the accuracy of data matching, thereby enhancing the ability to respond to disasters and manage risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0041] Figure 1 A flow chart of a disaster metadata automatic matching method based on a word2vec model provided in an embodiment of the present invention;

[0042] Figure 2 A flow chart of extracting structured metadata using data parsing rules provided in an embodiment of the present invention;

[0043] Figure 3 A technical roadmap for extracting unstructured metadata provided by an embodiment of the present invention;

[0044] Figure 4 The consistency, completeness and accuracy evaluation of the matching results provided by the embodiment of the present invention in the data table dimension;

[0045] Figure 5 The matching results provided by the embodiment of the present invention are evaluated for consistency, completeness and accuracy under the data item dimension. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. Obviously, the specific embodiments described herein are only used to explain the present invention and are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0047] Example 1

[0048] Embodiment 1 of the present invention provides a disaster metadata automatic matching method based on the word2vec model, such as Figure 1-3 As shown, the flow chart specifically includes:

[0049] Step 1: Metadata collection and preprocessing: Collect natural disaster metadata, divide the collected metadata into structured metadata (such as xlsx / xls, shp, tif, csv, etc.) and unstructured metadata (such as doc / docx, etc.), convert the unstructured metadata into text format data, and use the doccano tool to perform part-of-speech tagging on the text format data as the text sequence x.

[0050] Specifically, the data source selected by the present invention is the 2023 China summer flood disaster data jointly released by the National Earth Observation Science Data Center and the National Comprehensive Earth Observation Data Sharing Platform, which are responsible for the operation and maintenance of the China Remote Sensing Satellite Ground Station of the Institute of Space Information Innovation of the Chinese Academy of Sciences. This special service currently covers Hubei, Zhejiang, Jiangxi, Shandong, Heilongjiang, Henan, Gansu and other regions, covering typical flood disasters that occurred between 2013 and 2021. Including typical disaster-related areas such as Poyang Lake, Taihu Lake, Tangxun Lake, Longgan Lake and the Songhua River Basin, the disaster process includes pre-disaster, mid-disaster and post-disaster stages. The data set contains 40 disaster remote sensing data sets, involving observation data from 4 high-resolution satellites (Gaofen-1 series, Gaofen-2, Gaofen-3, Gaofen-4, Environmental-1) and 3 medium and low-resolution satellites (Sentinel-1, Sentinel-2, Landsat series). In addition, it also includes social media data on typhoon disaster events in China in 2021, and a dataset of future dam breach flood and inundation risks in Shanghai with a resolution of 50 meters covering the period from 2010 to 2100.

[0051] Specifically, the present invention uses GIS software and Python software for data preprocessing, converts unstructured metadata into text format data to facilitate data extraction, and uses doccano tools to perform part-of-speech tagging, which can improve the accuracy and semantic understanding of disaster metadata extraction.

[0052] Step 2: Model training: The CBOW architecture of the word2vec model is used and trained on the LCQMC corpus.

[0053] Specifically, given a fixed context window 2c, the objective function of CBOW can be written as:

[0054]

[0055] in,

[0056]

[0057] Among them, v i is the vector representation of the input word (obtained through the matrix V), u 0 is the output word w 0 The corresponding vector (obtained through matrix U), the context window 2c takes c words on the left and right of the current word as context.

[0058] Step 3: Determine the confidence threshold: Based on the training results, determine the mean cosine distance between matches and mismatches as the confidence threshold.

[0059] For example, based on the training results, the confidence threshold is set to 0.8 to ensure that high-confidence matching results are screened out in subsequent steps.

[0060] Step 4: Metadata extraction: Use data parsing rules to extract structured metadata, and use natural language processing technology NLP and information extraction framework UIE to extract unstructured metadata.

[0061] Furthermore, the structured metadata is data in the formats of shp, tif, xlsx / xls, csv, etc., and the structured metadata is extracted using data parsing rules such as Figure 2 As shown, specifically:

[0062] a.shp file metadata analysis

[0063] Use the PyShp library to read the shp file, obtain the geometry type information, the number of record data, and obtain the data range by calculating the minimum and maximum coordinate values ​​in the shp file; read the .prj file to obtain the coordinate system information; use the PyShp library to read the DBF file and obtain the field information, including field name, data type, field length, etc.; based on the xml.etree.ElementTree module, use the ET.parse() method to load the XML file and obtain the tags and attributes of the metadata, such as creation time, file directory, etc.

[0064] b.tif file metadata analysis

[0065] Use standard text file reading methods to read geographic coordinates and geographic transformation parameters from tfw files. Use the GDAL library to parse metadata in tif files, such as geometry type, and use the PIL library to read basic information of tif images and metadata that may exist in the tif file header. Use the XML parsing library (xml.etree.ElementTree) to parse the AUX.XML file and read geographic coordinates and projection information. Use the XML parsing library to obtain metadata tags and attributes, such as creation time, file directory, etc.

[0066] c.Excel file metadata analysis

[0067] Use the third-party library openpyxl in Python to read the basic information of the file. The first row of the Excel file usually contains the labels or headers of the columns. Read the metadata information of the first row of the Excel file as part of the column labels or headers, and then check these labels when parsing the file.

[0068] d.csv file metadata parsing

[0069] Check whether there are comments in the csv file. If there are comments, use regular expressions to parse the metadata in the comments. The first line of the csv file usually contains column labels or headers. The metadata information is used as part of the column labels or headers. These labels are checked when parsing the file.

[0070] Furthermore, the natural language processing technology (NLP) and information extraction framework (UIE) are used to extract unstructured metadata. Figure 3 As shown, it includes entity recognition, relationship extraction, and information extraction. Specifically, UIE takes the given structural pattern guide (s) and text sequence (x) as input to generate a linearized SEL (y)

[0071]

[0072] Where x=[x 1 , …, x |x| ] is a text sequence, s=[s 1 ,…,s |s| ] is the structural pattern indicator, y=[y 1 , …, y |y| ] is a SEL sequence that can be easily converted into an extracted information record.

[0073] Step 5: Word segmentation: Use the jieba tool to segment the text sequence (x).

[0074] Step 6: Word vector conversion: Convert the name of the data into the corresponding word vector.

[0075] Step 7: Calculate the cosine distance: By calculating the cosine distance between word vectors, if the cosine distance is greater than a preset confidence threshold, the metadata is determined to be matched.

[0076] Specifically, the formula for cosine distance is as follows:

[0077]

[0078] Among them, A and B are two word vectors respectively.

[0079] Example 2

[0080] Embodiment 2 of the present invention provides a disaster metadata automatic matching method based on the word2vec model, such as Figure 1-5 As shown, the flow chart specifically includes:

[0081] Step 1: Metadata collection and preprocessing: Collect natural disaster metadata, divide the collected metadata into structured metadata (such as xlsx / xls, shp, tif, csv, etc.) and unstructured metadata (such as doc / docx, etc.), convert the unstructured metadata into text format data, and use the doccano tool to perform part-of-speech tagging on the text format data as the text sequence x.

[0082] Specifically, the data source selected by the present invention is the 2023 China summer flood disaster data jointly released by the National Earth Observation Science Data Center and the National Comprehensive Earth Observation Data Sharing Platform, which are responsible for the operation and maintenance of the China Remote Sensing Satellite Ground Station of the Institute of Space Information Innovation of the Chinese Academy of Sciences. This special service currently covers Hubei, Zhejiang, Jiangxi, Shandong, Heilongjiang, Henan, Gansu and other regions, covering typical flood disasters that occurred between 2013 and 2021. Including typical disaster-related areas such as Poyang Lake, Taihu Lake, Tangxun Lake, Longgan Lake and the Songhua River Basin, the disaster process includes pre-disaster, mid-disaster and post-disaster stages. The data set contains 40 disaster remote sensing data sets, involving observation data from 4 high-resolution satellites (Gaofen-1 series, Gaofen-2, Gaofen-3, Gaofen-4, Environmental-1) and 3 medium and low-resolution satellites (Sentinel-1, Sentinel-2, Landsat series). In addition, it also includes social media data on typhoon disaster events in China in 2021, and a dataset of future dam breach flood and inundation risks in Shanghai with a resolution of 50 meters covering the period from 2010 to 2100.

[0083] Specifically, the present invention uses GIS software and Python software for data preprocessing, converts unstructured metadata into text format data to facilitate data extraction, and uses doccano tools to perform part-of-speech tagging, which can improve the accuracy and semantic understanding of disaster metadata extraction.

[0084] Step 2: Model training: The CBOW architecture of the word2vec model is used and trained on the LCQMC corpus.

[0085] Specifically, given a fixed context window 2c, the objective function of CBOW can be written as:

[0086]

[0087] in,

[0088]

[0089] Among them, v i is the vector representation of the input word (obtained through the matrix V), u 0 is the output word w 0The corresponding vector (obtained through matrix U), the context window 2c takes c words on the left and right of the current word as context.

[0090] Step 3: Determine the confidence threshold: Based on the training results, determine the mean cosine distance between matches and mismatches as the confidence threshold.

[0091] For example, based on the training results, the confidence threshold is set to 0.8 to ensure that high-confidence matching results are screened out in subsequent steps.

[0092] Step 4: Metadata extraction: Use data parsing rules to extract structured metadata, and use natural language processing technology (NLP) and information extraction framework (UIE) to extract unstructured metadata.

[0093] Furthermore, the structured metadata is data in the formats of shp, tif, xlsx / xls, csv, etc., and the structured metadata is extracted using data parsing rules such as Figure 2 As shown, specifically:

[0094] e.shp file metadata analysis

[0095] Use the PyShp library to read the shp file, obtain the geometry type information, the number of record data, and obtain the data range by calculating the minimum and maximum coordinate values ​​in the shp file; read the .prj file to obtain the coordinate system information; use the PyShp library to read the DBF file and obtain the field information, including field name, data type, field length, etc.; based on the xml.etree.ElementTree module, use the ET.parse() method to load the XML file and obtain the tags and attributes of the metadata, such as creation time, file directory, etc.

[0096] f.tif file metadata analysis

[0097] Use standard text file reading methods to read geographic coordinates and geographic transformation parameters from tfw files. Use the GDAL library to parse metadata in tif files, such as geometry type, and use the PIL library to read basic information of tif images and metadata that may exist in the tif file header. Use the XML parsing library (xml.etree.ElementTree) to parse the AUX.XML file and read geographic coordinates and projection information. Use the XML parsing library to obtain metadata tags and attributes, such as creation time, file directory, etc.

[0098] g.Excel file metadata analysis

[0099] Use the third-party library openpyxl in Python to read the basic information of the file. The first row of the Excel file usually contains the labels or headers of the columns. Read the metadata information of the first row of the Excel file as part of the column labels or headers, and then check these labels when parsing the file.

[0100] h.csv file metadata parsing

[0101] Check whether there are comments in the csv file. If there are comments, use regular expressions to parse the metadata in the comments. The first line of the csv file usually contains column labels or headers. The metadata information is used as part of the column labels or headers. These labels are checked when parsing the file.

[0102] Furthermore, the natural language processing technology (NLP) and information extraction framework (UIE) are used to extract unstructured metadata. Figure 3 As shown, it includes entity recognition, relationship extraction, and information extraction. Specifically, UIE takes the given structural pattern guide (s) and text sequence (x) as input to generate a linearized SEL (y)

[0103]

[0104] Where x=[x 1 , …, x |x| ] is a text sequence, s=[s 1 ,…,s |s| ] is the structural pattern indicator, y=[y 1 , …, y |y| ] is a SEL sequence that can be easily converted into an extracted information record.

[0105] Step 5: Word segmentation: Use the jieba tool to segment the text sequence (x).

[0106] Step 6: Word vector conversion: Convert the name of the data into the corresponding word vector.

[0107] Step 7: Calculate the cosine distance: By calculating the cosine distance between word vectors, if the cosine distance is greater than a preset confidence threshold, the metadata is determined to be matched.

[0108] Specifically, the formula for cosine distance is as follows:

[0109]

[0110] Among them, A and B are two word vectors respectively.

[0111] The present invention can obtain metadata extraction and matching results through steps 1 to 7. In order to further improve the accuracy of the metadata extraction and matching results, the present invention also includes an evaluation of the matching results obtained in step 7.

[0112] Step 8: Evaluate the matching results to ensure their completeness, consistency and accuracy.

[0113] Furthermore, the matching results are evaluated to ensure their completeness, consistency and accuracy. Specifically:

[0114] a. Integrity Assessment Algorithm

[0115] Given a relation R containing N tuples, the set of attributes on relation R is A = {A1, A2, ..., Am}, the primary key constraint is A1, the number of null values ​​is M1, the joint primary key constraint set is B = {B1, B2, ..., Bn}, the number of null values ​​is M2, the non-null constraint set is C = {C1, C2, ..., Ct}, the number of null values ​​is M3, where Cj, Bi∈A (i = 1, 2, ..., n; j = 1, 2, ..., t; t < m), and Cj is a single element set. At the same time, in all constraint rule metadata, violation of the non-null constraint rule will indicate a violation of data integrity. Therefore, the integrity measurement function F1 on relation R can be defined as:

[0116]

[0117] b. Consistency Assessment Algorithm

[0118] According to the consistency constraint rules such as foreign key constraint rules and function dependency constraint rules, the number of data that does not meet these rules is found, and then the consistency problem incidence rate is evaluated using the formula algorithm. Among all the constraint rule metadata, violation of the name, alias and dimension consistency constraint rules will indicate violation of data consistency. Therefore, the data consistency evaluation algorithm F2 is defined as:

[0119]

[0120] Among them, Scm represents the total number of data records; Scc represents the number of attribute columns; Sc1 represents the number of data that violates the equality consistency constraint rule; Sc2 represents the number of data that violates the existence consistency constraint rule; Sc3 represents the number of data that violates the logical consistency constraint rule; Sc4 represents the number of data that violates the foreign key constraint rule; Sc5 represents the number of data that violates the equality dependency constraint rule; Sc6 represents the number of data that violates the logical dependency constraint rule; Sc7 represents the number of data that violates the code constraint rule; Sid represents the number of empty data in the problem record; Scr represents the size of the problem record.

[0121] c. Accuracy Evaluation Algorithm

[0122] According to the accuracy constraint rules such as value range problem constraints and data type problem constraints, the number of data that do not meet these rules is found, and then the accuracy problem incidence rate is evaluated using the formula algorithm. Among all the constraint rule metadata, violations of length, precision, minimum value, maximum value and fixed value set constraint rules will indicate violations of data accuracy. Therefore, the data accuracy evaluation algorithm F3 is defined as:

[0123]

[0124] Among them, Sa1 represents the number of data that violates the value range constraint rules; Sa2 represents the number of data that violates the data type constraint rules; Sa3 represents the number of data that violates the data format constraint rules; Sa4 represents the number of data that violates the fixed value constraint rules; Sa5 represents the number of data problems that violate the precision constraint rules.

[0125] Furthermore, the present invention uses the above-mentioned integrity algorithm, consistency algorithm and accuracy algorithm to evaluate the matching results under the data table dimension. The evaluation results are as follows: Figure 4 As shown in the figure, the minimum consistency value is 0.58 of MD_limit, the maximum value is 1.0, and the mean value is 0.947; the minimum completeness value is 0.58 of MD_limit, the maximum value is 1.0, and the mean value is 0.908; the minimum accuracy value is 0.85 of MD_identification, the maximum value is 1.0, and the mean value is 0.970. The metadata table shows a high level of accuracy and is relatively stable; although there are some fluctuations in consistency and completeness, it shows good performance overall.

[0126] Furthermore, the present invention uses the above-mentioned integrity algorithm, consistency algorithm and accuracy algorithm to evaluate the matching results in the data item dimension. The evaluation results are as follows: Figure 5 As shown. Taking the data items in the metadata table as an example, in terms of consistency, the minimum value is 0.73 for character set information, the maximum value is 1.0, and the mean value is 0.942, indicating that the data items have relatively high consistency in character set information. In terms of completeness, the minimum value is 0.54 for restriction information, the maximum value is 1.0, and the mean value is 0.899. Although there are some fluctuations, the overall performance is good. In terms of accuracy, the minimum value is 0.97 for metadata reference information, extended information, and identification information, the maximum value is 1.0, and the mean value is 0.992, showing that the data items have very high accuracy in these aspects. These results show that the system can maintain a high level of consistency, completeness, and accuracy when processing specific data items, providing a solid foundation for the accurate extraction and matching of metadata.

[0127] Example 3

[0128] According to another aspect of an embodiment of the present invention, a disaster metadata automatic matching system based on a word2vec model is provided, comprising:

[0129] Metadata collection and preprocessing module: collect natural disaster metadata, divide the collected metadata into structured metadata and unstructured metadata, convert the unstructured metadata and perform part-of-speech tagging to form a text sequence x;

[0130] Model training module: The CBOW architecture of the word2vec model is used for training on the LCQMC corpus. Based on the training results, the mean cosine distance between metadata matches and metadata mismatches is determined as the confidence threshold.

[0131] Metadata extraction module: extract structured metadata using data parsing rules, and extract unstructured metadata using NLP and UIE;

[0132] Word segmentation module: Use the jieba tool to segment the extracted metadata;

[0133] Word vector conversion module: converts the name of metadata into the corresponding word vector;

[0134] Judgment module: Calculate the cosine distance between word vectors: If the calculated cosine distance between word vectors is greater than the set confidence threshold, the metadata is determined to be matched.

[0135] Furthermore, the system also includes an evaluation module for evaluating the matching results to ensure their completeness, consistency and accuracy.

[0136] Example 4

[0137] According to another aspect of an embodiment of the present invention, a storage medium is provided, on which program instructions are stored, wherein when the program instructions are executed, a device where the storage medium is located is controlled to execute any one of the above-mentioned disaster metadata automatic matching methods based on the word2vec model.

[0138] The above embodiments only express the preferred implementation mode of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A disaster metadata automatic matching method based on word2vec model, characterized in that: The following steps are involved: Step 1: Metadata collection and preprocessing: Collect natural disaster metadata, divide the collected metadata into structured metadata and unstructured metadata, convert the unstructured metadata and perform part-of-speech tagging to form a text sequence x; Step 2: Model training: Use the CBOW architecture of the word2vec model and train it on the LCQMC corpus; Step 3: Set the confidence threshold: Based on the training results, determine the mean cosine distance between metadata matches and metadata mismatches as the confidence threshold; Step 4: Metadata extraction: Use data parsing rules to extract structured metadata, and use NLP and UIE to extract unstructured metadata; Step 5: Word segmentation: Use the jieba tool to segment the extracted metadata; Step 6: Word vector conversion: convert the metadata name into the corresponding word vector; Step 7: Calculate the cosine distance between word vectors: If the calculated cosine distance between word vectors is greater than the confidence threshold set in step 3, the metadata is determined to be a match.

2. The disaster metadata automatic matching method based on the word2vec model according to claim 1 is characterized in that: The structured metadata includes data in shp, tif, xlsx / xls, and csv formats; the unstructured metadata includes data in doc / docx formats.

3. The disaster metadata automatic matching method based on the word2vec model according to claim 2 is characterized in that: The step of converting the unstructured metadata and then performing part-of-speech tagging specifically includes: using GIS software and python software to convert the unstructured metadata into text format data, and using doccano tool to perform part-of-speech tagging.

4. The disaster metadata automatic matching method based on the word2vec model according to claim 3 is characterized in that: The CBOW architecture using the word2vec model is trained on the LCQMC corpus as follows: Given a fixed context window 2c, the objective function of CBOW is: in, Among them, v i is the vector representation of the input word, u0 is the vector corresponding to the output word w0, and the context window 2c takes c words on the left and right of the current word as context.

5. The disaster metadata automatic matching method based on the word2vec model according to claim 4 is characterized in that: The extraction of unstructured metadata using NLP and UIE is specifically as follows: UIE takes a given structural pattern guide s and a text sequence x as input and generates a linearized SEL(y) Where x=[x1,…,x |x| ] is a text sequence, s=[s1,…,s |s| ] is the structural pattern indicator, y=[y1,…,y |y| ] is a SEL sequence that can be converted into record information for metadata extraction.

6. The disaster metadata automatic matching method based on the word2vec model according to claim 5 is characterized in that: The calculation formula of the cosine distance is as follows: Among them, A and B are two word vectors respectively.

7. The disaster metadata automatic matching method based on the word2vec model according to claim 6 is characterized in that: After step 7, the method further includes: evaluating the matching results to ensure their completeness, consistency and accuracy.

8. The disaster metadata automatic matching method based on the word2vec model according to claim 7 is characterized in that: The evaluation of the matching results to ensure their completeness, consistency and accuracy specifically includes: evaluating the matching results under the data table dimension.

9. The disaster metadata automatic matching method based on the word2vec model according to claim 7 is characterized in that: The evaluation of the matching results to ensure their completeness, consistency and accuracy specifically includes: evaluating the matching results under the data item dimension.

10. A disaster metadata automatic matching system based on word2vec model, characterized in that: include: Metadata collection and preprocessing module: collect natural disaster metadata, divide the collected metadata into structured metadata and unstructured metadata, convert the unstructured metadata and perform part-of-speech tagging to form a text sequence x; Model training module: uses the CBOW architecture of the word2vec model and trains on the LCQMC corpus; Based on the training results, the mean cosine distance between metadata matching and metadata mismatching is determined as the confidence threshold; Metadata extraction module: extract structured metadata using data parsing rules, and extract unstructured metadata using NLP and UIE; Word segmentation module: Use the jieba tool to segment the extracted metadata; Word vector conversion module: converts the name of metadata into the corresponding word vector; Judgment module: Calculate the cosine distance between word vectors: If the calculated cosine distance between word vectors is greater than the set confidence threshold, the metadata is determined to be matched; Evaluation module: Quality assessment is performed on the extracted and matched metadata tables and metadata items from the three dimensions of consistency, completeness, and accuracy.

Citation Information

Patent Citations

  • Similar case matching method based on semantic similarity

    CN110580281A

  • Geographic information service metadata text multi-level multi-label classification method

    CN110704624A