Database integration method, system and device for cross-domain collaboration
By marking the target entities and candidate associated words in the emergency management data and calculating the degree of fusion matching in different fields, the problem of data fusion conflict in emergency response is solved, the efficient integration and accuracy of multi-field data are achieved, and a stable emergency database is generated.
Patent Information
- Application Number
- CN202511106285.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-08-08
AI Technical Summary
In emergency response, matching and fusing data entities from different fields based on cosine similarity of feature semantics can easily lead to conflicts in fused data and data matching errors during event alignment, making it difficult to achieve efficient fusion of multi-field data.
By acquiring emergency management data from multiple fields, marking them as multiple target entities and determining candidate associated words, the degree of fusion matching between target entities in different fields is calculated based on the possible degree of association between the target entities and the candidate associated words, and data fusion processing is performed, and finally integrated into an emergency database.
It improves the accuracy of multi-domain data fusion, generates a more stable and reliable emergency management database, and supports real-time disaster emergency response.
Smart Images

Figure CN120610943B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a database integration method, system, and device for cross-domain collaboration. Background Art
[0002] Emergency response and intelligent decision-making support for sudden incidents require the support of a knowledge base. In the current big data context, the data in the field of emergency management is abundant, but emergency knowledge is relatively scarce. Without the condensation of emergency knowledge, it is difficult to make a rapid emergency response based on the actual disaster situation. Therefore, it is necessary to integrate and analyze multi-source heterogeneous data from different sources and different fields in each emergency processing link for actual decision-making.
[0003] The traditional data fusion method is to match and fuse data entities in different fields based on the cosine similarity of feature semantics. However, due to the inconsistent text description standards for emergency information in different links, multiple matching inconsistencies will occur during the data entity matching and fusion process, resulting in conflicts in the fused data during event alignment and data matching errors. Summary of the Invention
[0004] The main purpose of this application is to provide a database integration method, system and device for cross-domain collaboration, aiming to solve the technical problem in related technologies that matching and fusing data entities in different fields based on cosine similarity of feature semantics leads to conflicts in fused data during event alignment and data matching errors.
[0005] To achieve the above objectives, the present invention provides a database integration method for cross-domain collaboration, including:
[0006] Access emergency management data in multiple areas;
[0007] Marking the data belonging to emergency disaster events in the emergency management data as multiple target entities, and determining multiple candidate associated words corresponding to the target entities;
[0008] Based on the possible correlation between the target entity and the candidate associated words, the fusion matching degree between the target entities from different fields is calculated;
[0009] Based on the degree of fusion matching, the corresponding data of each target entity is fused to obtain fused data;
[0010] Emergency management data and fusion data are integrated and processed to obtain an emergency database.
[0011] In a possible implementation of the present application, the fusion matching degree between target entities from different fields is calculated based on the possible association degree between the target entities and the candidate association words, including:
[0012] For any field, the candidate association evaluation value between the candidate association word and the target entity in the current field is calculated based on the possible association degree between the target entity and the candidate association word;
[0013] Based on the candidate association evaluation value, the association degree value between the target entity and the candidate association word in different fields is calculated;
[0014] Based on the association degree value, the fusion matching degree between the target entities from different fields is calculated.
[0015] In a possible implementation of the present application, before the candidate association evaluation value between the candidate association word and the target entity in the current field is calculated based on the possible association degree between the target entity and the candidate association word, the method further includes:
[0016] The total number of first sentences in which the target entity exists in the emergency management data of the current field and the total number of second sentences in which the target entity and the candidate association word exist at the same time are determined;
[0017] Based on the total number of first sentences and the total number of second sentences, the possible association degree of each candidate association word for describing the target entity is calculated.
[0018] In a possible implementation of the present application, the candidate association evaluation value between the candidate association word and the target entity in the current field is calculated based on the possible association degree between the target entity and the candidate association word, including:
[0019] The text distance between the candidate association word and the target entity in the same sentence and the first number of times of occurrence of the candidate key word in the current sentence are determined;
[0020] Based on the text distance, the first number of times and the possible association degree, the candidate association evaluation value between the candidate association word and the target entity in the current field is calculated.
[0021] In a possible implementation of the present application, the association degree value between the target entity and the candidate association word in different fields is calculated based on the candidate association evaluation value, including:
[0022] The association comparison value of the candidate association evaluation value of the target entity corresponding to any two different fields is determined;
[0023] The association difference value between the candidate association evaluation value of the target entity of each field and the smallest candidate association evaluation value is calculated;
[0024] Based on the association comparison value and the association difference value, the association degree values between the target entities and the candidate association words in different fields are calculated.
[0025] In a possible implementation of the present application, the degree of fusion matching between target entities from different fields is calculated based on the association degree value, including:
[0026] Extract the word vector corresponding to each target entity;
[0027] Based on word vectors, calculate the semantic similarity between target entities;
[0028] Based on the semantic similarity and association degree values, the fusion matching degree between target entities is calculated.
[0029] In a possible implementation of the present application, the degree of fusion matching includes semantic similarity. Based on the degree of fusion matching, the data corresponding to each target entity is fused to obtain fused data, including:
[0030] Compare the semantic similarity between any two target entities with a preset similarity threshold;
[0031] If the comparison result shows that the semantic similarity is greater than the preset similarity threshold, the corresponding data of the two target entities are fused based on the degree of fusion matching until the data fusion of all target entities that meet the corresponding conditions of the preset similarity threshold is completed, and multiple fused data are obtained.
[0032] In a possible implementation of the present application, data belonging to emergency disaster events in the emergency management data are marked as multiple target entities, including:
[0033] Marking original data belonging to emergency disaster events in the emergency management data as a plurality of first entities;
[0034] Dividing the text data in the first entity into a plurality of vocabulary data;
[0035] Identify multiple attribute words representing attributes of disaster events in the vocabulary data, and mark each attribute word as an associated entity;
[0036] Based on the first entity and the associated entities, target entities are determined.
[0037] This application also provides a database integration system for cross-domain collaboration, which includes:
[0038] Acquisition module, which is used to acquire emergency management data in multiple fields;
[0039] The determining module is configured to mark data belonging to an emergency disaster event in the emergency management data as a plurality of target entities, and determine a plurality of candidate association words corresponding to the target entities;
[0040] The calculating module is configured to calculate a fusion matching degree between target entities from different fields based on a possible association degree between the target entities and the candidate association words;
[0041] The processing module is configured to perform fusion processing on data corresponding to each target entity based on the fusion matching degree, to obtain fusion data;
[0042] The integration processing module is configured to perform integration processing on the emergency management data and the fusion data, to obtain an emergency database.
[0043] The present application also provides a database integration device for cross-field collaboration, which is an entity node device. The database integration device for cross-field collaboration comprises a memory, a processor, and a program of a database integration method for cross-field collaboration stored in the memory and executable on the processor. When the program of the database integration method for cross-field collaboration is executed by the processor, the steps of the database integration method for cross-field collaboration can be implemented.
[0044] The present application provides a database integration method, system and device for cross-field collaboration. Compared with the related art, in which different field data entities are matched and fused based on feature semantics similarity, the present application can realize efficient fusion of multi-field data. In the present application, emergency management data of multiple fields are acquired, data belonging to an emergency disaster event in the emergency management data are marked as a plurality of target entities, candidate association words corresponding to the target entities are determined, the performance difference of the target entity context description relationship in different fields is quantified according to the possible association degree between the target entities and the candidate association words, the fusion matching degree between target entities from different fields is calculated, the data corresponding to each target entity is fused according to the fusion matching degree, to obtain fusion data, and the emergency management data and the fusion data are integrated, to obtain an emergency database. Thus, the matching degree of target entities in different fields is combined to correct the process of matching and fusing data between fields, thereby improving the accuracy of multi-field data fusion. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The flowchart of the first embodiment of the database integration method for cross-field collaboration of the present application is shown in the figure;
[0046] Figure 2 The device structure diagram of the hardware running environment involved in the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0047] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0048] The present application provides a database integration method for cross-domain collaboration. In the first embodiment of the database integration method for cross-domain collaboration, referring to Figure 1 , methods include:
[0049] Step S10, acquiring emergency management data in multiple fields;
[0050] Step S20: marking the data belonging to emergency disaster events in the emergency management data as multiple target entities, and determining multiple candidate associated words corresponding to the target entities;
[0051] Step S30, calculating the fusion matching degree between target entities from different fields based on the possible association degree between the target entity and the candidate association words;
[0052] Step S40: Based on the fusion matching degree, the corresponding data of each target entity is fused to obtain fused data;
[0053] Step S50: Integrate the emergency management data and the fusion data to obtain an emergency database.
[0054] The specific scenarios targeted by the embodiments of this application are:
[0055] The knowledge graph integrates emergency data from different fields by constructing a unified semantic model, and interconnects data from various fields in the emergency management process by creating a graph containing entities and relationships to form a structured emergency management database. In the process of integrating multi-field emergency data using knowledge graph technology, due to the lack of a unified narrative format between data in various fields, the representation of the same data entity in data from different sources may be very different. Therefore, in the process of data fusion, it is difficult to use expression features to extract and fuse knowledge. However, in different fields, the context information of similar entities also has descriptive similarities. Therefore, this application reconstructs the feature semantics of the entity by combining the context information to avoid mismatching caused by data conflicts during the integration process.
[0056] This embodiment aims to improve the accuracy of multi-domain data fusion, obtain a more stable and reliable emergency management database, and achieve emergency response to real-time disaster situations.
[0057] The specific steps are as follows:
[0058] Step S10, obtaining emergency management data in multiple fields.
[0059] As an example, a database integration method for cross-domain collaboration can be applied to a database integration device for cross-domain collaboration. The database integration device for cross-domain collaboration belongs to a database integration system for cross-domain collaboration, and the database integration system for cross-domain collaboration belongs to a database integration device for cross-domain collaboration.
[0060] As an example, emergency management data involves multiple fields, such as natural disasters, accidents and disasters, public health, and social security incidents. Data is collected from relevant data sources in these fields. The collected emergency management data includes data on monitoring, rescue and disposal, and recovery and reconstruction.
[0061] Specifically, emergency management data also includes disaster data, emergency resource data, and emergency management plans, as described in detail below:
[0062] Disaster data: text data describing the type, scale, impact range, etc. of disaster events.
[0063] Emergency resource data: When a disaster occurs, text data such as all emergency resources that need to be dispatched and confirmed, the disaster situation, and the number of people to be evacuated.
[0064] Emergency management plan: a series of disaster prevention and mitigation activities implemented during the disaster response process.
[0065] As an example, after obtaining emergency management data from various fields, due to the wide range of sources of emergency management data, these data may contain repeated, redundant, inconsistent or erroneous "interference data". The above collected data will be cleaned, including removing noise, filling missing values, and processing anomalies in the collected data to improve the quality of the collected data and provide strong data support for subsequent data analysis and fusion.
[0066] Specifically, duplicate data is removed by comparing related fields in data records; then missing values in the data summary are filled by calculating the mean; on this basis, the Z-score standardization method is used to standardize heterogeneous data in multiple fields, and data sets of different dimensions are converted into standard normal distributions to facilitate subsequent processing.
[0067] Step S20: Mark the data belonging to emergency disaster events in the emergency management data as multiple target entities, and determine multiple candidate associated words corresponding to the target entities.
[0068] As an example, when data integration is performed, the time series data of any emergency disaster time needs to be analyzed, the target entity is used to represent a specific object or concept with clear characteristics and attributes related to the event, such as the cause of the disaster event, the location of the event, the organization in the response process, the rescue team, etc. By identifying and defining these target entities, heterogeneous data from different fields can be effectively integrated. Since different fields may have different descriptions of the same disaster event, in order to provide structured and comparable representation features for subsequent entity matching and knowledge fusion process, it is necessary to first extract the associated feature entities from the collected text and construct the semantic features of the entities.
[0069] As an example, the extraction process of the candidate associated word can be: marking the words or phrases in the single-field text sentence other than the target entity as verbs and adjectives as candidate associated words. In this way, candidate associated words in multiple fields can be obtained.
[0070] In step S20 of marking the data belonging to the emergency disaster event in the emergency management data as a plurality of target entities, the following steps are included:
[0071] The original data belonging to the emergency disaster event in the emergency management data is marked as a plurality of first entities.
[0072] As an example, due to the particularity of the time attribute, when describing the disaster occurrence or end event, specific time description words are used. Therefore, the present application constructs knowledge graphs in different fields based on disaster events described under the same specific time.
[0073] As an example, in the process of marking entities, each disaster event is marked as an entity to obtain a plurality of entities belonging to the emergency disaster event in the emergency management data, and the time stamps of the emergency management data in different fields are aligned to determine the different field original data texts that may describe the same entity in the same time period.
[0074] The text data in the first entity is divided into a plurality of word data.
[0075] As an example, after dividing the plurality of first entities, the text data in each first entity is divided, the original data is segmented according to the punctuation marks (in this embodiment, the punctuation marks refer to period, exclamation mark, etc.), the string between any two punctuation marks is recorded as a complete sentence, and a plurality of independent sentences are obtained in the original data. Then, by using the accurate mode of the Chinese word segmentation tool, each independent sentence is cut into a plurality of words or phrases to obtain a plurality of word data.
[0076] A plurality of attribute words representing the attributes of the disaster event in the word data are identified, and each attribute word is marked as an associated entity.
[0077] As an example, the attribute vocabulary corresponding to the disaster event attribute can be the vocabulary or phrase of the semantics of "disaster type", "disaster location", "rescue team", "impact range", "rescue materials", etc. By using named entity recognition technology, these attribute vocabularies in the disaster event are identified, and these attribute vocabularies are marked as associated entities.
[0078] Based on the first entity and the associated entity, each target entity is determined.
[0079] As an example, all the first entities and associated entities collected are integrated into a target entity set: , obtaining a plurality of target entities.
[0080] Step S30, based on the possible association degree between the target entity and the candidate association word, the fusion matching degree between the target entities from different fields is calculated.
[0081] As an example, in the process of entity fusion of analyzing the text data of any emergency disaster event, the single entity word vector description may exist in different field information, which may cause expression ambiguity, resulting in fusion conflict in the alignment process using the similarity between the entity word vectors.
[0082] In actual unstructured original data text, the extracted entities are not isolated, and often have a connection with adjacent entities, sentences or paragraphs. In the process of different field entity fusion, in addition to the similar or identical semantics of the entity itself, there is also a description similarity between the context of the target entity in the field. By analyzing the connection between the target entity and the context information, the association words in a single field that are directly associated with the target entity are screened out.
[0083] As an example, after determining the candidate association words, the possible association degree between each candidate association word and the target entity is calculated, and then the association word set corresponding to each target entity in a single field is screened out according to the calculated possible association degree. The difference between the association evaluation of the association words in the cross-neighborhood association word set of the target entity is calculated, the difference between the context description relationship of the target entity in different fields is quantified, and the higher the description similarity between the two entities in the context is, the more likely the two entities represent the same disaster event. Therefore, the alignment and fusion process of the two entities is more accurate.
[0084] As an example, the fusion matching degree can be the matching degree of the entities in different fields. The greater the numerical value corresponding to the fusion matching degree, the stronger the association between the two, and the more likely it is to be used to describe the same emergency disaster event.
[0085] The step S30 of calculating the fusion matching degree between the target entities from different fields based on the possible association degree between the target entities and the candidate association words comprises:
[0086] The step S31 of calculating the candidate association evaluation value between the candidate association word and the target entity in the current field based on the possible association degree between the target entity and the candidate association word comprises:
[0087] As an example, the candidate association evaluation value is used to represent the association degree between the candidate association word and the target entity. The higher the candidate association evaluation value is, the stronger the association between the two is. The candidate association word set with strong association with the target entity is extracted by screening the candidate association word based on the candidate association evaluation value.
[0088] The step S31 of calculating the candidate association evaluation value between the candidate association word and the target entity in the current field based on the possible association degree between the target entity and the candidate association word further comprises:
[0089] determining the total number of first sentences in which the target entity exists in the emergency management data of the current field and the total number of second sentences in which the current target entity and the candidate association word exist at the same time;
[0090] As an example, before calculating the candidate association evaluation value, the possible association degree between the target entity and the candidate association word needs to be determined. The possible association degree is used to represent the possibility of the candidate association word for describing the target entity.
[0091] As an example, taking the xth candidate association word and the ith target entity as an example, the total number of first sentences represents the total number of sentences in which the current ith target entity exists in the data of the current field, and the total number of second sentences represents the total number of sentences in which the current ith target entity and the xth candidate association word exist at the same time in the field data.
[0092] Based on the total number of first sentences and the total number of second sentences, the possible association degree of each candidate association word for describing the target entity is calculated.
[0093] As an example, when the frequency of the candidate association word and the target entity appearing in the same sentence is higher, the possibility of the candidate association word for describing the target entity is higher. The possibility of each candidate association word for describing the target entity is calculated, and the possible association degree The calculation method of the possible association degree can be:
[0094]
[0095] wherein, represents the possible association degree of the xth candidate key word for describing the ith target entity; Indicates the total number of first sentences in the domain data where the current i-th target entity exists; Indicates the total number of second sentences in the domain data that contain both the current i-th target entity and the x-th candidate associated word.
[0096] As an example, by calculating the frequency of the target entity and each candidate associated word appearing together in a single domain text, the possibility of the current candidate associated word being used to describe the target entity is reflected.
[0097] The step S31 of calculating the candidate association evaluation value between the candidate association word and the target entity in the current domain based on the possible association degree between the target entity and the candidate association word includes:
[0098] Determine the textual distance between the candidate associated words and the target entity in the same sentence and the number of first occurrences of the candidate keyword in the current sentence;
[0099] As an example, the number of characters between the candidate associated words and the target entity in the same sentence is counted as the text distance between the candidate associated words and the target entity. , Represents the text distance between the xth candidate associated word and the ith target entity in the ath sentence.
[0100] As an example, the first number represents the number of times the candidate keyword appears in the current sentence.
[0101] Based on the text distance, the first number and the possible association degree, the candidate association evaluation value between the candidate association word and the target entity in the current field is calculated.
[0102] As an example, in a single field, the closer the text distance between the candidate associated words and the target entity, the greater the possibility that the candidate associated words are used to describe the current entity text. At the same time, combined with the possible degree of association between the target entity and each candidate associated word, the candidate association evaluation value between the candidate associated words and the target entity in the current field is calculated.
[0103] As an example, the candidate association evaluation value The calculation method can be:
[0104]
[0105] in, represents the candidate association evaluation value between the xth candidate keyword and the i-th target entity, Indicates the The number of times the xth candidate conjunction appears in a sentence, that is, the number of first occurrences; Indicates the possible relevance of the x-th candidate keyword to describe the i-th target entity. represents the text distance between the xth candidate association word and the ith target entity, represents the total number of all sentences containing the xth candidate association word and the ith target entity in the field data.
[0106] As an example, after calculating the candidate association evaluation value, the candidate association evaluation value is used to filter the key words from the current field candidate association words, and the association word set is constructed , wherein, represents the association word set corresponding to the ith target entity in the sth field, and the candidate association evaluation value of the corresponding candidate association word and the target entity is marked in the set. When constructing the association word set, the candidate association evaluation values of the target entity and each candidate key word are sorted from high to low, and the candidate key words corresponding to the first preset number of candidate association evaluation values are selected. In addition, the number of key words in the association word set of each target entity can be set to a fixed number (for example, 20). When the number of association words in all association word sets is not equal, the minimum value of the number of association words in each association word set is set as the number of key words in each set, and the selection of the association word set of each target entity is completed.
[0107] At this point, the association word set of the ith target entity in a single field is obtained .
[0108] Step S32, based on the candidate association evaluation value, the association degree value between the target entity and the candidate association word in different fields is calculated.
[0109] As an example, based on the candidate association evaluation value, the association degree value between the target entity and the filtered association word set in different fields is calculated.
[0110] As an example, according to the characteristics of the emergency management field, the attributes of disaster events mainly include: disaster-causing factors, disaster attributes, spatial attributes, emergency task attributes, etc. For entities with the same attributes, similar common features are used in the process of describing the entities, so that the rule matching of the entities can be performed according to the common features of the same type of attributes. For the association word set filtered according to the target entity in different fields, when the description similarity between the verbs, adjectives and other words in the context of the two entities across fields is higher, it means that the possibility of the two entities representing the same disaster event is greater. Therefore, after filtering the association word set corresponding to the target entity in a single field in the above step, the association degree value between the target entity and the same association word in different fields also needs to be calculated.
[0111] Wherein, the step S32 of calculating the association degree value between the target entity and the candidate association word in different fields based on the candidate association evaluation value, comprises:
[0112] determining a correlation contrast value of a candidate correlation evaluation value between any two different domains;
[0113] calculating a correlation difference value between the candidate correlation evaluation value and the minimum candidate correlation evaluation value of each domain target entity;
[0114] Based on the correlation contrast value and the correlation difference value, the correlation degree value between the target entity and the candidate correlation word in different domains is calculated.
[0115] As an example, when calculating the correlation degree value, the correlation degree value between the target entity in different domains and the correlation word in the correlation word set is calculated, and the correlation word belongs to part of the candidate correlation word. The parameters calculated using the correlation word will be more accurate.
[0116] As an example, the correlation contrast value between the candidate correlation evaluation values of two different domains is represented as , wherein represents the candidate correlation evaluation value between the ith target entity and the xth correlation word in the sth domain; represents the candidate correlation evaluation value between the ith target entity and the xth correlation word in the tth domain, and the tth domain represents any one of all other domains except the sth domain, wherein the ith target entity in different domains refers to the same entity, for example, fire is considered as a target entity, and in the whole emergency management process, multiple domains such as news reporting, emergency management scheme, post-disaster reconstruction and the like are involved. There are multiple texts about the target entity of fire description in these domains, which can be determined as the same target entity in different domains.
[0117] As an example, the correlation difference value is represented as , wherein represents the minimum candidate correlation evaluation value between the ith target entity and the xth correlation word in several domains.
[0118] As an example, the correlation degree value The calculation method can be:
[0119]
[0120] , wherein represents the correlation degree value between the ith target entity and the correlation word set in the sth domain; represents the correlation evaluation value between the ith target entity and the xth correlation word in the sth domain; represents the correlation evaluation value between the ith target entity and the xth correlation word in the tth domain; represents the minimum association evaluation value between the ith target entity and the xth association word in the s th field; nx represents the number of association words contained in the set. By calculating the difference between the minimum association evaluation between the target entity and the association word in several fields, the influence of the high association word related to the target entity on the association degree of the target entity is strengthened. T-1 represents the number of other fields except the s th field, T represents the total number of fields, and e represents a natural constant, represents a normalization function, such as the hyperbolic tangent function tanh.
[0121] Step S33, based on the association degree value, the fusion matching degree between the target entities from different fields is calculated.
[0122] As an example, the difference between the association evaluation values between the association words in the cross-neighborhood association word set of the target entity is calculated, the difference between the performance of the context description relationship of each target entity in different fields is quantified, and then the performance difference of the entity association word is combined to determine the matching degree between each target entity in the cross-field fusion process of the current target entity.
[0123] Among them, the step S33 of calculating the fusion matching degree between the target entities from different fields based on the association degree value includes:
[0124] Extracting the word vector corresponding to each target entity;
[0125] Based on the word vector, the semantic similarity between each target entity is calculated;
[0126] As an example, the semantic similarity between two target entities The way can be:
[0127]
[0128] Among them, , represents the word vector of the ith target entity in the given s th field and the word vector of the j th target entity in the t th field. represents the length of the two word vectors. Since the numerical range of is -1 to 1, the norm normalization method here can be: .
[0129] Based on the semantic similarity and the association degree value, the fusion matching degree between each target entity is calculated.
[0130] As an example, for semantic similarity, a preset similarity threshold of 0.75 is set, and data sources / target entities whose semantic similarity between target entities is greater than or equal to the preset similarity threshold are fused. If the semantic similarity between two target entities is less than the preset similarity threshold, they will not participate in the fusion process.
[0131] As an example, calculate the degree of fusion matching The way can be:
[0132]
[0133] in, represents the semantic similarity between the i-th target entity and the j-th target entity, The Jaccard coefficient (a statistical indicator used to measure the similarity between two sets) represents the set of associated words between the i-th target entity in the s-th field and the j-th target entity in the t-th field. The larger the value, the greater the similarity between the two target entity associated word sets. It represents the performance difference between the association degree value of the i-th target entity in the s-th field and the association degree value between the j-th target entity and the associated word set in the t-th field, that is, the performance difference between the association degree between the target entity and the associated word set in different fields; when the performance difference between the association degree between the associated word set and the target entity in different fields is smaller, it means that the connection between the currently determined associated word and the entity is closer. This is to avoid the denominator being 0 while ensuring that the denominator logic remains unchanged.
[0134] Step S40: Based on the fusion matching degree, the corresponding data of each target entity is fused to obtain fused data.
[0135] The step S40 of fusing the corresponding data of each target entity based on the degree of fusion matching to obtain fused data includes:
[0136] Step S41, comparing the semantic similarity between any two target entities with a preset similarity threshold;
[0137] As an example, after calculating the semantic similarity between any two target entities, the semantic similarity between the two target entities is compared with a preset similarity threshold to determine the degree of similarity between the two entities.
[0138] As an example, the preset similarity threshold may be 0.6, 0.75, etc., which is not specifically limited.
[0139] In step S42, if the comparison result shows that the semantic similarity is greater than the preset similarity threshold, the two target entity corresponding data are fused based on the fusion matching degree until the data of all target entities that meet the corresponding conditions of the preset similarity threshold are fused to obtain multiple fused data.
[0140] As an example, when the comparison result shows that the semantic similarity is greater than the preset similarity threshold, the two target entities are fused according to the weight corresponding to the fusion matching degree. The target entity with a high fusion matching degree is assigned a relatively high weight, and the target entity with a relatively low fusion matching degree is assigned a low weight value, until the data fusion of all target entities that meet the conditions corresponding to the preset similarity threshold is completed, and multiple fused data are obtained.
[0141] Step S50: Integrate the emergency management data and the fusion data to obtain an emergency database.
[0142] As an example, after obtaining the fused data, a new database is first generated based on the fused data, and then the collected emergency management data is imported into the new database through data mapping to obtain a complete emergency database.
[0143] As an example, due to the continuous updating of data, the constructed database needs to be continuously updated and maintained. For the new original data that appears in the field in real time, the above steps are used to generate new sorted data. At the same time, Cypher statements are used to perform corresponding modification operations on the nodes and edges in the graph through local updates to complete the real-time update of the database integration process.
[0144] The present application provides a database integration method for cross-domain collaboration. Compared with the related art, which cannot achieve efficient fusion of multi-domain data by matching and fusing data entities in different domains based on the similarity of feature semantics, in the present application, emergency management data in multiple domains are obtained, the data belonging to emergency disaster events in the emergency management data are marked as multiple target entities, and candidate associated words corresponding to the target entities are determined. According to the possible degree of association between the target entity and the candidate associated words, the performance difference of the context description relationship of the target entity in different domains is quantified, thereby calculating the degree of fusion matching between target entities from different domains, and then according to the degree of fusion matching, the data corresponding to each target entity is fused to obtain fused data, and then the emergency management data and the fused data are integrated to obtain an emergency database. Therefore, by combining the matching degree of target entities in different domains, the process of matching and fusing data between each domain is corrected, thereby achieving efficient fusion of multi-domain data.
[0145] The embodiment of the application further provides a database integration system for cross-domain collaboration, which comprises:
[0146] An acquisition module is configured to acquire emergency management data of multiple domains.
[0147] A determination module is configured to mark data belonging to emergency disaster events in the emergency management data as target entities, and determine candidate associated words corresponding to the target entities.
[0148] A calculation module is configured to calculate fusion matching degrees between the target entities from different domains based on possible association degrees between the target entities and the candidate associated words.
[0149] A processing module is configured to perform fusion processing on data corresponding to each target entity based on the fusion matching degrees, to obtain fusion data.
[0150] An integration processing module is configured to perform integration processing on the emergency management data and the fusion data, to obtain an emergency database.
[0151] Referring to Figure 2 , Figure 2 is a device structure diagram of a hardware running environment involved in the embodiment of the application.
[0152] As Figure 2 shown, the database integration device for cross-domain collaboration can comprise a processor 1001, a memory 1005, and a communication bus 1002. The communication bus 1002 is configured to realize connection and communication between the processor 1001 and the memory 1005.
[0153] Optionally, the database integration device for cross-domain collaboration can further comprise a user interface, a network interface, a camera, an RF (Radio Frequency, radio frequency) circuit, a sensor, a WiFi module, and the like. The user interface can comprise a display screen (Display), an input sub-module such as a keyboard (Keyboard), and the optional user interface can further comprise a standard wired interface, a wireless interface. The network interface can comprise a standard wired interface, a wireless interface (such as a WI-FI interface).
[0154] Those skilled in the art can understand that Figure 2 the database integration device structure for cross-domain collaboration shown in the embodiment of the application does not constitute a limitation on the database integration device for cross-domain collaboration, and can comprise more or fewer components than the diagram, or combine certain components, or different component arrangements.
[0155] As Figure 2As shown, memory 1005, which serves as a storage medium, may include an operating system, a network communication module, and a database integration program for cross-domain collaboration. The operating system is a program that manages and controls the hardware and software resources of the database integration device for cross-domain collaboration, supporting the operation of the database integration program for cross-domain collaboration and other software and / or programs. The network communication module is used to enable communication between the various components within memory 1005, as well as communication with other hardware and software in the database integration system for cross-domain collaboration.
[0156] exist Figure 2 In the cross-domain collaboration-oriented database integration device shown, the processor 1001 is configured to execute the cross-domain collaboration-oriented database integration program stored in the memory 1005 to implement any of the steps of the cross-domain collaboration-oriented database integration method described above.
[0157] The specific implementation of the database integration device for cross-domain collaboration in this application is basically the same as the various embodiments of the database integration method for cross-domain collaboration described above, and will not be repeated here.
[0158] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0159] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0160] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as mentioned above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of each embodiment of the present application.
[0161] The above are only preferred embodiments of the present application and do not limit the scope of application of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application description and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of protection of the present application.
[0162] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0163] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A database integration method for cross-domain collaboration, characterized in that: The method comprises: Access emergency management data in multiple areas; Marking the data belonging to emergency disaster events in the emergency management data as multiple target entities, and determining multiple candidate associated words corresponding to the target entities; Determine the total number of first sentences containing the target entity and the total number of second sentences containing both the current target entity and the candidate associated words in the emergency management data of the current domain; Based on the total number of the first sentences and the total number of the second sentences, calculating the possible association degree of each candidate associated word for describing the target entity; The calculation method of the possible association degree is: in, Indicates the possible degree of association of the xth candidate association word used to describe the i-th target entity; Indicates the total number of first sentences in the domain data where the current i-th target entity exists; Indicates the total number of second sentences in the domain data that contain both the current i-th target entity and the x-th candidate associated word; For any field, determining the text distance between the candidate associated word and the target entity in the same sentence and the number of first occurrences of the candidate associated word in the current sentence; Calculating a candidate association evaluation value between the candidate associated word and the target entity in the current domain based on the text distance, the first number, and the possible association degree; The candidate association evaluation value is calculated as follows: in, represents the candidate association evaluation value between the xth candidate association word and the i-th target entity, Indicates the The number of first occurrences of the x-th candidate conjunction in a sentence; Indicates the possible association degree of the xth candidate association word used to describe the i-th target entity, represents the text distance between the xth candidate associated word and the i-th target entity, represents the total number of all sentences in the domain data that contain the xth candidate associated word and the i-th target entity; Determine the correlation comparison value of the target entities between any two different fields corresponding to the candidate correlation evaluation values; Calculate the correlation difference between the candidate correlation evaluation value corresponding to the target entity in each field and the minimum candidate correlation evaluation value; Based on the association comparison value and the association difference value, calculating the association degree value between the target entity in different fields and the candidate association word; The calculation method of the correlation degree value is as follows: in, Indicates the association degree between the i-th target entity and the associated word set in the s-th domain; represents the association evaluation value between the i-th target entity and the x-th associated word in the s-th domain; represents the association evaluation value between the i-th target entity and the x-th associated word in the t-th domain; represents the minimum association evaluation value between the i-th target entity and the x-th associated word in several fields; nx represents the number of associated words contained in the set; By calculating the correlation difference of the minimum correlation evaluation between the target entity and the associated words in several fields, represents the correlation comparison value between candidate correlation evaluation values in different fields; T-1 represents the number of fields other than the sth field, then T represents the total number of fields, e represents a natural constant, represents the normalization function; Extracting word vectors corresponding to each target entity; Calculating the semantic similarity between the target entities based on the word vectors; Calculating the degree of fusion matching between the target entities based on the semantic similarity and the association degree value; The calculation method of the fusion matching degree is as follows: in, represents the semantic similarity between the i-th target entity and the j-th target entity, represents the Jaccard coefficient of the set of associated words between the i-th target entity in the s-th field and the j-th target entity in the t-th field, It represents the difference between the association degree value of the i-th target entity in the s-th domain and the association degree value between the j-th target entity and the associated word set in the t-th domain; Based on the fusion matching degree, the corresponding data of each target entity is fused to obtain fused data; The emergency management data and the fusion data are integrated and processed to obtain an emergency database.
2. The database integration method for cross-domain collaboration according to claim 1, characterized in that: The step of fusing the corresponding data of each target entity based on the fusion matching degree to obtain fused data includes: Comparing the semantic similarity between any two target entities with a preset similarity threshold; If the comparison result shows that the semantic similarity is greater than the preset similarity threshold, the data corresponding to the two target entities are fused based on the fusion matching degree until the data fusion of all target entities that meet the corresponding conditions of the preset similarity threshold is completed, and multiple fused data are obtained.
3. The database integration method for cross-domain collaboration according to claim 1, characterized in that: The step of marking the data belonging to the emergency disaster event in the emergency management data as a plurality of target entities includes: Marking original data belonging to emergency disaster events in the emergency management data as a plurality of first entities; Dividing the text data in the first entity into a plurality of vocabulary data; Identifying a plurality of attribute words representing attributes of a disaster event in the vocabulary data, and marking each of the attribute words as an associated entity; Based on the first entity and the associated entities, target entities are determined.
4. A database integration system for cross-domain collaboration, characterized by: The database integration method for cross-domain collaboration according to any one of claims 1 to 3, wherein the database integration system for cross-domain collaboration comprises: An acquisition module, the acquisition module is used to acquire emergency management data in multiple fields; a determination module, configured to mark data belonging to emergency disaster events in the emergency management data as a plurality of target entities, and determine a plurality of candidate associated words corresponding to the target entities; a calculation module configured to calculate a degree of fusion matching between target entities from different fields based on a possible degree of association between the target entity and the candidate associated words; A processing module, configured to perform fusion processing on the data corresponding to each target entity based on the fusion matching degree to obtain fused data; An integration processing module is used to integrate the emergency management data and the fusion data to obtain an emergency database.
5. A database integration device for cross-domain collaboration, characterized in that: The device includes: a memory, a processor, and a database integration program for cross-domain collaboration stored in the memory and executable on the processor, wherein the database integration program for cross-domain collaboration is configured to implement the steps of the database integration method for cross-domain collaboration as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Fusion method of household registration population data and real estate registration data
CN118277953A
Knowledge graph data fusion
US20240144032A1