A multi-source heterogeneous data warehousing service system for regional seismic safety evaluation

By building a multi-source heterogeneous data storage service system, the problems of data redundancy and repeated database entry in regional earthquake safety evaluation are solved, unified encoding and redundant identification of data are realized, and standardization and query efficiency of data storage are improved.

CN119917523BActive Publication Date: 2025-07-11HEBEI EARTHQUAKE ADMINISTRATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510404899.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-11
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

In the prior art, the management of multi-source heterogeneous data in regional seismic safety evaluation lacks a unified labeling system, structured matrix construction and dynamic identification mechanism, resulting in repeated data entry into the database, invalid information accumulation and resource waste, making it difficult to meet the needs of systematic and standardized data processing.

Method used

Build a multi-source heterogeneous data storage service system that includes evaluation behavior labels. Through triple determination modules, advanced matrix construction modules, evaluation label determination modules and database matrix construction modules, unified encoding representation and redundancy judgment of multi-source heterogeneous data are realized, and the three-element advanced matrix is formed to dynamically identify and eliminate redundant data.

Benefits of technology

It significantly improves the adaptability and accuracy of standardized in-store intake of multi-source heterogeneous data, dynamically identify and eliminate redundant data, improves data query and call efficiency, and supports structured feature extraction and precise query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917523B_ABST
    Figure CN119917523B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-source heterogeneous data warehousing service system for regional seismic safety evaluation, including: a data set acquisition module, a triple determination module, an advanced matrix construction module for constructing a three-element advanced matrix according to N double-label data triples; an evaluation label determination module for determining an evaluation behavior label from the three-element advanced matrix; wherein, the evaluation behavior label is a binary label used to mark the evaluation effectiveness of adjacent data pairs, and its values are: "reasonable evaluation" and "unnecessary evaluation"; a warehousing matrix construction module for constructing a multi-source heterogeneous data warehousing matrix according to the evaluation behavior label and the three-element advanced matrix; a server generation module for storing the multi-source heterogeneous data warehousing matrix in a server to generate a multi-source data server; the present invention can automatically judge the behavior of seismic safety evaluation result data with redundancy and effectively control the inflow of invalid information into the warehouse.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data warehousing service system, and more particularly to a multi-source heterogeneous data warehousing service system for regional seismic safety evaluation. Background Art

[0002] Currently, the data types involved in regional seismic safety evaluation work exhibit significant multi-source and heterogeneous characteristics, including but not limited to data such as geological surveys, geophysical surveys, drilling analyses, seismic activity assessments, and multi-probability ground motion parameters. These data vary greatly in terms of source, structure, expression form, and storage format. Typically, various types of data coexist, such as geospatial data (points, lines, surfaces), tabular data, map data, and raw data files, and they are usually managed using different carriers such as proprietary databases or file systems like ArcGIS and MapInfo. This results in inconsistent data structures, chaotic field naming, high information duplication rates, and great difficulty in structure comparison, severely restricting the subsequent data warehousing, structure integration, analysis and evaluation, and query and invocation efficiency.

[0003] The patent document with the patent publication number CN115964360B discloses a method and system for constructing a seismic safety evaluation database, which can obtain a classification result for representing the seismic safety level label of the area to be evaluated. However, in the prior art, the management method for earthquake-related multi-source data lacks dynamic identification and organization means for data structure differences and time series relationships, and also lacks a quantitative judgment mechanism for data redundancy and evaluation necessity, resulting in problems such as duplicate data warehousing, accumulation of invalid information, and waste of evaluation resources. In addition, there is currently a lack of a complete service method that can unify the label system, construct a structured matrix, and support behavior recognition and warehousing decision-making, making it difficult to meet the technical requirements for systematic, standardized, and automated processing of data in seismic safety evaluation. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides a multi-source heterogeneous data warehousing service system for regional seismic safety evaluation, which solves the technical problems raised in the background art by constructing a multi-source heterogeneous data warehousing matrix containing evaluation behavior labels.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0006] A multi-source heterogeneous data warehousing service system for regional seismic safety evaluation, comprising:

[0007] A data set acquisition module for acquiring a data set composed of N multi-source heterogeneous data;

[0008] A triple determination module for determining N double-label data triples according to the data set;

[0009] An advanced matrix construction module for constructing a three - element advanced matrix according to N double - label data triples;

[0010] An evaluation label determination module for determining an evaluation behavior label from the three - element advanced matrix; wherein, the evaluation behavior label is a binary label used to mark the evaluation validity of adjacent data pairs, and its values are: "reasonable evaluation" and "unnecessary evaluation";

[0011] An incoming - warehouse matrix construction module for constructing a multi - source heterogeneous data incoming - warehouse matrix according to the evaluation behavior label and the three - element advanced matrix;

[0012] A server generation module for storing the multi - source heterogeneous data incoming - warehouse matrix in a server to generate a multi - source data server;

[0013] When a client issues a data query request, the multi - source data server responds to the data query request to generate a query vector;

[0014] The multi - source data server receives the query vector as an input, extracts features from the multi - source heterogeneous data incoming - warehouse matrix, and outputs data features based on the data query request.

[0015] In some embodiments, the triple determination module is specifically used for:

[0016] S2 - 1. Determine the stage - continuous label and the structure - discrete label of the multi - source heterogeneous data from the dataset; wherein, the stage - continuous label is used to characterize the acquisition stage of the multi - source heterogeneous data in the time dimension, and its minimum value is defined as the origin label;

[0017] S2 - 2. Pair the multi - source heterogeneous data with its corresponding stage - continuous label and structure - discrete label to construct N double - label data triples.

[0018] In some embodiments, determining the stage - continuous label and the structure - discrete label of the multi - source heterogeneous data from the dataset includes:

[0019] S2 - 1 - 1. Extract the marked acquisition timestamps of N multi - source heterogeneous data in the dataset;

[0020] S2 - 2 - 2. Assign stage - continuous labels to N multi - source heterogeneous data according to the chronological order of the acquisition timestamps; wherein, the stage - continuous label is used to characterize the acquisition stage of the multi - source heterogeneous data in the time dimension, and its minimum value is defined as the origin label, corresponding to the multi - source heterogeneous data with the earliest acquisition time;

[0021] S2-2-3. Identify the data structures of N multi-source heterogeneous data to generate structural discrete labels, where the structural discrete labels are used to characterize the type attribution of multi-source heterogeneous data at the structural level;

[0022] In some of these embodiments, the advanced matrix construction module is specifically configured to:

[0023] S3-1. Construct a matrix coordinate system in a two-dimensional plane, where the origin of the matrix coordinate system is characterized as the origin label, and the coordinate Y-axis is characterized as the virtual connection axis of the phase continuous label;

[0024] S3-2. Perform matrix arrangement on N triple-label data triples in the matrix coordinate system to construct a structure element matrix and a phase element matrix;

[0025] S3-3. Extract an ordered sequence of N multi-source heterogeneous data from the structure element matrix;

[0026] S3-4. Construct a three-element advanced matrix according to the ordered sequence of the multi-source heterogeneous data and the phase element matrix.

[0027] In some of these embodiments, performing matrix arrangement on N triple-label data triples in the matrix coordinate system to construct a structure element matrix and a phase element matrix includes:

[0028] S3-2-1. Extract the phase continuous label and its corresponding multi-source heterogeneous data from N triple-label data triples;

[0029] S3-2-2. Fill the extracted phase continuous label and its corresponding multi-source heterogeneous data into the first quadrant of the matrix coordinate system, where the origin label is located on the X-axis, and the phase continuous label and its corresponding multi-source heterogeneous data are arranged equidistantly along the positive Y-axis direction;

[0030] S3-2-3. Define the phase continuous label and its corresponding multi-source heterogeneous data located in the first quadrant as the structure element matrix;

[0031] S3-2-4. Translate and copy the elements of the second column vector in the structure element matrix to the second quadrant of the matrix coordinate system as the second column vector in the second quadrant;

[0032] S3-2-5. Fill in the corresponding structural discrete labels as the first column vector in the second quadrant according to the elements of the second column vector in the second quadrant;

[0033] S3-2-6. Define the structural discrete labels and their multi-source heterogeneous data located in the second quadrant as the phase element matrix.

[0034] In some of these embodiments, extracting N ordered sequences of multi-source heterogeneous data from the structural element matrix includes:

[0035] S3-3-1. Defining the multi-source heterogeneous data corresponding to the origin label in the structural element matrix as reference data, and the multi-source heterogeneous data corresponding to non-origin labels as paired data;

[0036] S3-3-2. Associating each paired data with the reference data to obtain M heterogeneous data pairs; where M = N - 1;

[0037] A6-3. Calculating the content similarity of the M heterogeneous data pairs;

[0038] S3-3-4. Sorting them in descending order according to the content similarity values of the M heterogeneous data pairs to generate an ordered sequence of data pairs;

[0039] S3-3-5. Assigning sorting numbers to the ordered sequence of data pairs;

[0040] S3-3-6. Judging whether the sorting number of the heterogeneous data pair is the positive integer 1; if the sorting number is the positive integer 1, marking the corresponding heterogeneous data pair as the origin data pair, otherwise, marking it as a redundant data pair;

[0041] S3-3-7. Deleting all the reference data in the redundant data pairs, and defining the data sequence of the M paired data after deletion and the origin data pair as the ordered sequence of the multi-source heterogeneous data; where the ordered sequence of the multi-source heterogeneous data is characterized as an ordered sequence obtained by arranging the N multi-source heterogeneous data in descending order based on content similarity.

[0042] In some of these embodiments, constructing a three-element advanced matrix according to the ordered sequence of the multi-source heterogeneous data and the phase element matrix includes:

[0043] S3-4-1. Defining the ordered sequence of the N multi-source heterogeneous data as the first column vector of the three-element advanced matrix;

[0044] S3-4-2. Extracting the phase continuous labels corresponding to the N multi-source heterogeneous data from the phase element matrix;

[0045] S3-4-3. Rearranging the phase continuous labels according to the correspondence between the first column vector of the three-element advanced matrix and the phase continuous labels to generate phase non-continuous labels;

[0046] S3-4-4. Defining the phase non-continuous labels as the second column vector of the three-element advanced matrix;

[0047] S3-4-5. Fill in the content similarity according to the correspondence between the first column vector of the three-element advanced matrix and the content similarity, and define it as the third column vector of the three-element advanced matrix; among them, fill in zero at the matrix position corresponding to the reference data.

[0048] S3-4-6. Pair the first column vector, the second column vector and the third column vector to construct the three-element advanced matrix.

[0049] In some embodiments, the evaluation label determination module is specifically configured to:

[0050] S4-1. According to the phase continuous label in the three-element advanced matrix, match the first acquisition timestamp and the second acquisition timestamp of adjacent phase continuous labels;

[0051] S4-2. Define the interval duration between the first acquisition timestamp and the second acquisition timestamp as the evaluation interval coefficient between adjacent multi-source heterogeneous data after normalization;

[0052] S4-3. Define the product of the evaluation interval coefficient and the content similarity as the evaluation redundancy value;

[0053] S4-4. Compare the evaluation redundancy value with the threshold to generate a binary classification comparison result;

[0054] S4-5. According to the binary classification comparison result, label the evaluation behavior label of adjacent heterogeneous data pairs.

[0055] In some embodiments, the labeling condition for labeling the evaluation behavior label of adjacent heterogeneous data pairs according to the binary classification comparison result is:

[0056] If the binary classification comparison result is that the evaluation redundancy value is greater than the threshold, then mark the data row corresponding to the second acquisition timestamp in the three-element advanced matrix as "unnecessary evaluation", and mark the data row corresponding to the first acquisition timestamp as "reasonable evaluation"; otherwise, mark both the data rows corresponding to the first acquisition timestamp and the second acquisition timestamp as "reasonable evaluation".

[0057] In some embodiments, according to the evaluation behavior label and the three-element advanced matrix, construct a multi-source heterogeneous data storage matrix, including:

[0058] S5-1. Initialize the fourth column vector in the three-element advanced matrix;

[0059] S5-2. Fill in the evaluation behavior label into the fourth column vector to generate a multi-source heterogeneous data storage matrix.

[0060] The present invention provides a multi-source heterogeneous data storage service system for regional seismic safety evaluation, having the following beneficial effects:

[0061] The present invention realizes the unified coding representation of multi-source heterogeneous data in the time dimension and the structure dimension by constructing a dual-tag system of phase continuous tags and structure discrete tags, and performs quadrant mapping on the tag information based on a matrix coordinate system to form a structure element matrix and a phase element matrix. Furthermore, a three-element advanced matrix is constructed to complete the structural sorting, time classification, and semantic alignment of the data. This tag system and matrix construction method not only adapt to various structural forms of data such as geospatial data, maps, tables, and layers, but also can uniformly organize and abstractly classify data with non-standard fields and heterogeneous sources, significantly improving the adaptability and accuracy of data standardization for storage.

[0062] Based on data structure modeling, the present invention forms a reference index for judging repeatability and necessity by calculating the product of the content similarity and the acquisition time interval coefficient of adjacent multi-source heterogeneous data, and uses this to label adjacent items in the data sequence with "reasonable evaluation" or "unnecessary evaluation" behavior tags. Specifically, when the acquisition time interval between two data is extremely short and their content expressions are highly similar, the system automatically determines that there is redundant evaluation behavior, retains only the data with an earlier time, and eliminates the duplicate data with a later time. This mechanism can dynamically identify "recent high-duplication" type redundant data and effectively control the inflow of invalid information into the database. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is a structural block diagram of a multi-source heterogeneous data storage service system for regional seismic safety evaluation according to the present invention;

[0064] Figure 2 is a storage flow chart of a multi-source heterogeneous data storage service system for regional seismic safety evaluation according to the present invention;

[0065] Figure 3 is a schematic diagram of the annotation process of the evaluation behavior tag described in the present invention

[0066] Figure 4 is a schematic diagram of the storage of the evaluation results in the client in the embodiment of the present invention;

[0067] Figure 5 is a schematic diagram of the query of the query request in the client in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] Example 1: Please refer to Figures 1 - 3 , the present invention provides a multi-source heterogeneous data warehousing service system for regional seismic safety evaluation, including:

[0070] A dataset acquisition module for acquiring a dataset composed of N multi-source heterogeneous data;

[0071] Exemplarily, in this embodiment, the multi-source heterogeneous data refers to the area evaluation data for regional seismic safety evaluation.

[0072] According to the warehousing stage, it can be divided into actual material data and result data.

[0073] For the actual material data, its content may include: data on geological surveys, geophysical explorations, seismic activity analyses, drilling, and engineering geological condition analyses, as well as process data for seismic hazard analyses.

[0074] For the result data, its content may include: multi-probability seismic motion parameters for the bedrock and surface of the controlled borehole points, as well as the multi-probability seismic motion parameter zoning map of the target area.

[0075] According to the data type, the data can be divided into the following categories:

[0076] Geospatial data: including points (Point), polylines (Polyline), and polygons (Polygon);

[0077] Data tables: that is, tabular data (Table);

[0078] Graphics: such as charts and figures (Figure);

[0079] Original data: including data files and specialized formats (Data FileA, Apecialized FormatA), etc.

[0080] Among the geospatial data provided by different practicing units, there are both data in the ArcGIA format and data in other formats such as Mapinfo and MapGIA. These data have problems of non-standardization and non-uniformity in terms of feature coding and field rules. However, these data of different types, formats, and storage methods are interrelated, constituting a typical multi-source and heterogeneous data environment.

[0081] A triple determination module for determining N double-label data triples according to the dataset;

[0082] An advanced matrix construction module for constructing a three-element advanced matrix according to N double-label data triples;

[0083] An evaluation label determination module for determining an evaluation behavior label from the three-element advanced matrix; wherein the evaluation behavior label is a binary label used to mark the evaluation effectiveness of adjacent data pairs, and its values are: "reasonable evaluation" and "unnecessary evaluation";

[0084] An incoming library matrix construction module for constructing a multi-source heterogeneous data incoming library matrix according to the evaluation behavior label and the three-element advanced matrix;

[0085] A server generation module for storing the multi-source heterogeneous data incoming library matrix in the server to generate a multi-source data server;

[0086] When the client issues a data query request, the multi-source data server responds to the data query request to generate a query vector;

[0087] The multi-source data server receives the query vector as input, extracts features from the multi-source heterogeneous data incoming library matrix, and outputs data features based on the data query request.

[0088] Specifically, the data query request can be initiated based on any single or combined conditions such as a structure label, a phase label, an evaluation behavior label, etc., and supports exact query and fuzzy matching.

[0089] In this embodiment, by uniformly obtaining a data set composed of N multi-source heterogeneous data, an original data basis for regional earthquake safety evaluation tasks is constructed. Among them, the multi-source heterogeneous data is divided into actual material data and result data according to the incoming library phase, corresponding to the input support and result expression of the evaluation process respectively; the former includes original materials such as geological surveys, geophysical explorations, and seismic activity analyses, and the latter covers model output results such as seismic motion parameters and zoning maps. The data types cover geospatial data (points, lines, surfaces), tabular data, map data, and original data in special formats, and the sources include various system formats such as ArcGIS, Mapinfo, and MapGIS, showing typical problems of structural heterogeneity and inconsistent naming rules. This embodiment provides comprehensive data support for subsequent label extraction, structure mapping, and evaluation modeling through the normalized access and structural classification processing of the above complex heterogeneous data. Further, the collected data will be incorporated into a uniformly constructed incoming library matrix system to realize the semantic labeling organization of heterogeneous data, provide the ability to extract structured features for terminal query requests, and ultimately support the accurate or fuzzy retrieval of the multi-source data server based on phase labels, structure labels, and evaluation behavior labels.

[0090] In this embodiment, the triple determination module is specifically used for:

[0091] S2-1. Determine the phase continuous label and the structure discrete label of the multi-source heterogeneous data from the dataset; wherein, the phase continuous label is used to characterize the acquisition phase of the multi-source heterogeneous data in the time dimension, and its minimum value is defined as the origin label;

[0092] S2-2. Pair the multi-source heterogeneous data with its corresponding phase continuous label and structure discrete label to construct N double-label data triples.

[0093] In this embodiment, by extracting the phase continuous label and the structure discrete label of the multi-source heterogeneous data and pairing them with the original data ontology, a set of double-label data triples consisting of N elements is constructed. The phase continuous label is used to identify the acquisition phase of the data in the time dimension, reflecting the timeliness of the data; the structure discrete label is used to distinguish the attribution category of the data in the structural expression, such as types of tables, graphics, spatial data, etc. This step realizes the dual-attribute calibration of the original heterogeneous data in the time dimension and the structure dimension, forming a unified labeled representation, enabling the heterogeneous data to be organized and compared orderly before being warehoused, and providing a data basis for finally forming a queryable, matchable, and inferable warehousing matrix.

[0094] Further, determining the phase continuous label and the structure discrete label of the multi-source heterogeneous data from the dataset includes:

[0095] S2-1-1. Extract the marked acquisition timestamps of N multi-source heterogeneous data in the dataset; wherein, the acquisition timestamp is the time attribute attached to each multi-source heterogeneous data during construction, used to characterize its acquisition time point; specifically, during the generation or acquisition process of the multi-source heterogeneous data, the acquisition timestamp has been marked by the original system, device, or personnel. Therefore, in this step, there is no need to recalculate the acquisition time, and only this time attribute needs to be read and extracted.

[0096] S2-2-2. Assign phase continuous labels to N multi-source heterogeneous data according to the time sequence of the acquisition timestamps; wherein, the phase continuous label is used to characterize the acquisition phase of the multi-source heterogeneous data in the time dimension, and its minimum value is defined as the origin label, corresponding to the multi-source heterogeneous data with the earliest acquisition time; the phase continuous label is essentially the phase number obtained by standardizing and sorting the acquisition timestamps, used to uniformly describe the relative acquisition phase of the data in the time dimension.

[0097] S2-2-3. Perform structure recognition on the data structures of N multi-source heterogeneous data to generate structure discrete labels; wherein, the structure discrete label is used to characterize the type attribution of the multi-source heterogeneous data at the structural level;

[0098] Among them, structure recognition refers to the process of identifying, classifying, and categorizing the organizational structure, data representation form, and semantic features of multi-source heterogeneous data. Preferably, the structure type can be divided according to multi-dimensional features such as the storage format, spatial attributes, field rules, and data expression form of the data. The structure recognition can be implemented by means of rule matching, metadata parsing, data structure model comparison, etc., to perform structure standardization processing on data objects from different sources and different formats, and classify them into predefined structure categories, thereby generating structure discrete labels.

[0099] In specific implementation, the structure recognition can be divided in the following ways:

[0100] If the data is a layer file with spatial attributes, it can be classified as geospatial data (points, lines, surfaces);

[0101] If the data is in pure table form (such as Excel, CSV, DBF), it is classified as table data;

[0102] If the data is map or image data (such as JPG, PDF, chart screenshots, etc.), it is classified as map data;

[0103] If the data is in the original format (such as proprietary format DataFileA, ShapeFile, Mapinfo, etc.), the meta-structure can be extracted according to the format rules; through the above structure attribute feature recognition and classification, a corresponding structure discrete label is assigned to each piece of data.

[0104] Therefore, the structure discrete label belongs to a discrete identification variable; in the present invention, the "phase continuous label" is used to describe the sequential attribute of multi-source heterogeneous data in the time dimension and belongs to a continuous identification variable; the "structure discrete label" is used to represent the attribution category of multi-source heterogeneous data in the structural feature dimension and belongs to a discrete identification variable.

[0105] In this embodiment, the advanced matrix construction module is specifically used for:

[0106] S3-1. Construct a matrix coordinate system in a two-dimensional plane, where the coordinate origin of the matrix coordinate system is represented as the origin label, and the coordinate Y-axis is represented as the virtual connection axis of the phase continuous label;

[0107] S3-2. Perform matrix arrangement on N double-label data triples in the matrix coordinate system to construct a structure element matrix and a phase element matrix;

[0108] S3-3. Extract an ordered sequence of N multi-source heterogeneous data from the structure element matrix;

[0109] S3-4. Construct a three-element advanced matrix according to the ordered sequence of the multi-source heterogeneous data and the phase element matrix.

[0110] In this embodiment, by introducing a two-dimensional matrix coordinate system to perform spatial mapping on N double-label data triples, a partitioned organization of structural attributes and temporal attributes in the matrix space is achieved. Specifically, the stage continuous labels form a virtual connection axis along the Y-axis direction, identifying the sequential logic of the data acquisition stage. The origin label is located at the coordinate origin, and the data is aligned based on this. The structural discrete labels are classified in the opposite quadrants of the matrix coordinate system through structural mapping with reference to the data ontology. Thus, the constructed structural element matrix and stage element matrix respectively record the arrangement states of heterogeneous data in the structural and temporal dimensions. On this basis, the extracted ordered data sequence, as the result of structural similarity sorting, is combined with the stage information and label information to form a three-element advanced matrix, realizing the spatial sorting modeling of multi-source heterogeneous data.

[0111] Furthermore, performing matrix arrangement on N double-label data triples in the matrix coordinate system to construct a structural element matrix and a stage element matrix includes:

[0112] S3-2-1: Extract the stage continuous labels and their corresponding multi-source heterogeneous data from the N double-label data triples;

[0113] S3-2-2: Fill the extracted stage continuous labels and their corresponding multi-source heterogeneous data into the first quadrant of the matrix coordinate system; wherein, the origin label is on the X-axis, and the stage continuous labels and their corresponding multi-source heterogeneous data are arranged equidistantly along the positive Y-axis direction;

[0114] S3-2-3: Define the stage continuous labels and their corresponding multi-source heterogeneous data located in the first quadrant as the structural element matrix; that is, the first column vector is the stage continuous label, and the stage identification numbers start from the X-axis and are arranged in the positive Y-axis direction in the first quadrant; the second column vector is the corresponding multi-source heterogeneous data, and the multi-source heterogeneous data corresponding to the origin label is also on the X-axis.

[0115] S3-2-4: Translate and copy the elements of the second column vector in the structural element matrix to the second quadrant of the matrix coordinate system as the second column vector in the second quadrant;

[0116] S3-2-5: Fill in the corresponding structural discrete labels as the first column vector in the second quadrant according to the elements of the second column vector in the second quadrant;

[0117] S3-2-6: Define the structural discrete labels and their multi-source heterogeneous data located in the second quadrant as the stage element matrix.

[0118] In this embodiment, by arranging the dual-label data triples in a quadrant matrix in the matrix coordinate system, a spatial distributed mapping expression of the phase-continuous label and the structure-discrete label is realized. Specifically, the phase-continuous label and its corresponding data are arranged in the first quadrant, where the origin label is aligned with the starting point of the X-axis, and the remaining label data are arranged equidistantly along the positive direction of the Y-axis in sequence to form a structure element matrix; the first column of the matrix records the time phase, and the second column maps the multi-source heterogeneous data ontology, constituting the structure data distribution under the time main axis. Further, by copying and translating the data ontology in the structure element matrix to the second quadrant and introducing structure-discrete labels according to their corresponding relationships, the correspondence between the structure labels and the data is established, thereby constituting a phase element matrix. The matrix layout logic is bounded by quadrants and axis by labels, realizing a symmetric expression of the dual-label attributes in the coordinate system, which helps to independently model and cross-compare the structure distribution and the time phase in the same space.

[0119] Further, extracting an ordered sequence of N multi-source heterogeneous data from the structure element matrix includes:

[0120] S3-3-1: Defining the multi-source heterogeneous data corresponding to the origin label in the structure element matrix as reference data, and the multi-source heterogeneous data corresponding to the non-origin label as paired data;

[0121] S3-3-2: Associating each paired data with the reference data to obtain M heterogeneous data pairs; where M = N - 1;

[0122] A6-3: Calculating the content similarity of M heterogeneous data pairs; where the content similarity is used to measure the degree of similarity between two multi-source heterogeneous data in terms of information content, numerical distribution, or expression results;

[0123] In this embodiment, the content similarity is used to measure the degree of proximity between two multi-source heterogeneous data in terms of information content, numerical distribution, or expression characteristics. Preferably, after performing feature vectorization processing on various data objects, **Cosine Similarity** can be used for calculation, and its expression is as follows:

[0124] ;

[0125] where A and B respectively represent the feature expression vectors of two multi-source heterogeneous data. The similarity value ranges from [0, 1], and the higher the value, the higher the similarity degree of the two data in content expression. The content similarity is in scalar form and can be directly used for subsequent evaluation of redundancy value calculation. If the multi-source heterogeneous data includes unstructured forms such as images, layers, and charts, it can be uniformly processed after being transformed into vectors through feature extraction (such as image texture, layer vector features, table summary indicators, etc.).

[0126] S3-3-4. Sort the M heterogeneous data pairs in descending order according to their content similarity values to generate an ordered data pair sequence; wherein, the ordered data pair sequence is used to represent the order of all paired data arranged from high to low according to the similarity degree with the reference data.

[0127] S3-3-5. Assign sorting numbers to the ordered data pair sequence.

[0128] Among them, the sorting numbers are a group of consecutive positive integers starting from 1, and each multi-source heterogeneous data pair is assigned in turn according to the order from high to low similarity; number 1 corresponds to the maximum similarity value, and number M corresponds to the minimum similarity value; the "sorting number" in this embodiment is used to identify the position of the data pair arranged in descending order in the ordered sequence.

[0129] S3-3-6. Determine whether the sorting number of the heterogeneous data pair is the positive integer 1; if the sorting number is the positive integer 1, mark the corresponding heterogeneous data pair as the origin data pair, otherwise, mark it as a redundant data pair.

[0130] S3-3-7. Delete all reference data in the redundant data pairs, and define the data sequence of the M paired data after deletion and the origin data pair as the ordered sequence of the multi-source heterogeneous data; wherein, the ordered sequence of the multi-source heterogeneous data is characterized as an ordered sequence of N multi-source heterogeneous data arranged in descending order based on content similarity.

[0131] In this embodiment, by using the multi-source heterogeneous data corresponding to the origin label in the structure element matrix as the reference data, associating it with the remaining N - 1 paired data one by one to form a data pair set, and comparing and sorting according to the content similarity, the clustering reconstruction of heterogeneous data at the content level is realized. The content similarity is constructed based on the feature vector expression. When processing different structure data such as drawings, layers, charts, and tables, their texture features, vector features, or summary indicators are respectively extracted and transformed into standard vector representations, and then the cosine similarity is used for measurement and calculation to ensure comparability between different data forms. After descending order sorting, combined with the setting of sorting numbers and the redundant marking logic, the redundant reference data in multiple repeated evaluations are systematically removed, and only the most representative data in each group of content is retained, thereby generating an ordered sequence of multi-source heterogeneous data dominated by structural similarity.

[0132] Furthermore, according to the ordered sequence of the multi-source heterogeneous data and the stage element matrix, construct a three-element advanced matrix, including:

[0133] S3-4-1. Define the ordered sequence of N multi-source heterogeneous data as the first column vector of the three-element advanced matrix.

[0134] Extract the sorted ordered sequence of N multi-source heterogeneous data according to the descending order result of content similarity, and form the first column vector of the three-element advanced matrix.

[0135] S3-4-2. Extract the stage continuous labels corresponding to N multi-source heterogeneous data from the stage element matrix;

[0136] By traversing each multi-source heterogeneous data in the first column vector, find its corresponding stage continuous label in the stage element matrix to form a one-to-one correspondence.

[0137] S3-4-3. Rearrange the stage continuous labels according to the corresponding relationship between the first column vector of the three-element advanced matrix and the stage continuous labels to generate stage discontinuous labels;

[0138] S3-4-4. Define the stage discontinuous labels as the second column vector of the three-element advanced matrix;

[0139] S3-4-5. Fill in the content similarity according to the corresponding relationship between the first column vector of the three-element advanced matrix and the content similarity, and define it as the third column vector of the three-element advanced matrix; among them, zero is filled in the matrix position corresponding to the reference data;

[0140] Specifically, the content similarity is a representation of the content similarity degree between the reference data and the paired data. In the ordered sequence of multi-source heterogeneous data (the first column vector of the three-element advanced matrix), although the reference data is deleted, the remaining paired data still has a corresponding relationship with the content similarity.

[0141] S3-4-6. Pair the first column vector, the second column vector and the third column vector to construct the three-element advanced matrix.

[0142] In this embodiment, based on the ordered sequence of multi-source heterogeneous data and the stage element matrix extracted in the previous stage, a three-element advanced matrix integrating sorting information, time tags and structural correlation is constructed. Specifically, the first column vector carries the main sequence of data sorted by content similarity, reflecting the distribution logic of the structural dimension; the second column vector generates stage discontinuous labels consistent with the current structural sorting through the sequential mapping and rearrangement of the stage continuous labels, realizing the alignment and unification of time stage information and structural sequence; the third column vector is filled according to the similarity calculation result in the original data pair, and the numerical value expresses the content proximity degree between the two groups of data. The numerical value corresponding to the reference data position is set to zero to highlight its reference role. Through the ordered pairing of the three column vectors, a composite data matrix with structural sortability, time stageability and content similarity numericality is formed, providing a complete three-dimensional input structure for the evaluation behavior determination, redundancy recognition and label-driven warehousing based on sorting logic.

[0143] In this embodiment, the evaluation label determination module is specifically used to:

[0144] S4-1, matching the first acquisition timestamp and the second acquisition timestamp of the adjacent stage continuous labels according to the stage continuous labels in the three-element advanced matrix;

[0145] S4-2, defining the interval between the first acquisition timestamp and the second acquisition timestamp as an evaluation interval coefficient between adjacent multi-source heterogeneous data after normalization;

[0146] S4-3, defining the product of the evaluation interval coefficient and the content similarity as the evaluation redundancy value;

[0147] In this embodiment, content similarity is used to measure the closeness of two multi-source heterogeneous data at the content level, and is defined as a scalar value, usually in the interval [0,1]. The larger the value, the closer the content. This value can be multiplied by the evaluation interval coefficient to generate an evaluation redundancy value. Under the premise of a short time interval, the data content is highly repeated, which may be a resource-wasting evaluation behavior.

[0148] S4-4, comparing the evaluation redundancy value and the threshold value to generate a binary classification comparison result;

[0149] S4-5. According to the binary classification comparison results, label the evaluation behavior labels of adjacent heterogeneous data pairs.

[0150] In this embodiment, based on the existing stage continuous labels and content similarity data in the three-element advanced matrix, a bivariate calculation is performed on adjacent multi-source heterogeneous data pairs. Specifically, first, by extracting the acquisition timestamps of adjacent data pairs, the time interval is calculated and normalized into the evaluation interval coefficient, which is used to quantify the compactness of the evaluation behavior in the time dimension; then combined with the content similarity index in the structural dimension, the evaluation redundancy value reflecting the "recent high weight" feature is generated by multiplying the two. As a scalar indicator, this redundancy value reflects that if the content similarity is still high when the time difference is small, the evaluation behavior may be redundant or waste resources. By comparing the redundancy value with the set threshold, a binary comparison result of "reasonable evaluation" and "unnecessary evaluation" is generated, and finally a clear evaluation behavior label is marked for each pair of adjacent data. This label can not only characterize whether the data has the meaning of retention, but also provide a judgment basis for the data selection of the storage matrix, and is one of the important screening labels in the storage stage.

[0151] Among them, according to the binary classification comparison results, the labeling conditions for labeling the evaluation behavior labels of adjacent heterogeneous data pairs are:

[0152] If the binary classification comparison result is that the evaluation redundancy value is greater than the threshold, then mark the data row corresponding to the second acquisition timestamp in the three-element advanced matrix as "unnecessary evaluation", and mark the data row corresponding to the first acquisition timestamp as "reasonable evaluation"; otherwise, mark the data rows corresponding to both the first acquisition timestamp and the second acquisition timestamp as "reasonable evaluation".

[0153] In this embodiment, based on the binary classification comparison result of the evaluation redundancy value and the preset threshold, accurate annotation of the evaluation behavior labels for the data rows corresponding to adjacent data pairs in the three-element advanced matrix is carried out, realizing dynamic identification and marking control of the data retention value. Specifically, when the redundancy value is higher than the threshold, it indicates that there are highly repeated evaluation behaviors within a short time interval. The system defaults to preferentially retaining the data row earlier in time and marks the subsequent data rows as "unnecessary evaluation" to reduce the storage occupancy of redundant information; on the contrary, if the redundancy value is lower than or equal to the threshold, it is considered that the two pieces of data have sufficient distinguishability in content or acquisition time sequence and can both be retained, and are uniformly marked as "reasonable evaluation". This annotation mechanism combines the primary and secondary sorting of the time dimension with the numerical discrimination of the redundancy value, realizing an auxiliary storage decision-making method with logical controllability and clear identification, ensuring that the stored data not only has structural representativeness but also avoids redundant content accumulation.

[0154] In this embodiment, according to the evaluation behavior label and the three-element advanced matrix, a multi-source heterogeneous data storage matrix is constructed, including:

[0155] S5-1. Initialize the fourth column vector in the three-element advanced matrix;

[0156] S5-2. Fill the evaluation behavior label into the fourth column vector to generate a multi-source heterogeneous data storage matrix.

[0157] In this embodiment, by initializing the fourth column vector in the three-element advanced matrix and filling the marking result obtained by the evaluation behavior label determination into this column, the unified integration of structural sorting, time label, content similarity, and behavior determination result is realized, and a multi-source heterogeneous data storage matrix is constructed. The storage matrix not only retains all the characteristic information of the original data in the time sequence, structure, and similarity dimensions, but also introduces the behavior label as an explicit screening basis, providing a matrix-type data carrier with complete structure and clear semantics for subsequent data storage, invocation, and visual query. Through this step, the original data has completed the transformation from an unstructured, multi-source heterogeneous state to a storage standard format that can be directly invoked, with good system adaptability, label-drivenness, and structure query ability, significantly improving the organization efficiency and intelligent level of seismic safety evaluation data in the backend platform.

[0158] Example 2: Refer to Figure 4 and Figure 5As shown, Embodiment 2 discloses the specific application of the client in the warehousing process and the specific application of the query path of the data query request in Embodiment 1.

[0159] The data query path of the client is as Figure 4 shown, which is used to query various types of data features (i.e., the source heterogeneous data warehousing matrix) in the area evaluation results, and each module therein is an adaptive display of the client interface.

[0160] And in Figure 5 , for example, referring to Figure 5 , taking the query vector as the input, data features based on the data query request can be output; the query vector can represent an area quantity query or a preparation point query, and the data features can be the total area or the project information of the target point.

[0161] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.).

[0162] The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVD ), or semiconductor media. The semiconductor media can be a solid-state drive.

[0163] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division of an underwater topographic change analysis system and method for waterways. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0164] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A multi-source heterogeneous data warehousing service system for regional seismic safety evaluation, characterized in that Including: A data set acquisition module for acquiring a data set composed of N multi-source heterogeneous data; A triple determination module for determining N double-label data triples according to the data set; An advanced matrix construction module for constructing a three-element advanced matrix according to N double-label data triples; An evaluation label determination module for determining an evaluation behavior label from the three-element advanced matrix; wherein, the evaluation behavior label is a binary label used to mark the evaluation validity of adjacent data pairs, and its values are: "reasonable evaluation" and "unnecessary evaluation"; A storage matrix construction module for constructing a multi-source heterogeneous data storage matrix according to the evaluation behavior label and the three-element advanced matrix; A server generation module for storing the multi-source heterogeneous data storage matrix in a server to generate a multi-source data server; When a client issues a data query request, the multi-source data server responds to the data query request to generate a query vector; The multi-source data server receives the query vector as an input, extracts features from the multi-source heterogeneous data storage matrix, and outputs data features based on the data query request; The triple determination module is specifically used for: S2-1. Determining the stage continuous label and the structure discrete label of the multi-source heterogeneous data from the data set; wherein, the stage continuous label is used to characterize the acquisition stage of the multi-source heterogeneous data in the time dimension, and its minimum value is defined as the origin label; S2-2. Pairing the multi-source heterogeneous data with its corresponding stage continuous label and structure discrete label to construct N double-label data triples; The advanced matrix construction module is specifically used for: S3-1. Constructing a matrix coordinate system in a two-dimensional plane, wherein the coordinate origin of the matrix coordinate system is characterized as the origin label, and the coordinate Y-axis is characterized as the virtual connection axis of the stage continuous label; S3-2. Performing matrix arrangement on N double-label data triples in the matrix coordinate system to construct a structure element matrix and a stage element matrix; S3-3. Extracting an ordered sequence of N multi-source heterogeneous data from the structure element matrix; S3-4. Constructing a three-element advanced matrix according to the ordered sequence of the multi-source heterogeneous data and the stage element matrix; Constructing a three-element advanced matrix according to the ordered sequence of the multi-source heterogeneous data and the stage element matrix, including: S3-4-1. Defining the ordered sequence of N multi-source heterogeneous data as the first column vector of the three-element advanced matrix; S3-4-2. Extracting the stage continuous labels corresponding to N multi-source heterogeneous data from the stage element matrix; S3-4-3. Rearranging the stage continuous labels according to the corresponding relationship between the first column vector of the three-element advanced matrix and the stage continuous labels to generate stage non-continuous labels; S3-4-4. Defining the stage non-continuous labels as the second column vector of the three-element advanced matrix; S3-4-5. Filling in the content similarity according to the corresponding relationship between the first column vector of the three-element advanced matrix and the content similarity, and defining it as the third column vector of the three-element advanced matrix; wherein, zero is filled in the matrix position corresponding to the reference data; S3-4-6. Pair the first column vector, the second column vector, and the third column vector to construct the three-element advanced matrix.

2. The multi-source heterogeneous data warehousing service system for regional seismic safety evaluation according to claim 1, characterized in that, Determine the stage continuous label and the structure discrete label of the multi-source heterogeneous data from the dataset, including: S2-1-1. Extract the marked acquisition timestamps of N multi-source heterogeneous data in the dataset; S2-2-2. Assign stage continuous labels to the N multi-source heterogeneous data according to the chronological order of the acquisition timestamps; wherein, the stage continuous label is used to characterize the acquisition stage of the multi-source heterogeneous data in the time dimension, and its minimum value is defined as the origin label, corresponding to the multi-source heterogeneous data with the earliest acquisition time; S2-2-3. Perform structure recognition on the data structures of the N multi-source heterogeneous data to generate structure discrete labels; wherein, the structure discrete label is used to characterize the type attribution of the multi-source heterogeneous data at the structural level.

3. The multi-source heterogeneous data warehousing service system for regional seismic safety evaluation according to claim 2, characterized in that, Perform matrix arrangement on the N double-label data triples in the matrix coordinate system to construct a structure element matrix and a stage element matrix, including: S3-2-1. Extract the stage continuous label and its corresponding multi-source heterogeneous data from the N double-label data triples; S3-2-2. Fill the extracted stage continuous label and its corresponding multi-source heterogeneous data into the first quadrant of the matrix coordinate system; wherein, the origin label is located on the X-axis, and the multi-source heterogeneous data corresponding to the stage continuous label are arranged equidistantly along the positive direction of the Y-axis; S3-2-3. Define the stage continuous label and its corresponding multi-source heterogeneous data located in the first quadrant as the structure element matrix; S3-2-4. Translate and copy the elements of the second column vector in the structure element matrix to the second quadrant of the matrix coordinate system as the second column vector in the second quadrant; S3-2-5. Fill in the corresponding structure discrete label as the first column vector in the second quadrant according to the elements of the second column vector in the second quadrant; S3-2-6. Define the structure discrete label and its multi-source heterogeneous data located in the second quadrant as the stage element matrix.

4. A multi-source heterogeneous data warehousing service system for regional seismic safety evaluation according to claim 3, characterized in that Extract the ordered sequence of N multi-source heterogeneous data from the structure element matrix, including: S3-3-1. Define the multi-source heterogeneous data corresponding to the origin label in the structure element matrix as the reference data, and the multi-source heterogeneous data corresponding to the non-origin label as the paired data; S3-3-2. Associate each paired data with the reference data to obtain M heterogeneous data pairs; wherein, M = N - 1; S3-3-3. Calculate the content similarity of the M heterogeneous data pairs; S3-3-4. Sort them in descending order according to the content similarity values of the M heterogeneous data pairs to generate an ordered sequence of data pairs; S3-3-5. Assign a sorting number to the ordered sequence of data pairs; S3-3-6. Determine whether the sorting number of the heterogeneous data pair is the positive integer 1; if the sorting number is the positive integer 1, mark the corresponding heterogeneous data pair as the origin data pair, otherwise, mark it as the redundant data pair; S3-3-7. Delete all the reference data in the redundant data pairs, and define the data sequences of the M paired data after deletion and the origin data pair as the ordered sequence of the multi-source heterogeneous data; wherein, the ordered sequence of the multi-source heterogeneous data is characterized as an ordered sequence obtained by descendingly arranging the N multi-source heterogeneous data based on the content similarity.

5. The multi-source heterogeneous data warehousing service system for regional seismic safety evaluation according to claim 4, characterized in that The evaluation label determination module is specifically configured to: S4-1. According to the phase continuous labels in the three-element advanced matrix, match the first acquisition timestamp and the second acquisition timestamp of adjacent phase continuous labels; S4-2. Define the interval duration between the first acquisition timestamp and the second acquisition timestamp, after normalization, as the evaluation interval coefficient between adjacent multi-source heterogeneous data; S4-3. Define the product of the evaluation interval coefficient and the content similarity as the evaluation redundancy value; S4-4. Compare the evaluation redundancy value with the threshold to generate a binary classification comparison result; S4-5. According to the binary classification comparison result, label the evaluation behavior labels of adjacent heterogeneous data pairs.

6. The multi-source heterogeneous data warehousing service system for regional seismic safety evaluation according to claim 5, characterized in that, The labeling condition for labeling the evaluation behavior labels of adjacent heterogeneous data pairs according to the binary classification comparison result is: If the binary classification comparison result is that the evaluation redundancy value is greater than the threshold, then mark the data row corresponding to the second acquisition timestamp in the three-element advanced matrix as "unnecessary evaluation", and mark the data row corresponding to the first acquisition timestamp as "reasonable evaluation"; otherwise, mark both the data rows corresponding to the first acquisition timestamp and the second acquisition timestamp as "reasonable evaluation".

7. A multi-source heterogeneous data warehousing service system for regional seismic safety evaluation according to claim 6, characterized in that Construct a multi-source heterogeneous data storage matrix according to the evaluation behavior labels and the three-element advanced matrix, including: S5-1. Initialize the fourth column vector in the three-element advanced matrix; S5-2. Fill the evaluation behavior labels into the fourth column vector to generate a multi-source heterogeneous data storage matrix.

Citation Information

Patent Citations

  • Earthquake Safety Assessment Database Construction Methods and Systems

    CN115964360B

  • Multi-source heterogeneous data entity alignment method oriented to field of public security

    CN111753024A

  • Natural resource multi-source heterogeneous data aggregation and fusion service system

    CN115774861A