A Construction Method for a Structured Database of Multivariate Heterogeneous Technical Trade Measures
Through the construction of multivariate heterogeneous data acquisition, cleaning, mapping, distributed storage and knowledge graphs, the problem of insufficient data integration and analysis capabilities in the existing technology is solved, efficient integration and dynamic update of multivariate heterogeneous technology trade measures data is achieved, and data management and analysis capabilities are improved.
Patent Information
- Application Number
- CN202510680502.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-26
AI Technical Summary
When handling multivariate and heterogeneous technical trade measures data, the existing technology has problems such as single data collection methods, lack of professional classification systems for data cleaning and integration, simple data analysis models, and lack of dynamic update mechanisms, resulting in incomplete information coverage, insufficient analysis capabilities, and inability to effectively deal with trade barriers.
Multivariate heterogeneous data acquisition, data cleaning, data mapping, distributed storage and knowledge graph construction methods are adopted, including network crawlers, BERT model entity extraction, two-way mapping of HS encoding and national economic industry classification code, hybrid storage architecture, TransE algorithm modeling and dynamic update mechanism, to achieve efficient integration and dynamic update of multivariate heterogeneous data.
It realizes efficient integration and quality optimization of multivariate heterogeneous data, ensures the comprehensiveness, accuracy and standardization of data, improves data management efficiency and retrieval performance, supports dynamic real-time updates and reliable traceability of historical data, and provides efficient data analysis and early warning capabilities.
Smart Images

Figure CN120196686B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing for international trade technical barriers, and particularly to a method for constructing a structured database of diverse heterogeneous technical trade measures. Background Art
[0002] With the in-depth development of global trade integration, technical trade measures have become an important factor affecting international trade. In the early stage, the information processing of technical trade measures mainly relied on manual collection and analysis, with low efficiency and limited coverage. With the progress of information technology, countries have gradually established electronic databases based on a single data source, realizing the digital storage of basic data such as WTO-TBT / SPS notifications and product recalls. In recent years, the application of big data and artificial intelligence technologies has promoted the development of automated data collection. For example, public data is obtained through web crawlers, and duplicate data is identified using simple algorithms, initially constructing a structured data storage framework. However, the existing technology still remains at the stage of isolated processing of single-type data, lacking systematic integration of diverse heterogeneous data, and it is difficult to form an associated analysis ability covering multiple dimensions such as technical regulations, industrial standards, and market feedback.
[0003] The deficiencies of the existing technology are mainly reflected in three aspects: First, the data collection means are single, and the ability to obtain unstructured policy documents and enterprise dynamic data is insufficient, resulting in incomplete information coverage; second, there is a lack of a professional classification system for technical trade measures in data cleaning and integration, making it difficult to achieve accurate mapping of product classification and industrial fields, and the data availability is low; third, the data analysis model is simple, and it can only complete basic data retrieval and statistics, unable to intelligently judge the risk level and influence scope of technical barriers, and lacking the ability of differential warning for different enterprise entities. In addition, the existing systems generally lack a dynamic update mechanism, making it difficult to track the changes of foreign technical trade measures in real time, resulting in enterprises obtaining lagged information and being unable to effectively respond to the real-time impact of trade barriers. These defects cause enterprises to generally have problems such as unsmooth information acquisition channels, insufficient analysis capabilities, and lack of response means when facing foreign technical trade measures, and there is an urgent need for a more efficient and intelligent technical solution to solve the problems of processing and application of diverse heterogeneous data. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above problems and provide a method for constructing a structured database of diverse heterogeneous technical trade measures. To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] A method for constructing a structured database of diverse heterogeneous technical trade measures includes the following steps:
[0006] Step S1: Collect raw data of diverse heterogeneous technical trade measures;
[0007] Step S2: cleaning the original data;
[0008] Step S3: Establish a data mapping system to associate the original data with the preset basic classification system;
[0009] Step S4: construct a distributed storage architecture, store the cleaned and mapped structured data into a database cluster according to preset rules, and form a structured database;
[0010] Step S5: Construct a technical trade measures knowledge graph based on the structured database.
[0011] Further, in step S1, the multi-heterogeneous original data of technical trade measures include WTO-TBT notification data, WTO-SPS notification data, European and American product recall data, domestic and foreign technical regulations data and industrial standards data;
[0012] Furthermore, in step S1, the method of collecting raw data includes: deploying a web crawler program, configuring a dynamic IP proxy pool and a request interval parameter of 5 seconds, and performing real-time monitoring and crawling of public data sources; connecting to industry databases and enterprise data platforms through an API interface based on the OAuth2.0 protocol, with a data request frequency limit of 100 times per minute; and using a BERT-base model to extract entities from unstructured text, where entity types include product names and regulatory terms, and the extraction accuracy is ≥95%.
[0013] Further, in step S2, the cleaning process includes the following steps:
[0014] Step S21: Verify data based on data integrity rules, and remove invalid data when the missing rate of key fields is greater than 5%;
[0015] Step S22: Use the Levenshtein edit distance algorithm to merge duplicate records, and the calculation formula is: Step S21: Verify the data based on the data integrity rule, and remove invalid data when the missing rate of key fields is greater than 5%;
[0016] Step S22: Use the Levenshtein edit distance algorithm to merge duplicate records. The calculation formula is: ;
[0017] in, and Separate strings and The length index of Before and The minimum edit distance of character substrings; is an indicator function. When is not equal to , it takes 1; otherwise, it takes 0. When a certain substring is empty , the edit distance is the length of the other substring; otherwise, it takes the minimum value of deletion, insertion, or replacement operations;
[0018] Step S23: Standardize the data format through a preset dictionary library, which includes a standard term library and a unit conversion table.
[0019] Furthermore, in step S3, the preset basic classification system includes the HS coding system, the national economic industry classification code system, and the product attribute classification system.
[0020] Furthermore, in step S3, the establishment of the data mapping system includes the following steps:
[0021] Step S31: Construct a two-way mapping table between HS codes and national economic industry classification codes. Match the first 6 digits of the HS code through regular expressions, and the mapping accuracy requirement is ≥99%;
[0022] Step S32: Calculate the correlation strength between technical indicators and product attributes through the cosine similarity algorithm. The specific formula is: ;
[0023] where is the technical indicator word vector, is the product attribute word vector, is the vector dot product, which measures the consistency of semantic directions; and are the vector norms, which normalize the calculation results; when the similarity is ≥0.8, it is determined that there is a mapping relationship between the technical indicators and the product attributes;
[0024] Step S33: Generate a credibility score based on the authority level and update frequency of the data source. The calculation formula is: .
[0025] Furthermore, in step S4, the distributed storage architecture adopts a hybrid storage mode that combines a relational database and a non-relational database. Structured data is partitioned and stored in the relational database according to the first 6 digits of the HS code, and an index based on the product code and release date is established; semi-structured data is sharded and stored in the non-relational database according to the data source identifier, and redundant copies are configured.
[0026] Furthermore, in step S4, the preset rules include dividing the data sub-library according to the first 4 digits of the HS code, partitioning the table according to the data source type, and establishing a time series index based on the UTC timestamp field.
[0027] Further, in step S5, the construction of the knowledge graph of technical trade measures includes the following steps:
[0028] Step S51: Define the entity types of the knowledge graph. The entity types include technical regulations, products, enterprises, and standards. The entity attributes of technical regulations include promulgation date, scope of application, and constraint intensity;
[0029] Step S52: Extract the association relationships between entities. The association relationships are modeled by the TransE algorithm, with a vector dimension of 100 and boundary hyperparameters , and the loss function is: ;
[0030] where and are the vector representations of the head entity and the tail entity respectively, is the vector representation of the relationship, is the vector representation of the head entity in the negative sample, is the vector representation of the tail entity in the negative sample;
[0031] Step S53: Calculate the entity importance through the PageRank algorithm. The damping factor of the algorithm , and the number of iterations is 100 times. Select the Top-10% entities as the core nodes.
[0032] Further, it also includes step S6: Establish a data dynamic update mechanism, including the following steps:
[0033] Step S61: Set periodic data collection tasks to obtain incremental data from each data source;
[0034] Step S62: Identify the updated content through a data comparison algorithm. The algorithm includes Jaccard similarity calculation, and the formula is: ;
[0035] where is the incremental data set, is the baseline data set; when the difference rate exceeds the preset threshold, perform incremental updates on the structured database and the knowledge graph;
[0036] Step S63: Generate a data version identifier based on the hash algorithm. The hash algorithm uses SHA-256, and its output is a 256-bit hash value. Associate the identifier with the timestamp and store it in the blockchain to form a traceable historical data chain.
[0037] The advantages of the present invention are:
[0038] The present invention realizes the efficient integration and quality optimization of multi-source heterogeneous data such as WTO-TBT / SPS notifications, product recall data, and technical regulations through a multi-source acquisition method that deploys web crawlers in combination with API docking and entity extraction using the BERT model, and is paired with integrity verification, duplicate record merging, and format standardization processing in data cleaning, ensuring the comprehensiveness, accuracy, and standardization of the original data, and laying a high-quality data foundation for subsequent data mapping and storage.
[0039] The present invention realizes the precise mapping and efficient storage of multi-source data under a unified classification system by constructing a bidirectional mapping table between HS codes and national economic industry classification codes, calculating the correlation strength between technical indicators and product attributes using cosine similarity, and combining a hybrid storage architecture of relational and non-relational databases and strategies for database partitioning, sub-table partitioning, and index setting, which not only meets the complex query requirements of structured data but also adapts to the flexible storage of semi-structured data, improving the management efficiency of the database and the data retrieval performance.
[0040] The present invention realizes the dynamic real-time update of technical trade measure data and the reliable traceability of historical versions by setting periodic incremental data acquisition tasks, using the Jaccard similarity algorithm to identify updated content and trigger incremental updates, and combining the SHA-256 hash algorithm with blockchain technology to generate traceable version identifiers, ensuring that the database continuously reflects the latest trade measure dynamics, and at the same time providing an immutable historical data chain support for data auditing and compliance analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings forming a part of this application are used to provide a further understanding of this application, making other features, objects, and advantages of this application more obvious. The schematic embodiments of the drawings of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application.
[0042] In the drawings:
[0043] Figure 1 It is a flowchart of a method for constructing a multi-source heterogeneous structured database of technical trade measures in Embodiment 1.
[0044] Figure 2 It is an architecture diagram of data acquisition for a method for constructing a multi-source heterogeneous structured database of technical trade measures in Embodiment 1.
[0045] Figure 3 It is a schematic diagram of the data cleaning process for a method for constructing a multi-source heterogeneous structured database of technical trade measures in Embodiment 1.
[0046] Figure 4 It is a flowchart of the data dynamic update mechanism for a method for constructing a multi-source heterogeneous structured database of technical trade measures in Embodiment 1. Specific Embodiments
[0047] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0048] The present invention will be introduced in detail and specifically below through specific embodiments to better understand the present invention. However, the following embodiments do not limit the protection scope of the present invention.
[0049] Embodiment:
[0050] A method for constructing a structured database of multi-source heterogeneous technical trade measures, comprising the following steps:
[0051] Step S1: Collect the original data of multi-source heterogeneous technical trade measures;
[0052] Step S2: Clean and process the original data;
[0053] Step S3: Establish a data mapping system to associate the original data with a preset basic classification system;
[0054] Step S4: Construct a distributed storage architecture, and store the cleaned and mapped structured data in a database cluster according to preset rules to form a structured database;
[0055] Step S5: Construct a knowledge graph of technical trade measures based on the structured database.
[0056] Through a series of processes, the data of multi-source heterogeneous technical trade measures are processed to realize the transformation from the original data to the structured database and the knowledge graph. The technical effect produced is that it can systematically integrate data from different sources and different structures to form a structured database, and further construct a knowledge graph, providing an infrastructure for the data management, analysis and application of technical trade measures, and facilitating the subsequent in-depth mining and utilization of relevant data.
[0057] Further, in step S1, the original data of the multi-source heterogeneous technical trade measures includes WTO-TBT notification data, WTO-SPS notification data, European and American product recall data, domestic and foreign technical regulation data, and industrial standard data;
[0058] For example, during the data collection stage, the following specific data are explicitly included: TBT notification data and SPS notification data are obtained from the WTO official website; product recall data are captured from the U.S. Consumer Product Safety Commission official website; national standards issued by the Standardization Administration of China and the EU RoHS directive are collected; and industry standards issued by the International Organization for Standardization are simultaneously included to ensure that the data covers international, domestic and foreign multi-source heterogeneous scenarios.
[0059] It ensures that the collected data covers internationally accepted notification data, foreign product recall information, domestic and foreign technical regulations and industry standards, etc. The data types are rich and targeted, providing rich original data for the subsequent construction of a comprehensive and authoritative technical trade measures database, which can meet various analysis and research needs.
[0060] Furthermore, in step S1, the method of collecting raw data includes: deploying a web crawler program, configuring a dynamic IP proxy pool and a request interval parameter of 5 seconds, and performing real-time monitoring and crawling of public data sources; connecting to industry databases and enterprise data platforms through an API interface based on the OAuth2.0 protocol, with a data request frequency limit of 100 times per minute; and using a BERT-base model to extract entities from unstructured text, where entity types include product names and regulatory terms, and the extraction accuracy is ≥95%.
[0061] For example, data collection is implemented as follows:
[0062] Web crawler: Use Python's Scrapy framework to deploy crawlers, configure the dynamic IP pool of Xigua proxy, set the request interval to 5 seconds, conduct real-time monitoring of the WTO-TBT / SPS notification official website, crawl HTML pages and parse structured data;
[0063] API connection: Access the EBSCO industry database and a multinational enterprise quality data platform through the OAuth2.0 protocol, set a request frequency limit of 100 times per minute, and obtain standardized internal enterprise technical standards and recall records;
[0064] Entity extraction: Based on Hugging Face's BERT-base model, unstructured regulatory texts are processed and two types of entities, product names and regulatory clauses, are defined. By fine-tuning the model, the extraction accuracy rate reaches 96%, ensuring accurate extraction of key information from unstructured texts.
[0065] For different data sources and data forms, appropriate technical means are adopted to collect data efficiently and accurately. The dynamic IP proxy pool and the setting of request interval parameters can avoid being restricted from accessing the data source and ensure the stability of real-time monitoring and scraping; API docking realizes a reliable connection with industry databases and enterprise data platforms to obtain professional data; the entity extraction with high accuracy of the BERT-base model can effectively extract key information from unstructured texts, and the combination of multiple methods ensures the comprehensiveness, efficiency, and accuracy of the original data collection.
[0066] Further, in step S2, the cleaning process includes the following steps:
[0067] Step S21: Verify the data based on the data integrity rule. When the missing rate of the key fields > 5%, eliminate the invalid data;
[0068] Step S22: Use the Levenshtein edit distance algorithm to merge duplicate records. The calculation formula is: ;
[0069] where, and are the length indexes of the strings and respectively, represents the minimum edit distance of the first and character substrings; is an indicator function. When is not equal to , take 1, otherwise take 0. When a certain substring is empty , the edit distance is the length of the other substring; otherwise, take the minimum value of deletion, insertion, or replacement operations;
[0070] Step S23: Standardize the data format through a preset dictionary library, and the dictionary library includes a standard term library and a unit conversion table.
[0071] For example, the data cleaning process is as follows:
[0072] Integrity verification: For key fields such as "notification number", "effective date", "products involved", etc., if the field missing rate of a certain piece of data exceeds 5%, it is determined as invalid data and eliminated;
[0073] Merging of duplicate records: Use the Levenshtein algorithm to calculate the edit distance of the titles of two notification data. For example, "New EU Regulations on Toy Safety" and "Latest EU Toy Safety Regulations", when the edit distance is less than 20% of the character length, it is determined as duplicate records, and the complete information is merged and retained;
[0074] Format standardization: Unify terms and units through a preset dictionary library. For example, unify "kilogram", "KG", and "kg" into "kg", and unify date formats such as "January 2023" and "2023-01" into "YYYY-MM" to ensure consistent data formats for subsequent processing.
[0075] Optimize the quality of the collected raw data, remove invalid data, merge duplicate records, and standardize the data format. Removing invalid data ensures the integrity and availability of the data. Merging duplicate records avoids data redundancy. Standardizing the data format makes the data consistent, improves the quality of the data, provides a high-quality data foundation for subsequent data mapping and storage operations, and ensures the accuracy and standardization of the data in the database.
[0076] Furthermore, in step S3, the preset basic classification system includes the HS coding system, the national economic industry classification code system, and the product attribute classification system.
[0077] For example, the HS coding system: Adopt the HS coding of the World Customs Organization to classify the product-related data collected.
[0078] National economic industry classification code: Refer to GB / T 4754-2017 and classify the smartphone corresponding to "8517.62" into the industry code of "3962 Communication terminal equipment manufacturing".
[0079] Product attribute classification: Establish an attribute system under the category of "electronic products". For example, "smartphone" corresponds to attributes such as "screen size", "battery capacity", and "network mode" to form a three-level classification framework, providing a unified classification standard for data mapping and ensuring that data from different sources can be classified according to the same dimension.
[0080] Using these widely used and mature classification systems, diverse and heterogeneous data can be classified and managed according to a unified standard, enabling data from different sources and forms to be integrated under the same classification framework, facilitating data organization, retrieval, and analysis, and improving the structural and manageable nature of the data in the database.
[0081] Furthermore, in step S3, the establishment of the data mapping system includes the following steps:
[0082] Step S31: Construct a two-way mapping table between the HS coding and the national economic industry classification code. Match the first 6 digits of the HS coding through regular expressions, and the mapping accuracy requirement is ≥99%.
[0083] Step S32: Calculate the correlation strength between technical indicators and product attributes through the cosine similarity algorithm. The specific formula is: ;
[0084] Among them, is the technical index word vector, is the product attribute word vector, is the dot product of vectors, which measures the consistency of semantic directions; and are the vector norms, which are the results of normalization calculations; when the similarity ≥ 0.8, it is determined that there is a mapping relationship between the technical index and the product attribute.
[0085] Step S33: Generate a credibility score based on the authority level and update frequency of the data source. The calculation formula is: ;
[0086] For example, the two-way mapping table: By matching the first 6 digits of the HS code through regular expressions, a mapping relationship with the national economic industry classification code is established. For example, HS8517 corresponds to "396 Communication Equipment Manufacturing", and the mapping accuracy has been verified to reach 99.2%;
[0087] Association strength calculation: Convert the technical index "operating voltage (5V)" into word vector A, and convert the product attribute "voltage compatibility (5V / 9V)" into word vector B. Calculate through the cosine similarity formula. When the similarity is 0.85 (≥ 0.8), it is determined that there is an association between the two. For example, an association is established between "5V operating voltage" and "5V compatibility";
[0088] Credibility score: Set the authority level of WTO notification data to 100, and the update frequency is once a day. The authority level of the enterprise data platform is 80, and the update frequency is once a week. Calculate the credibility score through the formula, and give priority to using high-credibility data to ensure the reliability of the mapped data.
[0089] The high-precision two-way mapping table realizes the mutual association between the HS code and the national economic industry classification code. The cosine similarity algorithm ensures the reasonable association between the technical index and the product attribute. The credibility scoring mechanism provides a reliability reference for the use of data. The combination of the three enables the data to be accurately mapped and integrated under different classification systems, ensuring the accuracy of data association and the credibility of the data, and facilitating the comprehensive application and analysis of the data.
[0090] Furthermore, in step S4, the distributed storage architecture adopts a hybrid storage mode that combines a relational database and a non-relational database. The structured data is partitioned and stored in the relational database according to the first 6 digits of the HS code, and an index based on the product code and release date is established; the semi-structured data is sharded and stored in the non-relational database according to the data source identifier, and redundant copies are configured.
[0091] For example, the implementation of the distributed storage architecture is as follows:
[0092] Relational database: Use a MySQL cluster to store structured data, partition by the first 6 digits of the HS code, and establish a combined index of "product code" and "release date" to accelerate queries by product and time dimensions;
[0093] Non-relational database: Adopt a MongoDB cluster to store semi-structured data, shard by data source identifier, and configure 3 redundant replicas to ensure high data availability. The hybrid architecture meets the complex query requirements of structured data and the fast storage needs of semi-structured data.
[0094] The hybrid storage mode gives full play to the advantages of relational and non-relational databases. Partitioning and sharding storage improve the efficiency of data storage and retrieval. The establishment of indexes facilitates the rapid search for data. The configuration of redundant replicas enhances the reliability and availability of data, ensuring the efficient storage and management of large-scale data and meeting the requirements of high database availability and high performance.
[0095] Furthermore, in step S4, the preset rules include dividing data sub-databases by the first 4 digits of the HS code, creating tables by data source type, and establishing a time series index based on the UTC timestamp field.
[0096] For example, the storage rules are specifically applied as follows:
[0097] Sub-database strategy: Divide the database by the first 4 digits of the HS code. For example, fruit data starting with "08" is stored in the "db_hs08" database, and electronic product data starting with "85" is stored in the "db_hs85" database. A total of 99 sub-databases are established;
[0098] Table creation strategy: Create tables by data source type within each database. For example, in the "db_hs85" database, create a TBT notification table and a recall data table to distinguish data from different sources;
[0099] Time series index: Establish an index on the "effective date" field to support rapid queries of trade measure updates in the past year or quarter in chronological order and optimize the retrieval efficiency of time series data.
[0100] Dividing data sub-databases by the first 4 digits of the HS code and creating tables by data source type make the data storage structure clearer and facilitate data management and maintenance. Establishing a time series index based on the UTC timestamp is suitable for processing data with time series characteristics, can improve the query efficiency of time-related data, overall optimize the database storage structure, enhance the performance of data management and query, and enable the database to better adapt to the organization and usage requirements of data.
[0101] Further, in step S5, the construction of the knowledge graph of technical trade measures includes the following steps:
[0102] Step S51: Define the entity types of the knowledge graph. The entity types include technical regulations, products, enterprises, and standards. The entity attributes of technical regulations include promulgation date, scope of application, and constraint intensity.
[0103] Step S52: Extract the association relationships between entities. The association relationships are modeled by the TransE algorithm. The vector dimension is 100 dimensions, the boundary hyperparameter γ = 1.0, and the loss function is: ;
[0104] where and are the vector representations of the head entity and the tail entity respectively, is the vector representation of the relationship, is the vector representation of the head entity in the negative sample, is the vector representation of the tail entity in the negative sample;
[0105] Step S53: Calculate the entity importance through the PageRank algorithm. The algorithm damping factor , the number of iterations is 100 times, and the Top-10% entities are selected as the core nodes.
[0106] For example, the process of constructing the knowledge graph is as follows:
[0107] Define entities: Create "technical regulation" entities, "product" entities, "enterprise" entities, and "standard" entities.
[0108] Model association relationships: Use the TransE algorithm, set the vector dimension to 100, the boundary hyperparameter γ = 1.0, define the "applies to" relationship, such as the RoHS directive → applies to → laptop, and optimize the entity vector representation through the loss function to make the "regulation - product" relationship satisfy "h + r ≈ t" in the vector space.
[0109] Screen core nodes: Run the PageRank algorithm, the damping factor d = 0.85, iterate 100 times, calculate the entity importance, select the top 10% entities, such as EU RoHS, US CPSC recall regulations, ISO 9001, etc. as the core nodes, and construct a knowledge graph centered on core regulations and standards for visual analysis of the trade measure association network.
[0110] Convert structured data into a knowledge graph, define entities and their attributes, establish the association relationships between entities, and determine the importance of entities. The construction of the knowledge graph presents data in the form of a graph, clearly showing the relationships between entities such as technical regulations, products, enterprises, and standards. The TransE algorithm modeling accurately represents the associations between entities, and the PageRank algorithm screens core nodes to highlight important entities, facilitating the visualization and in-depth analysis of knowledge related to technical trade measures, providing a basis for intelligent applications based on the knowledge graph, such as complex relationship queries and decision support, etc.
[0111] Furthermore, it also includes step S6: Establish a data dynamic update mechanism, including the following steps:
[0112] Step S61: Set periodic data collection tasks to obtain incremental data from each data source;
[0113] Step S62: Identify updated content through a data comparison algorithm, and the algorithm includes Jaccard similarity calculation, and the formula is: ;
[0114] where, is the incremental data set, is the baseline data set; when the difference rate exceeds the preset threshold, perform incremental updates on the structured database and the knowledge graph;
[0115] Step S63: Generate a data version identifier based on the hash algorithm. The hash algorithm uses SHA-256, and its output is a 256-bit hash value, and associate the identifier with the timestamp and store it in the blockchain to form a traceable historical data chain.
[0116] For example, the dynamic update mechanism is implemented as follows:
[0117] Incremental collection: Set the data collection task to be executed at 2 am every day, obtain the newly added TBT / SPS notifications in the previous 24 hours from the WTO official website, and obtain the latest recall data from the CPSC to form an incremental data set ;
[0118] Difference identification: Use Jaccard similarity calculation to calculate the difference rate, where is the baseline data set. If the difference rate exceeds the preset threshold of 5%, for example, 100 new notification data are added and the difference rate reaches 8%, then trigger an incremental update and only update the newly added and changed records;
[0119] Version Traceability: Generate a 256-bit hash value for the updated data block using the SHA-256 algorithm, associate the hash value with the UTC timestamp, and store it in the Hyperledger Fabric blockchain to form an immutable historical data chain, supporting data version backtracking and auditing.
[0120] Periodically collecting incremental data ensures that the database can obtain the latest information in a timely manner. The difference rate calculation accurately identifies the updated content, enabling precise incremental updates and avoiding the resource waste of full-scale updates. The blockchain stores the version identifier and timestamp to form a traceable historical data chain, ensuring the transparency and immutability of data updates, enabling the database to dynamically adapt to data changes, maintaining the freshness and reliability of data, and providing support for the long-term management and auditing of data.
[0121] The specific embodiments of the present invention have been described in detail above, but they are only examples, and the present invention is not equivalent to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.
Claims
1. A method for constructing a structured database of multi - heterogeneous technical trade measures, characterized in that, It includes the following steps: Step S1: Collect the original data of diverse heterogeneous technical trade measures; Step S2: Clean and process the original data; Step S3: Establish a data mapping system to associate the original data with a preset basic classification system; The preset basic classification system includes the HS coding system, the national economic industry classification code system, and the product attribute classification system; The establishment of the data mapping system includes the following steps: Step S31: Construct a two-way mapping table between the HS coding and the national economic industry classification code. Match the first 6 digits of the HS coding through regular expressions, and the mapping accuracy requirement is ≥99%; Step S32: Calculate the association strength between technical indicators and product attributes through the cosine similarity algorithm. The specific formula is: Similarity = ; Among them, is the technical index word vector, is the product attribute word vector, is the dot product of vectors, measuring the consistency of semantic directions; and are the vector norms, normalizing the calculation results; when the similarity ≥ 0.8, it is determined that there is a mapping relationship between the technical index and the product attribute; Step S33: Generate a credibility score based on the authority level and update frequency of the data source. The calculation formula is: ; Step S4: Construct a distributed storage architecture, and store the cleaned and mapped structured data in the database cluster according to preset rules to form a structured database; Step S5: Construct a knowledge graph of technical trade measures based on the structured database; The construction of the knowledge graph of technical trade measures includes the following steps: Step S51: Define the entity types of the knowledge graph. The entity types include technical regulations, products, enterprises, and standards. The entity attributes of technical regulations include the promulgation date, scope of application, and constraint intensity; Step S52: Extract the association relationships between entities. The association relationships are modeled through the TransE algorithm. The vector dimension is 100 dimensions, the boundary hyperparameter γ = 1.0, and the loss function is: ; Among them, h and t are the vector representations of the head entity and the tail entity respectively, and r is the vector representation of the relation; is the vector representation of the head entity in the negative sample, is the vector representation of the tail entity in the negative sample; is the distance of the positive sample; is the distance of the negative sample; Step S53: Calculate the importance of entities through the PageRank algorithm. The algorithm damping factor d = 0.85, the number of iterations is 100 times, and the Top-10% entities are selected as core nodes; Step S6: Establish a data dynamic update mechanism, including the following steps: Step S61: Set periodic data collection tasks to obtain incremental data from each data source; Step S62: Identify the updated content through a data comparison algorithm. The algorithm includes Jaccard similarity calculation. The formula is: ; where A is the incremental data set and B is the baseline data set; when the difference rate exceeds the preset threshold, perform incremental updates on the structured database and the knowledge graph; Step S63: Generate a data version identifier based on the hash algorithm. The hash algorithm uses SHA-256, and its output is a 256-bit hash value. Associate the identifier with the timestamp and store it in the blockchain to form a traceable historical data chain.
2. The construction method of a structured database for a multi - heterogeneous technical trade measure according to claim 1, characterized in that, In step S1, the original data of diverse heterogeneous technical trade measures includes WTO-TBT notification data, WTO-SPS notification data, European and American product recall data, domestic and foreign technical regulation data, and industrial standard data.
3. The construction method of a structured database for a multi - heterogeneous technical trade measure according to claim 2, characterized in that, In step S1, the methods for collecting raw data include: deploying a web crawler program, configuring a dynamic IP proxy pool and a request interval parameter of 5 seconds, and performing real-time monitoring and crawling of public data sources; connecting to industry databases and enterprise data platforms through an API interface based on the OAuth2.0 protocol, with a data request frequency limit of 100 times per minute; and using a BERT-base model to extract entities from unstructured texts, with entity types including product names and regulatory terms, and an extraction accuracy rate of ≥95%.
4. The construction method of a structured database for multi - heterogeneous technical trade measures according to claim 3, characterized in that, In step S2, the cleaning process includes the following steps: Step S21: Verify data based on data integrity rules, and remove invalid data when the missing rate of key fields is greater than 5%; Step S22: Use the Levenshtein edit distance algorithm to merge duplicate records. The calculation formula is: ; Among them, and are the length indices of the strings and respectively, indicating the minimum edit distance of the first and character substrings; is an indicator function that takes 1 when is not equal to and 0 otherwise. When a certain substring is empty ( ), the edit distance is the length of the other substring; otherwise, it takes the minimum value of deletion, insertion, or replacement operations; Step S23: standardizing the data format through a preset dictionary library, wherein the dictionary library includes a standard term library and a unit conversion table.
5. The construction method of a structured database for a multi - heterogeneous technical trade measure according to claim 4, wherein, In step S4, the distributed storage architecture adopts a hybrid storage mode combining a relational database and a non-relational database. The structured data is partitioned and stored in the relational database according to the first 6 digits of the HS code, and an index based on the product code and release date is established; the semi-structured data is partitioned and stored in the non-relational database according to the data source identifier, and redundant copies are configured.
6. The construction method of a structured database for multi - heterogeneous technical trade measures according to claim 5, characterized in that, In step S4, the preset rules include dividing the data base according to the first 4 digits of the HS code, dividing the table according to the data source type, and establishing a time series index based on the UTC timestamp field.
Citation Information
Patent Citations
Production intelligent decision-making system and method for oil and gas field
CN114676978A
Multi-source heterogeneous data processing method and apparatus, computer device and storage medium
WO2023123182A1