Construction method of multivariate and heterogeneous technical trade measure structured database
Through methods such as multi-disciplinary collection, data cleaning and integration, data mapping and distributed storage, a structured database and knowledge graph of multi-various heterogeneous technology trade measures was constructed, solving the problem of insufficient data integration and analysis capabilities in the existing technology, and achieving efficient and intelligent data management and analysis.
Patent Information
- Application Number
- CN202510680502.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
It is difficult for the existing technology to systematically integrate multivariate and heterogeneous technical trade measures data, lack of correlation analysis capabilities for technical regulations, industrial standards, market feedback and other multi-dimensional correlation analysis capabilities, and the data collection methods are single, data cleaning and integration lack professional classification systems, and the data analysis model is simple, so it is impossible to intelligently analyze the risk level and impact scope of technical barriers.
Multiple collection methods are adopted, including network crawlers, API docking and BERT model entity extraction, data cleaning and integration, a data mapping system is established, the original data is associated with the preset basic classification system, a distributed storage architecture and knowledge graph are built, and a dynamic update mechanism is set.
It has realized efficient integration and quality optimization of data on diversified heterogeneous technology trade measures, ensured the comprehensiveness, accuracy and standardization of the data, improved the availability and analysis capabilities of data, and had the functions of dynamic updates and historical version traceability.
Smart Images

Figure CN120196686A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing for international trade technical barriers, and particularly to a method for constructing a structured database of diverse heterogeneous technical trade measures. Background Art
[0002] With the in-depth development of global trade integration, technical trade measures have become an important factor affecting international trade. In the early stage, the information processing of technical trade measures mainly relied on manual collection and analysis, with low efficiency and limited coverage. With the progress of information technology, countries have gradually established electronic databases based on a single data source, realizing the digital storage of basic data such as WTO-TBT / SPS notifications and product recalls. In recent years, the application of big data and artificial intelligence technologies has promoted the development of automated data collection. For example, public data is obtained through web crawlers, and duplicate data is identified using simple algorithms, initially constructing a framework for structured data storage. However, the existing technology still remains at the stage of isolated processing of single-type data, lacking systematic integration of diverse heterogeneous data and being difficult to form an associative analysis ability covering multiple dimensions such as technical regulations, industrial standards, and market feedback.
[0003] The deficiencies of the existing technology are mainly reflected in three aspects: First, the data collection methods are single, and the ability to obtain unstructured policy documents and enterprise dynamic data is insufficient, resulting in incomplete information coverage; Second, the data cleaning and integration lack a professional classification system for technical trade measures, making it difficult to achieve accurate mapping of product classification and industrial fields, and the data availability is low; Third, the data analysis model is simple, and it can only complete basic data retrieval and statistics, unable to intelligently judge the risk level and influence scope of technical barriers, and lacking the ability of differential early warning for different enterprise entities. In addition, the existing systems generally lack a dynamic update mechanism, making it difficult to track the changes of foreign technical trade measures in real time, resulting in enterprises obtaining lagged information and being unable to effectively respond to the real-time impact of trade barriers. These defects have led to problems such as unsmooth information acquisition channels, insufficient analysis capabilities, and lack of response means for enterprises when facing foreign technical trade measures, and there is an urgent need for a more efficient and intelligent technical solution to solve the problems of processing and application of diverse heterogeneous data. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above problems and provide a method for constructing a structured database of diverse heterogeneous technical trade measures. To achieve the above purpose, the present invention adopts the following technical solutions: A method for constructing a structured database of diverse heterogeneous technical trade measures, comprising the following steps: Step S1: Collect original data of diverse heterogeneous technical trade measures; Step S2: Clean and process the original data; Step S3: Establish a data mapping system to associate the original data with a preset basic classification system; Step S4: Construct a distributed storage architecture, and store the cleaned and mapped structured data in a database cluster according to preset rules to form a structured database; Step S5: Construct a knowledge graph of technical trade measures based on the structured database.
[0005] Further, in Step S1, the original data of diverse heterogeneous technical trade measures includes WTO-TBT notification data, WTO-SPS notification data, European and American product recall data, domestic and foreign technical regulation data, and industrial standard data; Further, in Step S1, the methods for collecting the original data include: deploying a web crawler program, configuring a dynamic IP proxy pool and a request interval parameter of 5 seconds, and performing real-time monitoring and scraping on public data sources; docking with industry databases and enterprise data platforms through API interfaces based on the OAuth2.0 protocol, with a data request frequency limit of 100 times per minute; using the BERT-base model to extract entities from unstructured texts, and the entity types include product names and regulatory clauses, with an extraction accuracy rate ≥ 95%.
[0006] Further, in Step S2, the cleaning process includes the following steps: Step S21: Verify the data based on data integrity rules, and eliminate invalid data when the missing rate of key fields > 5%; Step S22: Use the Levenshtein edit distance algorithm to merge duplicate records. The calculation formula is: Step S21: Verify the data based on data integrity rules, and eliminate invalid data when the missing rate of key fields > 5%; Step S22: Use the Levenshtein edit distance algorithm to merge duplicate records. The calculation formula is: ; where and are the length indexes of the strings and respectively, represents the minimum edit distance of the first and character substrings; is an indicator function, which takes 1 when is not equal to , otherwise it takes 0. When a certain substring is empty , the edit distance is the length of the other substring; otherwise, take the minimum value of deletion, insertion, or replacement operations; Step S23: Standardize the data format through a preset dictionary library, which includes a standard term library and a unit conversion table.
[0007] Further, in step S3, the preset basic classification system includes the HS coding system, the national economic industry classification code system, and the product attribute classification system.
[0008] Further, in step S3, the establishment of the data mapping system includes the following steps: Step S31: Construct a two-way mapping table between HS codes and national economic industry classification codes. Match the first 6 digits of the HS code through regular expressions, and the mapping accuracy requirement is ≥99%; Step S32: Calculate the correlation strength between technical indicators and product attributes through the cosine similarity algorithm. The specific formula is: ; Among them, is the technical indicator word vector, is the product attribute word vector, is the vector dot product, which measures the consistency of semantic directions; and are the vector norms, which are the results of normalization calculations; when the similarity ≥0.8, it is determined that there is a mapping relationship between the technical indicator and the product attribute; Step S33: Generate a credibility score based on the authority level and update frequency of the data source. The calculation formula is: .
[0009] Further, in step S4, the distributed storage architecture adopts a hybrid storage mode that combines a relational database and a non-relational database. Structured data is partitioned and stored in the relational database according to the first 6 digits of the HS code, and an index based on the product code and release date is established; semi-structured data is sharded and stored in the non-relational database according to the data source identifier, and redundant copies are configured.
[0010] Further, in step S4, the preset rules include dividing data sub-libraries according to the first 4 digits of the HS code, dividing tables according to the data source type, and establishing a time series index based on the UTC timestamp field.
[0011] Further, in step S5, the construction of the technical trade measures knowledge graph includes the following steps: Step S51: Define the entity types of the knowledge graph. The entity types include technical regulations, products, enterprises, and standards. The entity attributes of technical regulations include promulgation date, scope of application, and constraint strength; Step S52: Extract the association relationships between entities. The association relationships are modeled through the TransE algorithm, and the vector dimension is 100 dimensions. The boundary hyperparameter , the loss function is as follows: ; Among them, and are the vector representations of the head entity and the tail entity respectively, is the relational vector representation, is the head entity vector representation in the negative sample, is the tail entity vector representation in the negative sample;
[0012] Step S53: Calculate the entity importance through the PageRank algorithm, with the algorithm damping factor , the number of iterations is 100 times, and the Top-10% entities are selected as the core nodes.
[0013] Furthermore, it also includes Step S6: Establish a data dynamic update mechanism, including the following steps: Step S61: Set periodic data collection tasks to obtain incremental data from each data source; Step S62: Identify the updated content through a data comparison algorithm, and the algorithm includes Jaccard similarity calculation, and the formula is: ; Among them, is the incremental data set, is the baseline data set; when the difference rate exceeds the preset threshold, perform incremental updates on the structured database and the knowledge graph; Step S63: Generate a data version identifier based on the hash algorithm. The hash algorithm uses SHA-256, and its output is a 256-bit hash value, and associate the identifier with the timestamp and store it in the blockchain to form a traceable historical data chain.
[0014] The advantages of the present invention are as follows: Through the multi-source acquisition method of deploying web crawlers combined with API docking and entity extraction of the BERT model, and matching the integrity verification, duplicate record merging and format standardization processing in data cleaning, the present invention realizes the efficient integration and quality optimization of multi-source heterogeneous data such as WTO-TBT / SPS notifications, product recall data, and technical regulations, ensuring the comprehensiveness, accuracy and standardization of the original data, and laying a high-quality data foundation for subsequent data mapping and storage.
[0015] By constructing a bidirectional mapping table between HS codes and national economic industry classification codes, calculating the association strength between technical indicators and product attributes through cosine similarity, and combining the hybrid storage architecture of relational and non-relational databases and the strategies of database partitioning, table partitioning and index setting, the present invention realizes the accurate mapping and efficient storage of multi-source data under a unified classification system, which not only meets the complex query requirements of structured data but also adapts to the flexible storage of semi-structured data, improving the management efficiency of the database and the data retrieval performance.
[0016] The present invention realizes the dynamic real - time update of technical trade measure data and the reliable traceability of historical versions by setting periodic incremental data collection tasks, using the Jaccard similarity algorithm to identify updated content and trigger incremental updates, and combining the SHA - 256 hash algorithm with blockchain technology to generate traceable version identifiers, ensuring that the database continuously reflects the latest trade measure dynamics, and at the same time providing an immutable historical data chain support for data auditing and compliance analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings forming a part of this application are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments and their descriptions of this application are used to explain this application and do not constitute an improper limitation of this application.
[0018] In the drawings: Figure 1 It is a flowchart of a method for constructing a multi - heterogeneous structured database of technical trade measures in Embodiment 1.
[0019] Figure 2 It is an architecture diagram of data collection for a method for constructing a multi - heterogeneous structured database of technical trade measures in Embodiment 1.
[0020] Figure 3 It is a schematic diagram of the data cleaning process for a method for constructing a multi - heterogeneous structured database of technical trade measures in Embodiment 1.
[0021] Figure 4 It is a flowchart of the data dynamic update mechanism for a method for constructing a multi - heterogeneous structured database of technical trade measures in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated in the drawings here can be arranged and designed in various different configurations.
[0023] The present invention will be described in detail and specifically below through specific embodiments to better understand the present invention. However, the following embodiments do not limit the protection scope of the present invention.
[0024] Embodiment:
[0025] A method for constructing a multi - heterogeneous structured database of technical trade measures includes the following steps: Step S1: Collecting multivariate and heterogeneous original data on technical trade measures; Step S2: cleaning the original data; Step S3: Establish a data mapping system to associate the original data with the preset basic classification system; Step S4: construct a distributed storage architecture, store the cleaned and mapped structured data into a database cluster according to preset rules, and form a structured database; Step S5: Construct a technical trade measures knowledge graph based on the structured database.
[0026] Through a series of processes, the multi-dimensional and heterogeneous technical trade measures data is processed to achieve the transformation from raw data to structured databases and knowledge graphs. The technical effect produced is that data from different sources and structures can be systematically integrated to form a structured database, and further construct a knowledge graph, which provides an infrastructure for the data management, analysis and application of technical trade measures, and facilitates the subsequent in-depth mining and utilization of relevant data.
[0027] Further, in step S1, the multi-heterogeneous original data of technical trade measures include WTO-TBT notification data, WTO-SPS notification data, European and American product recall data, domestic and foreign technical regulations data and industrial standards data; For example, during the data collection stage, the following specific data are explicitly included: TBT notification data and SPS notification data are obtained from the WTO official website; product recall data are captured from the U.S. Consumer Product Safety Commission official website; national standards issued by the Standardization Administration of China and the EU RoHS directive are collected; and industry standards issued by the International Organization for Standardization are simultaneously included to ensure that the data covers international, domestic and foreign multi-source heterogeneous scenarios.
[0028] It ensures that the collected data covers internationally accepted notification data, foreign product recall information, domestic and foreign technical regulations and industry standards, etc. The data types are rich and targeted, providing rich original data for the subsequent construction of a comprehensive and authoritative technical trade measures database, which can meet various analysis and research needs.
[0029] Furthermore, in step S1, the method of collecting raw data includes: deploying a web crawler program, configuring a dynamic IP proxy pool and a request interval parameter of 5 seconds, and performing real-time monitoring and crawling of public data sources; connecting to industry databases and enterprise data platforms through an API interface based on the OAuth2.0 protocol, with a data request frequency limit of 100 times per minute; and using a BERT-base model to extract entities from unstructured text, where entity types include product names and regulatory terms, and the extraction accuracy is ≥95%.
[0030] For example, data collection is implemented as follows: Web crawler: Use Python's Scrapy framework to deploy crawlers, configure the dynamic IP pool of Xigua proxy, set the request interval to 5 seconds, conduct real-time monitoring of the WTO-TBT / SPS notification official website, crawl HTML pages and parse structured data; API connection: Access the EBSCO industry database and a multinational enterprise quality data platform through the OAuth2.0 protocol, set a request frequency limit of 100 times per minute, and obtain standardized internal enterprise technical standards and recall records; Entity extraction: Based on Hugging Face's BERT-base model, unstructured regulatory texts are processed and two types of entities, product names and regulatory clauses, are defined. By fine-tuning the model, the extraction accuracy rate reaches 96%, ensuring accurate extraction of key information from unstructured texts.
[0031] According to different data sources and data forms, appropriate technical means are used for efficient and accurate data collection. Dynamic IP proxy pool and request interval parameter settings can avoid access restrictions by data sources and ensure the stability of real-time monitoring and crawling; API docking realizes reliable connection with industry databases and enterprise data platforms to obtain professional data; BERT-base model's high-accuracy entity extraction can effectively extract key information from unstructured text. The combination of multiple methods ensures the comprehensiveness, efficiency and accuracy of original data collection.
[0032] Further, in step S2, the cleaning process includes the following steps: Step S21: Verify data based on data integrity rules, and remove invalid data when the missing rate of key fields is greater than 5%; Step S22: Use the Levenshtein edit distance algorithm to merge duplicate records. The calculation formula is: ; in, and Separate strings and The length index of Before and The minimum edit distance of character substrings; is the indicator function, when and If they are not equal, it takes 1, otherwise it takes 0. When , the edit distance is the length of another substring; otherwise, the minimum value of the deletion, insertion or replacement operation is taken; Step S23: Standardize the data format through a preset dictionary library, which includes a standard term library and a unit conversion table.
[0033] For example, the data cleaning process is as follows: Integrity check: For keyword fields such as "notification number", "effective date", "products involved", etc., if the field missing rate of a certain piece of data exceeds 5%, it is determined as invalid data and excluded; Duplicate record merging: Use the Levenshtein algorithm to calculate the edit distance of the titles of two notification data. For example, for "New EU Regulations on Toy Safety" and "Latest EU Toy Safety Regulations", when the edit distance is less than 20% of the character length, it is determined as a duplicate record, and the complete information is merged and retained; Format standardization: Unify terms and units through a preset dictionary library. For example, unify "kilogram", "KG", "kg" to "kg", and unify date formats such as "January 2023", "2023-01" to "YYYY-MM" to ensure consistent data formats for subsequent processing.
[0034] Optimize the quality of the collected raw data, remove invalid data, merge duplicate records, and unify the data format. Removing invalid data ensures the integrity and availability of the data. Merging duplicate records avoids data redundancy. Standardizing the data format makes the data consistent, improves the quality of the data, provides a high-quality data foundation for subsequent data mapping and storage operations, and ensures the accuracy and standardization of the data in the database.
[0035] Furthermore, in step S3, the preset basic classification system includes the HS coding system, the national economic industry classification code system, and the product attribute classification system.
[0036] For example, the HS coding system: Adopt the HS coding of the World Customs Organization to classify the collected product-related data; National economic industry classification code: Refer to GB / T 4754-2017 and classify the smartphone corresponding to "8517.62" into the industry code of "3962 Communication Terminal Equipment Manufacturing"; Product attribute classification: Establish an attribute system under the category of "electronic products". For example, "smartphone" corresponds to attributes such as "screen size", "battery capacity", "network mode", etc., forming a three-level classification framework to provide a unified classification standard for data mapping and ensure that data from different sources can be classified according to the same dimension.
[0037] Using these widely applied and mature classification systems, heterogeneous data can be classified and managed according to a unified standard, enabling data from different sources and in different forms to be integrated within the same classification framework, facilitating data organization, retrieval, and analysis, and enhancing the structure and manageability of the data in the database.
[0038] Further, in step S3, the establishment of the data mapping system includes the following steps: Step S31: Construct a two-way mapping table between HS codes and national economic industry classification codes. Match the first 6 digits of the HS code through regular expressions, with a mapping accuracy requirement of ≥99%. Step S32: Calculate the correlation strength between technical indicators and product attributes through the cosine similarity algorithm. The specific formula is: ; where, is the technical indicator word vector, is the product attribute word vector, is the vector dot product, measuring the consistency of semantic directions; and are the vector norms, the results of normalized calculation; when the similarity ≥0.8, it is determined that there is a mapping relationship between the technical indicator and the product attribute.
[0039] Step S33: Generate a credibility score based on the authority level and update frequency of the data source. The calculation formula is: ; For example, two-way mapping table: Match the first 6 digits of the HS code through regular expressions to establish a mapping relationship with the national economic industry classification code. For example, HS8517 corresponds to "396 Communication Equipment Manufacturing", and the mapping accuracy is verified to reach 99.2%; Calculation of correlation strength: Convert the technical indicator "operating voltage (5V)" into word vector A, and convert the product attribute "voltage compatibility (5V / 9V)" into word vector B. Calculate through the cosine similarity formula. When the similarity is 0.85 (≥0.8), it is determined that there is a correlation between the two, such as establishing a correlation between "5V operating voltage" and "5V compatibility"; Credibility score: Set the authority level of WTO notification data to 100, with an update frequency of once a day, and the authority level of the enterprise data platform to 80, with an update frequency of once a week. Calculate the credibility score through the formula, and give priority to using high-credibility data to ensure the reliability of the mapped data.
[0040] A high-precision two-way mapping table realizes the mutual association between HS codes and the classification codes of industries in the national economy. The cosine similarity algorithm ensures the reasonable association between technical indicators and product attributes. The credibility scoring mechanism provides a reliable reference for the use of data. The combination of the three enables data to be accurately mapped and integrated under different classification systems, ensuring the accuracy of data association and the credibility of data, and facilitating the comprehensive application and analysis of data.
[0041] Further, in step S4, the distributed storage architecture adopts a hybrid storage mode combining a relational database and a non-relational database. Structured data is partitioned and stored in the relational database according to the first 6 digits of the HS code, and an index based on the product code and the release date is established; semi-structured data is sharded and stored in the non-relational database according to the data source identifier, and redundant copies are configured.
[0042] For example, the distributed storage architecture is implemented as follows: Relational database: Use a MySQL cluster to store structured data, partition it according to the first 6 digits of the HS code, and establish a combined index of "product code" and "release date" to accelerate queries by product and time dimensions; Non-relational database: Adopt a MongoDB cluster to store semi-structured data, shard it according to the data source identifier, and configure 3 redundant copies to ensure high data availability. The hybrid architecture meets the complex query requirements of structured data and the fast storage requirements of semi-structured data.
[0043] The hybrid storage mode gives full play to the advantages of relational databases and non-relational databases. Partitioning and sharding storage improve the efficiency of data storage and retrieval. The establishment of indexes facilitates the rapid search for data. The configuration of redundant copies enhances the reliability and availability of data, ensuring the efficient storage and management of large-scale data and meeting the requirements of high database availability and high performance.
[0044] Further, in step S4, the preset rules include dividing data sub-databases according to the first 4 digits of the HS code, creating tables according to the data source type, and establishing a time series index based on the UTC timestamp field.
[0045] For example, the storage rules are specifically applied as follows: Sub-database strategy: Divide the database according to the first 4 digits of the HS code. For example, fruit data starting with "08" is stored in the "db_hs08" database, and electronic product data starting with "85" is stored in the "db_hs85" database. A total of 99 sub-databases are established; Table creation strategy: Create tables according to the data source type in each database. For example, in the "db_hs85" database, create a TBT notification table and a recall data table to distinguish data from different sources; Time series index: Index the "Effective Date" field to support fast chronological queries for trade measure updates in the past year or quarter, optimizing the retrieval efficiency of time series data.
[0046] Divide the data sub-library by the first 4 digits of the HS code and create separate tables according to the data source type, making the data storage structure clearer and facilitating data management and maintenance; establish a time series index based on UTC timestamps, which is suitable for processing data with time series characteristics, can improve the query efficiency of time-related data, optimize the overall database storage structure, enhance the performance of data management and query, and enable the database to better adapt to the organization and usage requirements of data.
[0047] Furthermore, in step S5, the construction of the technology trade measure knowledge graph includes the following steps: Step S51: Define the entity types of the knowledge graph, which include technical regulations, products, enterprises, and standards. The entity attributes of technical regulations include promulgation date, scope of application, and constraint intensity; Step S52: Extract the association relationships between entities. The association relationships are modeled by the TransE algorithm, with a vector dimension of 100, a boundary hyperparameter γ = 1.0, and the loss function is: ; Among them, and are the vector representations of the head entity and the tail entity respectively, is the vector representation of the relationship, is the vector representation of the head entity in the negative sample, is the vector representation of the tail entity in the negative sample;
[0048] Step S53: Calculate the entity importance through the PageRank algorithm, with the algorithm damping factor , and the number of iterations is 100 times. Screen the top 10% entities as core nodes.
[0049] For example, the process of constructing the knowledge graph is as follows: Define entities: Create "technical regulation" entities, "product" entities, "enterprise" entities, and "standard" entities; Model the association relationships: Use the TransE algorithm, set the vector dimension to 100, the boundary hyperparameter γ = 1.0, define the "applies to" relationship, such as the RoHS directive → applies to → laptops, and optimize the entity vector representations through the loss function so that the "regulation - product" relationship satisfies "h + r ≈ t" in the vector space; Core node screening: Run the PageRank algorithm with a damping factor d = 0.85 and iterate 100 times to calculate the importance of entities. Select the top 10% of entities, such as EU RoHS, US CPSC recall regulations, ISO 9001, etc. as core nodes, and construct a knowledge graph centered on core regulations and standards to facilitate visual analysis of the trade measure association network.
[0050] Convert structured data into a knowledge graph, define entities and their attributes, establish the association relationships between entities, and determine the importance of entities. The construction of the knowledge graph presents the data in the form of a graph, clearly showing the relationships between entities such as technical regulations, products, enterprises, and standards. The TransE algorithm modeling accurately represents the associations between entities, and the PageRank algorithm screens core nodes to highlight important entities, facilitating the visualization and in-depth analysis of technical trade measure-related knowledge, providing a basis for intelligent applications based on the knowledge graph, such as complex relationship queries and decision support.
[0051] Furthermore, it also includes step S6: Establish a data dynamic update mechanism, including the following steps: Step S61: Set periodic data collection tasks to obtain incremental data from each data source; Step S62: Identify updated content through a data comparison algorithm, and the algorithm includes Jaccard similarity calculation, and the formula is: ; Among them, is the incremental data set, is the baseline data set; when the difference rate exceeds the preset threshold, perform incremental updates on the structured database and the knowledge graph; Step S63: Generate a data version identifier based on the hash algorithm. The hash algorithm uses SHA-256, and its output is a 256-bit hash value, and associate the identifier with the timestamp and store it in the blockchain to form a traceable historical data chain.
[0052] For example, the dynamic update mechanism is implemented as follows: Incremental collection: Set the data collection task to be executed at 2 am every day, obtain the newly added TBT / SPS notifications in the previous 24 hours from the WTO official website, and obtain the latest recall data from CPSC to form an incremental data set ; Difference identification: Use Jaccard similarity to calculate the difference rate, where is the baseline data set. If the difference rate exceeds the preset threshold of 5%, such as adding 100 notification data and the difference rate reaches 8%, then trigger incremental updates and only update the newly added and changed records; Version tracing: Generate a 256-bit hash value for the updated data block using the SHA-256 algorithm, associate the hash value with the UTC timestamp, and store it in the Hyperledger Fabric blockchain to form an immutable historical data chain, supporting data version backtracking and auditing.
[0053] Periodically collecting incremental data ensures that the database can obtain the latest information in a timely manner. The difference rate calculation accurately identifies the updated content, realizes precise incremental updates, and avoids the waste of resources caused by full-scale updates. The blockchain stores version identifiers and timestamps to form a traceable historical data chain, ensuring the transparency and immutability of data updates, enabling the database to dynamically adapt to data changes, maintaining the freshness and reliability of data, and providing support for the long-term management and auditing of data.
[0054] The specific embodiments of the present invention have been described in detail above, but they are only examples, and the present invention is not equivalent to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.
Claims
1. A method for constructing a structured database of multi - heterogeneous technical trade measures, characterized in that, The following steps are involved: Step S1: Collecting multivariate and heterogeneous original data on technical trade measures; Step S2: cleaning the original data; Step S3: Establish a data mapping system to associate the original data with the preset basic classification system; Step S4: construct a distributed storage architecture, store the cleaned and mapped structured data into a database cluster according to preset rules, and form a structured database; Step S5: Construct a technical trade measures knowledge graph based on the structured database.
2. The construction method of a structured database for a multi - heterogeneous technical trade measure according to claim 1, characterized in that, In step S1, the multi-heterogeneous original data on technical trade measures include WTO-TBT notification data, WTO-SPS notification data, European and American product recall data, domestic and foreign technical regulations data and industrial standards data.
3. The construction method of a structured database for multi - heterogeneous technical trade measures according to claim 2, characterized in that, In step S1, the methods for collecting raw data include: deploying a web crawler program, configuring a dynamic IP proxy pool and a request interval parameter of 5 seconds, and performing real-time monitoring and crawling of public data sources; connecting to industry databases and enterprise data platforms through an API interface based on the OAuth2.0 protocol, with a data request frequency limit of 100 times per minute; and using a BERT-base model to extract entities from unstructured texts, with entity types including product names and regulatory terms, and an extraction accuracy rate of ≥95%.
4. A method for constructing a structured database of multi - heterogeneous technical trade measures according to claim 3, characterized in that, In step S2, the cleaning process includes the following steps: Step S21: Verify data based on data integrity rules, and remove invalid data when the missing rate of key fields is greater than 5%; Step S22: Merge duplicate records using the Levenshtein edit distance algorithm, and the calculation formula is: ; Among them, and are the length indexes of the strings and respectively, indicating the minimum edit distance of the first and character substrings; is an indicator function, taking 1 when is not equal to and 0 otherwise. When a certain substring is empty , the edit distance is the length of the other substring; otherwise, take the minimum value of deletion, insertion, or replacement operations; Step S23: standardizing the data format through a preset dictionary library, wherein the dictionary library includes a standard term library and a unit conversion table.
5. The construction method of a structured database for multi - heterogeneous technical trade measures according to claim 4, characterized in that, In step S3, the preset basic classification system includes the HS coding system, the national economic industry classification code system and the product attribute classification system.
6. The construction method of a structured database for multi - heterogeneous technical trade measures according to claim 5, characterized in that, In step S3, the establishment of the data mapping system includes the following steps: Step S31: construct a bidirectional mapping table between HS codes and national economic industry classification codes, match the first 6 digits of the HS codes through regular expressions, and the mapping accuracy is required to be ≥ 99%; Step S32: Calculate the correlation strength between technical indicators and product attributes through the cosine similarity algorithm. The specific formula is: ; Among them, is the technical index word vector, is the product attribute word vector, is the dot product of vectors, measuring the semantic direction consistency; and are the vector norms, normalizing the calculation results; when the similarity ≥ 0.8, it is determined that there is a mapping relationship between the technical index and the product attribute; Step S33: Generate a credibility score based on the authority level and update frequency of the data source. The calculation formula is as follows: .
7. A method for constructing a structured database of multi - heterogeneous technical trade measures according to claim 6, characterized in that, In step S4, the distributed storage architecture adopts a hybrid storage mode combining a relational database and a non-relational database. The structured data is partitioned and stored in the relational database according to the first 6 digits of the HS code, and an index based on the product code and release date is established; the semi-structured data is partitioned and stored in the non-relational database according to the data source identifier, and redundant copies are configured.
8. The construction method of a structured database for multi - heterogeneous technical trade measures according to claim 7, characterized in that, In step S4, the preset rules include dividing the data base according to the first 4 digits of the HS code, dividing the table according to the data source type, and establishing a time series index based on the UTC timestamp field.
9. The construction method of a structured database for multi - heterogeneous technical trade measures according to claim 8, characterized in that, In step S5, the construction of the technical trade measures knowledge graph includes the following steps: Step S51: defining entity types of the knowledge graph, wherein the entity types include technical regulations, products, enterprises and standards, and the technical regulations entity attributes include promulgation date, scope of application and binding strength; Step S52: Extract the association relationships between entities. The association relationships are modeled by the TransE algorithm, with a vector dimension of 100 and boundary hyperparameters , and the loss function is: ; Among them, and are the vector representations of the head entity and the tail entity respectively, is the vector representation of the relationship, is the vector representation of the head entity in the negative sample, is the vector representation of the tail entity in the negative sample; Step S53: Calculate the importance of entities through the PageRank algorithm, and the damping factor of the algorithm , with the number of iterations being 100 times, and screen the top 10% of entities as core nodes.
10. The construction method of a structured database for multi - heterogeneous technical trade measures according to claim 9, characterized in that, The step S6 is also included: establishing a data dynamic update mechanism, including the following steps: Step S61: Setting a periodic data collection task to obtain incremental data from each data source; Step S62: Identify the updated content through a data comparison algorithm, and the algorithm includes Jaccard similarity calculation, and the formula is: ; Among them, is the incremental data set, is the baseline data set; when the difference rate exceeds the preset threshold, the structured database and the knowledge graph are incrementally updated; Step S63: Generate a data version identifier based on a hash algorithm. The hash algorithm uses SHA-256, and its output is a 256-bit hash value. Then, associate the identifier with a timestamp and store them in the blockchain to form a traceable historical data chain.
Citation Information
Patent Citations
Production intelligent decision-making system and method for oil and gas field
CN114676978A
Multi-source heterogeneous data processing method and apparatus, computer device and storage medium
WO2023123182A1