Data transmission method and device, equipment and medium
By improving the data transmission method between relational and analytical databases, acquiring and converting log types, generating target database tables, and performing partitioning, the problems of data transmission efficiency and accuracy are solved, achieving efficient and stable data synchronization, which is suitable for real-time data analysis in the medical and financial fields.
Patent Information
- Application Number
- CN202511060281.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing relational databases are inefficient and inaccurate in data transmission, making it difficult to meet the needs of scenarios such as real-time data analysis and emergency medical decision-making. In particular, in the medical and financial fields, data transmission delays may lead to the loss of critical information.
By obtaining the log type of the source database, converting it to the log type of the target database, generating the target database table, and performing data parsing and partitioning, synchronous data transmission is achieved. Intelligent parsing and dynamic partitioning technologies, combined with priority-driven and fuzzy compatibility processing, ensure data integrity and accuracy.
It significantly improves the efficiency and accuracy of data transmission, ensures that critical data is processed first, reduces transmission latency and error rate, and provides efficient and stable data synchronization support for distributed systems and real-time data analysis.
Smart Images

Figure CN120873085A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a data transmission method, apparatus, device, and medium. Background Technology
[0002] In today's digital age, relational databases, as key data storage tools, are widely used in various IT systems such as e-commerce, banking, and healthcare. They have advantages such as data consistency and fast read and write speeds, but they are difficult to conduct large-scale joint data analysis. Therefore, data needs to be transferred from relational databases to dedicated analytical databases during data transmission, which affects the efficiency and quality of data transmission.
[0003] For example, in the process of digitalizing healthcare, relational databases have become an important tool for storing critical data such as patient information and medical records. However, medical data is massive and complex. When integrating medical data from different departments and hospitals for disease research or medical quality assessment, the data needs to be transmitted to a specialized analytical database, which is difficult to meet the needs of scenarios such as real-time disease monitoring and emergency medical decision-making. At the same time, when the fields of a relational database table are changed, the synchronization transmission program needs to be rewritten. If the synchronization task fails, it cannot be automatically restarted and requires manual intervention. In emergency medical situations, this may delay the data transmission process and affect the effectiveness of medical treatment.
[0004] For example, in the fintech field, relational databases are the cornerstone for storing core data such as customer information and transaction records. However, financial data is massive and growing rapidly, making it difficult for relational databases to perform large-scale joint analysis. This makes them unable to meet the needs of scenarios such as real-time risk control and high-frequency trading, which may lead to delayed analysis of key data and increase financial risks.
[0005] Therefore, improving the accuracy and efficiency of data transmission has become an urgent problem to be solved. Summary of the Invention
[0006] This invention provides a data transmission method, apparatus, device, and medium, the main purpose of which is to solve the problems of low data transmission efficiency and inaccuracy.
[0007] In a first aspect, to achieve the above objectives, the present invention provides a data transmission method comprising: Obtain the source database and identify the source log type of the source log data in the source database; Convert the source log data type to the preset target log type of the target database; Obtain the target table configuration information corresponding to the target database, and generate the target database table according to the target log type and the target table configuration information; The source log data is parsed to obtain multiple target data, and partitions are created for the target database based on the multiple target data to obtain target data blocks; Data is synchronously transmitted to the target database based on the target database table and the target data block to obtain target synchronous transmission data.
[0008] In a second aspect, the present invention also provides a data transmission device, comprising: The log type identification module is used to obtain the source database and identify the source log type of the source log data in the source database; The log type conversion module is used to convert the source log data type into a preset target log type of the target database. The database table generation module is used to obtain the target table configuration information corresponding to the target database, and generate the target database table according to the target log type and the target table configuration information. The data block construction module is used to parse the source log data to obtain multiple target data, and create partitions for the target database based on the multiple target data to obtain target data blocks; The data synchronization and transmission module is used to perform data synchronization and transmission on the target database according to the target database table and the target data block to obtain target synchronization and transmission data.
[0009] Thirdly, the present invention also provides an electronic device, the electronic device comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the data transmission method described above.
[0010] Fourthly, the present invention also provides a computer-readable storage medium storing at least one computer program, which is executed by a processor in an electronic device to implement the above-described data transmission method.
[0011] In this embodiment of the invention, source data is directly extracted by accurately locating the log storage area, avoiding information loss due to intermediate conversion, ensuring data integrity, prioritizing the conversion of critical data, and improving processing accuracy. A priority-driven matching and search mechanism combined with fuzzy compatibility processing enhances the flexibility of type conversion, avoids data loss due to format differences, and ensures the uniqueness and consistency of the target log type. Dynamic table generation technology significantly improves the efficiency and flexibility of database log management. Deep parsing of the target log type can automatically extract structured features, and the normalization generated by the target table configuration can accurately adapt to the storage specifications of different databases, ensuring compatibility. Intelligent log parsing and dynamic partitioning technology significantly improve log management efficiency. To improve efficiency and data availability, the partitioning strategy is dynamically adapted based on field statistical analysis during the partitioning creation phase. Combined with priority sorting and capacity threshold control, this ensures balanced data distribution and prioritizes critical data processing, optimizing query performance and storage utilization. This provides efficient and reliable underlying support for log analysis, troubleshooting, and security auditing. Through structured synchronization and intelligent transmission strategies, the efficiency and reliability of database data synchronization are significantly improved, ensuring priority transmission of critical data and efficient use of network bandwidth. Structured synchronization control enables automatic alignment between the source and target ends, and checksum verification ensures data consistency, effectively reducing transmission latency and error rates. This provides efficient and stable data synchronization support for distributed systems, real-time data analysis, and other scenarios. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of an application environment for a data transmission method according to an embodiment of the present invention; Figure 2 This is a schematic flowchart of a data transmission method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the process of synchronizing and transmitting data to a target database based on a target database table and a target data block, according to an embodiment of the present invention. Figure 4 This is a schematic diagram of a data transmission device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device that implements a data transmission method according to an embodiment of the present invention; Figure 6This is another schematic diagram of an electronic device for implementing a data transmission method according to an embodiment of the present invention.
[0014] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0015] To enable those skilled in the art to better understand the technical solutions of this disclosure, and to fully understand and implement the process of how this disclosure applies technical means to solve technical problems and achieve corresponding technical effects, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, not all embodiments. The embodiments of this disclosure and the various features within them can be combined with each other without conflict, and the resulting technical solutions are all within the protection scope of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort should fall within the protection scope of this disclosure.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0017] This application provides a data transmission method. The execution subject of the data transmission method includes, but is not limited to, at least one of electronic devices that can be configured to execute the device provided in this application, such as a server or a terminal. In other words, the data transmission method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0018] This invention discloses a data transmission method that can be applied to applications such as... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain the source and target databases from the client, directly extracting source data by accurately locating the log storage area, avoiding information loss due to intermediate conversion and ensuring data integrity; using unified encoding to convert logs into standardized data strings eliminates multi-format compatibility issues and significantly reduces manual classification costs; through hierarchical splitting and priority marking, complex log types are refined into manageable units, ensuring that key data is converted first, improving processing accuracy; the priority-driven matching and search mechanism combined with fuzzy compatibility processing enhances the flexibility of type conversion, avoids data loss due to format differences, and ensures the uniqueness and consistency of the target log type; dynamic table generation technology significantly improves the efficiency and flexibility of database log management, and deep analysis of the target log type can automatically extract structured features, and the normalization generated by the target table configuration can accurately adapt to the storage specifications of different databases, ensuring compatibility. To avoid partition failures due to null values and reduce development costs for log migration and querying, intelligent log parsing and dynamic partitioning technologies significantly improve log management efficiency and data availability. During partition creation, the partitioning strategy is dynamically adapted based on field statistical analysis, combined with priority sorting and capacity threshold control to ensure balanced data distribution and prioritized processing of critical data, optimizing query performance and storage utilization. This provides efficient and reliable underlying support for log analysis, troubleshooting, and security auditing. Furthermore, structured synchronization and intelligent transmission strategies significantly improve the efficiency and reliability of database data synchronization. The system automatically parses the target table structure and accurately locates the data blocks to be synchronized, avoiding performance losses from full table scans. Checksums are generated based on data records and encapsulated in transmission packets. Combined with dynamic routing and priority scheduling strategies, critical data is prioritized for transmission and network bandwidth is used efficiently. Structured synchronization control enables automatic alignment between the source and target ends, and checksum verification ensures data consistency, effectively reducing transmission latency and error rates. This provides efficient and stable data synchronization support for distributed systems, real-time data analysis, and other scenarios. Finally, the target synchronized data is output back to the client. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0019] Reference Figure 2 The diagram shown is a flowchart illustrating a data transmission method according to an embodiment of the present invention. In this embodiment, the data transmission method includes: S1. Obtain the source database and identify the source log type of the source log data in the source database.
[0020] In this embodiment of the invention, the source database is the starting point for data migration or synchronization, stores the original data, and is the database currently used by the business system.
[0021] In detail, confirm the purpose of the synchronization transmission (such as system upgrade, cloud migration), and sort out the source database type (such as MySQL, Oracle), version, data volume and network environment. At the same time, determine the architecture (such as distributed, columnar storage) and compatibility requirements of the target database. Connect directly to the source database through database management tools (such as DBeaver, Navicat) or command line (such as the mysql client of MySQL), and use command lines to view the database list and table structure. For the target database, create an empty instance in advance and obtain its connection string, port and authentication information.
[0022] Furthermore, ETL tools (such as Informatica and Talend) or database built-in export functions (such as MySQL's mysqldump and Oracle's Data Pump) can be used to extract the source data structure and data, and generate SQL scripts or binary files. For large-scale data, table partitioning or parallel extraction techniques (such as Spark SQL) can be used to improve efficiency.
[0023] For example, in a healthcare scenario, acquiring source and target databases requires balancing data sensitivity and business continuity. For instance, a top-tier hospital plans to migrate patient records, examination reports, and other data from its old HIS system (source database: Oracle) to a new cloud HIS (target database: MySQL cluster). During implementation, structured data is first exported via the hospital's intranet using Oracle Data Pump. Simultaneously, ETL tools (such as Informatica) are used to clean and anonymize sensitive fields (such as patient ID numbers). Then, encrypted data is imported into the cloud MySQL via a dedicated line. Unstructured data (such as DICOM images) is transferred in batches to cloud object storage using distributed storage migration tools (such as AWS Snowball). Finally, a relational index is established in the target database to ensure real-time data synchronization during the parallel operation of the old and new systems. After the migration, the integrity of 100,000 medical records is compared through sampling.
[0024] For example, in fintech scenarios, data migration must meet compliance and millisecond-level latency requirements. For instance, a bank migrated account information and transaction logs from its core transaction system (source database being DB2) to a distributed database (such as TiDB). During implementation, the table structure was first exported using DB2's EXPORT command, and incremental data was captured by parsing transaction logs using a self-developed middleware (such as IBMCDC) and synchronized to a Kafka message queue. The target database used the TiDB Lightning tool to batch load historical data, while simultaneously synchronizing incremental changes in real time based on Binlog. Before the migration, the balances of 5 million accounts were verified, and a dual-write mechanism was used during the migration to ensure zero transaction loss. Finally, the performance of the target database was verified through TPC-C benchmark testing to ensure a seamless switch for financial services.
[0025] In this embodiment of the invention, identifying the source log type of the source log data in the source database includes: Extract source log data from the log storage area in the source database; The source log data is encoded to obtain a log data string; Obtain a sliding window, and perform window truncation on the log data string according to the sliding window to obtain a log data substring; The log data substring is pattern matched according to a preset log type template library to obtain the source log data type corresponding to the source log data.
[0026] In detail, the log storage location can be located through the database management system (such as MySQL's error log or Oracle's alert log) or file system path (such as Linux's / var / log directory). The contents of the log file can be directly extracted using built-in database tools (such as using the SELECT statement to query system tables) or file reading interfaces (such as Python's open() function) to ensure that the original data structure is not damaged.
[0027] Specifically, non-text data (such as binary logs) in log files are converted to a unified encoding format (such as UTF-8) to avoid garbled characters. If the logs are structured data (such as JSON format), they are converted into a string sequence using a parsing tool (such as Python's json.loads()). If they are unstructured text, they are directly concatenated line by line into a continuous string.
[0028] Furthermore, a fixed-length window (e.g., 100 characters) is set, and the window slides gradually from the beginning of the log data string, moving one character or a fixed step (e.g., 10 characters) each time to extract the substring within the window; if the length of the log data string is less than the window size, the remaining part is directly extracted as the substring to avoid data loss.
[0029] Specifically, feature patterns for different log types are predefined (e.g., error logs contain the keyword ERROR, access logs contain GET / POST requests), stored as regular expressions or keyword sets, and the substring structure is validated using regular expression rules in the template. If a substring matches a template, it is marked as the corresponding log type; otherwise, it is classified as an "unknown type" or the default type.
[0030] In this embodiment of the invention, source data is directly extracted by accurately locating the log storage area, avoiding information loss due to intermediate conversion and ensuring data integrity; unified encoding is used to convert logs into standardized data strings, eliminating multi-format compatibility issues and improving subsequent processing efficiency; the sliding window mechanism dynamically truncates substrings to flexibly adapt to log segments of different lengths, reducing unnecessary calculations; combined with pattern matching technology from a preset template library, log types (such as error, access, and operation logs) can be quickly identified, significantly reducing manual classification costs.
[0031] S2. Convert the source log data type to the preset target log type of the target database.
[0032] In this embodiment of the invention, the target database is the endpoint for data migration or synchronization. It is used to receive and store the migrated data. It is a new system, cloud database, or backup library, ensuring that data smoothly transitions from the old environment to the new environment.
[0033] In detail, the source log data type is converted into a target log type compatible with the target database through hierarchical splitting, priority marking, mapping rule matching, and conflict detection.
[0034] In this embodiment of the invention, converting the source log data type into a preset target database target log type includes: The source log data type is hierarchically split to obtain multiple log-level units, and type conversion priority markers are added to the multiple log-level units; Based on the source log data type and the preset target database log type mapping rule base, an initial type mapping relationship is constructed; Based on the initial type mapping relationship, the log-level units with added type conversion priority tags are sequentially searched for target types to obtain the initial compatible types; A bidirectional conflict type detection is performed on the initial compatible type to obtain the target log type of the target database.
[0035] The bidirectional conflict detection includes: forward verification (whether the target type meets the source data constraints) and reverse verification (whether the converted data is compatible with the source library).
[0036] In detail, the source log types are decomposed into structural levels to facilitate fine-grained processing. Priorities are assigned to each unit according to business needs or transformation complexity to ensure that critical data is transformed first. The correspondence between source log types (such as MySQL's slow query logs) and target database types (such as PostgreSQL's execution timeout logs) is predefined and stored as key-value pairs or rule tables. By querying the rule database, the source log types are matched with the types supported by the target database to generate an initial mapping table (such as "error logs → exception record table").
[0037] Furthermore, log-level units are processed in priority order (e.g., high-priority error codes are processed first, followed by low-priority user IDs), and compatible types supported by the target library are searched level by level. If strict matching fails, candidate compatible types are generated using approximate matching rules in the rule base.
[0038] This involves checking whether the same source log unit matches multiple target types, or whether different source units are mapped to the same target type. Based on a preset strategy (such as highest priority or field uniqueness constraint), a unique target type is selected to ensure that data is written to the target database without ambiguity.
[0039] For example, in the field of healthcare, hospital electronic medical record systems generate massive amounts of log data every day, covering different business processes such as patient information entry, diagnosis record updates, and medical order issuance. These log data types are diverse and complex in structure. When it is necessary to convert the source log data types of the existing system into the target log types of the new medical big data platform (preset target database), the source log data types are first split into hierarchical types.
[0040] For example, logs related to patient diagnosis records are broken down into log-level units such as patient basic information units, diagnosis disease information units, and diagnosis time information units. Type conversion priority markers are added to these units, with the patient basic information unit having the highest priority to ensure that key information is converted first. Next, based on the source log data type and the log type mapping rule library of the new medical big data platform, an initial type mapping relationship is constructed to clarify the corresponding type of different log-level units in the new platform. Then, according to the initial type mapping relationship, the log-level units with priority markers are sequentially matched with the target type to obtain the initial compatible type.
[0041] For example, patient diagnosis information units are matched to the corresponding disease coding types in the new platform. Finally, bidirectional conflict type detection is performed on the initial compatible types to check whether there are conflicts of the same information under different type definitions. For example, disease codes may have slight differences in different departments. Through detection and correction, the accurate target log type of the new medical big data platform is finally obtained, ensuring the accurate migration and subsequent analysis and utilization of patient data.
[0042] In this embodiment of the invention, complex log types are refined into manageable units through hierarchical splitting and priority marking, ensuring that key data is converted first and improving processing accuracy. Initial type relationships are automatically constructed based on a predefined mapping rule base, reducing manual configuration costs and supporting rapid adaptation to different database environments. The priority-driven matching and search mechanism combined with fuzzy compatibility processing enhances the flexibility of type conversion and avoids data loss due to format differences. Bidirectional conflict detection technology effectively solves many-to-one or one-to-many mapping conflicts, ensuring the uniqueness and consistency of the target log type. Ultimately, it achieves efficient and reliable conversion of log data during cross-database migration, providing a standardized data foundation for system operation and maintenance and data analysis.
[0043] S3. Obtain the target table configuration information corresponding to the target database, and generate the target database table according to the target log type and the target table configuration information.
[0044] In this embodiment of the invention, in computer technology, obtaining the target table configuration information corresponding to the target database requires a combination of automated tools and the database's native interface to ensure information integrity and real-time performance.
[0045] In detail, you can directly connect to the target database through database management tools (such as MySQL Workbench, SQL Server Management Studio) or programming interfaces (such as JDBC / ODBC) and use system tables or metadata query statements to extract table structure information, including field names, data types, primary key constraints, index definitions, etc.
[0046] For cloud databases (such as AWS RDS and Alibaba Cloud PolarDB), the metadata query interface can be called through the SDK provided by the cloud platform to obtain the configuration details of the table; if the target database supports the information schema, the view can be queried directly to obtain cross-database metadata in a standardized manner.
[0047] For distributed databases (such as TiDB and CockroachDB), it is necessary to use their management console or distributed coordination service (such as etcd) to obtain distributed characteristic information such as table sharding rules and replica configurations; for systems that have deployed database change management tools (such as Liquibase and Flyway), the current configuration status of the target table can be extracted by parsing its change scripts (such as XML / YAML format change sets).
[0048] Furthermore, to ensure information accuracy, a multi-source verification mechanism can be adopted to compare the direct query results with database backup files (such as DDL scripts) or cached metadata snapshots, correcting differences caused by permission restrictions or caching delays, and encapsulating the obtained configuration information into structured data (such as JSON / YAML format) for use in subsequent data migration, synchronization, or table structure verification tasks.
[0049] For example, in the fintech field, bank transaction systems generate a large number of transaction logs, including customer account operations, fund transfers, and transaction times. When upgrading the transaction system or connecting to a new regulatory platform, log type conversion is required. Similarly, the source log data types are first hierarchically split, then priority tags are added, an initial type mapping relationship is constructed, and a matching search is performed. In the two-way conflict type detection stage, because financial transactions have extremely high requirements for data accuracy and compliance, the focus is on detecting type conflicts involving key information such as fund amounts and the identities of the transacting parties. This ensures that the converted log types meet the requirements of the new system and regulatory standards, providing strong support for the security and stability of financial transactions.
[0050] In this embodiment of the invention, generating the target database table based on the target log type and the target table configuration information includes: The target log type is parsed to obtain the log structure characteristics; Generate a table structure paradigm based on the log structure characteristics and the target table configuration information; Dynamic partitioning parameters are injected into the table structure normalization to obtain table partitioning parameters; Extract the partition field from the table partition parameters, add default value constraints to the partition field, and generate the target database table based on the default value constraints, the table partition parameters, and the table structure normalization.
[0051] In detail, the structural characteristics of the target log type are parsed through a predefined log type template library (such as a rule file in JSON / XML format), including field names, data types (such as strings and integers), hierarchical relationships between fields (such as subfields in nested JSON), and whether key attributes such as timestamps and unique identifiers are included.
[0052] If the log type is semi-structured (such as key-value pairs in log text), regular expressions or natural language processing (NLP) techniques are used to extract field boundaries and semantic features; if it is completely unstructured text, the underlying structure is inferred through keyword matching or machine learning models (such as named entity recognition).
[0053] Furthermore, the fields in the log features are matched with the column names and data types in the target table configuration. If there are type incompatibilities (such as the strings in the log needing to be mapped to the enumeration type in the target table), they are handled using type conversion rules (such as string truncation and enumeration value mapping tables). If the log features contain fields not defined in the target table (such as fields added in the new version of the log), the table structure is dynamically expanded, new columns are added and default values are set or null values are allowed, the field order is adjusted according to the log access pattern (such as frequently queried fields), or denormalization techniques (such as redundant storage of commonly used calculated fields) are used to improve query performance.
[0054] Specifically, if the log contains a timestamp field, partitioning rules are automatically generated based on time ranges (such as days and months) (e.g., PARTITION BY RANGE (timestamp_column)); hash partitioning is used to balance data distribution for high-cardinality segments (such as user ID and device type), or list partitioning is used to improve query efficiency for low-cardinality segments (such as log level); partitioning parameters (such as the number of partitions and the interval unit) are dynamically replaced through configuration files or environment variables to avoid hard coding and adapt to different sizes of data volume (e.g., small partitions for test environments and large partitions for production environments).
[0055] If the partition field is a timestamp and is not explicitly assigned a value, then the default value is set to the current time to ensure that data is automatically assigned to the correct partition when it is inserted. Set business-related default values for partition fields (such as log level) to avoid partition failure due to null values. Concatenate the table structure normalization, partition parameters and default value constraints into a complete SQL statement to generate the target database table.
[0056] For example, during the upgrade of a hospital's information system, it is necessary to integrate log data generated by different business systems into a new medical big data platform. Taking the log processing of the electronic medical record system as an example, firstly, the target log type is parsed. For example, the patient medical record log contains structured features such as basic patient information, diagnosis details, and medical orders. Next, based on the preset target table configuration information, such as stipulating that basic patient information must include fields such as name, age, and gender, and diagnosis details must record disease name, diagnosis time, etc., a table structure normalization is generated to determine the basic framework and field types of the table. Then, dynamic partitioning parameters are injected into the table structure normalization. Considering that data is stored in partitions by time for easy querying and management, "year-month" is set as the dynamic partitioning parameter. After extracting the partitioning fields "year" and "month", default value constraints are added to prevent the partitioning fields from being empty when inserting data, such as the default value being the current year and month. Finally, the target database table is generated based on the default value constraints, table partitioning parameters, and table structure normalization. The resulting table can clearly store patient medical record log data, and improve data query efficiency through partitioning, making it convenient for doctors to quickly retrieve patient historical medical records for analysis and diagnosis, thereby improving the quality of medical services.
[0057] In this embodiment of the invention, dynamic table generation technology significantly improves the efficiency and flexibility of database log management. Deep analysis of the target log type automatically extracts structured features, eliminating errors from manual modeling. The generated normalization pattern, combined with the target table configuration, accurately adapts to the storage specifications of different databases, ensuring compatibility. The dynamic partitioning parameter injection mechanism automatically adjusts the partitioning strategy (such as time-based or hash-based partitioning) according to the log size, optimizing query performance and storage balance. Default value constraints on partitioning fields further ensure data integrity, preventing partition failures due to null values and reducing the development costs of log migration and queries.
[0058] S4. Parse the source log data to obtain multiple target data, and create partitions for the target database based on the multiple target data to obtain target data blocks.
[0059] In this embodiment of the invention, source log data is divided into three categories—structured, semi-structured, and unstructured—through classification processing. Key-value pairs, regular expression matching, and natural language entities are extracted respectively. Contextual association analysis is performed using timestamps to generate a data lineage graph, and finally, the target data is extracted. This achieves unified parsing and efficient partitioning of heterogeneous logs, ensuring data correlation and storage optimization.
[0060] In this embodiment of the invention, the step of parsing the source log data to obtain multiple target data includes: The source log data is classified to obtain structured log data, semi-structured log data, and unstructured log data; The structured key-value pairs of the structured log data are extracted, and regular expression matching is performed on the semi-structured log data to obtain semi-structured target data. Natural language recognition processing is performed on the unstructured log data to obtain unstructured log entities. Obtain the log timestamp corresponding to the source log data, and perform context association analysis on the structured key-value pairs, the semi-structured target data, and the unstructured log entities based on the log timestamp to generate a log data lineage graph; Data extraction is performed on the pedigree graph of the log data to obtain multiple target data.
[0061] In detail, the structured log data is identified by detecting predefined formats (such as JSON, CSV) or matching database table structures (such as strict correspondence between field names and data types); the semi-structured logs are identified by data containing partial structure markers (such as XML tags, key-value separators "=" or ":" in the log header); and the unstructured logs are determined by text features (such as no fixed separators, free text paragraphs) or machine learning models (such as training a classifier to identify plain text paragraphs).
[0062] Specifically, fields in structured logs are directly parsed, or field names and values are extracted through database metadata mapping; for semi-structured logs, key-value pairs are extracted using predefined regular expression templates; entities such as time, location, and error type are extracted using Named Entity Recognition (NER) models (such as BERT and Spacy), and domain dictionaries (such as "Exception" and "Timeout") are used to assist in identifying key information.
[0063] Furthermore, time fields can be directly extracted from structured or semi-structured data, or time expressions can be identified from unstructured text using NLP models. Time slots can be divided according to fixed intervals (such as minutes or hours) or event-driven (such as the same transaction ID) to group log data. Dependencies between data can be established based on timestamps, unique identifiers (such as request IDs) or causal relationships (such as "error log → subsequent retry logs"). Graph databases such as Neo4j can be used to store lineage relationships, with nodes representing log entries and edges representing association types (such as "trigger" or "follow").
[0064] Specifically, logs for specific paths are extracted using depth-first search (DFS) or breadth-first search (BFS). Fields of associated logs (such as error count and processing time) are counted and summed. Key nodes and edges in the lineage graph are converted into tables (such as CSV) or JSON formats for downstream analysis, thereby obtaining multiple target data.
[0065] In this embodiment of the invention, the step of creating partitions for the target database based on multiple target data to obtain target data blocks includes: Extract the partition fields of the target data, perform statistical analysis on the partition fields, and obtain the partition field analysis results; The partition field analysis results are transformed using target partition parameters to obtain the target partition identifier for each target data. The target data is grouped into multiple temporary datasets according to the target partition identifier, and the target data in each temporary dataset is sorted according to the priority of the target data to obtain a sorting result; Based on the sorting results and the preset block capacity threshold, each temporary dataset is divided into one or more target data blocks.
[0066] In detail, candidate partition fields are extracted directly from the structured target data or the parsed semi-structured data. The number and frequency distribution of unique values of the partition fields are calculated. The cardinality (total number of unique values) of the hash partition candidate fields (such as user ID) is calculated to ensure that the number of partitions is reasonable.
[0067] If the field is a timestamp or a numerical value, a partition key is generated according to the range (e.g., one partition per day). Low cardinality numeric segments are directly mapped to fixed partitions, while high cardinality numeric segments (e.g., user IDs) are distributed evenly using a hash function (e.g., modulo). The partitioning parameters are adjusted based on the statistical analysis results (e.g., the high-frequency value "East China" is divided into a separate partition, and the remaining values are merged into the "Other" partition), and a partition label is attached to each target data entry.
[0068] Furthermore, the hash value of the partition identifier is used to quickly allocate data to temporary buckets in memory. Sorting fields are defined according to business needs (such as error logs taking precedence over ordinary logs, or high-value user data taking precedence). Data in each temporary dataset is sorted by priority and timestamp (such as stable sorting to ensure that data with the same priority is processed in chronological order).
[0069] Specifically, the preset block size (e.g., 100MB or 100,000 records per block) is dynamically adjusted to adapt to different data densities. Data is filled into the current block according to the sorting results. When the threshold is exceeded, a new block is created to ensure that data from the same transaction is not split into different blocks (e.g., grouped by transaction ID and then divided into blocks). The actual block size is adjusted according to the data compression rate (e.g., text data may be smaller than the threshold after compression).
[0070] In this embodiment of the invention, intelligent log parsing and dynamic partitioning technologies significantly improve log management efficiency and data availability. The log parsing stage employs multimodal technology, combining regular expression matching and natural language processing to accurately extract key information, and constructs a lineage graph through time slot association to achieve complete traceability of log context. In the partitioning stage, the partitioning strategy is dynamically adapted based on field statistical analysis, combined with priority sorting and capacity threshold control to ensure balanced data distribution and priority processing of key data, optimize query performance and storage utilization, and provide efficient and reliable underlying support for log analysis, fault diagnosis, and security auditing.
[0071] S5. Perform data synchronization transmission on the target database according to the target database table and the target data block to obtain target synchronization transmission data.
[0072] In this embodiment of the invention, by extracting records from the target data block and generating a checksum, encapsulating it into a synchronous transmission data packet, and formulating a synchronization strategy, a structural synchronization control is performed in the target database based on the strategy to achieve reliable data transmission and generate target synchronous transmission data.
[0073] For example, in the fintech field, when a bank upgrades its core system, it needs to synchronize data from the old system database to the new system database. The old system database stores basic customer information, account balances, transaction records, etc. First, the target database table structure information is obtained, the data blocks to be synchronized are determined, the data records of the target data blocks are extracted and a checksum is generated, and these are encapsulated into a synchronization transmission data packet. The synchronization transmission strategy is determined based on the importance and real-time requirements of the data. For critical data such as account balances, a real-time synchronization strategy is adopted to ensure data consistency between the old and new systems. Structural synchronization control is performed within the target database, and the checksum ensures accurate data transmission, resulting in the target synchronized transmission data. This ensures a smooth data transition and business continuity during the bank's core system upgrade process.
[0074] like Figure 3 As shown in this embodiment of the invention, the step of performing data synchronization transmission on the target database based on the target database table and the target data block to obtain target synchronization transmission data includes: Obtain the structure information of the target database table and determine the data blocks to be synchronized in the target database table; Extract data records from the target data block and generate corresponding data verification codes based on the data records; The data verification code and the data record are encapsulated into a synchronous transmission data packet, and a synchronous transmission strategy is determined based on the data block to be synchronized and the synchronous transmission data packet. Within the target database, the synchronization transmission strategy is used to perform structural synchronization control on the data block to be synchronized and the synchronization transmission data packet to obtain the target synchronization transmission data.
[0075] In detail, the target table's field names, data types, primary keys, indexes, and other structural information are obtained through database metadata interfaces (such as MySQL's INFORMATION_SCHEMA or Hive's DESCRIBE TABLE). The schema differences between the source table and the target table are compared, and a compatibility report is generated (such as marking warnings when fields are missing or types do not match).
[0076] Based on the partition field value or hash value range, a subset of data to be synchronized is selected from the target data block. Incremental synchronization (e.g., synchronizing only newly added partitions) or full synchronization (overwriting all data in the target table) is supported. Data is read row by row from the target data block. Batch read optimization is supported (e.g., reading 1000 records at a time). Serialization processing is performed on special fields (e.g., BLOB type) to ensure transmission integrity.
[0077] Furthermore, data records and checksums are packaged into a structured format (such as JSON or Protocol Buffers), with additional metadata (such as data source and generation time), and compressed transmission (such as Gzip compressed data packets) is supported to reduce network bandwidth usage. The nearest database node is selected based on the partition identifier of the data block to be synchronized (such as region=East China). Real-time transmission channels are enabled for high-priority data (such as error logs), while ordinary data goes through batch channels. The checksums of the transmitted data are recorded, and retransmission is performed from the breakpoint in case of failure.
[0078] Specifically, missing fields are automatically created or data types are adjusted to ensure that the partitioning strategy of the target table (such as monthly partitioning) is consistent with the partitioning granularity of the data blocks to be synchronized; the checksum of the received data is recalculated on the target side and compared with the checksum in the transmission packet. Inconsistent data is discarded, and atomicity is ensured by batch commit (such as committing once every 1000 records) or two-phase commit (2PC). For data with primary key conflicts, the strategy is to choose to overwrite, ignore, or record it in the conflict table for manual processing.
[0079] For example, in the scenario of building a medical data sharing platform in a large hospital, it is necessary to synchronize the database data of different departmental subsystems to the central database. Taking the radiology department and clinical departments as examples, the radiology department's database stores various imaging examination data and reports of patients, while the clinical department's database contains detailed information such as patients' diagnoses and treatments.
[0080] First, obtain the structural information of the target database tables (such as patient image information tables and clinical diagnosis information tables), clarifying the table fields, data types, primary keys, etc. By analyzing the data update frequency and importance, determine the data blocks to be synchronized, for example, recently entered patient image data and diagnostic data as the data blocks to be synchronized. Next, extract the data records from the target data blocks, such as patient image file names, examination times, and diagnostic results, and use a hash algorithm to generate a corresponding data checksum for each data record to ensure the integrity and accuracy of the data during transmission. Then, encapsulate the data checksum and data records into a synchronization transmission data packet, and determine the synchronization transmission strategy based on the size of the data block to be synchronized, the data volume, and network bandwidth. Finally, within the target database (central database), use the synchronization transmission strategy to perform structural synchronization control on the data blocks to be synchronized and the synchronization transmission data packets. During the synchronization process, check the data integrity based on the data checksum; if data corruption or loss occurs, retransmit the corresponding data packet.
[0081] Through this series of operations, the target synchronous transmission data is obtained, realizing accurate synchronization of data between the radiology department and clinical departments in the central database. This allows doctors to comprehensively view patient information and improve diagnostic efficiency and accuracy.
[0082] In this embodiment of the invention, structured synchronization and intelligent transmission strategies significantly improve the efficiency and reliability of database data synchronization. The system automatically parses the target table structure and accurately locates the data blocks to be synchronized, avoiding performance losses caused by full table scans. Secondly, it generates checksums based on data records and encapsulates them into transmission packets. Combined with dynamic routing and priority scheduling strategies, it ensures that critical data is transmitted first and network bandwidth is used efficiently. Finally, through structured synchronization control, it achieves automatic schema alignment between the source and target ends, and uses checksum verification to ensure data consistency, effectively reducing transmission latency and error rates. This provides efficient and stable data synchronization support for distributed systems, real-time data analysis, and other scenarios.
[0083] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0084] like Figure 4 The diagram shown is a functional block diagram of a data transmission device provided in an embodiment of the present invention.
[0085] In this embodiment of the disclosure, a data transmission device is provided, which corresponds one-to-one with the data transmission method of the above embodiments. For example... Figure 4As shown, the data transmission device 100 can be installed in an electronic device. Depending on its functions, the data transmission device 100 includes a log type identification module 101, a log type conversion module 102, a database table generation module 103, a data block construction module 104, and a data synchronization and transmission module 105. Detailed descriptions of each functional module are as follows: Log type identification module 101 is used to obtain the source database and identify the source log type of the source log data in the source database; Log type conversion module 102 is used to convert the source log data type into a preset target log type of the target database; The database table generation module 103 is used to obtain the target table configuration information corresponding to the target database, and generate the target database table according to the target log type and the target table configuration information. The data block construction module 104 is used to parse the source log data to obtain multiple target data, and create partitions for the target database based on the multiple target data to obtain target data blocks; The data synchronization transmission module 105 is used to perform data synchronization transmission on the target database according to the target database table and the target data block to obtain target synchronization transmission data.
[0086] In one embodiment, when the log type identification module 101 performs the task of identifying the source log type of the source log data in the source database, it is used to: Extract source log data from the log storage area in the source database; The source log data is encoded to obtain a log data string; Obtain a sliding window, and perform window truncation on the log data string according to the sliding window to obtain a log data substring; The log data substring is pattern matched according to a preset log type template library to obtain the source log data type corresponding to the source log data.
[0087] In one embodiment, when the log type conversion module 102 performs the conversion of the source log data type to a preset target log type of the target database, it is used to: The source log data type is hierarchically split to obtain multiple log-level units, and type conversion priority markers are added to the multiple log-level units; Based on the source log data type and the preset target database log type mapping rule base, an initial type mapping relationship is constructed; Based on the initial type mapping relationship, the log-level units with added type conversion priority tags are sequentially searched for target types to obtain the initial compatible types; A bidirectional conflict type detection is performed on the initial compatible type to obtain the target log type of the target database.
[0088] In one embodiment, when the database table generation module 103 generates a target database table based on the target log type and the target table configuration information, it is used to: The target log type is parsed to obtain the log structure characteristics; Generate a table structure paradigm based on the log structure characteristics and the target table configuration information; Dynamic partitioning parameters are injected into the table structure normalization to obtain table partitioning parameters; Extract the partition field from the table partition parameters, add default value constraints to the partition field, and generate the target database table based on the default value constraints, the table partition parameters, and the table structure normalization.
[0089] In one embodiment, when the data block construction module 104 performs log data parsing on the source log data to obtain multiple target data, it is used to: The source log data is classified to obtain structured log data, semi-structured log data, and unstructured log data; The structured key-value pairs of the structured log data are extracted, and regular expression matching is performed on the semi-structured log data to obtain semi-structured target data. Natural language recognition processing is performed on the unstructured log data to obtain unstructured log entities. Obtain the log timestamp corresponding to the source log data, and perform context association analysis on the structured key-value pairs, the semi-structured target data, and the unstructured log entities based on the log timestamp to generate a log data lineage graph; Data extraction is performed on the pedigree graph of the log data to obtain multiple target data.
[0090] In one embodiment, when the data block construction module 104 performs the operation of creating partitions for the target database based on a plurality of target data to obtain target data blocks, it is configured to: Extract the partition fields of the target data, perform statistical analysis on the partition fields, and obtain the partition field analysis results; The partition field analysis results are transformed using target partition parameters to obtain the target partition identifier for each target data. The target data is grouped into multiple temporary datasets according to the target partition identifier, and the target data in each temporary dataset is sorted according to the priority of the target data to obtain a sorting result; Based on the sorting results and the preset block capacity threshold, each temporary dataset is divided into one or more target data blocks.
[0091] In one embodiment, when the data synchronization transmission module 105 performs data synchronization transmission of the target database based on the target database table and the target data block to obtain target synchronization transmission data, it is used to: Obtain the structure information of the target database table and determine the data blocks to be synchronized in the target database table; Extract data records from the target data block and generate corresponding data verification codes based on the data records; The data verification code and the data record are encapsulated into a synchronous transmission data packet, and a synchronous transmission strategy is determined based on the data block to be synchronized and the synchronous transmission data packet. Within the target database, the synchronization transmission strategy is used to perform structural synchronization control on the data block to be synchronized and the synchronization transmission data packet to obtain the target synchronization transmission data.
[0092] In this invention, specific limitations regarding the data transmission device can be found in the limitations regarding the data transmission method described above, and will not be repeated here. Each module in the aforementioned data transmission device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in a computer device, or stored in software in the memory of a computer device, so that the processor can call and execute the operations corresponding to each module.
[0093] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the data transmission method on the server side.
[0094] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the data transmission method on the client side.
[0095] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain the source database and identify the source log type of the source log data in the source database; Convert the source log data type to the preset target log type of the target database; Obtain the target table configuration information corresponding to the target database, and generate the target database table according to the target log type and the target table configuration information; The source log data is parsed to obtain multiple target data, and partitions are created for the target database based on the multiple target data to obtain target data blocks; Data is synchronously transmitted to the target database based on the target database table and the target data block to obtain target synchronous transmission data.
[0096] In the several embodiments provided by this invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0097] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0098] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0099] In some embodiments of this example, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that the computer program, when executed by a processor, implements the steps of the method described in the above embodiments.
[0100] The readable storage medium of the present invention stores a computer program, which, when executed by a processor of an electronic device, can perform the following: Obtain the source database and identify the source log type of the source log data in the source database; Convert the source log data type to the preset target log type of the target database; Obtain the target table configuration information corresponding to the target database, and generate the target database table according to the target log type and the target table configuration information; The source log data is parsed to obtain multiple target data, and partitions are created for the target database based on the multiple target data to obtain target data blocks; Data is synchronously transmitted to the target database based on the target database table and the target data block to obtain target synchronous transmission data.
[0101] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0102] Computer-readable storage media may also store at least one computer-executable program / instruction, such as computer-readable instructions. Computer-readable storage media include, but are not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Computer-readable storage media may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer, and then, when the computing device executes the computer-readable instructions stored on the computer-readable storage medium, the various methods described above can be performed.
[0103] In addition, the computer device may include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (e.g., keyboard, mouse, speakers, etc.).
[0104] In one embodiment, the at least one computer-executable instruction may also be compiled into or comprise a software product / computer program product, wherein one or more computer-executable instructions are executed by a processor to perform the steps of the various functions and / or methods in the embodiments described herein.
[0105] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Furthermore, any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory.
[0106] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0107] In the embodiments provided in this disclosure, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0108] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
[0109] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.
Claims
1. A data transmission method, characterized in that, The method includes: Obtain the source database and identify the source log type of the source log data in the source database; Convert the source log data type to the preset target log type of the target database; Obtain the target table configuration information corresponding to the target database, and generate the target database table according to the target log type and the target table configuration information; The source log data is parsed to obtain multiple target data, and partitions are created for the target database based on the multiple target data to obtain target data blocks; Data is synchronously transmitted to the target database based on the target database table and the target data block to obtain target synchronous transmission data.
2. The data transmission method as described in claim 1, characterized in that, The identification of the source log type of the source log data in the source database includes: Extract source log data from the log storage area in the source database; The source log data is encoded to obtain a log data string; Obtain a sliding window, and perform window truncation on the log data string according to the sliding window to obtain a log data substring; The log data substring is pattern matched according to a preset log type template library to obtain the source log data type corresponding to the source log data.
3. The data transmission method as described in claim 1, characterized in that, The step of converting the source log data type to the preset target database target log type includes: The source log data type is hierarchically split to obtain multiple log-level units, and type conversion priority markers are added to the multiple log-level units; Based on the source log data type and the preset target database log type mapping rule base, an initial type mapping relationship is constructed; Based on the initial type mapping relationship, the log-level units with added type conversion priority tags are sequentially searched for target types to obtain the initial compatible types; A bidirectional conflict type detection is performed on the initial compatible type to obtain the target log type of the target database.
4. The data transmission method as described in claim 1, characterized in that, The step of generating the target database table based on the target log type and the target table configuration information includes: The target log type is parsed to obtain the log structure characteristics; Generate a table structure paradigm based on the log structure characteristics and the target table configuration information; Dynamic partitioning parameters are injected into the table structure normalization to obtain table partitioning parameters; Extract the partition field from the table partition parameters, add default value constraints to the partition field, and generate the target database table based on the default value constraints, the table partition parameters, and the table structure normalization.
5. The data transmission method as described in claim 1, characterized in that, The source log data is parsed to obtain multiple target data, including: The source log data is classified to obtain structured log data, semi-structured log data, and unstructured log data; The structured key-value pairs of the structured log data are extracted, and regular expression matching is performed on the semi-structured log data to obtain semi-structured target data. Natural language recognition processing is performed on the unstructured log data to obtain unstructured log entities. Obtain the log timestamp corresponding to the source log data, and perform context association analysis on the structured key-value pairs, the semi-structured target data, and the unstructured log entities based on the log timestamp to generate a log data lineage graph; Data extraction was performed on the pedigree graph of the log data to obtain multiple target data.
6. The data transmission method as described in claim 1, characterized in that, The step of creating partitions for the target database based on multiple target data to obtain target data blocks includes: Extract the partition fields of the target data, perform statistical analysis on the partition fields, and obtain the partition field analysis results; The analysis results of the partition fields are transformed using target partition parameters to obtain the target partition identifier for each target data. The target data is grouped into multiple temporary datasets according to the target partition identifier, and the target data in each temporary dataset is sorted according to the priority of the target data to obtain a sorting result; Based on the sorting results and the preset block capacity threshold, each temporary dataset is divided into one or more target data blocks.
7. The data transmission method as described in claim 1, characterized in that, The step of synchronizing and transmitting data to the target database based on the target database table and the target data block to obtain target synchronized transmission data includes: Obtain the structure information of the target database table and determine the data blocks to be synchronized in the target database table; Extract data records from the target data block and generate corresponding data verification codes based on the data records; The data verification code and the data record are encapsulated into a synchronous transmission data packet, and a synchronous transmission strategy is determined based on the data block to be synchronized and the synchronous transmission data packet. Within the target database, the synchronization transmission strategy is used to perform structural synchronization control on the data block to be synchronized and the synchronization transmission data packet to obtain the target synchronization transmission data.
8. A data transmission device, characterized in that, The device includes: The log type identification module is used to obtain the source database and identify the source log type of the source log data in the source database; The log type conversion module is used to convert the source log data type into a preset target log type of the target database. The database table generation module is used to obtain the target table configuration information corresponding to the target database, and generate the target database table according to the target log type and the target table configuration information. The data block construction module is used to parse the source log data to obtain multiple target data, and create partitions for the target database based on the multiple target data to obtain target data blocks; The data synchronization and transmission module is used to perform data synchronization and transmission on the target database according to the target database table and the target data block to obtain target synchronization and transmission data.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data transmission method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the data transmission method as described in any one of claims 1 to 7.