Data acquisition method and system based on low-code development platform
By implementing automatic data source identification and mapping, predefined rules to define data acquisition logic and real-time trigger data acquisition operations on a low-code development platform, the problems of low data quality and large storage burden in the existing technology are solved, and high-quality data acquisition and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202510049096.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing data acquisition systems cannot effectively process data that develops towards a bad trend but still meets the requirements, resulting in low data quality and increasing data storage burden.
By automatically identifying and connecting data sources on a low-code development platform, automatically inferring the relationship between data sources for data mapping and transformation, using predefined rules to define data acquisition logic, including filtering conditions and sorting rules, and triggering data acquisition operations in real time when the data source changes, processing encryption and access control of sensitive data.
Effectively screen out low-quality data, improve data quality, and reduce the burden of data storage, improve data processing efficiency through comprehensive analysis of multiple indicators, and achieve more comprehensive and accurate data quality analysis.
Smart Images

Figure CN120067080A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data acquisition, and specifically relates to a data acquisition method and system based on a low-code development platform. Background Art
[0002] A data acquisition system for a low-code development platform refers to a set of systems used to acquire, integrate, and process data in a low-code development environment. The low-code development platform aims to accelerate the application development process by minimizing manual coding. On such a platform, developers can use a graphical interface and a small amount of coding to create applications without having to delve into the underlying programming languages and details;
[0003] The prior art has the following defects: The existing acquisition system usually sets corresponding indicators for data. When the data to be acquired does not meet any of the indicators, the data is not acquired. However, if multiple indicators of the acquired data all show a bad trend but still meet the requirements, it will still result in low data quality. The acquisition system cannot effectively handle such problems, easily leading to low data quality and increasing the data storage burden. Summary of the Invention
[0004] The purpose of the present invention is to provide a data acquisition method and system based on a low-code development platform to solve the deficiencies in the background art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions: A data acquisition method based on a low-code development platform, the acquisition method comprising the following steps:
[0006] The acquisition system detects and identifies data sources by automatically scanning the network, API catalog, or metadata. The data sources include databases, file systems, and cloud services. After discovering the data sources, connections are automatically established;
[0007] Automatically identify and adapt to the connection methods of various data sources using predefined connectors. The acquisition system automatically infers the relationships between different data sources and performs data mapping and transformation;
[0008] Automatically define data acquisition logic using predefined rules, including filtering conditions and sorting rules, automatically execute data acquisition operations according to a schedule, and immediately trigger data acquisition operations when the data source changes, and handle encryption and access control of sensitive data;
[0009] Automatically test the acquired data after acquisition, verify whether the data acquisition process is working properly, and automatically record monitoring data, and automatically handle abnormal situations and take measures during the data acquisition process.
[0010] Preferably, automatically defining data acquisition logic using predefined rules comprises the following steps:
[0011] After obtaining data from the low-code development platform, obtain the integrity, duplication, and update frequency of the data;
[0012] Comprehensively calculate the condition coefficient of the data by combining the integrity, duplication, and update frequency. The expression is:
[0013] ,
[0014] In the formula, TJS is the condition coefficient, WZD, GXP, and CFD are the integrity, update frequency, and duplication respectively, n is the number of data segments, WZDi represents the integrity of the i-th data segment, GXPi represents the update frequency of the i-th data segment, CFDi represents the duplication of the i-th data segment, and α, β, and γ are the proportionality coefficients of the integrity, update frequency, and duplication respectively, and α, β, and γ are all greater than 0.
[0015] Preferably, the filtering conditions and sorting rules are:
[0016] After obtaining the condition coefficient of the data, compare the condition coefficient with the preset condition threshold. If the condition coefficient is less than the condition threshold, analyze that the data quality is low and filter out the data;
[0017] After obtaining all the data, sort all the data in descending order according to the condition coefficient, generate a data list, and send the data list to the administrator.
[0018] Preferably, the integrity acquisition logic is: query the missing values in the data set, understand the frequency and distribution of the missing values of each field, and use the statistical functions of the database or data analysis tool to calculate the number of non-missing values in each column or row;
[0019] The update frequency acquisition logic is: use a monitoring tool or query the update log of the data source to obtain the update frequency of the data in real time;
[0020] The duplication acquisition logic is: use an aggregation function to calculate the number of unique values of the field or row to obtain the duplication of the data.
[0021] Preferably, after discovering the data source, automatically establish a connection, and use a predefined connector to automatically identify and adapt to the connection methods of various data sources, including the following steps:
[0022] Build a connector library in the acquisition system, which contains predefined connectors. Each connector is used for a specific type of data source. The connectors include database connectors, API connectors, cloud service connectors, etc. When the acquisition system detects a data source, it matches according to the type of the data source to determine which predefined connector to use. The acquisition system automatically selects a connector that matches the detected data source type, automatically extracts the information required for connection from the metadata of the data source, and dynamically configures the connection parameters for different data sources. Before establishing a connection, the acquisition system automatically conducts a connection test.
[0023] Preferably, the acquisition system automatically infers the relationships between different data sources and performs data mapping and transformation, including the following steps:
[0024] Extract metadata from each data source, including table structure, field type, and key relationship information, conduct statistics and analysis on the data, use heuristic methods to infer the relationships between data sources, and based on the analysis results, automatically generate mapping rules between data sources, including field mapping, value transformation, and data type conversion rules. Use intelligent matching algorithms to automatically generate data type conversion logic based on the mapping rules. For similar but different data values in different data sources, generate corresponding mapping and transformation logic.
[0025] Preferably, after obtaining the data, automatically conduct tests to verify whether the data acquisition process is working properly and automatically record monitoring data, including the following steps:
[0026] Use an automated testing tool to execute the designed automated test script, simulate changes in the data source during the test, integrate a performance monitoring tool, record the performance metrics of the data acquisition system, record detailed error logs, including exception situations, error codes, and stack traces, automatically generate a detailed automated test report, including test coverage, pass rate, failed test cases, and execution time, and integrate a log analysis tool to analyze the running logs of the acquisition system.
[0027] Preferably, during the data acquisition process, automatically handle exception situations, including the following steps:
[0028] Real-time monitor the running status of the data acquisition system, including the data source connection status, response time, and data acquisition success rate. For an unavailable data source, select to retry the connection. For a connection failure, trigger an alarm and perform automatic repair. When a connection failure or data source unavailability is detected, automatically conduct a certain number of retries and automatically execute repair actions. When an exception occurs, generate a detailed log record, including the type of exception, timestamp, and related data. Define error codes for different types of exceptions. If an exception causes some operations in the data acquisition process to go wrong, the acquisition system automatically performs a rollback operation to restore the acquisition system status to the normal working state.
[0029] The present invention also provides a data acquisition system based on a low-code development platform, including an identification module, a mapping and conversion module, a data screening module, a timing module, and a testing module;
[0030] Identification module: Detect and identify data sources by automatically scanning the network, API catalog, or metadata. The data sources include databases, file systems, and cloud services. The identification results of the data sources are sent to the connection module;
[0031] Mapping and conversion module: Automatically establish a connection after discovering the data source, and use predefined connectors to automatically identify and adapt to the connection methods of various data sources, automatically infer the relationships between different data sources, and perform data mapping and conversion;
[0032] Data screening module: Automatically define data acquisition logic using predefined rules, including filtering conditions and sorting rules;
[0033] Timing module: Automatically execute data acquisition operations according to a schedule, and immediately trigger data acquisition operations when the data source changes, handling encryption and access control of sensitive data;
[0034] Testing module: Automatically perform tests after obtaining data, verify whether the data acquisition process is working properly, and automatically record monitoring data, automatically handling abnormal situations and taking measures during the data acquisition process.
[0035] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0036] 1. By using predefined connectors to automatically identify and adapt to the connection methods of various data sources, the acquisition system automatically infers the relationships between different data sources, performs data mapping and conversion, automatically defines data acquisition logic using predefined rules, including filtering conditions and sorting rules, automatically executes data acquisition operations according to a schedule, and immediately triggers data acquisition operations when the data source changes, handling encryption and access control of sensitive data, automatically performing tests after obtaining data, verifying whether the data acquisition process is working properly, and automatically recording monitoring data, and automatically handling abnormal situations and taking measures during the data acquisition process. This acquisition method can comprehensively analyze multiple indicators of the acquired data, thereby effectively screening out low-quality data, improving data quality while reducing the data storage burden.
[0037] 2. After obtaining data from the low-code development platform, the present invention divides the data into multiple data segments, and obtains the integrity, duplication degree, and update frequency of the multiple data segments. The condition coefficient of the data is obtained by comprehensively calculating the integrity, duplication degree, and update frequency. After obtaining the condition coefficient of the data, the condition coefficient is compared with a preset condition threshold, so as to analyze the quality of the overall data. Through the comprehensive analysis method, not only the processing efficiency of the data is improved, but also the analysis is more comprehensive and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0039] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0041] Embodiment 1: Please refer to Figure 1 As shown, the data acquisition method based on the low-code development platform in this embodiment includes the following steps:
[0042] The acquisition system detects and identifies available data sources by automatically scanning the network, API directories, or metadata. The data sources include databases, file systems, and cloud services. After discovering the data sources, it automatically establishes connections and uses predefined connectors to automatically identify and adapt to the connection methods of various data sources, such as addresses, credentials, etc. The acquisition system has the ability of intelligent mapping and transformation, automatically infers the relationships between different data sources, and performs data mapping and transformation to ensure data consistency and the correctness of applications. It automatically defines data acquisition logic using predefined rules, including filtering conditions, sorting rules, etc., so as to dynamically adapt to changes in the data source without manual adjustment. It automatically executes data acquisition operations according to a schedule and immediately triggers data acquisition operations when the data source changes. It processes the encryption and access control of sensitive data to ensure that only authorized users or systems can obtain specific data. After acquiring the data, it automatically conducts tests to verify whether the data acquisition process is working properly and automatically records monitoring data to promptly discover and solve potential problems. It automatically handles abnormal situations during the data acquisition process, such as unavailable data sources, connection failures, etc., and takes appropriate measures, such as triggering an alarm or performing automatic repair.
[0043] In this application, by using predefined connectors to automatically identify and adapt to the connection methods of various data sources, the acquisition system automatically infers the relationships between different data sources, performs data mapping and transformation, automatically defines data acquisition logic using predefined rules, including filtering conditions and sorting rules, automatically executes data acquisition operations according to a schedule and immediately triggers data acquisition operations when the data source changes, processes the encryption and access control of sensitive data, automatically conducts tests after acquiring the data to verify whether the data acquisition process is working properly, and automatically records monitoring data. It automatically handles abnormal situations during the data acquisition process and takes measures. This acquisition method can comprehensively analyze multiple indicators of the acquired data, thereby effectively screening out low-quality data, improving data quality while reducing the data storage burden.
[0044] Embodiment 2: A. The acquisition system detects and identifies available data sources by automatically scanning the network, API directories, or metadata. The data sources include databases, file systems, and cloud services, and it includes the following steps:
[0045] Network scanning: Target definition: The system first needs to define the network scope to be scanned, including IP address ranges, subnets, domain names, etc.
[0046] Scanning tool: Use a network scanning tool, such as Nmap or ZMap, to scan the target network and identify active hosts and open ports.
[0047] API Directory Scanning: Collect API Endpoint Information: Use automated tools or web crawler techniques to collect API endpoint information in the target system. This may include the URLs of APIs, parameters, supported HTTP methods, etc.
[0048] Analyze API Documentation: For documented APIs, the system can parse the documentation to obtain detailed API information, including supported resources, data formats, etc.
[0049] Metadata Analysis: Database Metadata Extraction: For databases, the system can extract database metadata, including table structures, fields, indexes, and other information. This can be done by querying the system tables of the database system or using metadata query statements specific to the database.
[0050] File System Metadata Analysis: For file systems, the system can examine the metadata of files and folders, including file types, sizes, creation dates, and other information.
[0051] Automatically Identify Data Source Types: Pattern Matching: Through predefined pattern matching rules, the system can identify the types of data sources. For example, a specific URL structure or port number can indicate a database, and a specific file path or file extension can indicate a file system.
[0052] Feature Analysis: Use heuristic methods or machine learning techniques to perform feature analysis on the collected information to identify the types of data sources. For example, databases typically exhibit specific connection and query patterns.
[0053] Establish Connections: Extract Connection Information: Once the data source type is determined, the system needs to extract the information required for connection, such as the database server address, API endpoint URL, file system path, etc.
[0054] Test Connections: The system can attempt to establish connections and perform simple tests to ensure that the data sources can be accessed successfully.
[0055] Generate a Data Source Inventory: Organize Information: Organize the detected data source information into an inventory, including data source types, connection information, etc.
[0056] Store Metadata: Store the obtained metadata for future reference. This can include API documentation, database metadata, file system information, etc.
[0057] Visualization and Management: User Interface Display: To enable developers or administrators to better understand the detected data sources, the system can provide a visual user interface to display the data source inventory and related information.
[0058] Management Tools: Provide management tools that allow users to configure or further investigate detected data sources, which may include manually adding, editing connection information, etc.
[0059] B. Automatically establish a connection after discovering the data source, and use predefined connectors to automatically identify and adapt to the connection methods of various data sources, such as addresses, credentials, etc., including the following steps:
[0060] Predefined Connectors: Connector Library Construction: Build a connector library in the system, which contains predefined connectors, and each connector is used for a specific type of data source. Connectors may include database connectors, API connectors, cloud service connectors, etc.
[0061] Connector Matching after Data Source Identification: Data Source Type Matching: When the system detects a data source, match according to the type of the data source to determine which predefined connector to use.
[0062] Automatically Select Connector: The system automatically selects a connector that matches the detected data source type.
[0063] Connection Information Extraction: Automatically Extract Connection Information: Automatically extract the information required for connection from the metadata of the data source, which may include addresses (URLs, IPs, etc.), credentials (usernames, passwords, API keys, etc.), port numbers, etc.
[0064] Dynamically Configure Connection Parameters: Adaptive Connection Configuration: For different data sources, the connector can dynamically configure connection parameters to adapt to specific connection methods and requirements. This may include protocols, security settings, authentication methods, etc.
[0065] Connection Testing: Automatically Test the Connection: Before establishing a connection, the system can automatically perform a connection test to ensure that the provided connection information is valid. This includes verifying the username and password, checking network reachability, etc.
[0066] Error Handling and Fallback Mechanism: Exception Handling: If an error occurs during the connection test or connection establishment process, the system should be able to capture the exception and perform appropriate error handling. This may include recording error information, triggering an alarm, or automatically falling back to an alternative connection method.
[0067] Connection Pool Management: Connection Pool Configuration: For situations that require frequent connections, the system can configure a connection pool to reduce the overhead of connection establishment.
[0068] Automatic Connection Pool Management: The system can automatically manage the connection pool, including connection creation, recycling, and timeout handling.
[0069] Dynamic connection update: Dynamic update: If the connection information of the data source changes (such as password reset, address change, etc.), the system should be able to dynamically update the connection information without stopping the application.
[0070] Security considerations: Encryption and security protocols: The system should consider the security of data transmission, support encrypted communication and use secure protocols.
[0071] Credential management: For sensitive information (such as credentials), the system should manage them securely, which may include encrypted storage and access control.
[0072] C. The acquisition system has the ability of intelligent mapping and conversion, automatically infers the relationships between different data sources, and performs data mapping and conversion to ensure data consistency and application correctness, including the following steps:
[0073] Automatic data source analysis: Metadata extraction: Extract metadata from each data source, including information such as table structure, field type, key relationships, etc.
[0074] Data statistics and analysis: Perform statistics and analysis on the data to understand the data distribution, data types, null values, etc.
[0075] Automatic relationship discovery: Intelligent relationship inference: Using heuristic methods or machine learning techniques, the system can intelligently infer the relationships between data sources, such as primary key - foreign key relationships, similarities, etc.
[0076] Pattern matching: Using predefined rules or pattern matching algorithms, the system can discover potential relationships between data sources.
[0077] Automatic generation of mapping rules: Automatic mapping generation: Based on the analysis results, the system can automatically generate mapping rules between data sources. This may include rules such as field mapping, value conversion, data type conversion, etc.
[0078] Intelligent matching algorithm: Use an intelligent matching algorithm to ensure that the generation of mapping rules is accurate, consistent and reliable.
[0079] User intervention and calibration: User visualization tool: Provide user - friendly graphical tools that allow developers or administrators to manually adjust and calibrate mapping rules when needed.
[0080] Conflict resolution strategy: If there are conflicts or uncertainties, the system should be able to provide conflict resolution strategies, allowing users to intervene and adjust.
[0081] Automatic generation of conversion logic: Data type conversion: Based on the mapping rules, the system can automatically generate data type conversion logic to ensure data compatibility between different data sources.
[0082] Value mapping and conversion: For similar but different data values in different data sources, the system can generate corresponding mapping and conversion logic to ensure data consistency.
[0083] Real-time or batch conversion strategy: Real-time conversion: If the application requires real-time data, the system can generate a real-time conversion strategy to ensure that the conversion is carried out immediately when the data changes.
[0084] Batch conversion: For large amounts of data, the system can generate a batch conversion strategy to effectively process large-scale data sets.
[0085] Exception handling and logging: Exception detection: The system should be able to automatically detect and handle exception situations such as mapping rule mismatches, data type inconsistencies, etc.
[0086] Logging: Record the logs during the conversion process for subsequent review and troubleshooting.
[0087] Automated testing: Automated testing tools: Provide automated testing tools to ensure the accuracy and reliability of the data mapping and conversion logic.
[0088] Testing with simulated data: Use simulated data for testing to verify the correctness of the conversion logic in different situations.
[0089] D. Automatically define the data acquisition logic using predefined rules, including filtering conditions, sorting rules, etc., so as to dynamically adapt to changes in the data source without manual adjustment, including the following steps:
[0090] The processing logic of the predefined rules is as follows: After obtaining data from the low-code development platform, obtain the integrity, duplication, and update frequency of the data;
[0091] Comprehensively calculate the condition coefficient of the data by combining the integrity, duplication, and update frequency. The expression is:
[0092] ,
[0093] In the formula, TJS is the condition coefficient, WZD, GXP, and CFD are the integrity, update frequency, and duplication respectively, n is the number of data segments, WZDi represents the integrity of the i-th data segment, GXPi represents the update frequency of the i-th data segment, CFDi represents the duplication of the i-th data segment, and α, β, and γ are the proportionality coefficients of the integrity, update frequency, and duplication respectively, and α, β, and γ are all greater than 0;
[0094] After obtaining data from a low-code development platform, this application divides the data into multiple data segments, and obtains the completeness, duplication, and update frequency of the multiple data segments. By comprehensively calculating the completeness, duplication, and update frequency, the condition coefficient of the data is obtained. After obtaining the condition coefficient of the data, the condition coefficient is compared with a preset condition threshold to analyze the quality of the overall data. Through the comprehensive analysis method, not only the data processing efficiency is improved, but also the analysis is more comprehensive and accurate.
[0095] After obtaining the condition coefficient of the data, the condition coefficient is compared with a preset condition threshold. If the condition coefficient is less than the condition threshold, it is analyzed that the data quality is low, and the data is screened out.
[0096] After obtaining all the data, all the data is sorted from largest to smallest according to the condition coefficient to generate a data list, and the data list is sent to the administrator.
[0097] Completeness:
[0098] Detect missing values: Query the missing values in the dataset to understand the frequency and distribution of missing values in each field.
[0099] Use statistical functions: Use statistical functions (such as COUNT, SUM) of the database or data analysis tool to calculate the number of non-missing values in each column or row, so as to evaluate the integrity of the data.
[0100] Data dictionary: If there is a data dictionary or metadata, view the definition and expected values of the fields, and compare them with the actual data to evaluate the completeness of the data.
[0101] Timeliness:
[0102] View timestamps: For data containing timestamps, observe the latest timestamp to understand the update frequency of the data.
[0103] Analyze historical data: Check historical data to understand the data update pattern, such as daily, weekly, or monthly updates.
[0104] Monitor data sources: Use monitoring tools or query the update logs of data sources to understand the update frequency of data in real time.
[0105] Duplication:
[0106] Check uniqueness constraints: For database tables, check whether there are uniqueness constraints to ensure the uniqueness of the data.
[0107] Use aggregation functions: Use aggregation functions (such as COUNT, DISTINCT) to calculate the number of unique values of fields or rows to evaluate the duplication of the data.
[0108] Data quality tools: Use data quality tools, which typically provide functions for duplicate data detection and deduplication.
[0109] E. Automatically execute data acquisition operations according to a schedule and immediately trigger data acquisition operations when the data source changes. Process the encryption and access control of sensitive data to ensure that only authorized users or systems can obtain specific data, including the following steps:
[0110] Scheduling and triggering mechanism: Scheduled task setting: Configure scheduled tasks to automatically execute data acquisition operations according to a schedule. This may involve timed triggering, such as executing once a day, once a week, or once a month.
[0111] Real-time triggering mechanism: Implement a real-time triggering mechanism to immediately trigger data acquisition operations when the data source changes. An event-driven mechanism, such as Webhooks or message queues, can be used.
[0112] Encryption of sensitive data: Data encryption algorithm: Select an appropriate encryption algorithm to encrypt sensitive data. Common encryption algorithms include AES, RSA, etc.
[0113] Encryption key management: Manage the generation, storage, and rotation of encryption keys to ensure the security of the keys.
[0114] Field-level encryption: During the data acquisition process, perform field-level encryption on sensitive fields to ensure that only authorized users can decrypt and access them.
[0115] Access control and authentication: User authentication: Configure a user authentication mechanism to ensure that only authorized users can access the system.
[0116] Access Control List (ACL): Use ACL to define access control rules for data sources and specify which users or systems are authorized to access specific data.
[0117] Role-based access control: Use role-based access control to assign users to different roles and grant corresponding permissions to the roles.
[0118] Key management and secure storage: API key or credential management: For API access, manage API keys or credentials and ensure their secure storage and regular rotation.
[0119] Secure storage: Ensure that all information such as credentials and keys for accessing sensitive data is stored in a secure manner. A dedicated key management service or Hardware Security Module (HSM) can be used.
[0120] Audit and Monitoring: Audit Logs: Record audit logs for all data acquisition operations, including information such as who, when, and what data was acquired from which data source.
[0121] Real-time Monitoring: Monitor the access activities of the system in real time, detect abnormal behaviors in a timely manner, and take appropriate response measures.
[0122] Automatic Fault Handling: Anomaly Detection and Handling: During the data acquisition process, the system should be able to detect abnormal situations, such as connection failures, permission errors, etc., and automatically perform corresponding fault handling.
[0123] Alarm and Notification: Specify the alarm and notification mechanism for anomalies and notify relevant personnel or system administrators in a timely manner.
[0124] Regular Security Reviews: Regular Reviews: Regularly review the security of the data acquisition system, including aspects such as permission settings, encryption implementation, access control, etc.
[0125] Vulnerability Management: Promptly repair the vulnerabilities found in the system to ensure the security of the system.
[0126] F. Automatically perform tests after data acquisition, verify whether the data acquisition process is working properly, and automatically record monitoring data to promptly detect and resolve potential problems, including the following steps:
[0127] Automated Test Script Design:
[0128] Data Acquisition Test Cases: Design automated test cases to cover different data acquisition scenarios, including different data sources, mapping rules, transformation logics, etc.
[0129] Abnormal Situation Testing: Write test cases to simulate possible abnormal situations, such as unavailable data sources, incorrect data formats, permission issues, etc.
[0130] Automated Test Execution:
[0131] Test Automation Tools: Use appropriate automated test tools, such as Selenium, JUnit, Postman, etc., to execute the designed automated test scripts.
[0132] Simulate Data Changes: Simulate changes in the data source during testing to ensure that the real-time trigger mechanism can work properly.
[0133] Monitoring Data Recording:
[0134] Performance Monitoring: Integrate performance monitoring tools to record the performance metrics of the data acquisition system, such as response time, throughput, etc.
[0135] Error Log: Record detailed error logs, including exceptions, error codes, stack traces, etc., for subsequent analysis and troubleshooting.
[0136] Data Quality Monitoring: Monitor data quality metrics, such as duplicate data, missing data, outliers, etc.
[0137] Automated Test Report Generation:
[0138] Test Report: Automatically generate a detailed automated test report, including information such as test coverage, pass rate, failed test cases, execution time, etc.
[0139] History Record: Retain historical test reports to compare the performance and stability of data acquisition processes in different versions.
[0140] Real-time Alarms and Notifications:
[0141] Anomaly Detection: Real-time monitor the running status of the system, detect anomalies, and trigger alarms.
[0142] Notification Mechanism: Set up a notification mechanism to promptly notify relevant operation and maintenance personnel or developers of anomalies for timely response and problem resolution.
[0143] Log Analysis and Tracing:
[0144] Log Analysis Tool: Integrate a log analysis tool to analyze the running logs of the system and identify potential problems and abnormal patterns.
[0145] Distributed Tracing: Implement distributed tracing in the system to track each step in the data acquisition process and find potential performance bottlenecks or error sources.
[0146] Automated Recovery Mechanism:
[0147] Automated Rollback: When a problem is detected, the system should be able to automatically perform a rollback operation to restore the system state to a normal working state.
[0148] Error Fix: Problems discovered in automated testing should be promptly fixed and verified in the next version.
[0149] Regular Performance Optimization and Tuning:
[0150] Performance Optimization Strategy: Based on monitoring data and test results, formulate regular performance optimization strategies to ensure the system runs in the best state.
[0151] Resource Adjustment: Based on system load and performance monitoring results, perform necessary resource adjustments, such as adding servers, optimizing database indexes, etc.
[0152] G. Automatically handle exceptions during data acquisition, such as unavailable data sources, connection failures, etc., and take appropriate measures, such as triggering alerts or performing automatic repairs, including the following steps:
[0153] Exception detection mechanism:
[0154] Monitor system status: Real-time monitor the running status of the data acquisition system, including data source connection status, response time, data acquisition success rate, etc.
[0155] Define exception metrics: Determine the metrics for exception situations, such as timeouts, connection errors, data format exceptions, etc.
[0156] Design exception handling strategies:
[0157] Formulate exception handling strategies: For different exception situations, design corresponding handling strategies. For example, for an unavailable data source, you can choose to retry the connection; for a connection failure, trigger an alarm and perform automatic repair.
[0158] Automated exception handling:
[0159] Automatic retry: When a connection failure or unavailable data source is detected, the system can automatically perform a certain number of retries to attempt to restore a normal connection.
[0160] Automatically switch to backup data source: If there is a backup data source, the system can automatically switch to the backup data source to ensure that the data acquisition process is not interrupted.
[0161] Automatic repair: For problems that can be automatically repaired, the system should be able to automatically execute repair actions, such as restarting services, reloading configurations, etc.
[0162] Alarms and notifications:
[0163] Trigger an alarm: When an exception situation is detected, the system should trigger an alarm to notify the relevant operations and maintenance team or administrator.
[0164] Notification mechanism: Set up a notification mechanism to notify relevant personnel in a timely manner, such as by email, SMS, instant messaging, etc.
[0165] Logging:
[0166] Detailed logging: When an exception occurs, the system should generate detailed log records, including the type of exception, timestamp, relevant data, etc., for subsequent analysis and troubleshooting.
[0167] Error code identification: Define error codes for different types of exceptions so that the system can identify the problem type based on the error code.
[0168] Automated rollback:
[0169] Rollback mechanism: If certain operations during data acquisition go wrong due to abnormal situations, the system should be able to automatically perform rollback operations to restore the system state to the normal working state.
[0170] Transaction management: For operations that support transactions, ensure rollback in case of exceptions to maintain data consistency.
[0171] System self-healing function:
[0172] Self-healing mechanism: For problems that can be automatically repaired, the system should implement a self-healing function, such as automatically restarting affected services or modules.
[0173] Health check: Regularly perform system health checks, automatically repair damaged components or services to improve system availability.
[0174] Monitoring system adjustment:
[0175] Dynamic adjustment of monitoring strategy: According to the experience of exception handling and the system operation situation, dynamically adjust the monitoring strategy to improve the accuracy and sensitivity of exception detection.
[0176] Predictive maintenance: Based on historical data and exception patterns, implement predictive maintenance to detect potential problems in advance and take preventive measures.
[0177] Embodiment 3: The data acquisition system based on the low-code development platform described in this embodiment includes an identification module, a mapping and conversion module, a data screening module, a timing module, and a testing module;
[0178] Identification module: Detect and identify data sources by automatically scanning the network, API catalog, or metadata. The data sources include databases, file systems, and cloud services. The data source identification results are sent to the connection module;
[0179] Mapping and conversion module: Automatically establish a connection after discovering the data source, and use predefined connectors to automatically identify and adapt to the connection methods of various data sources, automatically infer the relationships between different data sources, and perform data mapping and conversion;
[0180] Data screening module: Automatically define data acquisition logic using predefined rules, including filtering conditions and sorting rules;
[0181] Timing module: Automatically execute data acquisition operations according to the plan, and immediately trigger data acquisition operations when the data source changes, handling encryption and access control of sensitive data;
[0182] Testing module: Automatically perform tests after obtaining data, verify whether the data acquisition process is working properly, automatically record monitoring data, and automatically handle exception situations and take measures during data acquisition.
[0183] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0184] In the description of this specification, the description referring to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0185] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to only the specific implementation manners. Obviously, according to the content of this specification, many modifications and variations can be made. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A data acquisition method based on a low-code development platform, characterized in that: The acquisition method comprises the following steps: The acquisition system detects and identifies data sources by automatically scanning the network, API directory or metadata. Data sources include databases, file systems, and cloud services. Once a data source is found, a connection is automatically established. Use predefined connectors to automatically identify and adapt to the connection methods of various data sources, and obtain the system to automatically infer the relationship between different data sources, perform data mapping and conversion; Automatically define data acquisition logic using predefined rules, including filtering conditions and sorting rules, automatically execute data acquisition operations according to schedule, and immediately trigger data acquisition operations when data sources change, handle encryption and access control of sensitive data; After acquiring the data, the system automatically performs tests to verify whether the data acquisition process is working properly, and automatically records the monitoring data. It also automatically handles abnormal situations and takes measures during the data acquisition process.
2. The data acquisition method based on the low-code development platform according to claim 1 is characterized in that: Automatically define data acquisition logic using predefined rules, including the following steps: After obtaining data from the low-code development platform, obtain the completeness, duplication, and update frequency of the data; The condition coefficient of the data is obtained by comprehensively calculating the completeness, repetition and update frequency. The expression is: , Where TJS is the conditional coefficient, WZD, GXP, and CFD are completeness, update frequency, and repetition, respectively, n is the number of data segments, WZDi represents the completeness of the i-th data segment, GXPi represents the update frequency of the i-th data segment, CFDi represents the repetition of the i-th data segment, α, β, and γ are the proportional coefficients of completeness, update frequency, and repetition, respectively, and α, β, and γ are all greater than 0.
3. The data acquisition method based on the low-code development platform according to claim 2 is characterized in that: The filtering conditions and sorting rules are: After obtaining the condition coefficient of the data, the condition coefficient is compared with the preset condition threshold. If the condition coefficient is less than the condition threshold, the quality of the analyzed data is low and the data is screened out; After obtaining all the data, all the data are sorted from large to small according to the condition coefficient, a data list is generated, and the data list is sent to the administrator.
4. The data acquisition method based on the low-code development platform according to claim 3 is characterized in that: The logic of obtaining completeness is as follows: query the missing values in the data set, understand the frequency and distribution of missing values in each field, and use the statistical functions of the database or data analysis tool to calculate the number of non-missing values in each column or row; The logic for obtaining the update frequency is: use monitoring tools or query the update log of the data source to obtain the data update frequency in real time; The logic for obtaining the duplication is: use aggregate functions to calculate the number of unique values of a field or row to obtain the duplication of the data.
5. The data acquisition method based on the low-code development platform according to claim 4 is characterized in that: Automatically establish connections after discovering data sources. The connection method of automatically identifying and adapting to various data sources using predefined connectors includes the following steps: A connector library is built in the acquisition system, which contains predefined connectors. Each connector is used for a specific type of data source. Connectors include database connectors, API connectors, cloud service connectors, etc. When the acquisition system detects the data source, it matches it according to the type of data source to determine which predefined connector to use. The acquisition system automatically selects the connector that matches the detected data source type and automatically extracts the information required for the connection from the metadata of the data source. For different data sources, the connector dynamically configures the connection parameters. Before establishing the connection, the acquisition system automatically performs a connection test.
6. The data acquisition method based on the low-code development platform according to claim 5 is characterized in that: The acquisition system automatically infers the relationship between different data sources, and performs data mapping and conversion, including the following steps: Extract metadata from each data source, including table structure, field type, and key relationship information, perform statistics and analysis on the data, use heuristic methods to infer the relationship between data sources, and automatically generate mapping rules between data sources based on the analysis results, including field mapping, value conversion, and data type conversion rules. Use intelligent matching algorithms to automatically generate data type conversion logic based on mapping rules, and generate corresponding mapping and conversion logic for similar but different data values in different data sources.
7. The data acquisition method based on the low-code development platform according to claim 6 is characterized in that: After acquiring the data, the test is automatically performed to verify whether the data acquisition process is working properly and automatically record the monitoring data. The following steps are included: Use automated testing tools to execute designed automated testing scripts, simulate changes in data sources during testing, integrate performance monitoring tools, record performance indicators of the data acquisition system, record detailed error logs, including exceptions, error codes, stack traces, and automatically generate detailed automated testing reports, including test coverage, pass rate, failed cases, execution time, and integrate log analysis tools to analyze the operation logs of the acquisition system.
8. The data acquisition method based on the low-code development platform according to claim 7 is characterized in that: Automatically handle exceptions during data acquisition, including the following steps: Monitor the running status of the data acquisition system in real time, including data source connection status, response time, and data acquisition success rate. If the data source is unavailable, choose to retry the connection. If the connection fails, trigger an alarm and perform automatic repairs. When a connection failure or data source is found to be unavailable, automatically retry a certain number of times and automatically perform repair actions. When an abnormal situation occurs, generate detailed log records, including the type of exception, timestamp, and related data, and define error codes for different types of exceptions. If the abnormal situation causes certain operations in the data acquisition process to fail, the acquisition system automatically performs a rollback operation to restore the acquisition system status to normal working status.
9. A data acquisition system based on a low-code development platform, used to implement the acquisition method according to any one of claims 1 to 8, characterized in that: It includes recognition module, mapping conversion module, data screening module, timing module and testing module; Identification module: detects and identifies data sources by automatically scanning the network, API directory or metadata. Data sources include databases, file systems, and cloud services. The data source identification results are sent to the connection module. Mapping and conversion module: automatically establishes a connection after discovering a data source, and uses predefined connectors to automatically identify and adapt to the connection methods of various data sources, automatically infer the relationship between different data sources, and perform data mapping and conversion; Data filtering module: automatically defines data acquisition logic using predefined rules, including filtering conditions and sorting rules; Timing module: automatically executes data acquisition operations according to the plan, triggers data acquisition operations immediately when the data source changes, and handles encryption and access control of sensitive data; Test module: automatically conducts tests after acquiring data to verify whether the data acquisition process is working properly, and automatically records monitoring data, automatically handles abnormal situations and takes measures during the data acquisition process.
Citation Information
Cited By
Data source fault processing system based on AI server
CN120315968A
Data monitoring system and method based on low-code platform
CN121387677A
A data monitoring system and method based on a low-code platform
CN121387677B