Adaptive data knitting performance optimization method based on artificial intelligence
By generating a logical data carrier containing multi-dimensional metadata and monitoring resource load characteristics in real time, the problem of insufficient adaptability of data weaving solutions in existing technologies is solved, and dynamic optimization of data weaving paths and efficient utilization of resources are achieved.
Patent Information
- Application Number
- CN202511284546.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-09-10
Smart Images

Figure CN120780876A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial technology, and in particular to an adaptive data weaving performance optimization method based on artificial intelligence. Background Art
[0002] In the process of enterprise digital transformation, the integration and efficient utilization of multi-source, heterogeneous data has become a core requirement for supporting business decision-making and improving operational efficiency. Currently, enterprise data sources exhibit significant heterogeneity, encompassing relational databases (such as MySQL and Oracle), NoSQL databases (such as MongoDB and Redis), unstructured files (such as CSV and JSON documents), and real-time data streams from IoT devices. This data requires cross-source integration through data weaving technology to meet the access requirements for a unified data view in diverse business scenarios such as real-time transaction analysis, batch data statistics, and static report queries. This is particularly true in sectors such as finance and e-commerce, which place higher demands on the real-time, stability, and resource adaptability of data weaving.
[0003] Existing data weaving solutions mostly adopt a static implementation approach: first, they rely on manual sorting of the basic attributes of the data source (such as storage location, data format, and other technical information), and establish the correspondence between the data source and the data access interface through manual configuration; then, they build the data weaving path based on preset fixed path rules (such as allocating fixed transmission protocols and processing nodes according to the data source type); at the resource scheduling level, fixed computing and storage resources are usually pre-allocated according to business scenarios, such as a fixed allocation of 50% of memory resources for batch data processing scenarios, and the path and resource allocation strategy are rarely dynamically adjusted after they are determined, lacking the ability to adaptively respond to real-time business changes and system status.
[0004] However, in traditional solutions, on the one hand, the metadata processing dimension is single, focusing only on technical metadata, and failing to integrate key information such as business semantics (such as the business domain to which the data belongs, business priority), operation records (such as historical access frequency, data processing time) and management specifications (such as data sensitivity level, compliance requirements). As a result, the constructed weaving path is difficult to adapt to complex and changing business needs. For example, it is impossible to dynamically adjust the data processing priority according to business urgency; on the other hand, resource scheduling relies on static rules and cannot monitor the load changes of the data access task queue (such as a sudden increase in task volume, high-priority tasks jumping in line) and system resource indicators (such as CPU occupancy, memory usage) in real time. When resource bottlenecks or task load fluctuations occur, resource allocation and path planning cannot be adjusted in time, which can easily lead to increased data access latency, low resource utilization, and even interruption of the data weaving process. It is difficult to meet the actual needs of enterprises for efficient, stable and adaptive integration of multi-source heterogeneous data. Summary of the Invention
[0005] To solve the above technical problems, the application provides an adaptive data weaving performance optimization method based on artificial intelligence to at least alleviate the above technical problems.
[0006] The technical scheme provided by the embodiments of the application is as follows: An adaptive data weaving performance optimization method based on artificial intelligence, comprising: Step 1, acquiring multi-source heterogeneous data based on a data source list and parsing the multi-source heterogeneous data to generate a metadata set containing at least technical metadata, business metadata, operation metadata and management metadata; Step 2, based on a virtual table interface, performing virtual table abstraction and encapsulation processing on the metadata set to generate a logical data carrier; Step 3, determining the operation attribute of the metadata set to generate a data weaving path of the logical data carrier based on the operation attribute; Step 4, monitoring a data access task queue and access index feedback in real time, extracting resource load correlation features based on a constructed resource scheduling model for the data access task queue and the access index feedback, and generating a resource scheduling strategy according to the resource load correlation features; Step 5, optimizing the data weaving path of the logical data carrier based on the resource scheduling strategy to obtain an optimized data weaving path to perform data weaving on the multi-source heterogeneous data.
[0007] In the application, multi-source heterogeneous data is acquired based on a data source list and parsed to generate a metadata set containing technical metadata, business metadata, operation metadata and management metadata. Compared with traditional single technical metadata manually sorted, the metadata set supplements information such as the business domain to which the data belongs, business priority (business metadata), historical access frequency, data processing time (operation metadata), and data sensitivity level and compliance requirements (management metadata), thereby providing more comprehensive decision basis for subsequent generation of data weaving paths. For example, based on the “real-time transaction business priority” in the business metadata and the “high-frequency access record” in the operation metadata, a more optimal processing resource can be allocated to data in a real-time transaction scenario, avoiding the defect that the traditional scheme cannot dynamically adjust according to business urgency due to the lack of business and operation dimension information, so that the subsequently generated path is more in line with actual business requirements.
[0008] In addition, in the present application, the metadata set is abstracted and encapsulated based on a virtual table interface to generate a logical data carrier. The corresponding relationship between the data source and the interface does not need to be manually established, but the metadata field and the virtual table structure are dynamically mapped through the virtual table, and the physical differences (such as database types and file formats) of the underlying multi-source heterogeneous data are shielded at the logical layer. Compared with the traditional manual configuration, this method not only reduces the labor cost, but also forms a unified logical data access layer, so that the subsequent data weaving path generation does not need to pay attention to the physical properties of the specific data source, and only needs to be based on the standardized logical carrier planning. The problem of poor interface adaptability and high maintenance cost in the traditional interface is solved, and a standardized basic carrier is provided for path generation.
[0009] In the present application, the operation attribute of the metadata set is determined (the operation attribute can be derived based on the operation metadata and business metadata in step 1, such as determining the operation attribute of “real-time read-write type” and “batch processing type” by combining the historical data processing frequency and the business SLA requirement), and then the data weaving path of the logical data carrier is generated based on the operation attribute. Compared with the traditional fixed path, the generated path can be differentially adapted according to the operation attribute, for example, the “real-time read-write type” operation attribute is assigned a low-delay transmission protocol and a high-frequency processing node, and the “batch processing type” operation attribute is assigned a high-throughput node, which solves the problem that the traditional path cannot adapt to different business scenario requirements, and makes the resource consumption of the path more matched with the business demand.
[0010] In the present application, by monitoring the data access task queue and access index feedback in real time, the resource scheduling model is constructed to extract the resource load correlation characteristics (such as correlating “high-priority task proportion 30%” with “CPU occupancy rate 85%”), and then a resource scheduling strategy is generated. Compared with the traditional static resource allocation, this method can real-time perceive the system state and task load fluctuation, for example, when it is monitored that the real-time transaction task queue length increases sharply and the CPU occupancy rate rises, the resource bottleneck can be identified in advance and the strategy of “allocating CPU resources to real-time tasks in priority” is generated, which solves the problem that the traditional method cannot respond to load changes in real time, and provides timely and accurate scheduling basis for subsequent path optimization.
[0011] In the present application, the data weaving path of the logical data carrier is optimized based on the generated resource scheduling strategy, which can adjust the path in a targeted manner when the resource bottleneck or load changes, for example, when the load of a certain node is too high, the tasks on the node are migrated to a low-load node; when the batch tasks occupy too much memory, the memory occupancy mode of the batch data processing path is optimized. Through dynamic optimization, the data access delay can be reduced, the resource utilization rate can be improved, and the process interruption problem caused by the traditional path without adjustment can be avoided, so that the multi-source heterogeneous data can maintain good performance and stability in different load scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 This is a flow chart of the adaptive data weaving performance optimization method based on artificial intelligence according to an embodiment of the present application. DETAILED DESCRIPTION
[0013] like Figure 1 As shown, an embodiment of the present application provides an artificial intelligence-based adaptive data knitting performance optimization method, the method comprising: Step 1: Based on the data source list, obtain multi-source heterogeneous data and parse it to generate a metadata set containing at least technical metadata, business metadata, operational metadata, and management metadata; Step 2: Based on the virtual table interface, perform virtual table abstraction and encapsulation processing on the metadata set to generate a logical data carrier; Step 3: determining the operational attributes of the metadata set to generate a data weaving path of the logical data carrier based on the operational attributes; Step 4: Monitor the data access task queue and access indicator feedback in real time, extract resource load correlation features from the data access task queue and access indicator feedback based on the constructed resource scheduling model, and generate a resource scheduling strategy based on the resource load correlation features; Step 5: Optimize the data weaving path of the logical data carrier based on the resource scheduling strategy to obtain the optimized data weaving path to weave multi-source heterogeneous data.
[0014] Optionally, step 1 specifically includes the following steps: Step 11: Perform type parsing on the multi-source heterogeneous data sources in the preset data source list to identify the storage engine type and transmission protocol type and generate a detailed list of data source types accordingly; Step 12: Based on the detailed list of data source types, perform protocol adaptation on the multi-source heterogeneous data sources to obtain multi-source heterogeneous data; Step 13: Perform data field segmentation processing on the multi-source heterogeneous data to obtain a structured field set, and perform metadata parsing on the structured field set to generate a metadata set that at least includes technical metadata, business metadata, operational metadata, and management metadata.
[0015] Optionally, step 11 specifically includes the following steps: Step 111: Based on the set data source matching rule library, a type parser is constructed. Based on the constructed type parser, type parsing is performed on the multi-source heterogeneous data sources in the preset data source list to obtain a data source type description; Step 112: extract features from the data source type description to generate a data source type feature coding table; Step 113: According to the set storage-computing protocol coding mapping rules, type matching and verification processing are performed on the data source type feature coding table to identify the storage engine type and transmission protocol type and generate a detailed list of data source types accordingly.
[0016] Preferably, the specific implementation process of step 111 is as follows: first, the set data source matching rule library is structured and initialized, and the rule library contains core identification rules for multi-source heterogeneous data sources - the technical rules cover the data source identification prefix (such as "mysql: / / " corresponding to relational database, "mongodb: / / " corresponding to NoSQL database, "mqtt: / / " corresponding to IoT stream data source), default port range (MySQL default port 3306, MongoDB default port 27017, MQTT default port 1883), feature field identification (such as the "information_schema" system table field of the relational database, the "collection" collection field of the NoSQL database), and the business rules cover the business domain to which the data source belongs (such as "order system-relational" and "device monitoring-IoT stream") to generate a structured data source matching rule library; then, an extensible rule engine type parser is constructed based on the structured rule library. The parser is different from the traditional fixed rule parser and supports dynamic addition of new data source rules through JSON configuration (such as adding "clickhouse : / / " identifier corresponding to the time series database rules), without modifying the underlying code to adapt to the expansion needs of enterprise data source types; then, the multi-source heterogeneous data sources in the preset data source list (such as "mysql: / / 192.168.1.100:3306 / order_db" and "mqtt: / / 192.168.1.200:1883 / device_data") are input into the type parser. The parser performs parsing in the order of "identifier prefix matching → port verification → feature field detection". For example, for "mysql: / / 192.168.1.100:3306 / order_db", the "mysql: / / " prefix is first matched, then the 3306 port is checked to see if it is within the MySQL default port range, and finally the "information_schema" field is detected. Finally, a data source type description containing the data source ID, identifier prefix, port, and feature field detection results is obtained (such as "DS001, identifier prefix: mysql: / / , port: 3306, feature field: information_schema exists, preliminary judgment: relational database").
[0017] Preferably, in the enterprise multi-source heterogeneous data source batch parsing scenario, when step 112 is specifically implemented: first, the data source type description generated in step 111 is used as the processing object, and multi-dimensional feature extraction processing is performed on each data source type description - extracting technical features (storage type: determined as relational / non-relational / file / stream data type based on the identification prefix and feature field, initial judgment of transmission protocol: determined as TCP / UDP / MQTT / HTTP based on the identification prefix), attribute features (whether transactions are supported: relational databases support by default, non-relational databases do not support by default; whether it is real-time streaming: MQTT identifier corresponds to streaming data, file identifier corresponds to non-streaming data), to obtain a multi-dimensional feature set for each data source (such as "DS001, storage type: relational, initial judgment of transmission protocol: TCP, support transactions: yes, real-time streaming: no"); then, a multi-dimensional feature hierarchical coding scheme is used to encode the multi-dimensional feature set. This scheme is different from the traditional single coding and divides the features into 3 coding levels, with a total of 32-bit binary code - the first 8 bits are storage type code (00000001=relational The middle 8 bits are the initial judgment code of the transmission protocol (00000001=TCP, 00000010=UDP, 00000100=MQTT, 00001000=HTTP), and the last 16 bits are the attribute feature code (000000000000001=support transaction, 00000000000001=real-time stream, and the remaining bits are reserved for extension); finally, each data source is The encoding results are organized according to the column structure of "data source ID, storage type code, transmission protocol preliminary judgment code, and attribute feature code" to generate a data source type feature coding table. Each row in the table corresponds to a data source, and the intersection of rows and columns is the specific coding value. For example, the data source ID "DS001" corresponds to the storage type code "00000001", the transmission protocol preliminary judgment code "00000001", and the attribute feature code "0000000000000001", which clearly maps the characteristics of "relational database, TCP transmission, and transaction support".
[0018] Preferably, in the specific technical implementation of step 113: first, load the preset storage and computing protocol coding mapping rule, which is a two-way mapping table - the forward mapping is "feature coding → specific type" (storage type code 00000001 corresponds to relational storage engines such as MySQL / Oracle, 00000010 corresponds to non-relational storage engines such as MongoDB / Redis; the transmission protocol initial judgment code 00000001 corresponds to JDBC / ODBC protocol, 00000100 corresponds to MQTT protocol), and the reverse mapping is "specific type → feature coding" (MySQL corresponds to storage type code 00000001, JDBC protocol corresponds to transmission protocol code 00000001), and also includes verification rules (such as when the storage engine is MySQL, the transmission protocol must support JDBC / ODBC, and if it matches MQTT, it is judged to be abnormal); then, based on the mapping rule, perform type matching processing on the data source type feature coding table generated in step 112 - convert the data source type feature coding table in the coding table to the data source type feature coding table. The storage type code is forward matched with the mapping rule to obtain the storage engine type candidate (such as the code 00000001 matches "MySQL, Oracle"), and the transmission protocol initial judgment code is forward matched to obtain the transmission protocol type candidate (such as the code 00000001 matches "JDBC, ODBC"); then, a two-way verification process is performed: one is to verify the coding accuracy based on the reverse mapping (such as the reverse mapping code of the storage engine candidate "MySQL" is 00000001, which is consistent with the coding table, then the verification passes); the other is to verify the type compatibility based on the verification rule (such as the storage engine "MySQL" is compatible with the transmission protocol "JDBC", the verification passes; if the transmission protocol candidate is "MQTT", it is marked as a compatibility exception), and the abnormal code is completed by rules or manually confirmed (such as adding a special mapping rule for "MySQL+MQTT"); finally, the storage engine type and transmission protocol type that have passed the verification are associated with the data source ID, and organized into a list containing "data source ID, storage engine type (such as MySQL 8.0), transport protocol type (such as JDBC 4.2), verification status (passed), and remarks (no exceptions).
[0019] Optionally, step 12 specifically includes the following steps: Step 121: Perform structured parsing on the detailed list of data source types to generate a data source adaptation basic information table; Step 122: The data source adaptation basic information table is matched with a preset "storage engine type - transmission protocol type" two-dimensional pre-configured protocol adaptation template library to generate a protocol adaptation template. Step 123: Based on the protocol adaptation template, protocol handshake and data reading processing are performed on the multi-source heterogeneous data sources to obtain multi-source heterogeneous data, and log writing processing is performed on the adaptation key information to generate an adaptation log table.
[0020] Preferably, the specific implementation process of step 121 is as follows: first, the detailed list of data source types generated in step 113 is used as the processing object, and the list includes fields such as "data source ID, storage engine type (such as MySQL 8.0, MongoDB 6.0), transmission protocol type (such as JDBC 4.2, MQTT 3.1.1), and verification status". A structured parsing process is performed on the list to extract the core adaptation elements of each type of data source, including the data source unique identifier (data source ID), storage engine detailed model (to distinguish different versions, such as the adaptation parameters of MySQL 8.0 and MySQL 5.7 are different), and transmission protocol version (to clarify the protocol version number to match the corresponding connection parameters, such as the SSL configuration supported by JDBC 4.2 and JDBC 4.0), data source physical address (IP address + port number + target library / collection / topic name, such as "192.168.1.100:3306 / order_db" and "192.168.1.200:1883 / device_topic"), authentication information identifier (marking whether username and password authentication is required, key file path or token validity period); then, the extracted core adaptation elements are sorted according to the column structure of "data source ID, storage engine model, transmission protocol version, physical address, authentication information identifier" to generate a data source adaptation basic information table. Each row in the table corresponds to a data source, and the intersection of rows and columns is the specific adaptation element value (such as the data source ID "DS001" corresponds to the storage engine model "MySQL 8.0", the transmission protocol version "JDBC 4.2", physical address "192.168.1.100:3306 / order_db", authentication information identifier "Username and password required, path: / conf / auth / mysql.xml"); this technology differs from the traditional parsing method that only extracts the basic "storage engine-protocol" combination. It adds the "protocol version" and "authentication information identifier" elements. Because the connection parameters of different protocol versions (such as the "serverTimezone" parameter in JDBC 4.2, which does not exist in JDBC 4.0) and authentication methods differ, which directly affects the adaptation success rate. Parsing these elements in advance can reduce the number of parameter debugging times in the subsequent adaptation process and improve adaptation efficiency.
[0021] Preferably, in the batch protocol adaptation scenario of enterprise multi-source heterogeneous data sources (including relational databases, NoSQL databases, IoT stream data sources), when step 122 is specifically implemented: first, load the preset "storage engine type-transmission protocol type" two-dimensional pre-configured protocol adaptation template library. The special design of this template library is that the row dimension is the storage engine type (subdivided into subtypes such as MySQL, Oracle, MongoDB, Kafka, MQTT device, etc.), the column dimension is the transmission protocol type (subdivided into subtypes such as JDBC, ODBC, MongoDB Driver, Kafka Consumer API, MQTT Client, etc.), and the intersection of rows and columns is the exclusive protocol adaptation template for the corresponding combination. Each template contains three core parameters: connection parameters (such as the MySQL+JDBC template's "url format: jdbc:mysql: / / {IP}:{Port} / {DB}?useSSL={Bool}&serverTimezone={Timezone}", the timeout default value "30s"), data reading parameters (such as MongoDB+MongoDB The Driver template's "Batch size: 1000 records / batch", the field projection rule "Return all fields by default, optionally excluding the _id field", and exception handling parameters (such as the MQTT device + MQTT Client template's "Retry times: 3 times" and "Retry interval: 5s"). Next, using the data source adaptation basic information table generated in step 121 as the matching basis, extract the "Storage Engine Model" and "Transport Protocol Version" for each row in the table. First, match the template library's row dimension by "Storage Engine Model" (for example, MySQL 8.0 matches the "MySQL" row), and then match the column dimension by "Transport Protocol Version" (for example, JDBC 4.2 matches the "JDBC" column). If a fully matching template exists, it is directly called. If only the storage engine type matches but the protocol version does not match (for example, there is no dedicated template for JDBC 4.2), the protocol general template for that storage engine type (such as the MySQL+JDBC general template) is called and version compatibility parameters are automatically added (for example, adding "serverTimezone=UTC" to the general template to adapt to JDBC). 4.2); Finally, a protocol adaptation template containing exclusive parameters is generated for each data source. This technology differs from traditional single-dimensional template libraries (based only on storage engine or protocol). The dual-dimensional design can accurately cover the "storage engine-protocol" combination scenario, avoiding adaptation failures caused by traditional single-dimensional templates that do not distinguish between protocol types (such as the difference in adaptation parameters between JDBC and ODBC for MySQL), thereby improving template matching accuracy.
[0022] Preferably, in the specific technical implementation of step 123: first, based on the protocol adaptation template generated in step 122, a protocol handshake process is initiated for each multi-source heterogeneous data source - the connection parameters in the template are called (such as the URL, username and password of the MySQL+JDBC template) to build a connection request, and a protocol version negotiation instruction is sent to the data source (such as the JDBC connection sends the "handshake_v10" instruction, and the MQTT connection sends the "CONNECT" message). After receiving the response message returned by the data source, the response is verified to see whether it meets the success identifier preset in the template (such as the JDBC connection returns "OK_Packet", the MQTT connection returns the "CONNACK" message and the return code is 0). If the verification passes, the protocol handshake is completed. If it fails, the exception handling parameters in the template are triggered (such as retries 3 times with an interval of 5s). After the protocol handshake is successful, the data reading parameters in the template are called to perform data reading processing - for relational databases (such as MySQL), table data is read batch by batch according to the "batch size", and for NoSQL databases (such as MongoDB), the data is read batch by batch according to the "batch size". The data source (DB) filters fields according to the "field projection rule" and then reads them. It receives streaming data in real time from IoT streaming data sources (such as MQTT devices) by "subscribing to topics" and records the read time and data volume for each batch of data. Subsequently, key adaptation information (data source ID, protocol handshake start time, handshake end time, handshake result, data read start time, read end time, data volume, read time, and exception information (blank if no exceptions)) is written to a log file in a preset format. This generates an adaptation log table. Each row in this table corresponds to a single adaptation process for a data source, and each column corresponds to the aforementioned key information. This facilitates subsequent operation metadata statistics (such as extracting the "historical access latency" of operation metadata in step 13). This technology differs from traditional adaptation methods that only complete data reading. It integrates protocol handshake parameter verification, data read performance statistics, and logging. This not only ensures the standardization of the adaptation process but also provides basic data for subsequent metadata analysis and resource scheduling (for example, read time can be used as the basis for calculating resource load correlation characteristics in step 4), avoiding the redundant operation of collecting this information later.
[0023] Optionally, step 13 specifically includes the following steps: Step 131: perform field splitting on structured data in multi-source heterogeneous data according to preset field delimiters, extract nested fields from semi-structured data according to JSON / XML tag paths, and locate key fields and perform field separation on unstructured data to generate structured field units. Step 132: Perform metadata extraction and association mapping processing on the structured field unit to generate a metadata association ledger; Step 133: Perform metadata completion and integration verification on the metadata association ledger to generate a metadata set that at least includes technical metadata, business metadata, operation metadata, and management metadata.
[0024] Preferably, in the field standardization scenario of enterprise multi-source heterogeneous data (including structured transaction data such as MySQL order tables, semi-structured log data such as system logs in JSON format, and unstructured document data such as customer contracts in PDF format), when step 131 is specifically implemented: first, the multi-source heterogeneous data obtained in step 123 is used as the processing object, and the structured data therein (such as CSV files with commas between fields and MySQL data stored in a fixed table structure) is first filtered out, and the preset field delimiter configuration library (including common delimiters such as commas, tabs, and vertical bars, which can be automatically matched according to the data source type) is called. The default characters for CSV and TSV are commas, and the default characters for TSV are tabs. The structured data is split into two separate fields: "order ID", "user ID", "order amount", and "order time". The structured split fields are obtained. Then, for semi-structured data (such as JSON data with nested structures `{"order":{"id":"O001","user":{"id":"U001","name":"张三"}}}`, XML data` <order> <id> O001< / id> <user> <id> U001< / id> < / user> < / order>`), extract nested fields according to the preset JSON / XML tag path rules (JSON paths such as "$.order.id" and "$.order.user.id", XML paths such as " / order / id" and " / order / user / id"), convert nested "order.id" and "order.user.id" fields into flat "order_id" and "user_id" fields to obtain semi-structured extracted fields; then, for unstructured data (such as text containing key fields such as "contract number", "contracting party", and "validity period" in PDF contracts), call the key field positioning module based on the BERT pre-trained model (this module has been fine-tuned through the company's historical contract text and can recognize the target field after prefixes such as "contract number:" and "contracting party:"). Text parsing is performed on unstructured data (PDF is converted into recognizable text). The positioning module then matches the prefix of key fields and extracts the field content after the prefix (e.g., "Contract Number: HT20240801" is extracted as "Contract Number = HT20240801") to obtain unstructured extracted fields. Finally, the structured split fields, semi-structured extracted fields, and unstructured extracted fields are integrated in the format of "data source ID-field name-field value-field source type (structured / semi-structured / unstructured)" to generate structured field units. This technology differs from the traditional method of processing only structured data. It covers three types of heterogeneous data through targeted field processing logic, especially the pre-trained model positioning solution for unstructured data, which solves the problem of traditional unstructured data having difficulty extracting fixed fields and improves the structuring rate of multi-source data.
[0025] Preferably, in the specific technical implementation of step 132: first, the structured field unit generated in step 131 is used as the processing object, and multi-dimensional metadata extraction processing is performed on each structured field unit - extracting technical metadata (field data type such as int / varchar / date, field length such as order ID is 10 digits, whether it is the primary key / foreign key such as "order ID" is the primary key), business metadata (field business meaning such as "order amount" corresponds to "customer actual payment amount", data belongs to the business domain such as "order domain", business person in charge such as "order system operation and maintenance group"), operation metadata (field historical access frequency such as "user ID" with an average daily access of 1000 times, the last update time such as 2024-08-27 10:30), manage metadata (field sensitivity level, such as "user ID" is highly sensitive, data retention period, such as 3 years), and obtain a single-field metadata set; then, load the company's preset business term association library (including business field semantic mapping relationships such as "cust_id-customer ID" and "order_amt-order amount"), and perform semantic matching processing on the field names in the single-field metadata set and the business term association library. For example, establish a mapping relationship between the "cust_id" field in the structured field unit and the "customer ID" in the association library, and at the same time associate this field with the same source fields in different data sources (such as "cust_id" in the MySQL order table and "customer_id" in the MongoDB customer table) to obtain the field association mapping result; Finally, the single-field metadata set and field association mapping results are organized into a column structure of "field unique identifier (data source ID + field name) - technical metadata - business metadata - operational metadata - management metadata - associated field identifier" to generate a metadata association ledger. Each row in the ledger corresponds to the metadata and associated information of a structured field unit, and the intersection of rows and columns is the specific metadata value (for example, the business meaning of the field unique identifier "DS001-Order ID" is "the number that uniquely identifies the customer's order record"). This technology differs from the traditional method of extracting only technical metadata. It simultaneously covers four types of metadata and adds business semantic associations, providing basic information on the business and operational dimensions for subsequent operation attribute determination (step 31), avoiding the semantic deviation caused by traditional metadata associations relying solely on field names.
[0026] Preferably, the specific implementation process of step 133 is as follows: First, the metadata associated ledger generated in step 132 is used as the processing object, and metadata completion processing is performed - the metadata completion rule library is loaded (including "If the field is 'amount type' (business metadata determination), then the technical metadata needs to be supplemented with 'precision such as 2 decimal places'" "If the field sensitivity level is highly sensitive (management metadata), then it is necessary to supplement 'encryption algorithm type such as SM4'" "If the field historical access frequency is greater than 500 times / day (operation metadata), then it is necessary to supplement 'cache validity period such as 1h'" and other rules), traverse each record in the ledger, and for missing metadata items (such as a " The precision of the "amount" field is not filled in, and the encryption algorithm of a highly sensitive field is not filled in) is automatically supplemented according to the rule base. For example, the "precision = 2 decimal places" of the technical metadata is completed for the "order amount" field, and the "encryption algorithm type = SM4" of the management metadata is completed for the "user ID" field to obtain the completed metadata associated ledger; then, the integrated verification processing is performed on the completed metadata associated ledger. The verification is divided into three categories: one is the consistency verification across metadata dimensions (for example, whether the "field type = date" of the technical metadata matches the "business meaning = order time" of the business metadata. If the field type is int, it is judged to be inconsistent); the second is the consistency verification. The first is the integrity check of the associated fields (for example, if the "order ID" in the ledger is associated with the "user ID", it is necessary to verify whether the metadata of the "user ID" is complete. If the sensitivity level of the "user ID" is missing, it is marked as abnormal). The third is the compliance check of the management metadata (for example, whether the highly sensitive fields are all supplemented with encryption algorithms, whether the data retention period meets the corporate compliance requirements, such as not less than 2 years). For the verification abnormal items (such as the field type is inconsistent with the business meaning, and the highly sensitive field has no encryption algorithm), the abnormality repair module is called (such as automatically correcting the int type "order time" field to the date type, and supplementing the default SM4 for the highly sensitive field without encryption algorithm). The algorithm is used to obtain a verified and repaired metadata association ledger. Finally, the verified and repaired metadata association ledger is reorganized according to the four dimensions of "technical metadata, business metadata, operational metadata, and management metadata," and redundant fields (such as duplicate association field identifiers) are deleted to generate a metadata set containing at least four types of metadata. This technology differs from traditional manual completion and verification methods. Through automatic completion based on a rule base and multi-dimensional verification, it improves the integrity and consistency of metadata and avoids subsequent path planning deviations caused by manual omissions in traditional metadata (for example, the lack of sensitivity level information when generating the path in step 3 makes it impossible to add encrypted nodes).
[0027] Optionally, step 2 specifically includes the following steps: Step 21: Based on the virtual table interface, dynamically map metadata fields to virtual table structures on the metadata set to obtain a metadata-virtual table field mapping relationship. Step 22: Perform virtual table structure abstraction construction processing on the metadata-virtual table field mapping relationship to generate a virtual table structure definition file; Step 23: Perform interface adaptation and function encapsulation processing on the virtual table structure definition file to generate a logical data carrier.
[0028] Optionally, step 21, based on the virtual table interface, performs dynamic mapping configuration processing on the metadata set between metadata fields and virtual table structures to obtain a metadata-virtual table field mapping relationship, specifically comprising the following steps: Step 211: Based on the metadata parsing specification of the virtual table interface, perform field attribute extraction processing on the metadata set to generate metadata field attribute features; Step 212: Dynamically match the metadata field attributes with the virtual table fields to generate a virtual table mapping candidate set; Step 213: Generate metadata-virtual table field mapping relationships based on the virtual table initial mapping candidate set.
[0029] Preferably, in the scenario where enterprise multi-source heterogeneous metadata (including metadata of multiple business domains such as order domain, customer domain, equipment domain, etc.) is adapted to the virtual table interface, when step 211 is specifically implemented: first, the metadata set generated in step 133 is used as the processing object, and the metadata set includes technical metadata, business metadata, operation metadata and management metadata, and at the same time, the metadata parsing specification preset by the virtual table interface is loaded - this specification is different from the traditional basic specification that only defines "field name-data type", and three new parsing dimensions are added: one is the field basic attribute dimension (requires the extraction of field name, data type, length, precision, and whether it is allowed to be empty), the second is the business association attribute dimension (requires the extraction of the business domain to which the field belongs, business term label, and associated business process ID), and the third is the Operational attribute dimension (requires extraction of field average access frequency, read-write operation ratio, and last update time); then, perform attribute extraction processing on each field in the metadata set according to the specification: extract basic field attributes from technical metadata (such as the name of the "order amount" field = order amount, data type = decimal, length = 10, precision = 2, allow null = no), extract business-related attributes from business metadata (such as business domain = order domain, business term label = transaction amount, associated business process ID = P001 order process), and extract operational attribute attributes from operational metadata (such as average access frequency = 500 times / day, read-write operation ratio = read:write = 8:2, last update time = 2024-08-27 Finally, the three extracted attributes are integrated according to the structure of "field unique identifier (data source ID + field name) - basic attributes - business-related attributes - operational characteristic attributes" to generate a metadata field attribute feature table. Each row in this table corresponds to a metadata field, and the intersection of rows and columns is the specific attribute value (for example, the field unique identifier "DS001-order amount" corresponds to the business domain = order domain, and the average access frequency = 500 times / day). This technology provides richer decision-making basis for subsequent virtual table field matching by extracting attributes from the newly added business and operational dimensions, avoiding the traditional matching deviation of "fields with the same name have different business meanings" caused by relying solely on basic attributes (for example, the business meaning of "customer ID" in the order domain and logistics domain is different, and needs to be distinguished by business-related attributes).
[0030] Preferably, in the specific technical implementation of step 212: first, the virtual table field structure library preset by the virtual table interface is loaded, and the structure library contains multiple virtual tables (such as "virtual order table", "virtual customer table" and "virtual device status table"). Each virtual table field has preset adaptation features (including field identifier, adaptation data type range (such as int / long compatibility), business domain label, and metadata field source allowed to be associated). For example, the adaptation features of the "virtual order table-order amount" field are "field identifier = VT001_OrderAmt, adaptation data type range = decimal(10,2) / decimal(12,2), business domain label = order domain, allowed source = MySQL order table / ERP order table"; then, the metadata field attribute feature table generated in step 211 is used as the matching basis, and multi-dimensional dynamic matching processing is performed - first, preliminary screening is performed according to the "business domain label" (such as only matching the metadata field with the business domain = order domain with the "virtual order table" field), and then verification is performed according to the "adaptation data type range" (such as the metadata field "order amount" d ecimal(10,2) types pass if they fall within the matching range of VT001_OrderAmt. Finally, the matching degree is calculated based on the "operational characteristic attributes" (for example, if the average access frequency of a metadata field matches the preset "high-frequency access / low-frequency access" label of a virtual table field, the matching degree is increased by 10%). This yields the matching results for each metadata field and virtual table field (including the matching virtual table ID, virtual table field identifier, and multi-dimensional matching degree). The matching results are then organized based on "metadata field unique identifier - matching virtual table information - matching degree", filtering out results with a matching degree below 60% (the threshold can be adjusted through configuration). This generates a candidate set of virtual table mappings, for example, "DS001 - Order Amount, matching virtual table ID: VT001, virtual table field identifier: VT001_OrderAmt, matching degree: 85%". This technology differs from traditional single-dimensional matching based on "field name + data type". By using multi-dimensional matching based on "business domain - data type - operational characteristics", it significantly reduces mismatches caused by similar field names but different business meanings, thereby improving the rationality of matching results.
[0031] Preferably, the specific implementation process of step 213 is as follows: first, the candidate set of virtual table mapping generated in step 212 is taken as the processing object, and the candidate set conflict detection and resolution processing is performed - if there are multiple virtual table fields corresponding to the same metadata field unique identifier (conflict scenario, such as "DS002-Customer ID" matches "VT002_Virtual Customer Table-CustID" and "VT003_Virtual Logistics Table-CustomerID" at the same time), the conflict resolution rule library is called: rule 1 preferentially selects the virtual table field with the same business domain label (such as the business domain of "DS002-Customer ID" = customer domain, and only the business domain label of "VT002_Virtual Customer Table" = customer domain, then the field is preferentially matched); rule 2, if there are still multiple matching results in the same business domain, the "access frequency" of the operation metadata is combined for weighting (such as the daily access of "DS002-Customer ID" is 1000 times, and the "high-frequency access" label of "VT002_CustID" is matched, and the matching degree is weighted to be higher than that of "VT003_CustomerID" "medium-frequency access" label, then the former is selected); if rules 1 and 2 still cannot be resolved, it is marked as "to be manually confirmed" and a conflict prompt log is generated; then, the candidate set after resolution is subjected to uniqueness verification to ensure that each metadata field corresponds to only one virtual table field, and each virtual table field can correspond to multiple homogenous metadata fields (such as the "order amount" fields of different data sources can all be mapped to "VT001_OrderAmt"); finally, the candidate set that passes the verification is arranged according to the column structure of "metadata field unique identifier-virtual table ID-virtual table field identifier-matching degree-matching state (automatic matching / manual confirmation)", and a metadata-virtual table field mapping relationship table is generated, which can be directly used for subsequent virtual table structure construction (step 22); this technology solves the "one-to-many" conflict problem in traditional automatic matching through the combination of conflict resolution rules and manual confirmation supplement, while retaining the manual intervention entrance, taking into account the automation efficiency and the adaptive flexibility in special scenarios, and ensuring that the generated mapping relationship has high usability.
[0032] Optionally, step 22 specifically includes the following steps: Step 221, according to the metadata-virtual table field mapping relationship and the technical metadata in the metadata set, performing attribute completion and constraint judgment processing on the virtual table field to generate a virtual table field attribute detail; Step 222, based on the virtual table field attribute detail, performing multi-source field association logic combing processing on the metadata-virtual table field mapping relationship to generate a virtual table field association rule table; Step 223, based on the virtual table field attribute detail and the virtual table field association rule table, performing structure framework construction processing on the virtual table to generate a virtual table structure framework; Step 224 : Perform format verification and standardization on the virtual table structure framework according to the structure definition specification preset by the virtual table interface to generate a virtual table structure definition file.
[0033] Preferably, in the scenario of building a unified virtual table with enterprise multi-source heterogeneous metadata (such as MySQL order table, MongoDB customer table, IoT device log table), when step 221 is specifically implemented: first, the metadata-virtual table field mapping relationship generated in step 213 is used as the core processing object, and the mapping relationship includes the corresponding relationship of "metadata field unique identifier-virtual table ID-virtual table field identifier", and the technical metadata (including field primary key / foreign key identifier, default value, data verification rule) in the metadata set of step 133 is loaded at the same time; then, group by virtual table ID, extract all virtual table fields under each virtual table, and perform attribute completion processing on each virtual table field - extract and integrate multi-source technology from the metadata field technical metadata associated with the mapping relationship Attributes (such as the "Order ID" field of the virtual table "VT001_Order Table", which is associated with the primary key identifier of the MySQL order table "order_id" and the index identifier of the MongoDB order log table "order_id", and is completed as "virtual table field primary key identifier = yes, index identifier = yes"), and virtual table-specific attributes are supplemented (such as the display order of fields in the virtual table, whether it is a virtual calculated field (such as "total order amount = unit price × quantity")); then, constraint judgment processing is performed - based on the field verification rules of technical metadata (such as "order amount > 0") and business constraints of business metadata (such as "when the customer level is VIP, the lower limit of the order amount = 100 yuan"), a combination constraint condition of the virtual table field (such as "order amount > 0 And (customer level ≠ VIP or order amount ≥ 100 yuan)"); Finally, the completed attributes and determined constraints are organized according to "virtual table ID - virtual table field identifier - basic attributes (data type / length, etc.) - associated metadata attributes (primary key / foreign key, etc.) - virtual table-specific attributes - combined constraints" to generate a detailed virtual table field attribute list. This technology differs from traditional metadata completion methods based solely on single-source technology. By integrating multi-source metadata attributes and business constraints, it ensures that virtual table field attributes are compatible with multi-source data source characteristics and meet actual business needs, avoiding the usage restrictions caused by the single attribute of traditional virtual table fields.
[0034] Preferably, in the specific technical implementation of step 222: first, the virtual table field attribute details generated in step 221 are used as the processing object, all fields in the same virtual table are screened out, and the metadata association ledger generated in step 132 is loaded at the same time (including the business semantic association relationship between fields, such as "order ID-order item ID" and "customer ID-order ID"); then, multi-source field association logic combing processing is performed, which is divided into two types of association combing: one is technical association combing - based on the primary key / foreign key identifier in the virtual table field attribute details, the technical dependency relationship between the fields in the virtual table is identified (for example, "order ID" of "VT001_order table" is the primary key, and "order ID" of "VT002_order item table" is the foreign key, which is combed into "VT002_order item table.order ID"). The second is business association sorting - based on the business semantic mapping of metadata-related ledgers, identify fields that have no technical foreign keys but have business associations (for example, "Customer ID" in "VT001_Order Table" and "Customer Number" in "VT003_Customer Table" have no technical foreign keys, but the business semantics are both "Customer Unique Identifier", sorted into "VT001_Order Table.Customer ID"). Then, define association rules for each association logic, clarifying the association type (technical foreign key association / business semantic association), association field pair (source field-target field), association matching condition (such as "exact match" or "fuzzy match (customer number = the first 8 digits of the customer ID)"), and association direction (one-way association / two-way association). Finally, organize the association rules according to "association rule ID-virtual table combination (such as VT001+VT002)-association type-association field pair-association matching condition-association direction" to generate a virtual table field association rule table. By distinguishing between technical and business association logic, this technology solves the problem of "business-related fields cannot be associated" caused by traditional sorting of only technical foreign key associations, providing a more comprehensive association basis for subsequent data linkage between virtual tables.
[0035] Preferably, the specific implementation process of step 223 is as follows: first, based on the virtual table field attribute details of step 221 and the virtual table field association rule table of step 222, a basic framework is constructed for a single virtual table - taking the virtual table ID as the identifier, the fields in the virtual table field attribute details are arranged in "display order" to form a virtual table field list, and at the same time, the basic attributes and constraint conditions (such as "order ID: decimal (10, 0), primary key, non-empty") are marked for each field, and virtual table basic information (such as virtual table name, business domain, creation time, and update time) is added; then, the association logic between virtual tables is integrated - based on the virtual table field association rule table, different virtual tables are connected through association fields to form an association topology structure between virtual tables (such as "VT001_order table ← (order ID association) → VT002_order item table ← (customer ID association) → VT003_customer table"), and the association rule ID and matching conditions are marked in the topology structure; then, index and query optimization design is performed - based on the "average access frequency" attribute of the virtual table field attribute details in step 221 (such as "order ID" daily access 1000 times, "customer name" daily access 300 times), virtual indexes are automatically added for high-frequency access fields (such as adding B-tree index for "order ID" and hash index for "customer name"), and the field association order of common query templates (such as "query order item and customer information by order ID") is preset; finally, the single virtual table basic framework, virtual table inter-association topology structure, virtual index, and query template are integrated to generate a virtual table structure framework containing "virtual table basic information-field list and attributes-table inter-association topology-virtual index-query template"; this technology designs virtual indexes and query templates by combining field access frequency, which is different from the traditional virtual table method of only constructing a basic framework, and provides support for subsequent data access optimization in advance, reducing the performance consumption of subsequent queries.
[0036] Preferably, in the scenario where the virtual table structure definition file needs to be compatible with multiple system interfaces (such as BI tools and data integration platforms), when step 224 is specifically implemented: first, load the virtual table interface preset structure definition specification - this specification is an XML format standard, which clearly defines the root node of the file (such as "virtual table structure definition") and the child nodes (such as "virtual table ID", "field list", "field association topology", "virtual index", and "query template"); <virtualtabledefinition>), child node level ( <virtualtableinfo>Virtual table basic information, <fields>Field list, <relations>Relationships between tables, <indexes>Virtual index, <querytemplates>query template) and the required attributes of each node (such as <virtualtableinfo>It must contain id / name / businessDomain attributes), and define the format requirements of the attribute values (such as the date format is "yyyy-MM-dd HH:mm:ss", and the field length is a positive integer); then, convert the virtual table structure framework generated in step 223 according to the format specification - map the "virtual table basic information" to <virtualtableinfo>attribute values (e.g., id=VT001, name=virtual order table, businessDomain=order domain) of the node, mapping "field list with attributes" to <fields>under the node <field>Child nodes (each <field>Contains sub-attributes such as id / name / dataType / length), and maps the "inter-table association topology" to <relations>Under the node <relation>Subnodes (including attributes such as ruleId / sourceField / targetField); then, perform format verification processing - call the XML Schema validator to verify whether the converted XML file conforms to the structure definition specification: first, node level verification (such as <field>The node must be located at <fields>node), and secondly, required attribute verification (e.g. <virtualtableinfo>Whether the name attribute is missing), thirdly, attribute format verification (such as whether the field length is a positive integer, whether the date format is correct), for verification anomalies (such as missing field dataType attribute, wrong date format), automatically locate the abnormal node and complete or correct it according to the specification (such as adding the default value "varchar(50)" for the field with missing dataType, correcting "2024.08.27" to "2024-08-27"); finally, perform standardization processing on the XML file that passes the verification - unify the field naming format (such as unify "order_amt" and "OrderAmt" into "order_amount"), standardize the constraint condition expression (such as unify "order amount>0" into "constraint: order_amount>0"), add the file version number and generation timestamp (such as version number = V1.0, generation time = 2024-08-27) 14:30:00), generating a standard virtual table structure definition file; this technology uses automated format conversion and verification, which is different from the traditional manual writing and checking method, significantly improving the standardization and accuracy of virtual table structure definition files and reducing system connection issues caused by format incompatibility.
[0037] Optionally, step 23 specifically includes the following steps: Step 231: extract the field access attributes in the virtual table structure definition file; Step 232: Perform multi-protocol adaptation layer construction based on field access attributes to generate "interface path / parameter name" mapping rules; Step 233: According to the "interface path / parameter name" mapping rule, perform interface adaptation and function encapsulation processing on the virtual table structure definition file.
[0038] Preferably, in a scenario where multiple enterprise systems (such as BI reporting tools, e-commerce order management systems, and logistics data integration platforms) access virtual tables, when step 231 is specifically implemented: first, the virtual table structure definition file generated in step 224 is used as the processing object, and the file contains core fields such as "virtual table ID-virtual table field identifier-field basic attributes-access frequency-sensitivity level"; then, field access attribute extraction processing is performed, and the extraction dimensions are divided into three categories: one is the access operation type - based on the operation metadata collected in step 133 metadata (field read-write ratio, such as "order ID" read:write = 9:1, "order status" read:write = 4:6), it is determined whether the field is "read-only", "write-only", or "read-write" (read-write ratio > 8:2 is determined as read-only, < 2:8 is determined as write-only, and the rest are read-write); the second is the access frequency level - based on the "average daily number of visits" of the operation metadata (such as "order amount" average 1,200 times per day, "remarks information" average 80 times per day). The first step is to classify access permissions into "high frequency (>1000 times / day)", "medium frequency (100-1000 times / day)", and "low frequency (<100 times / day)". The third step is access permission constraints. Based on the "sensitivity level" of management metadata (e.g., "customer phone number" is highly sensitive, and "order number" is medium sensitive), the permission verification requirements that fields must meet are determined (high-sensitivity fields require "user role + token dual verification", medium-sensitivity fields require "user role verification", and non-sensitive fields require "no special verification"). Finally, the three types of extraction results are integrated in the format of "virtual table field unique identifier (virtual table ID + field identifier) - access operation type - access frequency level - access permission constraints" to generate a field access attribute table. This technology extracts multi-dimensional access attributes by combining operational metadata with management metadata. This differs from the traditional method of extracting only the "read and write" attribute. This provides a basis for "on-demand adaptation" for subsequent multi-protocol adaptation, avoiding resource waste caused by over-adaptation of low-frequency fields.
[0039] Preferably, in the specific technical implementation of step 232: first, based on the field access attribute table generated in step 231, load the enterprise's preset "access attribute-protocol type" mapping rule library (the rule library defines that "read-only + high frequency" fields are adapted to SQL query protocols, and "read-write + medium frequency" fields are adapted to RESTful The API protocol and "read-write + low-frequency" fields are adapted to the GRPC protocol, while also supporting expansion based on business needs, such as the addition of MQTT protocol mapping for "IoT device data"; then, the multi-protocol adaptation layer is constructed and divided into the SQL adaptation sublayer, the RESTful adaptation sublayer, and the GRPC adaptation sublayer according to protocol type. For the SQL adaptation sublayer, for "read-only + high-frequency" fields (such as "order amount"), a mapping rule is generated: "interface path = / sql / virtual table ID / field identifier" and "parameter name = field identifier (lowercase)" (for example, for the "order amount" field in the virtual table VT001, the interface path is / sql / vt001 / order_amount and the parameter name is order_amount); for the RESTful adaptation sublayer, for "read-write + medium-frequency" fields (such as "order status"), mapping is differentiated by HTTP method (GET method path = / api / vt001 / order_status?order_id=xxx, PUT method path = / sql / vt001 / order_status?order_id=xxx, = / api / vt001 / order_status, parameter name = order_status); for the GRPC adapter sublayer, for "read-write + low-frequency" fields (such as "remarks"), a mapping rule is generated: "service name = virtual table ID + FieldService", "method name = Get / Set + field identifier", and "parameter type = field data type" (for example, service name = VT001FieldService, method name = GetRemarkInfo, parameter type = string). Finally, the mapping rules for each sublayer are organized according to "protocol type - virtual table field unique identifier - interface path - parameter name - HTTP method / GRPC method" to generate an "interface path / parameter name" mapping rule table. This technology's unique design avoids the traditional "one table, one protocol" fixed adaptation model. Instead, it assigns differentiated protocols based on field access attributes. This allows high-frequency read-only fields to achieve higher query efficiency using the SQL protocol, medium-frequency read-write fields to achieve flexibility through the RESTful API, and low-frequency fields to reduce connection overhead through GRPC, significantly improving protocol adaptation efficiency in multiple scenarios.
[0040] Preferably, the specific implementation process of step 233 is as follows: first, with the "interface path / parameter name" mapping rule table generated in step 232 and the virtual table structure definition file of step 224 as dual input, perform interface adaptation processing - bind the "interface path / parameter name" in the mapping rule table with the "virtual table field" in the virtual table structure definition file, and generate a dedicated interface description for each virtual table field (such as "VT001_Order Amount" binds the SQL interface description "path= / sql / vt001 / order_amount, supports SELECT operation, parameter order_amount must match decimal(10,2) type"); then, perform function encapsulation processing, and integrate three types of core functional modules: one is the protocol conversion module - based on the "interface path / parameter name" mapping rule, automatically converts access requests from different systems into protocols recognizable by the virtual table (such as the SQL query request of the BI tool is converted into a virtual table internal data reading instruction, the order system's RESTful interface is converted into a virtual table internal data reading instruction, and the order system's RESTful interface is converted into a virtual table internal data reading instruction). The PUT request is converted into a virtual table field update instruction); the second is the permission verification module - the "access permission constraint" in the associated step 231 field access attribute table automatically triggers the "role + token" verification for highly sensitive field requests (for example, a customer mobile phone number query request needs to verify that the user role is "customer service administrator" and the token is not expired), and the medium-sensitive field triggers role verification to ensure access compliance; the third is the access cache module - for "high-frequency access" fields (such as order amount), the cache strategy is configured based on the access frequency level (cache validity period = 5 minutes, and when the cache is full, it is eliminated according to the LRU rule) to reduce the overhead of reading duplicate data; then, the interface adaptation results are integrated with the three types of functional modules to generate a A structured encapsulation package consisting of "virtual table basic information - field interface list - protocol conversion logic - permission verification rules - cache strategy" is created. Finally, the package is subjected to availability verification (simulating scenarios such as BI tool SQL queries and order system RESTful updates to verify whether the interface responses are normal, permission verification is effective, and cache hits are detected). Once the verification passes, a logical data carrier that can be directly called by multiple systems is generated. This technology differs from the traditional method of only encapsulating basic interfaces. By integrating protocol conversion, permission verification, and cache strategies, the logical data carrier can meet the differentiated access requirements of multiple systems without relying on external components, solving the architectural complexity problem caused by the need to deploy additional adapter tools for traditional virtual tables.
[0041] Optionally, step 3 specifically includes the following steps: Step 31: extracting operational attribute features from the metadata set to determine the operational attributes of the metadata set; Step 32: Based on the operation attributes, perform multi-target path mapping on the logical data carrier to establish a corresponding relationship of "operation attribute-logical carrier node-data interaction protocol"; Step 33: Based on the correspondence between "operation attribute-logical carrier node-data interaction protocol", generate a data weaving path of the logical data carrier.
[0042] Optionally, step 31, extracting operational attribute features from the metadata set to determine the operational attributes of the metadata set, specifically includes the following steps: Step 311: Perform structured analysis on the operational metadata and business metadata in the metadata set to determine basic information associated with the metadata set's operational attributes. Step 312: Based on the determined basic information, perform feature extraction on the metadata set for "average response time," "peak access frequency," "business process urgency," and "data update timeliness requirements" to obtain operational attribute features. Step 313: Cluster and label the operational attribute features to determine the operational attributes of the metadata set.
[0043] Preferably, in the metadata operation attribute determination scenario of multiple business scenarios of an enterprise (such as real-time transaction settlement, daily batch data statistics, and monthly static report query), when step 311 is specifically implemented: first, the metadata set generated in step 133 is used as the processing object, and the operation metadata and business metadata therein are screened out - the operation metadata includes "historical data read and write records (each record includes access timestamp, read and write type, and response time)", "data processing task execution log (including task ID, processed data volume, and execution time)", and "data access frequency statistics (number of visits in hourly / day dimensions)"; the business metadata includes "business process to which the data belongs (such as 'order payment process' and 'inventory monthly settlement process')", "business SLA requirements (such as 'real-time transaction response ≤ 500ms' and 'batch statistics completion time limit ≤ 2 hours')", and "business priority tags (such as 'core transactions' and 'non-core reports')"; then, a structured analysis is performed on the operation metadata: group by "data unique identifier (data source ID + field identifier)", and extract the "nearly 7 The system then analyzes the data metadata for the following: average daily read and write times, maximum / minimum single response time, and high-frequency access time periods (e.g., 9:00-11:00 AM). It also performs structured parsing on business metadata: by associating data with the same unique identifier, it extracts the business process ID and name, SLA response time, and business priority level (core / important / general). Finally, the parsed results of the operational metadata and business metadata are integrated based on the "data unique identifier - operational dimension information (average read and write times, response time range, high-frequency time periods) - business dimension information (business process, SLA time limit, priority)" to generate a metadata set operational attribute association basic information table. This technology differs from traditional methods that only parse operational metadata by associating business metadata with business metadata to supplement the "business requirement dimension" information. This is because operational attributes (e.g., real-time requirements) depend not only on the frequency of technical operations but also directly on the urgency of the business process. For example, the SLA time limit for the "order payment process" is much shorter than that for the "monthly inventory settlement process." This pre-emptive association ensures that subsequent feature extraction is more closely aligned with business realities.
[0044] Preferably, in the specific technical implementation of step 312: first, the operation attribute associated basic information table generated in step 311 is used as input, and is processed one by one according to the "data unique identifier", and four feature extractions are performed: the first is average response time extraction - based on the historical response time data in the "response time range" in the basic information table (such as 1,000 response records in the past 7 days), the arithmetic mean is calculated (formula: total response time / number of records). If there are abnormal values (such as extreme values exceeding 3 times the standard deviation of the mean), they are eliminated and then calculated again to obtain the average response time feature of the data (such as 200ms); the second is peak access frequency extraction - based on the "high-frequency access time period" and the corresponding "hourly level access times", the maximum number of access times in the time period is screened out (such as 200 access times in the 9:00-10:00 period, the highest in the whole day) as the peak access frequency feature; the third is business process urgency extraction - loading the company's preset "business process-urgency" mapping rules (such as "order payment process = urgent", "inventory monthly settlement process = general", "static report query = low" Fourth, data update timeliness requirement extraction: Based on the "Business SLA Requirements" in the basic information table, the "Data Update Interval Requirements" are extracted (e.g., "SLA requirement 'Data delay ≤ 1 minute' corresponds to 'High Update Timeliness', and "SLA requirement 'Update once a day' corresponds to 'Low Update Timeliness'") as the data update timeliness requirement feature. Subsequently, the four features are organized according to the formula "Data Unique Identifier - Average Response Time - Peak Access Frequency - Business Process Urgency - Data Update Timeliness Requirements" to generate an operational attribute feature set. This technology is unique in that it combines business features such as "Business Process Urgency" and "Data Update Timeliness Requirements" with operational features such as "Average Response Time" and "Peak Access Frequency". This distinguishes it from the traditional method of extracting only technical operational features. This allows the features to better reflect the nature of "business demand-driven operational attributes." For example, given the same peak access frequency, the operational attribute determination for "urgent" businesses should prioritize real-time performance.
[0045] Preferably, the specific implementation process of step 313 is as follows: First, the operation attribute feature set generated in step 312 is used as the processing object, and the features are first standardized - the numerical features such as "average response time" (unit: ms) and "peak access frequency" (unit: times / hour) are standardized to the interval [0,1] using Min-Max (for example, the average response time of 200ms is standardized to 0.2, and the average response time of 500ms is standardized to 0.5); the categorical features such as "business process urgency" (urgent / general / low urgent) and "data update timeliness requirement" (high / medium / low) are standardized to the interval [0,1]. Use one-hot encoding to convert it into a numerical vector (e.g. "urgent" is encoded as [1,0,0], "high update timeliness" is encoded as [1,0,0]), and obtain a standardized operational attribute feature matrix; the row dimension of the matrix is "data unique identifier" (each row corresponds to a data object), the column dimension is "standardized feature item" (such as average response time, peak access frequency, urgency coding bit 1, urgency coding bit 2, etc.), and the intersection of rows and columns is the specific standardized value; then, load the preset K-means clustering model (the number of clusters K=3, determined by the company's historical business scenarios, The standardized feature matrix is input into the model, and the model clusters according to feature similarity (for example, the features of "average response time ≤ 0.3, peak access frequency ≥ 0.8, urgency code [1,0,0], and update timeliness code [1,0,0]" are grouped into one category); after clustering, labeling is performed: a "clustering result-operation attribute label" mapping is established based on the enterprise business rules (the first category "short response, high frequency, urgent, high update" corresponds to the "real-time read and write" label, the second category "long response, medium frequency, general, medium update" corresponds to the "batch processing" label). The first category is labeled "real-time type", and the third category "no real-time response requirements, low frequency, low urgency, and low update" corresponds to the "static query type" label; finally, the "data unique identifier-clustering category-operation attribute label" is integrated to generate a metadata set operation attribute determination result table; this technology differs from the traditional method of manually labeling operation attributes. It automatically clusters through multi-dimensional standardized features, which not only reduces labor costs but also dynamically adjusts labels according to feature changes (for example, when the peak access frequency of a certain data increases from 0.5 to 0.9, the clustering result automatically changes from "batch type" to "real-time type") to adapt to dynamic changes in business scenarios.
[0046] Optionally, step 32 specifically includes the following steps: Step 321: Perform node classification processing on the logical data carrier, synchronously collect the data interaction protocols supported by each type of node, and generate a logical carrier node-protocol adaptation representation; Step 322: associate and match the operation attribute with the logical carrier node-protocol adaptation representation to establish a corresponding relationship of "operation attribute-logical carrier node-data interaction protocol".
[0047] Preferably, the specific implementation process of step 321 is as follows: First, the logical data carrier generated in step 2 (including virtual table structure definition, interface adaptation rules and field access attributes) is used as the processing object, and node disassembly and function analysis processing is performed on it - the core logical nodes for data processing, transmission and storage in the logical data carrier (such as data receiving nodes, computing processing nodes, cache nodes, transmission nodes, result output nodes) are extracted, and the core function parameters of each node (processing delay limit, data throughput limit, cache capacity, and number of concurrent connections supported) are obtained; then, based on the metadata set operation attribute type determined in step 313 (real-time read and write type, batch processing type, static query type), node classification processing is performed: "processing delay limit ≤ 100ms, concurrent connection number ≥ 5" are classified into the following types: 00" are classified as "low-latency processing nodes" (adapting to real-time read and write operation attributes), nodes with "data throughput limit ≥100MB / s and support for batch data block processing" are classified as "high-throughput storage nodes" (adapting to batch processing operation attributes), and nodes with "cache capacity ≥10GB and cache hit rate ≥90%" are classified as "high cache hit rate query nodes" (adapting to static query operation attributes); then, data interaction protocol collection and processing are performed on each type of node - by calling the node's built-in protocol detection interface, all data interaction protocols and protocol configuration parameters currently supported by the node are obtained (for example, low-latency processing nodes support MQTT protocol (QoS level = 1) and GRPC protocol (compression method = gzip), and high-throughput storage nodes support Apache Arrow protocol (batch size = 1024 entries), high cache hit rate query nodes support the Redis protocol (data structure = Hash)); finally, the node classification results and collected protocol information are integrated according to the "node type (low latency / high throughput / high cache) - node unique identifier - core function parameters - supported protocol list and configuration" structure to generate a logical carrier node-protocol adaptation representation table. This technology differs from the traditional classification method based on "hardware type" and reversely defines the node classification standard based on operational attribute requirements. At the same time, protocol configuration parameters are simultaneously collected to avoid adaptation failures caused by unclear protocol parameters during subsequent matching, thereby improving the rationality of matching nodes and operational attributes.
[0048] Preferably, in the scenario where multiple operation attributes (real-time read-write, batch processing, static query) of an enterprise are matched with logical carrier nodes, when step 322 is specifically implemented: first, the metadata set operation attribute determination result table generated in step 313 is loaded (including "data unique identifier-operation attribute type-operation attribute characteristics (such as real-time requirements, data processing volume)"), and the logical carrier node-protocol adaptation representation table generated in step 321 is loaded at the same time; then, feature extraction processing is performed on each operation attribute type - extracting the "maximum allowed delay (≤200ms), concurrent access requirements (≥300 times / second)" of the real-time read-write operation attribute, and the "single data processing volume (≥100,000)" of the batch processing type. ), throughput requirements (≥80MB / s)", "query response delay tolerance (≤1s), cache validity requirements (≥85% hit rate)" for static query type, generate operation attribute feature vector; then, perform association matching processing: perform multi-dimensional matching on the operation attribute feature vector and the node core function parameters and supported protocol list in the logical carrier node-protocol adaptation representation table - the real-time read and write feature vector is matched with the "processing delay upper limit and concurrent connections" of the "low latency processing node", and the "MQTT / GRPC protocol" supported by the node is filtered; the batch processing type is matched with the "throughput upper limit and batch processing capability" of the "high throughput storage node", and the "Apache The static query type is matched with the cache capacity and hit rate of the "high cache hit rate query node" and the "Redis protocol" is selected. After the matching is completed, a secondary protocol compatibility check is performed (for example, confirming that the MQTT protocol QoS level associated with the real-time read and write operation attributes is consistent with the protocol configuration of the low-latency node), and incompatible matching results are eliminated. Finally, the "operation attribute type - logical carrier node type - data interaction protocol" that have passed the verification are sorted according to the "operation attribute identifier - node type - node unique identifier - protocol name and configuration parameters" to generate an "operation attribute - logical carrier node - data interaction protocol" correspondence table. The unique design of this technology lies in the use of a three-layer matching logic of "operation attribute characteristics - node function parameters - protocol configuration", which is different from the traditional two-layer hard binding of "operation attribute - node". When adding new operation attribute types (such as quasi-real-time analysis) or logical node types, matching can be completed by simply adding the corresponding feature vectors or node parameters, without modifying the underlying matching logic, which has high scalability.
[0049] Optionally, step 33 specifically includes the following steps: Step 331: Perform node flow logic reasoning on the "operation attribute - logical carrier node - data interaction protocol" correspondence to determine the data flow direction between the logical carrier nodes and generate node flow logic accordingly; Step 332: Generate a data weaving path for the logical data carrier according to the node flow logic.
[0050] Preferably, in the specific technical implementation of step 331: first, the "operation attribute-logical carrier node-data interaction protocol" correspondence generated in step 322 is used as the processing object, and the enterprise's preset "business process-node dependency" mapping rule library is synchronously loaded - the rule library defines the pre-dependence logic between nodes based on multiple business scenarios (such as the "order real-time transaction process" must first complete data legitimacy verification before performing storage operations, and the "customer data batch synchronization process" must first complete data format unification before performing aggregation operations), and also includes "inter-node protocol compatibility judgment rules" (such as when the output protocol of the preceding node is MQTT, the subsequent nodes must support the MQTT protocol or the associated protocol conversion sub-node to avoid data transmission interruption due to protocol incompatibility); then, the corresponding relationships are grouped according to the operation attribute type (real-time read-write type, batch processing type, static query type), and node flow logic reasoning is performed for each group of operation attributes: taking the real-time read-write operation attribute as an example, the associated logical carrier nodes ( The system then verifies that all three nodes support the MQTT protocol, eliminating the need for additional protocol conversion modules. The system then performs a logical conflict check on each set of inference results (for example, if an abnormal dependency of "storage node precedes verification node" is detected, the rule base automatically adjusts the dependency to "verification node precedes" based on the rule base). It also adds trigger conditions for data flows between nodes (for example, "when a data receiving node receives 10 pieces of data, a transmission to the data verification node is triggered"). Finally, the node flow logic table is generated by integrating the verified "operation attribute type - node sequence - protocol connection relationship - trigger condition" into a table. Each row in the table corresponds to the complete flow logic for an operation attribute, for example, "Real-time read / write type: data receiving node (MQTT) → Data verification node (MQTT) → low-latency storage node (MQTT), trigger condition: transmission is triggered every 10 data items. This technology differs from the traditional fixed node order reasoning method. By dynamically combining business process dependencies and protocol compatibility determination, the node flow logic not only meets actual business needs but also ensures smooth data interaction, avoiding traditional path execution failures caused by ignoring business dependencies or protocol compatibility.
[0051] Preferably, in the data weaving path generation scenario of multiple business domains (order domain, customer domain, logistics domain), when step 332 is specifically implemented: first, the node flow logic table generated in step 331 is used as the processing object, and structured path conversion is performed for each node flow logic - a unique identifier is assigned to each logical carrier node in the logic (such as "DRN_001" represents a real-time data receiving node, "DVN_001" represents a real-time data verification node, and "LLSN_001" represents a low-latency storage node), and at the same time, the "data interaction protocol configuration parameters" of each node are extracted from the corresponding relationship in step 322 (such as MQTT protocol QoS level = 1, transmission port = 1883, message timeout = 5s, GRPC protocol compression method = gzip, connection timeout = 3s); then, the connection relationship between nodes is defined: with node identifier For indexing, the mapping relationship between the preceding and succeeding nodes is clarified (e.g., "DRN_001→DVN_001" and "DVN_001→LLSN_001"). Data transmission rules are also added for each connection relationship (e.g., "DVN_001 to DVN_001 transmission uses batch packaging mode, with a single packet size of ≤1KB" and "DVN_001 to LLSN_001 transmission uses real-time push mode, with data transmitted immediately after verification."). Subsequently, path fault tolerance configuration is integrated: Based on the operational metadata (historical node failure records and fault recovery time) in the metadata set, fault tolerance mechanisms are added for key nodes (e.g., data verification nodes and storage nodes) (e.g., failure retry count = 3, retry interval = 2s, and automatic failover to the backup node when the primary node fails, e.g., failover to DVN_002 when "DVN_001" fails).Finally, the node identification sequence, protocol configuration parameters, connection relationship, and fault tolerance configuration are integrated according to the path definition format (such as JSON format) preset by the virtual table interface to generate a data weaving path file of the logical data carrier. The core content of the "real-time read-write type" path in the file is as follows: {"operation attribute": "real-time read-write type", "path ID": "RWP_001", "node sequence": ["DRN_001", "DVN_001", "LLSN_001"], "protocol configuration": {"DRN_001": {"protocol": "MQTT", "port": 1883, "QoS": 1, "timeout": 5}, "DVN_001": {"protocol": "MQTT", "port": 1884, "QoS": 1, "timeout": 5}, "LLSN_001": {"protocol": "MQTT", "port": 1885, "QoS": 1, "timeout": 5}}, "fault tolerance configuration": {"retryCount": 3, "retryInterval": 2, "backupNodes": {"DVN_001": "DVN_002", "LLSN_001": "LLSN_002"}}}; The special design of this technology is that the generated path file not only contains node sequence, but also integrates protocol specific configuration and fault tolerance mechanism, which is different from the traditional path generation method that only records node sequence. This makes it unnecessary to query protocol parameters or fault tolerance rules during subsequent path execution, significantly improving the execution efficiency and stability of data weaving, and facilitating the optimization and adjustment of the path by subsequent resource scheduling strategies (step 4).
[0052] Optionally, step 4 specifically includes the following steps: Step 41, real-time collection of data access task queue information by deploying a distributed task monitoring agent, and collection of access index data by an index collection module to monitor the data access task queue and access index feedback; Step 42, based on the constructed resource scheduling model including the LSTM time series prediction network and the random forest classification network, preliminary feature extraction processing is performed on the data access task queue and access index feedback to obtain task resource demand features and index load trend features. The task resource demand features and index load trend features are associated and fused to extract resource load association features, and the resource load association features are matched with a preset resource scheduling rule library to generate a resource scheduling strategy.
[0053] Preferably, the specific implementation process of step 41 is as follows: First, for the full-link nodes of the data access task (including data access task scheduling nodes, business domain computing nodes such as orders / inventory / customers, and data source access nodes such as MySQL / MongoDB / IoT devices), deploy a lightweight distributed task monitoring agent (each node independently deploys one agent instance to avoid single point failure affecting overall monitoring). The agent establishes a TCP long connection with the node local task scheduler to collect data access task queue information in real time - the specific collection fields include the task unique identifier (task ID, in the format of "Task_node IP_timestamp"), task business type (real-time transaction query / batch data synchronization / static report generation), target data identifier (data source ID The data is collected by combining virtual table fields, such as "DS001_VT001_Order ID"), task priority level (1 is the highest and 5 is the lowest, preset by business metadata), task submission timestamp (accurate to milliseconds), and the current execution status of the task (waiting for scheduling / executing / failed / completed). The collection trigger mechanism is set to "queue change trigger" (that is, collection is performed immediately when a task queue is added, its status changes, or it is deleted) to avoid information lags caused by periodic collection and generate node-level task queue details. Next, the indicator collection module is synchronously deployed on the above-mentioned full-link nodes. This module calls the node operating system's native API (such as the Linux system's proc file system to read CPU / memory data, the Windows system's Performance Counter interface to obtain disk I / O data) and the node's built-in resource monitoring interface (such as the cache node's Redis INFO command and the database node's SHOW command). STATUS command) to collect access metric data. Specific metrics include node CPU utilization (in %), taking the average utilization of all CPU cores, memory utilization (in %), calculating the ratio of used memory to total memory, disk I / O throughput (in MB / s, distinguishing between read and write throughput), network transmission latency (in ms, taking the average of three TCP connection times between the node and the target data source), and cache hit rate (in %), for cache nodes only, calculating the ratio of hits to total queries. The collection frequency is fixed at 1 second per time, balancing real-time performance and resource consumption, and generating node-level metric data units. Subsequently, the distributed task monitoring agent and metric collection module transmit the generated task queue details and metric data units to the central monitoring node via the lightweight MQTT protocol (QoS level set to 1 to ensure data transmission without loss or duplication). The central monitoring node deploys a data cleaning module to first format the received data (for example, verifying whether the task ID conforms to the preset format and whether the metric value is within a reasonable range (CPU utilization 0-100%)) and eliminate data with incorrect format or abnormality.Finally, the central monitoring node uses "node IP - collection timestamp" as the association key to correlate and integrate task queue details and indicator data units within the same 1-second time window for the same node. This generates a real-time monitoring dataset consisting of "node basic information (IP address, node type, such as scheduling node / compute node) - collection timestamp (accurate to the second) - task queue aggregate information (total number of tasks, number of tasks by business type, number of pending tasks by priority level) - indicator data aggregate information (average and maximum values of each indicator within 1 second)." This dataset directly serves as input to the resource scheduling model in step 42, providing task and indicator linkage information for extracting resource load correlation features. This technology differs from traditional monitoring solutions (which often collect task information at the scheduling node and indicator data at the server level, resulting in unrelated information and incomplete monitoring node coverage). This technology implements full-link node monitoring through distributed agents, simultaneously correlating task business attributes (type, priority) with resource indicators. This ensures that the monitoring data reflects the "business task type-resource consumption" relationship, avoiding the problem of traditional monitoring, which suffers from data fragmentation and prevents subsequent resource scheduling from accurately matching business needs.
[0054] Optionally, step 42 specifically includes the following steps: Step 421: perform standardized preprocessing on the data access task queue and access indicator feedback to generate a task data matrix and an indicator time series data matrix respectively; Step 422: Input the task data matrix into the random forest classification network for feature extraction to obtain the task resource requirement characteristics; Step 423: Input the indicator time series data matrix into the LSTM time series prediction network to capture the time dependency of the indicator data, output the indicator load prediction value, and predict the load increase rate and load fluctuation amplitude based on the prediction value to obtain the indicator load trend characteristics; Step 424: construct a feature weight mapping table based on the operation metadata in the metadata set to perform weighted fusion of the task resource demand feature and the indicator load trend feature to obtain the resource load correlation feature; Step 425: traverse and match the resource load association characteristics with the preset resource scheduling rule library, and generate a specific resource scheduling strategy based on the matching results.
[0055] Preferably, in the data access monitoring data processing scenario of enterprise multi-business scenarios (including real-time transactions, batch statistics, and static queries), when step 421 is specifically implemented: first, the real-time monitoring data set generated in step 41 is used as the processing object, and the data set is split into two types of core data - data access task queue data (including task ID, task business type, priority, target data identifier, submission timestamp) and access indicator feedback data (including node IP, collection timestamp, CPU occupancy, memory usage, disk I / O throughput, network transmission delay, cache hit rate); then, standardized preprocessing is performed on the data access task queue data: for any task, the data access task queue data is processed by the data access task queue data, and the data access task queue data is processed by the data access task queue data. The classification features such as business type (real-time transaction / batch statistics / static query) and priority (level 1-5) are converted into numerical vectors using one-hot encoding (e.g., "real-time transaction" is encoded as [1,0,0], and priority level 1 is encoded as [1,0,0,0,0]); for the task submission timestamp (milliseconds), the time difference with the current time is calculated (in seconds) and normalized to the interval [0,1] using Min-Max; the encoded classification features and the normalized time difference features are integrated according to "task ID" to generate a task data matrix - the row dimension of the matrix is the number of tasks (each row corresponds to a data access task), and the column dimension is the number of features (e.g., business The intersection of rows and columns is the specific standardized value (for example, "1" in the first row and the first column represents that the task is a real-time transaction type). Then, the access indicator feedback data is subjected to standardized preprocessing: grouping by "node IP + 1-minute sliding window" (each window contains 60 1-second collection points), calculating the mean of each indicator in each window (such as the mean CPU occupancy rate and the mean memory usage rate), and using Z-Score standardization (subtracting the mean and dividing by the standard deviation) to eliminate the dimension effect; the standardized window indicators are integrated according to "node IP - window start timestamp" to generate the indicator time series data matrix - this The matrix row dimension represents the number of time windows (each row corresponds to a 1-minute window), the column dimension represents the number of metrics (e.g., CPU, memory, disk I / O, network latency, cache hit rate, totaling 5 columns), and the intersection of rows and columns represents the specific standardized metric value (e.g., "0.8" in the second row, first column represents the standardized CPU utilization value for a certain window on a certain node). This technology differs from traditional methods that only standardize numerical features by incorporating task business attributes (type, priority) into the preprocessing scope. This allows subsequent models to extract features based on business semantics, avoiding misjudgments of resource requirements caused by ignoring business attributes (e.g., treating high-priority real-time tasks and low-priority batch tasks with the same standard).
[0056] Preferably, in the specific technical implementation of step 422: first, load the preset random forest classification network model. The special structural design of this model is adapted to the resource demand classification scenario of the data access task - it contains 15 CART decision trees (determined by cross-validation, too many will easily overfit, too few will have low classification accuracy), the maximum depth of each tree is set to 8 (to limit complexity and avoid learning noise), the split feature selection adopts the Gini coefficient (for classification tasks to calculate feature impurity), and the split feature priority is configured according to "task business type > priority > time difference" (because the business type directly determines the resource demand type, such as real-time transactions require high CPU, batch statistics require high I / O); the model training phase has completed training based on the historical task-resource consumption data set (containing more than 100,000 records, each record contains task features and annotated "high CPU demand / high I / O demand / balanced demand" labels), 5-fold cross-validation is used during the training process, and the classification accuracy is stable at more than 85%; then, the generated in step 421 is used. The task data matrix is input into the random forest classification network, which then performs parallel classification using multiple decision trees. Each tree traverses the task's characteristics starting from the root node (for example, if "Business Type = Real-Time Transaction," it enters the left subtree, while "Priority = Level 1" results in further filtering). The network then outputs its prediction for the task's resource requirement type. The network then uses a majority vote on the predictions of the 15 trees to determine the final task resource requirement type and calculate the confidence level for that type (number of votes divided by the total number of trees). Finally, the network combines the "task ID - resource requirement type (High CPU / High I / O / Balanced) - confidence level" to generate a task resource requirement feature table. This technology differs from the traditional random forest design of splitting features indiscriminately. By configuring split feature priorities based on business importance, the model prioritizes business type in determining resource requirements. For example, for a "Real-Time Transaction" task, even if the time difference feature differs, it will still be classified as "High CPU Requirement," which better reflects the resource consumption patterns of actual enterprise businesses.
[0057] Preferably, the specific implementation process of step 423 is as follows: first, load the preset LSTM time series prediction network model, the structure of which is optimized for the time series characteristics of access indicators - including 1 input layer (the number of neurons = the number of indicators, i.e. 5, corresponding to CPU / memory / disk I / O / network latency / cache hit rate), 2 hidden layers (64 LSTM units in each layer, using ReLU activation function to alleviate the gradient disappearance problem; Dropout layer is added to each layer, dropout rate = 0.2, to prevent overfitting), 1 fully connected output layer (the number of neurons = the number of indicators, i.e. 5, using the Linear activation function to output the indicator prediction value); the model training phase has completed training based on the historical indicator time series data of 1 month (1440 1-minute windows per node per day), the loss function uses the mean square error (MSE), the optimizer uses Adam (learning rate = 0.001), and the indicator prediction error (MAE) after training is controlled within 5%; then, the indicator time series data matrix generated in step 421 is input into the LSTM network in chronological order, and the input layer converts the standardized indicator value into a feature vector and then passes it to the first hidden layer; the LSTM unit captures the time dependency of the indicator through the forget gate (controls the discarding of historical information), the input gate (updates the cell state), and the output gate (generates the current output) (such as the rising trend of CPU occupancy as the amount of real-time tasks increases, and the fluctuation law of memory usage with batch data processing); the second layer The hidden layer further processes the output of the first layer to enhance the extraction of long-term time series features. The output layer outputs the predicted indicator load value for the next one-minute window (such as CPU utilization and memory usage). Subsequently, based on the predicted value and the actual indicator values of the last three windows, the indicator load increase rate (for example, CPU utilization increase rate = (predicted value - actual value three windows ago) / 3 minutes) and indicator load fluctuation range (for example, memory utilization fluctuation range = predicted value - absolute value of the actual value in the most recent window) are calculated. Finally, the "node IP - predicted time window - indicator load predicted value - increase rate - fluctuation range" formula is integrated to generate an indicator load trend feature table. This technology adds a business-oriented time series window design to the traditional LSTM (the one-minute window matches the decision cycle of enterprise resource scheduling) and strengthens the capture of time series dependencies through two hidden layers. Compared with a single-hidden-layer LSTM, the trend prediction accuracy for sudden loads (such as a sudden increase in real-time task volume) is improved by approximately 15%.
[0058] Preferably, in the multi-node multi-task resource load correlation feature fusion scenario, when step 424 is specifically implemented: first, based on the operation metadata in the metadata set of step 133, a feature weight mapping table is constructed - the row dimension of the table is the task resource demand type (high CPU / high I / O / balanced), the column dimension is the indicator load trend feature (CPU increase rate / memory fluctuation range / disk I / O throughput / network delay / cache hit rate), and the intersection of the row and column is the weight value (the sum is 1); the weight determination logic is based on the "task type" in the historical operation metadata. Type - Indicator Sensitivity" statistics (for example, the CPU indicator sensitivity of high CPU demand tasks is 0.6, and the disk I / O indicator sensitivity of high I / O demand tasks is 0.5). For example, the corresponding weights for "High CPU demand" are [0.6, 0.1, 0.1, 0.1, 0.1] (CPU ramp rate has the highest weight), the corresponding weights for "High I / O demand" are [0.1, 0.1, 0.5, 0.1, 0.1] (Disk I / O throughput has the highest weight), and the corresponding weights for "Balanced demand" are [0.2, 0.2, 0.2, 0.2, 0.2] Next, the task resource demand characteristics of step 422 are associated with the indicator load trend characteristics of step 423 based on the "node IP + task association" (for example, a node's real-time trading task (high CPU demand) is associated with the node's CPU ramp rate characteristics). A weighted fusion calculation is then performed: for each association combination, the weight vector corresponding to the task resource demand type is element-wise multiplied with the indicator load trend characteristic vector (normalized ramp rate, fluctuation range, etc.), and the sum is calculated to obtain the resource load association characteristic value (for example, the association characteristic value for a high CPU demand task = 0.6 × CPU ramp rate + 0.1 × memory fluctuation range + ... + 0.1 × cache hit rate). Finally, the "node IP - associated task type - resource load association characteristic value" is integrated to generate a resource load association characteristic table. This technology differs from traditional fixed-weight fusion methods by dynamically adjusting weights based on operational metadata, allowing the feature fusion results to reflect the "task demand - indicator status" correlation. For example, when the proportion of high I / O tasks is high, the weight of the disk I / O indicator is automatically increased, avoiding the disconnection between the association characteristics caused by fixed weights and actual resource bottlenecks.
[0059] Preferably, in the specific technical implementation of step 425: first, a preset resource scheduling rule base is loaded. The rule base is categorized into "resource constraints - task priority constraints - security constraints". Each rule includes a trigger condition (based on resource load correlation characteristic value), an execution action, and a priority (priority is executed in case of conflict between rules). For example, a resource constraint rule: "CPU load correlation characteristic value of a certain node > 0.8 (corresponding to actual CPU occupancy > 85%) → trigger CPU resource allocation + suspend low-priority (level 4-5) batch tasks of the node"; a task priority constraint rule: "The proportion of real-time transaction tasks > 30% and the node network delay correlation characteristic value > 0.7 → prioritize the allocation of low-latency transmission links to real-time tasks"; a security constraint rule: "The cache hit rate associated with highly sensitive data tasks (management metadata tags) < 0.6 → trigger the allocation of low-latency transmission links to the node"; Prohibit cache node resource reduction"; then, using the resource load association feature table generated in step 424 as the matching basis, perform rule base traversal matching on each feature: extract the "node IP-associated task type-feature value" in the feature, and compare it with the trigger conditions of each rule in the rule base (for example, if the CPU feature value of a certain node is 0.85>0.8, the CPU resource increase rule is matched); for multiple matched rules, sort them by rule priority (security constraint>task priority constraint>resource constraint), and select the rule combination with the highest priority; then, generate a specific resource scheduling strategy based on the filtered rules: clarify the scheduling object (node IP), scheduling action (such as adding 2 core CPUs, Suspend 5 level 4 priority batch tasks, switch transmission links), execution time (immediate execution / delayed execution for 30 seconds, emergency rules are executed immediately); finally, the scheduling policy is organized according to "policy ID-node IP-scheduling action-execution time-triggering rule" to generate an executable resource scheduling policy file; in this application, the rule base design incorporates security constraints of management metadata (such as protection of highly sensitive data tasks), and the rule priority matches the actual business needs, avoiding the traditional security risks caused by focusing only on resource indicators (such as reducing cache nodes for highly sensitive data to save resources). At the same time, the comprehensiveness of policy generation is ensured through rule traversal matching, reducing the probability of missing key scheduling actions.
[0060] Optionally, step 5 specifically includes the following steps: Step 51: Generate path optimization constraints based on resource scheduling strategy; Step 52: Analyze and evaluate the current resource consumption of the data weaving path of the logical data carrier, and identify bottleneck logical nodes in the data weaving path that do not match the path optimization constraints; Step 53: Optimize the data weaving path of the logical data carrier based on the bottleneck logical node to obtain the optimized data weaving path to weave the multi-source heterogeneous data.
[0061] Optionally, step 51 specifically includes the following steps: Step 511: Perform structured parsing on the core instructions in the resource scheduling policy to generate policy instruction parsing factors; Step 512: Based on the operational metadata and management metadata in the metadata set and according to the policy instruction parsing factors, determine the constraint dimension and the corresponding constraint index; Step 513: Based on the constraint dimensions and constraint indicators, generate path optimization constraint conditions classified and integrated according to "resource constraints - business constraints - security constraints".
[0062] Preferably, the specific implementation process of step 511 is as follows: First, the specific resource scheduling policy generated in step 425 is used as the processing object, and the policy includes "resource adjustment instructions (such as 'increase 20% CPU resources for real-time transaction nodes' and 'reduce 15% memory resources for batch statistics nodes')", "task priority adjustment instructions (such as 'increase the priority of order payment tasks to level 1' and 'reduce the priority of monthly report tasks to level 4')", "node load distribution instructions (such as 'limit the single node load of cache nodes to ≤80%), "replace IoT stream data nodes with the resource allocation instructions" and "replace ... Node load diversion to backup node'); then, perform structured parsing processing, define the parsing dimensions as "instruction type-resource type-adjustment range-associated node / task type-effective condition": for resource adjustment instructions, extract "resource type (CPU / memory / disk I / O)", "adjustment range (increase 20% / decrease 15%)", "associated node (real-time transaction node / batch statistics node)", "effective condition (effective when task queue length > 50)"; for task priority adjustment instructions, extract "task type (order priority)", "task priority" and "task priority"; for task priority adjustment instructions, extract "task type (order priority)". Single payment / monthly report)", "Target priority (Level 1 / Level 4)", "Effective conditions (immediate effect without special conditions)"; for node load distribution instructions, extract "Resource type (load rate)", "Adjustment range (≤80% / diversion)", "Associated node (cache node / IoT streaming data node)", "Effective conditions (effective when the load rate is > 90% for 1 minute)"; finally, organize the parsing results according to "Parsing factor ID-Instruction type-Core parsing item (resource type / task type, etc.)-Specific value" to generate a policy instruction parsing factor table. For example, "Parsing factor ID = F001, Instruction type = Resource adjustment, Resource type = CPU, Adjustment range = 20% increase, Associated node = Real-time transaction node, Effective condition = Task queue length > 50"; this technology differs from the traditional simple parsing method that only extracts "resource-range". By supplementing "Associated object (node / task)" and "Effective condition", subsequent constraints can accurately locate application scenarios, avoiding the generalization of constraints caused by incomplete information in traditional parsing (such as applying CPU increase instructions to all nodes indiscriminately).
[0063] Preferably, in the scenario of formulating data weaving path constraint conditions in multiple business domains (order domain, customer domain, logistics domain) of an enterprise, when step 512 is specifically implemented: first, load the operation metadata and management metadata in the metadata set of step 133 - the operation metadata includes "historical data weaving path resource consumption log (each record contains path ID, node ID, CPU / memory occupancy average, data transmission delay)" and "path execution frequency statistics (number of executions in hourly / daily dimensions)"; the management metadata includes "field sensitivity level (high / medium / low sensitivity)" and "data transmission security specifications (such as high-sensitivity fields need to be added)". Encrypted transmission, transmission protocol must comply with TLS1.3)" "Business process SLA requirements (such as real-time path response delay ≤ 500ms, batch path execution time ≤ 2 hours)"; then, using the policy instruction parsing factor table generated in step 511 as the matching basis, the metadata is associated one by one according to the "parsing factor ID": for the "CPU increase by 20% (associated with real-time transaction node)" parsing factor, the historical average CPU usage of the node (such as 70%) is extracted from the operation metadata, and the "resource constraint dimension" is determined in combination with the increase range. The corresponding constraint indicator is "real-time transaction node CPU usage ≤ (70% × (1 +20%)) = 84%"; for the "order payment task priority is raised to level 1" parsing factor, the SLA response delay of the order payment business (≤500ms) is extracted from the management metadata to determine the "business constraint dimension", and the corresponding constraint indicator is "order payment task data weaving path response delay ≤500ms"; for the "highly sensitive field associated node" parsing factor (implicit in the security requirements of the resource scheduling strategy), the field sensitivity level is extracted from the management metadata to determine the "security constraint dimension", and the corresponding constraint indicator is "the path where the highly sensitive field is located must integrate the national secret SM4 encryption module and transmission protocol TLS 1.3 is required. Finally, the "analysis factor ID - constraint dimension (resource / business / security) - constraint indicator name - constraint indicator threshold / requirement - associated metadata basis" are integrated to generate a constraint dimension and indicator correspondence table. The unique design of this technology is that it does not rely on a single resource scheduling instruction to formulate constraints. Instead, it combines the historical consumption patterns of operation metadata with the security / business requirements of management metadata to ensure that the constraint indicators meet both resource scheduling goals and the actual operating patterns of data weaving (for example, determining thresholds based on historical CPU usage to avoid resource waste caused by too high thresholds or performance deficiencies caused by too low thresholds).
[0064] Preferably, in the specific technical implementation of step 513: first, the constraint dimension and indicator correspondence table generated in step 512 is used as input, and classified and integrated according to the three dimensions of "resource constraint-business constraint-security constraint": the resource constraint dimension integrates the "CPU / memory / disk I / O occupancy rate", "node load rate", and "resource adjustment range" related indicators, such as "real-time transaction node CPU occupancy rate ≤ 84%", "cache node load rate ≤ 80%", and "batch statistics node memory occupancy rate after reduction ≤ 65%"; the business constraint dimension integrates the "path response delay", "task priority matching", and "path execution time" related indicators. Indicators, such as "Order payment task path response delay ≤ 500ms," "Level 1 priority task path execution priority is higher than Level 4," and "Monthly report path execution time ≤ 2 hours." The security constraint dimension integrates related indicators such as "Field sensitivity level matching," "Transmission security specifications," and "Encryption module requirements," for example, "Highly sensitive field paths must integrate the SM4 encryption module," "All path data transmission protocols must use TLS 1.3," and "Medium sensitive field paths must retain permission verification nodes." Next, standardize the indicators for each constraint type: for numerical indicators (such as CPU usage ≤ 84%, latency ≤ 500ms), , clarify "Indicator Name - Threshold - Unit - Effective Node / Task Type"; for rule-based indicators (such as integrated encryption modules, protocol requirements), clarify "Indicator Name - Specific Requirements - Effective Object (Highly Sensitive Field Path / All Paths)"; then, perform constraint conflict check: detect whether different constraints of the same node / task type are contradictory (for example, "Real-time transaction node CPU occupancy ≤ 84%" and "CPU must meet peak load after increase" have no conflict, if "CPU occupancy ≤ 84%" and "CPU occupancy ≥ 90%" appear, mark the conflict), in case of conflict, use the business / security requirements of management metadata Finally, the constraints that have passed the verification are organized according to "constraint category - constraint ID - indicator name - specific requirements - effective object - effective condition" to generate a path optimization constraint table classified and integrated by "resource constraint - business constraint - security constraint". This technology differs from the traditional unclassified constraint generation method. Through classification and integration, the subsequent bottleneck node identification (step 52) can be matched item by item by category. At the same time, the conflict verification mechanism ensures the rationality of the constraints, avoiding traditional optimization failures caused by constraint conflicts (such as requiring both high and low CPU usage).
[0065] Optionally, step 52 specifically includes the following steps: Step 521: Perform node disassembly processing on the data weaving path of the logical data carrier, extract all logical nodes in the path, and obtain real-time resource consumption data of each node to generate node resource consumption details; Step 522: Match and evaluate the node resource consumption details with the path optimization constraints one by one, and mark the nodes that do not meet the constraints to generate a target node sequence; Step 523: Perform influence range analysis on the nodes in the target node sequence, and select nodes whose influence weight on the overall performance of the path is greater than the set weight threshold as abnormal nodes and bottleneck logic nodes in the data weaving path that do not match the path optimization constraints.
[0066] Preferably, the specific implementation process of step 521 is as follows: first, the data weaving path of the logical data carrier generated in step 332 is taken as the processing object, and the path includes multiple scenario paths such as real-time transaction path (such as order payment data weaving path), batch statistical path (such as daily inventory data weaving path), and static report path (such as monthly revenue data weaving path); node disassembly processing is performed on each path according to the link logic of "data source access-data processing-data transmission-result output", all logical nodes in the path are extracted, and the dependency relationship between the nodes is recorded (such as "data source access node A→data verification node B→real-time calculation node C→result output node D"), so as to avoid the traditional disassembly that only extracts nodes without retaining dependencies, resulting in the inability to analyze the impact scope later; then, the distributed resource monitoring interface is called (the interface is deployed on the server where each node is located and supports real-time data query), and the real-time resource consumption data of each logical node is collected at a frequency of 1 second / time. The collection dimensions include: CPU usage (unit: %), taking the CPU of the server to which the node belongs The system then integrates the disassembled node information with the collected resource consumption data according to the structure of "node unique identifier (path ID + node ID) - node type (access / processing / transmission / output / security) - real-time resource consumption data (CPU / memory / throughput / latency / encryption status) - collection timestamp" to generate a detailed node resource consumption breakdown. This technology differs from the traditional method of simply collecting only CPU and memory by supplementing business-related resource dimensions (throughput, latency) and security (encryption status). The performance of the data weaving path depends not only on the underlying resources but also on business processing efficiency (throughput), transmission real-time performance (latency), and security compliance (encryption). Comprehensive collection ensures that no omissions are missed in subsequent evaluations.
[0067] Preferably, in the node compliance verification scenario of multiple constraint dimensions (resource constraint, business constraint, security constraint) of an enterprise, when step 522 is specifically implemented: first, load the path optimization constraint conditions of the "resource constraint-business constraint-security constraint" classification integration generated in step 513, and clarify the verification dimensions and thresholds / requirements of each type of constraint - resource constraints include "CPU occupancy ≤ 84% (real-time node)", "memory usage ≤ 75% (batch node)", "throughput ≥ 50MB / s (processing node)"; business constraints include "real-time path transmission delay ≤ 500ms" and "batch path execution time ≤ 2 hours"; security constraints include "highly sensitive field node encryption module status = true" and "transmission node protocol = TLS1.3"; then, using the node resource consumption details generated in step 521 as the verification basis, associate the corresponding path optimization constraint conditions one by one according to the "node unique identifier" (such as the CPU constraint and delay constraint of the real-time path associated with the real-time computing node C); then, perform item-by-item matching evaluation: for resource constraints, compare the node real-time CPU / memory / throughput with the constraint threshold (such as the node For example, the CPU usage of point C is 88% > 84%, which is considered unsatisfied. For business constraints, the transmission node latency is compared with the constraint threshold (for example, the latency of transmission node E is 580ms > 500ms, which is considered unsatisfied). For security constraints, the encryption module status and protocol type are verified (for example, the encryption module status of highly sensitive field node F is false, which is considered unsatisfied). A "Constraint Satisfaction Status" field (Satisfied / Unsatisfied) is added to each node, and the unsatisfied constraint item and deviation value are recorded (for example, "CPU Usage: 88% > 84%, Deviation 4%" and "Latency: 580ms > 500ms, Deviation 80ms"). Finally, nodes with "Constraint Satisfaction Status = Unsatisfied" are filtered out and sorted according to the "node unique identifier - unsatisfied constraint category (resource / business / security) - specific constraint item - deviation value" to generate a target node sequence. The unique design of this technology is that the matching evaluation is prioritized according to "security constraints > business constraints > resource constraints". If a node fails to meet multiple constraint categories simultaneously, the higher-priority constraint item is marked first, avoiding the traditional non-priority marking method that can lead to loss of focus during subsequent analysis.
[0068] Preferably, in the specific technical implementation of step 523: first, take the target node sequence generated in step 522 as the processing object, and synchronously load the output results of the random forest classification network in step 422 and the LSTM time series prediction network in step 423 - the random forest classification network (containing 15 CART decision trees, with the split features configured as "task business type > priority > resource demand") has output the "resource demand weight" of each node associated task (for example, the resource demand weight of the real-time transaction task associated node is 0.7, and the batch statistical task associated node is 0.3), which reflects the resource importance of the node associated task; the LSTM time series prediction network (containing 2 layers of hidden layers, 64 LSTM units in each layer, and a Dropout rate of 0.2) has output the "load influence coefficient" of each node in the next 5 minutes (for example, the load influence coefficient of node C is 0.8, representing that the influence degree of its load anomaly on downstream nodes is higher), which is generated based on node dependency relationship and historical load conduction law; then, a node influence range analysis model is constructed: based on the node dependency graph recorded in step 521, the "direct influence node number" (for example, node C directly influences output node D, and the direct influence number = 1) and the "indirect influence node number" (for example, node C indirectly influences report generation node G, and the indirect influence number = 1) of each target node are calculated, combined with the resource demand weight of the random forest and the load influence coefficient of the LSTM, and calculated according to the formula "total influence weight = (direct influence number x 0.4 + indirect influence number x 0.2) x resource demand weight x load influence coefficient" (the weight coefficient is calibrated through historical fault data); for example, the direct influence number of target node C is 1, the indirect influence number is 1, the resource demand weight is 0.7, and the load influence coefficient is 0.8, the total influence weight is (1 x 0.4 + 1 x 0.2) x 0.7 x 0.8 = 0.336; then, a preset influence weight threshold (for example, 0.3, which can be adjusted according to the business scenario) is loaded, and the target nodes with total influence weight > 0.3 are screened out; finally, the screened-out nodes are sorted according to "node unique identifier-total influence weight-associated task type-influence node list", and the list of bottleneck logical nodes that do not match the path optimization constraint conditions in the data weaving path is generated; this technology is different from the traditional method of screening bottlenecks based only on resource consumption deviation value, and by combining the task importance evaluation of the random forest and the load influence prediction of the LSTM, the screened-out bottleneck nodes not only meet the resource constraint requirements, but also reflect their actual influence on the overall performance of the business task and path, avoiding the problem of misjudging low-influence nodes as bottlenecks in the traditional method.
[0069] Optionally, step 53 specifically includes the following steps: Step 531, analyze the abnormal types of the bottleneck logical nodes to formulate targeted path optimization schemes; Step 532, based on the path optimization scheme, fragment disassembly is performed on the data weaving path of the logical data carrier to locate the path fragment where the bottleneck logical node is located; Step 533, based on the specific adjustment rules in the path optimization scheme, the path fragment where the bottleneck logical node is located is optimized according to the process of "node parameter adjustment - module integration - link adaptation", and an optimized target path fragment is obtained; Step 534, calling the protocol adaptation module, the optimized target path fragment and the non-optimized fragment in the original path are subjected to protocol compatibility verification and splicing processing to obtain a spliced path, and meanwhile, it is verified whether the spliced path satisfies the path optimization constraint condition, if yes, the whole path is taken as the optimized data weaving path, otherwise, steps 531-534 are repeatedly executed until the optimized data weaving path is obtained; Step 535, calling the optimized data weaving path, field association, format conversion and secure transmission processing are performed on the multi-source heterogeneous data to complete the data weaving of the multi-source heterogeneous data.
[0070] Preferably, the specific implementation process of step 531 is as follows: First, the bottleneck logical node list generated in step 523 is processed as the object, which includes "node unique identifier - total impact weight - associated task type - affected node list", and the output results of the random forest classification network in step 422 and the LSTM time series prediction network in step 423 are simultaneously loaded. The random forest classification network (containing 15 CART decision trees, with split features configured according to "task business type > priority > resource requirement") has output the "resource requirement type (high CPU / high I / O / balanced)" of the bottleneck node associated task. For example, the resource requirement type of the real-time transaction associated node is "high CPU"; the LSTM time series prediction network (containing 2 hidden layers, 64 LSTM units in each layer, and a dropout rate of 0.2) The "load trend (increasing / stable / decreasing)" of the bottleneck node for the next 5 minutes has been output. For example, the load trend of a high-CPU node is "continuously increasing." Next, an anomaly type analysis is performed on the bottleneck logical node, and three types of anomalies are classified: resource adaptation anomalies (for example, the CPU usage of a high-CPU demand node continuously exceeds the constraint threshold, and the LSTM predicted load increases); function loss anomalies (for example, the encryption module is not integrated into the node associated with a highly sensitive field, and the security constraint is not met); and link adaptation anomalies (for example, the transmission node protocol is incompatible with the downstream node, resulting in data transmission delay exceeding the business constraint); then, based on the anomaly type and model output, a targeted path optimization plan is formulated: for resource adaptation anomalies (high CPU + load increase), the plan is "node CPU parameter adjustment (increasing the number of cores by 20%) + Load offloading (migrating 30% of low-priority tasks to standby nodes)"; for function loss anomalies (lack of encryption module integration), the solution is "module integration (installing the national secret SM4 encryption module) + encryption parameter configuration (key rotation period set to 24 hours)"; for link adaptation anomalies (protocol incompatibility), the solution is "protocol conversion submodule deployment (supporting GRPC and MQTT protocol conversion) + link latency monitoring (collecting transmission latency at 1 second / time)"; finally, the "bottleneck node unique identifier - anomaly type - optimization plan details (adjusting parameters / modules / submodules) - association model basis (random forest resource demand / LSTM load trend)" are integrated to generate a path optimization plan table. This technology differs from traditional "one-size-fits-all" optimization solutions. By combining random forest resource demand judgment with LSTM load trend prediction, the solution not only resolves current anomalies but also copes with future load changes. For example, for nodes with "high CPU and increasing load", not only current resources are increased, but tasks are also offloaded in advance to avoid bottlenecks in a short period of time.
[0071] Preferably, in the fragment disassembly scenario of enterprise multi-scenario data weaving paths (real-time transaction path, batch statistics path, static report path), when step 532 is specifically implemented: first, based on the data weaving path of the logical data carrier generated in step 332, according to the business link logic of "data source access-data preprocessing-core processing-data transmission-result output", perform fragment disassembly processing on each path containing a bottleneck node - for example, the real-time transaction path "data source access node A→data verification node B (bottleneck node)→real-time computing node C→transmission node D→output node E" is disassembled into "access fragment (A)-preprocessing fragment (B)-core processing fragment (C)-transmission fragment (D)-output fragment (E)", and the node dependency relationship of each fragment is recorded at the same time (for example, the output of the preprocessing fragment is the input of the core processing fragment); then, locate the target path fragment where the bottleneck logical node is located: from the path of step 531 The unique identifier of the bottleneck node is extracted from the optimization plan table and matched with the list of disassembled fragment nodes to determine the fragment to which the bottleneck node belongs (for example, bottleneck node B belongs to the preprocessing fragment). Next, the key attributes of the target path fragment are extracted, including the fragment's associated business scenario (real-time / batch / static), the number and type of nodes within the fragment, the fragment's input and output data format (for example, JSON / Parquet), and the fragment's current average resource consumption (CPU / memory / latency). Finally, the target path fragment location table is generated by combining the "path ID - target path fragment (for example, preprocessing fragment) - fragment key attributes - bottleneck node position within the fragment." This technology differs from traditional direct optimization without disassembly. By disassembling fragments by business link, subsequent optimization is performed only on the target fragment, avoiding interference with non-bottleneck fragments (for example, optimizing only the preprocessing fragment does not affect the normal operation of the core processing and transmission fragments), thereby improving optimization efficiency.
[0072] Preferably, in the specific technical implementation of step 533: first, with the path optimization solution table of step 531 and the target path segment positioning table of step 532 as dual inputs, the "node parameter adjustment-module integration-link adaptation" process optimization is performed for the target path segment; the first step is to perform node parameter adjustment: taking the pre-processing segment of "high CPU demand + load increase" as an example, the node parameter configuration interface is called to increase the number of CPU cores of the bottleneck node B from 4 cores to 5 cores (increase by 20%), and at the same time adjust the CPU scheduling strategy to "real-time task priority scheduling", and based on the load trend predicted by LSTM, configure the load threshold warning for node B in advance (triggering the warning when the CPU occupancy rate exceeds 80%); the second step is to perform module integration: for the transmission segment of "functional loss exception", the module management interface is called to install the national secret SM4 encryption module for the bottleneck transmission node D, automatically load the encryption key preset in the management metadata (path: / conf / security / sm4.key), and configure the operation parameters of the encryption module (encryption block size = 128 bits, key rotation period = 24 hours), and after integration, the module is verified through the interface The block's running status is checked (if it returns to "normal operation," proceed to the next step). The third step is link adaptation: For the core processing segment with "link adaptation exception," a GRPC-MQTT protocol conversion submodule is deployed. The "interface path / parameter name" mapping rule generated in step 232 is read and the conversion submodule's parameters are configured (input interface path = / grpc / core / process, output interface path = / mqtt / transmit, parameter name mapping = "process_id → transmit_id"). This ensures compatibility between the GRPC protocol of the optimized core processing node C and the MQTT protocol of the downstream transmission node D. Finally, real-time resource consumption data (CPU utilization, latency, and encryption module status) is collected for the optimized target path segment, generating a detailed evaluation of the optimized target path segment. This technology's unique design prioritizes parameter adjustment, module integration, and link adaptation in the order of "basic resource adjustment first, then function completion, and finally link compatibility" to avoid resource conflicts during parameter adjustment caused by module integration first. Furthermore, a status check is performed after each optimization step to ensure effective optimization.
[0073] Preferably, the specific implementation process of step 534 is as follows: first, call the multi-protocol adaptation module (supporting SQL, RESTful API, MQTT, GRPC and other multi-protocol conversions) constructed in step 232, and use the optimized target path segment generated in step 533 and the non-optimized segment in the original path as processing objects to perform protocol compatibility verification: extract the output protocol of the optimized target path segment (such as the output protocol of the optimized preprocessing segment is GRPC) and the input protocol of the non-optimized segment in the original path (such as the input protocol of the core processing segment is GRPC). If the protocols are consistent, they are directly spliced; if the protocols are inconsistent (such as the output protocol of the optimized transmission segment is MQTT, and the input protocol of the original output segment is HTTP), call the protocol conversion submodule to perform protocol conversion according to the "interface path / parameter name" mapping rule (convert the MQTT message into an HTTP request, and map the parameter name "transmit_data" to "http_data") to generate protocol-compatible segment connection data; then, splice the optimized target path segment and the non-optimized segment according to the node dependency relationship of the original path to form a complete splicing path; then, load the generated in step 513 Based on the path optimization constraints, the concatenated path is checked for resource constraints (e.g., whether the CPU / memory usage of all nodes in the concatenated path is ≤ a threshold), business constraints (e.g., whether the total path latency is ≤ 500ms), and security constraints (e.g., whether the encryption module of the highly sensitive field node is operating normally). If the checks pass, the concatenated path is used as the optimized data weaving path. If the checks fail (e.g., the latency of the core processing node after concatenation is 550ms > 500ms), the process returns to step 531 to reanalyze the anomaly type (determined to be "link adaptation anomaly not fully resolved"), adjust the optimization plan (e.g., add a single CPU core to the core processing node), and repeat steps 531-533 until the concatenated path meets all constraints. This technology differs from the traditional method of not verifying concatenation after concatenation. Through a closed-loop process of "protocol verification-concatenation-constraint verification," it ensures that the optimized path not only resolves the original bottleneck but also meets the overall constraint requirements, thus avoiding new performance or security risks after optimization.
[0074] Preferably, in the enterprise multi-source heterogeneous data (MySQL order table, MongoDB customer table, JSON format IoT device log) integration scenario, when step 535 is specifically implemented: first, call the optimized data weaving path generated by step 534, load the metadata set (including technical / business / operational / management metadata) generated by step 133, and determine the processing requirements of multi-source heterogeneous data - perform field association on the MySQL order table (structured data) (associate "order ID" with the "order ID" of the customer table), perform format conversion on the MongoDB customer table (semi-structured data) (convert to Parquet format to improve transmission efficiency), perform secure transmission on the IoT device log (JSON format) (based on high-sensitivity field annotation of management metadata, and perform SM4 encryption on the "device key" field); then, perform processing according to the node flow logic of the optimized path: the data source access node reads the multi-source heterogeneous data, the preprocessing node performs field cleaning (eliminating empty value order records in the order table), the core processing node performs field association (associate the order table and customer table data by "order ID"), format conversion The node uniformly converts the associated data from JSON / MySQL format to Parquet format. The encryption transmission node performs SM4 encryption on data containing highly sensitive fields. The output node writes the processed data to the target data warehouse (such as Hive). During this process, key metrics for path execution (field association success rate, format conversion time, and encryption transmission latency) are collected in real time. Historical execution metrics are correlated with the operation metadata to calculate the performance improvement percentage for this execution (for example, encryption transmission latency dropped from 600ms to 350ms, an improvement of approximately 42%). Finally, a data weaving processing report is generated, including "optimized path ID - multi-source data processing volume - key metrics (association success rate 98%+, conversion time 20s, transmission latency 350ms) - performance improvement percentage), completing the data weaving of multi-source heterogeneous data. This technology differs from traditional manually written processing scripts. By invoking optimized adaptive paths, it automatically completes the association, conversion, and secure transmission of multi-source data. The processing process also incorporates business and security requirements of metadata to ensure that the output data conforms to business semantics and meets security specifications, while significantly improving performance compared to pre-optimization methods.
[0075] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.< / virtualtableinfo> < / fields> < / field> < / relation> < / relations> < / field> < / field> < / fields> < / virtualtableinfo> < / virtualtableinfo> < / querytemplates> < / indexes> < / relations> < / fields> < / virtualtableinfo> < / virtualtabledefinition>
Claims
1. An artificial intelligence-based adaptive data knitting performance optimization method, characterized in that: include: Step 1: Based on the data source list, obtain multi-source heterogeneous data and parse it to generate a metadata set containing at least technical metadata, business metadata, operational metadata, and management metadata; Step 2: Based on the virtual table interface, perform virtual table abstraction and encapsulation processing on the metadata set to generate a logical data carrier; Step 3: determining the operational attributes of the metadata set to generate a data weaving path of the logical data carrier based on the operational attributes; Step 4: Monitor the data access task queue and access indicator feedback in real time, extract resource load correlation features from the data access task queue and access indicator feedback based on the constructed resource scheduling model, and generate a resource scheduling strategy based on the resource load correlation features; Step 5: Optimize the data weaving path of the logical data carrier based on the resource scheduling strategy to obtain the optimized data weaving path to weave multi-source heterogeneous data.
2. The method according to claim 1, characterized in that Step 1 specifically includes the following steps: Step 11: Perform type parsing on the multi-source heterogeneous data sources in the preset data source list to identify the storage engine type and transmission protocol type and generate a detailed list of data source types accordingly; Step 12: Based on the detailed list of data source types, perform protocol adaptation on the multi-source heterogeneous data sources to obtain multi-source heterogeneous data; Step 13: Perform data field segmentation processing on the multi-source heterogeneous data to obtain a structured field set, and perform metadata parsing on the structured field set to generate a metadata set that at least includes technical metadata, business metadata, operational metadata, and management metadata.
3. The method according to claim 1, characterized in that Step 2 specifically includes the following steps: Step 21: Based on the virtual table interface, dynamically map metadata fields to virtual table structures on the metadata set to obtain a metadata-virtual table field mapping relationship. Step 22: Perform virtual table structure abstraction construction processing on the metadata-virtual table field mapping relationship to generate a virtual table structure definition file; Step 23: Perform interface adaptation and function encapsulation processing on the virtual table structure definition file to generate a logical data carrier.
4. The method according to claim 3, characterized in that Step 21 specifically includes the following steps: Step 211: Based on the metadata parsing specification of the virtual table interface, perform field attribute extraction processing on the metadata set to generate metadata field attribute features; Step 212: Dynamically match the metadata field attributes with the virtual table fields to generate a virtual table mapping candidate set; Step 213: Generate metadata-virtual table field mapping relationships based on the virtual table initial mapping candidate set.
5. The method according to claim 1, wherein Step 3 specifically includes the following steps: Step 31: extracting operational attribute features from the metadata set to determine the operational attributes of the metadata set; Step 32: Based on the operation attributes, perform multi-target path mapping on the logical data carrier to establish a corresponding relationship of "operation attribute-logical carrier node-data interaction protocol"; Step 33: Based on the correspondence between "operation attribute-logical carrier node-data interaction protocol", generate a data weaving path of the logical data carrier.
6. The method according to claim 5, characterized in that Step 31 specifically includes the following steps: Step 311: Perform structured analysis on the operational metadata and business metadata in the metadata set to determine basic information associated with the metadata set's operational attributes. Step 312: Based on the determined basic information, perform feature extraction on the metadata set for "average response time," "peak access frequency," "business process urgency," and "data update timeliness requirements" to obtain operational attribute features. Step 313: Cluster and label the operational attribute features to determine the operational attributes of the metadata set.
7. The method according to claim 1, characterized in that Step 4 specifically includes the following steps: Step 41: Deploy a distributed task monitoring agent to collect data access task queue information in real time, and collect access indicator data through an indicator collection module to monitor the data access task queue and access indicator feedback; Step 42: Based on the constructed resource scheduling model including the LSTM time series prediction network and the random forest classification network, preliminary feature extraction and processing are performed on the data access task queue and the access indicator feedback to obtain the task resource demand features and the indicator load trend features, and the task resource demand features and the indicator load trend features are correlated and fused to extract the resource load correlation features, and the resource load correlation features are matched with the preset resource scheduling rule library to generate a resource scheduling strategy.
8. The method according to claim 7, characterized in that Step 42 specifically includes the following steps: Step 421: perform standardized preprocessing on the data access task queue and access indicator feedback to generate a task data matrix and an indicator time series data matrix respectively; Step 422: Input the task data matrix into the random forest classification network for feature extraction to obtain the task resource requirement characteristics; Step 423: Input the indicator time series data matrix into the LSTM time series prediction network to capture the time dependency of the indicator data, output the indicator load prediction value, and predict the load increase rate and load fluctuation amplitude based on the prediction value to obtain the indicator load trend characteristics; Step 424: construct a feature weight mapping table based on the operation metadata in the metadata set to perform weighted fusion of the task resource demand feature and the indicator load trend feature to obtain the resource load correlation feature; Step 425: traverse and match the resource load association characteristics with the preset resource scheduling rule library, and generate a specific resource scheduling strategy based on the matching results.
9. The method according to claim 1, characterized in that Step 5 specifically includes the following steps: Step 51: Generate path optimization constraints based on resource scheduling strategy; Step 52: Analyze and evaluate the current resource consumption of the data weaving path of the logical data carrier, and identify bottleneck logical nodes in the data weaving path that do not match the path optimization constraints; Step 53: Optimize the data weaving path of the logical data carrier based on the bottleneck logical node to obtain the optimized data weaving path to weave the multi-source heterogeneous data.
10. The method according to claim 9, characterized in that Step 51 specifically includes the following steps: Step 511: Perform structured parsing on the core instructions in the resource scheduling policy to generate policy instruction parsing factors; Step 512: Based on the operational metadata and management metadata in the metadata set, and according to the policy instruction parsing factors, determine the constraint dimension and the corresponding constraint index; Step 513: Based on the constraint dimensions and constraint indicators, generate path optimization constraint conditions classified and integrated according to "resource constraints - business constraints - security constraints".
Citation Information
Patent Citations
Multi-source computing power data integration and intelligent scheduling system and method
CN118916147A
Data sharing method for digital twin cities and related device
CN119415607A
Integrated metadata management method and system based on data source expandability
CN119576863A
Database load optimization method and device based on dynamic fragmentation, equipment and medium
CN119847739A
Method and system for identifying potential information in combination with heterogeneous network environment
CN120416092A
Cited By
Security isolation system for USB (Universal Serial Bus) mass storage equipment and use method of security isolation system
CN121051809A
Distributed data unified management and intelligent scheduling method based on data braiding
CN121092336A
Distributed data unified management and intelligent scheduling method based on data weaving
CN121092336B
Data processing method and device, equipment, storage medium and program product
CN121188012A
Instruction interaction method and device of multi-task computing platform and electronic equipment
CN121326585A