Multi-plane data infrastructure system and bidirectional optimization method based on same

By constructing a multi-plane data infrastructure system, we have achieved full-link bidirectional collaboration and self-optimization of generative artificial intelligence, which solves the problem that existing data infrastructure cannot support diverse data needs and improves the training efficiency and inference performance of GAI.

CN122019154APending Publication Date: 2026-05-12PURPLE MOUNTAIN LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PURPLE MOUNTAIN LAB
Filing Date
2026-01-21
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The existing data infrastructure cannot effectively support the diverse data needs of generative artificial intelligence (GAI), resulting in low training efficiency and unstable inference performance. It lacks systematic optimization throughout the entire lifecycle and lacks a two-way optimization mechanism, which limits the comprehensive application of GAI.

Method used

Construct a multi-plane data infrastructure system, including an edge device plane, a network orchestration plane, a computing power processing plane, a data transaction plane, and a generative intelligence plane. Through bidirectional data flow, a self-optimizing closed loop is formed, realizing full-link bidirectional collaboration and self-optimization.

Benefits of technology

It achieves targeted and controllable data perception for GAI needs, provides efficient and deterministic data transmission and standardized preprocessing, accurately matches the high-quality data requirements of GAI models, promotes the bidirectional symbiotic improvement of data and models, and forms a self-optimizing and iterative data support system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019154A_ABST
    Figure CN122019154A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a multi-plane data infrastructure system and a bidirectional optimization method based on the multi-plane data infrastructure system. The plurality of function planes comprise an edge device plane, a network arrangement plane, a computing power processing plane, a data transaction plane and a generative intelligent plane, and each function plane forms a self-optimization closed loop through bidirectional data flow. According to the technical scheme provided by the invention, a multi-plane data infrastructure system capable of realizing bidirectional collaboration can be provided, and the technical effects of full-link bidirectional collaboration and self-optimization iteration are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a multi-plane data infrastructure system and a bidirectional optimization method based thereon. Background Technology

[0002] In recent years, the rapid development of Generative Artificial Intelligence (GAI) has placed higher demands on data infrastructure, especially given the significant differences in data requirements between general-purpose large models and small-scale specialized models. However, with the depletion of publicly available internet data and the high degree of fragmentation of data within industries, acquiring high-quality, timely data has become increasingly difficult, posing a bottleneck to the sustainable development of GAI. Current data infrastructure is mostly designed for traditional analytical tasks, lacking unified support for diverse data, resulting in low training efficiency and unstable inference performance. Furthermore, existing systems often fail to systematically consider the data lifecycle and neglect the optimization of data management through model feedback, limiting the comprehensive application of GAI.

[0003] Therefore, in response to the specific needs of GAI, how to build a multi-plane, bidirectional collaborative data infrastructure system has become an urgent technical problem to be solved. Summary of the Invention

[0004] This application provides a multi-plane data infrastructure system and a bidirectional optimization method based thereon, which achieves the technical effect of providing a multi-plane, bidirectionally collaborative data infrastructure system, realizing full-link bidirectional collaboration and self-optimization iteration.

[0005] To achieve the above objectives, the main technical solutions adopted in this application include: In a first aspect, embodiments of this application provide a multi-plane data infrastructure system, the system comprising multiple functional planes that work in sequence and collaboratively, including an edge device plane, a network orchestration plane, a computing power processing plane, a data transaction plane, and a generative intelligence plane, each functional plane forming a self-optimizing closed loop through bidirectional data flow; wherein... The edge device plane is used to collect sensing data from multi-source heterogeneous devices and adjust the collection parameters according to a self-optimization strategy. The network orchestration plane is used to plan transmission links, allocate bandwidth, and schedule transmission priorities for the sensed data, and output the transmitted sensed data to the computing power processing plane. The computing power processing plane includes a computing power resource pool encapsulated through virtualization and containerization technologies. It is used to select computing power units from the computing power resource pool based on business task requirements, preprocess the perceived data, and output standardized preprocessed data. The data transaction platform is used to realize the asset circulation and value distribution of the preprocessed data based on blockchain and smart contracts, and to filter high-quality data that is compatible with the GAI model through value scoring. The generative intelligent plane includes a cross-modal knowledge base and a GAI model, which are used to train the GAI model using the high-quality data, perform inference tasks based on the trained GAI model, and generate self-optimization strategies adapted to each functional plane based on the inference results and the data quality characteristics of the high-quality data. Each functional plane optimizes based on the corresponding self-optimization strategy and outputs updated high-quality data back to the generative intelligent plane to form a self-optimization closed loop.

[0006] Secondly, embodiments of this application provide a bidirectional optimization method based on the system described above, the method including a forward data flow process and a reverse optimization flow process.

[0007] In one implementation, the forward data flow process includes: Sensing data from multi-source heterogeneous devices is collected through the edge device plane and transmitted to the network orchestration plane; The network orchestration plane performs transmission link planning and resource allocation on the sensed data, and then transmits it to the computing power processing plane. The computing power processing plane selects suitable computing power units from its computing power resource pool, preprocesses the perceived data, outputs standardized preprocessed data, and transmits it to the data transaction plane. The preprocessed data is value-scored through the data transaction plane, and high-quality data that is suitable for the GAI model is selected and transmitted to the generative intelligence plane. The generative intelligent plane uses the high-quality data to train a GAI model, performs inference tasks based on the trained GAI model, generates self-optimization strategies adapted to each functional plane based on the inference results and the data quality characteristics of the high-quality data, and stores the high-quality data in the cross-modal knowledge base after structuring.

[0008] In one implementation, the preprocessing of the sensed data to output standardized preprocessed data includes: The sensed data is cleaned and denoised to remove redundant and abnormal data; A multimodal encoder is used to perform semantic-level fusion of perceptual data from different modalities to generate feature vectors of a unified dimension. The feature vectors are aligned to output standardized preprocessed data.

[0009] In one implementation, the step of scoring the preprocessed data through the data transaction plane to filter out high-quality data suitable for the GAI model includes: The preprocessed data is divided into several data units; Obtain the data usage frequency score, data quality score, and model performance contribution score for each data unit; Based on the data usage frequency score, the data quality score, and the model performance contribution score, the value score of each data unit is determined. Data units with value scores greater than a preset score threshold are selected as high-quality data for the GAI model.

[0010] In one implementation, training a GAI model using the high-quality data and performing an inference task based on the trained GAI model includes: The high-quality data is used to train the GAI model, resulting in a well-trained GAI model. The trained GAI model generates a first semantic vector for each data unit in the cross-modal knowledge base and a second semantic vector for the input sample; wherein the first semantic vector and the second semantic vector have the same dimension. Determine the semantic similarity between the first semantic vector and the second semantic vector, and obtain the retrieval results based on the similarity filtering results; The input sample and the retrieval result are concatenated and then input into the trained GAI model to output the inference result.

[0011] In one implementation, the step of generating a self-optimization strategy adapted to each functional plane based on the inference results and the data quality characteristics of the high-quality data includes: The prediction confidence of the input sample is obtained through the trained GAI model; Determine the change in inference confidence between a pre-set target confidence threshold and the predicted confidence level; The data coverage gap is determined based on the change in inference confidence. Based on the sum of the inference confidence change and the data coverage gap, a self-optimizing strategy is generated to guide the edge device plane to adjust the acquisition frequency and acquisition range, and to guide the network orchestration plane to optimize transmission priority.

[0012] In one implementation, the step of generating a self-optimization strategy adapted to each functional plane based on the inference results and the data quality characteristics of the high-quality data includes: The trained GAI model is used to obtain the data quality score of each data unit in the high-quality data. Based on the data quality score, a self-optimizing strategy is generated to guide the data trading platform in adjusting data pricing, sorting, and value screening rules.

[0013] In one implementation, the reverse optimization workflow includes: The inference task is performed by the trained GAI model of the generative intelligent plane. The inference results and the data quality characteristics of the high-quality data are combined to generate a self-optimizing strategy that is adapted to each functional plane. Each functional plane adjusts its operating parameters based on its corresponding self-optimization strategy. After each functional plane adjusts its operating parameters, it re-executes the forward data flow process and outputs updated, high-quality data back to the generative intelligent plane to form a self-optimizing closed loop.

[0014] In one embodiment, the method further includes a synthetic data generation and reflow process, comprising: The generative intelligent plane obtains scarce seed samples, which are real scarce scene data obtained from the edge device plane or auxiliary samples retrieved from the cross-modal knowledge base. Using the trained GAI model, several synthetic data are generated based on the features of the scarce seed samples. The trained GAI model is fine-tuned using the synthetic data to determine the changes in model performance metrics before and after fine-tuning. If the amount of change is greater than a preset change threshold, the synthesized data is determined to be valid data, and the valid data is returned to the data transaction plane.

[0015] In one embodiment, the method further includes an objective function optimization step, comprising: The generative intelligent plane determines the current dataset state and the current model state at a preset period. Determine the data-side resource consumption cost corresponding to the current dataset state, and determine the model performance improvement benefit corresponding to the current model state; The objective function is the difference between the performance improvement benefit of the model and the resource consumption cost on the data side. The net benefits of multiple cycles are compared, and an iterative optimization strategy is used to approximate the optimal solution. The target dataset state and target model state corresponding to the optimal solution are used as the steady-state benchmark of the system to guide the generation of the self-optimization strategy.

[0016] Thirdly, embodiments of this application provide a computer device, including: The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes these computer instructions to perform the bidirectional optimization method described above.

[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions, which are used to cause a computer to execute the bidirectional optimization method described above.

[0018] The technical solutions provided by one or more embodiments of this application, in which the edge device plane collects multi-source heterogeneous sensing data and dynamically adjusts the collection parameters according to a self-optimization strategy, realizes targeted and controllable data sensing for GAI needs from the data source, ensuring the supply of multimodal data foundation for GAI training and inference; the network orchestration plane, through transmission link planning, bandwidth allocation and priority scheduling of sensing data, realizes efficient and deterministic transmission of multi-source data, providing stable network support for the flow of high-quality data required by GAI; the computing power processing plane, relying on the computing power resource pool formed by virtualization and containerization, completes the standardized preprocessing of sensing data, transforming large-scale, multi-source heterogeneous raw data into standardized data that can be directly used by GAI models, significantly reducing the preprocessing cost of GAI models; and the data transaction plane, based on blockchain and intelligent... The system enables the assetization and fair distribution of preprocessed data through contracts. Simultaneously, it accurately selects high-quality data suitable for the GAI model through value scoring, achieving precise supply and demand matching between data and GAI needs. This provides high-value, high-information-density core data for GAI model training. The generative intelligent plane utilizes high-quality data to complete the training and inference service output of the GAI model. Based on inference results and data quality characteristics, it reverse-engineers self-optimization strategies for each plane, driving improvements in the output of data collection, transmission, preprocessing, and filtering stages. Furthermore, the optimized, updated high-quality data flows back to the generative intelligent plane to feed back into the GAI model iteration. This self-optimizing closed loop allows the entire data infrastructure system to continuously adapt to the training and inference needs of the GAI model, achieving a two-way symbiotic improvement in data quality and GAI model performance. Ultimately, this multi-plane data infrastructure system becomes a dedicated data support system capable of accurately matching and dynamically adapting to the specific needs of GAI, and enabling end-to-end two-way collaboration and self-optimization iteration. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 A schematic diagram of a multi-plane data infrastructure system provided in this application embodiment; Figure 2 A flowchart of the forward data transfer process provided in the embodiments of this application; Figure 3 A flowchart of step S5 provided in an embodiment of this application; Figure 4 A flowchart of step S7 provided in an embodiment of this application; Figure 5 A flowchart illustrating the execution of an inference task based on a trained GAI model, provided in an embodiment of this application; Figure 6 A flowchart illustrating a self-optimization strategy for generating adaptable functional planes, provided as an embodiment of this application; Figure 7 A flowchart illustrating another self-optimization strategy for generating adaptations to various functional planes, provided as an embodiment of this application; Figure 8 A flowchart of the reverse optimization workflow provided in this application embodiment; Figure 9 A flowchart of the synthetic data generation and reflow process provided in the embodiments of this application; Figure 10 A flowchart of the objective function optimization steps provided in the embodiments of this application; Figure 11 This application provides a schematic diagram illustrating the entire lifecycle of data flow in a multi-plane data infrastructure system. Figure 12 A block diagram of a bidirectional optimization device for the above-described system provided in an embodiment of this application; Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In recent years, the rapid development of Generative Artificial Intelligence (GAI) has brought profound changes to many industries, which has also placed higher demands on data infrastructure. The two main technical approaches to GAI—large-scale general-purpose models and small-scale specialized models—have significantly different data requirements. Large-scale general-purpose models rely on massive and diverse data to achieve broad intelligent capabilities, making data coverage and multimodality key considerations. In contrast, small-scale specialized models targeting specific industries require high-quality, timely data to accurately adapt to the needs of specific business scenarios.

[0023] However, with the gradual depletion of publicly available internet data, especially the highly fragmented nature of high-value industry-specific data, acquiring such data has become increasingly difficult. In this context, how to systematically organize and supply data has become a bottleneck for the sustainable development of GAI (Government-Assisted Institutional Analyzer). Therefore, there is an urgent need to establish an effective data ecosystem capable of centrally providing data from different sources and types to better support the training and application of GAI. Simultaneously, this system must possess the ability to dynamically adjust to adapt to ever-changing business needs and technological advancements, thereby achieving a virtuous cycle between data and models.

[0024] Current solutions for improving data infrastructure, such as blockchain technology and edge-cloud collaborative architectures, while addressing issues related to data collection, labeling, and circulation to some extent, are mostly plug-and-play enhancements for single aspects, lacking a systematic consideration of the entire data lifecycle. This prevents data infrastructure from effectively utilizing feedback from GAI models, thus missing opportunities for intelligent development.

[0025] Against this technological backdrop, the development of generative artificial intelligence (GAI) faces multiple challenges, particularly inadequate data infrastructure. These challenges not only affect the training efficiency and inference performance of GAI models but also hinder their widespread adoption in practical applications. Firstly, existing data infrastructure is often designed for traditional analytical tasks or single-type AI services, lacking overall uniformity and flexibility, and thus failing to effectively support GAI's demand for diverse data. For example, while storage and computing resources can be easily expanded for general-purpose large models, the lack of consistency guarantees and quality control mechanisms for multi-source, multi-modal data leads to data quality instability and delayed updates during model training, severely impacting model performance and reliability.

[0026] Secondly, in the application scenarios of small-scale, specialized models, existing data infrastructure struggles to meet the needs of specific business processes. In many industrial internet examples, scarce critical data such as fault images and extreme condition logs often cannot be collected and managed in a timely and refined manner. In this case, the mismatch between the data types required by the model and the data types provided by the infrastructure directly leads to low training efficiency and unstable inference performance.

[0027] Furthermore, traditional data infrastructure generally lacks a holistic consideration of the entire GAI lifecycle in its design. Existing data processing workflows are typically limited to "collection-storage-query-simple processing," failing to optimize for the characteristics of GAI tasks. This limitation manifests in the lack of intelligent information filtering and priority control during the data collection phase, resulting in a large amount of low-value data consuming valuable link and storage resources; in the data preprocessing phase, most adopt batch processing methods, failing to match the specific structure and optimization requirements of GAI models; and in the data utilization and feedback phase, almost no feedback on model performance changes is incorporated into the infrastructure design, causing the data lifecycle to exhibit a one-time use characteristic, making dynamic adjustment and reuse impossible.

[0028] Finally, the current infrastructure architecture does not treat GAI as an active component and lacks a two-way optimization mechanism, resulting in a weak interaction between the model and the data. In existing mainstream architectures, the data platform is often only responsible for data production and supply, while the model only acts as the data consumer. This one-way interaction mode limits the reverse optimization of the data infrastructure by GAI's intelligent capabilities.

[0029] In summary, given the specific needs of GAI, how to construct a multi-plane, bidirectional collaborative data infrastructure system has become an urgent technical problem to be solved.

[0030] To address the aforementioned technical problems, this application provides an embodiment of a multi-plane data infrastructure system. Figure 1 A schematic diagram of a multi-plane data infrastructure system provided in this application embodiment is shown below. Figure 1 As shown, the system comprises multiple functional planes that work in tandem. These functional planes include an edge device plane, a network orchestration plane, a computing power processing plane, a data transaction plane, and a generative intelligence plane. Each functional plane forms a self-optimizing closed loop through bidirectional data flow. The edge device plane is used to collect sensing data from multi-source heterogeneous devices and adjust the collection parameters according to a self-optimization strategy.

[0031] Specifically, the edge device plane is not a single device node. The device types are adapted to the needs of various scenarios such as industry, IoT, intelligent transportation, and aerospace. These include ground terminals: IoT smart terminals, industrial production equipment, various sensors (temperature, humidity, pressure, fault detection, etc.), and intelligent vehicle onboard units, primarily collecting business data, environmental data, status data, and multimedia data from industrial and civilian scenarios. Edge processing terminals: Edge computing servers (MEC), which do not directly participate in raw sensing but perform preliminary local processing (caching, semantic extraction, and simple cleaning) on ​​data collected by surrounding terminals, serving as the "front-end computing core" of the edge device plane. Aerospace terminals: Satellites and drones, primarily collecting environmental data and image / video data from wide-area geographic spaces and special scenarios (such as remote areas and hazardous conditions), filling the coverage gaps of ground terminals.

[0032] The edge device plane serves as the underlying data entry point and sensing terminal for the entire multi-plane data infrastructure system. It undertakes four core functions: raw data acquisition, heterogeneous device access, local lightweight processing, and strategic dynamic acquisition. It is the fundamental data source for all data flow, model training, and reverse optimization within the system. Compared to the passive sensing nodes in traditional data acquisition architectures, the edge device plane in this solution overcomes the limitations of "fixed acquisition and indiscriminate transmission." Through bidirectional interaction with the top-level generative intelligent plane, data acquisition transforms from "passive, indiscriminate acquisition" to "proactive supply tailored to model needs," ensuring the matching degree between data and upper-layer GAI model training and inference tasks from the source. This achieves proactive and targeted data acquisition based on model needs. Simultaneously, combined with local lightweight preprocessing and semantic communication technologies, it reduces transmission costs and improves data effectiveness from the data source, laying a high-quality data foundation for the processing flows of upper-layer planes. Furthermore, based on locally extracted semantic information, wireless semantic communication transmission is achieved. Compared with the traditional "full transmission of raw data", semantic communication only transmits the core semantic features of the data, which greatly reduces the amount of data transmitted, reduces the requirements for network bandwidth, and improves data transmission efficiency, thus reducing the burden on the transmission scheduling of the network orchestration plane.

[0033] The network orchestration plane is used to plan transmission links, allocate bandwidth, and schedule transmission priorities for the sensed data, and output the transmitted sensed data to the computing power processing plane.

[0034] Specifically, the network orchestration plane is not a single collection of network devices. It integrates technologies such as Software-Defined Networking (SDN), Network Functions Virtualization (NFV), Deterministic Networking (DetNet) / Time-Sensitive Networking (TSN), and Computing Network (CPN), and is combined with access devices such as smart gateways and API gateways to form an integrated intelligent network orchestration system. Each technology and device performs its own function and works together to realize the core functions of data transmission and resource scheduling.

[0035] As an access bridge between wireless edge devices and the network orchestration plane, the intelligent gateway performs protocol conversion and format standardization of data collected by edge devices, converting non-standardized data from heterogeneous wireless devices into standardized data that can be transmitted in the network, while also completing the initial data aggregation. The API gateway, as an access bridge for wired data sources (such as web data providers and cloud platform data sources), enables secure access, access control, and data aggregation of wired data through standardized API interfaces, ensuring the flexibility and security of wired data source access.

[0036] Software-defined networking (SDN) adopts a "control plane and data plane separation" architecture. Through an SDN controller, it centrally controls and globally schedules multiple data paths across the entire network, breaking the limitations of traditional networks' "distributed control, fixed paths, and difficulty in dynamic adjustment." It supports dynamic planning of transmission links and allocation of bandwidth resources based on data type and transmission requirements. Network function virtualization (NFV) encapsulates traditional hardware network functions such as firewalls, load balancers, and routers into virtualized network function instances, deploying them on general-purpose servers. This enables elastic scaling and on-demand deployment of network functions, significantly improving network resource utilization and reducing hardware deployment costs. The combination of these two technologies forms the core technical support for achieving "intelligent link planning and dynamic resource allocation" in the network orchestration plane.

[0037] Time-Sensitive Networking (TSN) primarily targets LAN scenarios such as industrial Ethernet. Through technologies like time synchronization, traffic scheduling, and preemptive transmission, it provides microsecond-level latency and jitter guarantees for data transmission, meeting the extremely high real-time requirements of scenarios such as industrial control and intelligent vehicle connectivity. Deterministic Networking (DetNet), primarily targeting WAN scenarios, uses technologies like path planning, resource reservation, and traffic shaping to provide end-to-end deterministic latency, jitter, and reliability guarantees for data transmission over a wide area, compensating for the shortcomings of TSN in WAN applications. The two complement each other, achieving deterministic transmission guarantees across all scenarios from LAN to WAN, solving the problems of "uncontrollable transmission latency and jitter, and easy loss of critical data" in traditional networks.

[0038] The network orchestration plane introduces a computing power network (CPN) to encode the computing resource status (such as computing power utilization, remaining computing power, and computing power type) of network nodes (such as edge nodes and core network nodes) into the network transmission protocol or data packets. This data is then transmitted to the upper-layer computing power processing plane along with the original data stream. This provides the upper-layer computing power processing plane with accurate resource status input for "selecting distributed preprocessing computing power units and planning computing power allocation strategies based on the computing power status of network nodes," thus avoiding blind computing power allocation by the computing power processing plane.

[0039] The computing power processing plane includes a computing power resource pool encapsulated through virtualization and containerization technologies. It is used to select computing power units from the computing power resource pool based on business task requirements, preprocess the perceived data, and output standardized preprocessed data.

[0040] Specifically, the computing power foundation of the computing power processing plane is a unified computing power resource pool encapsulated through virtualization and containerization technologies. This resource pool integrates heterogeneous computing power resources from edge servers, cloud servers, and third-party computing platforms, providing computing power support for all data processing operations on the plane. It breaks down the computing power boundaries between edge, cloud, and third-party platforms, incorporating physical computing resources of different brands, configurations, and deployment locations into a unified management system. Specifically, it uses virtualization technologies such as KVM and VMware to virtualize physical server resources, and containerization technologies such as Docker and Kubernetes to encapsulate virtualized computing power into elastically scalable and on-demand scheduling computing power units, shielding the hardware differences of the underlying physical computing power. Based on the processing requirements of business tasks (such as low-latency edge data processing, high-computing-power cloud data distillation, and high-concurrency multimodal data fusion), suitable computing power units are selected from the computing power resource pool and allocated to corresponding data processing tasks, achieving fine-grained scheduling of computing power resources.

[0041] The selected computing units address the "dirty, messy, and scattered" problems of the original data by performing missing value completion, outlier filtering, redundant data deletion, format unification conversion, and time-series data alignment, ensuring data integrity and standardization. For multi-source heterogeneous modal data such as text, images, videos, and time-series sensor data, different modalities are transformed into feature vectors of a unified dimension, and semantic alignment of multi-modal data is achieved in a unified semantic space, solving the problem of "inconsistent features and incompatible semantics" in multi-modal data, thus adapting to the multi-modal training requirements of GAI models. Based on the real-time requirements of generative intelligent planar GAI model training and inference, real-time data distillation is performed on large-scale data after semantic-level processing. By selecting core samples, compressing the data size, and retaining key semantic information, large-scale original data is transformed into small-scale distilled datasets, reducing the computing power and time costs of subsequent model training without losing key information for model training.

[0042] Finally, the dataset, after being cleaned, standardized, feature extracted, semantically aligned, and possibly distilled, will be uniformly structured and packaged to form high-information-density, high-availability, and standardized preprocessed data. This preprocessed data will then be transmitted as the core output to the downstream data trading platform, providing a foundation for data value screening and asset circulation.

[0043] The data transaction platform is used to realize the asset circulation and value distribution of pre-processed data based on blockchain and smart contracts, and to filter high-quality data that is compatible with the GAI model through value scoring.

[0044] Specifically, the data transaction plane is the core platform responsible for the assetization and value circulation of data. It structures and stores standardized pre-processed data on the blockchain for ownership verification, resolving issues such as unclear data ownership, susceptibility to tampering, and lack of traceability, thus laying a legal and compliant asset foundation for data circulation. Through blockchain and smart contracts, it achieves transparency, automation, and traceability throughout the entire data transaction process, addressing the problems of centralized monopoly, cumbersome transaction processes, chaotic pricing, and delayed settlement in traditional data transactions. It integrates the core requirements of GAI training / inference into the data value evaluation system, achieving a dual screening of data value and model suitability. This outputs accurately adapted, high-value data to the generative intelligence plane, while simultaneously achieving a virtuous cycle of data supply through value incentives.

[0045] The data transaction platform comprises a global database and an AIGC database. The global database stores a full range of standardized preprocessed data of various types, serving as the "basic resource pool" for data transactions and covering multimodal data such as text, images, videos, and time-series feature vectors. The AIGC database, designed for GAI model training / inference, performs preliminary classification and storage of data from the global database (such as multimodal semantic alignment data and distillation datasets), serving as a "dedicated resource pool" for GAI data screening, improving the efficiency of subsequent value scoring and selection. The two databases work together to achieve "full storage + accurate classification" of preprocessed data, ensuring data integrity while meeting the specific needs of GAI models.

[0046] The generative intelligent plane, including a cross-modal knowledge base and a GAI model, is used to train the GAI model with high-quality data, perform inference tasks based on the trained GAI model, and generate self-optimization strategies adapted to each functional plane based on the inference results and the data quality characteristics of the high-quality data. Each functional plane optimizes based on the corresponding self-optimization strategy and outputs updated high-quality data back to the generative intelligent plane to form a self-optimization closed loop.

[0047] Specifically, the generative intelligent plane is the top-level core brain and bidirectional optimization closed-loop engine of the entire multi-plane data infrastructure system. As the final consumer of system data flow and the sole output of strategy feedback, it undertakes functions such as full lifecycle management of GAI models, multi-scenario generative inference services, cross-plane self-optimization strategy generation based on model feedback, rare scenario synthetic data generation, and construction of a data-model symbiotic optimization closed loop. Through a technical architecture of cross-modal knowledge base + multi-type GAI models, the generative intelligent plane transforms high-quality data input from the data transaction plane into model performance and inference capabilities. At the same time, based on model inference results, performance characteristics, and data quality characteristics, it reverse-generates self-optimization strategies adapted to the four underlying functional planes, promoting the optimized data output of each plane and forming a high-quality data feedback loop. Ultimately, it constructs a full-system self-optimization closed loop of "data feeding the model - model generating strategies - strategies optimizing data - optimized data feeding back to the model," which is the core key to achieving adaptive evolution, optimal benefit and cost, and data-model symbiotic improvement of the entire multi-plane system.

[0048] Specifically, the generative intelligence plane serves as the lifecycle management center for GAI models, enabling pre-training, fine-tuning, and iterative optimization of various GAI models. It transforms the value of high-quality data into the model's inference capabilities and performance. Based on the trained and optimized GAI models, it provides real-time / near-real-time generative services to downstream vertical businesses such as industrial production, autonomous driving, and AR / VR, serving as the external output port of the system's intelligent capabilities. Based on model inference results, performance characteristics, and data quality characteristics, it identifies data issues at each underlying stage, generates targeted optimization strategies, and distributes them to the four functional planes, acting as the core feedback node for the system's bidirectional optimization. Furthermore, through generative capabilities, it recreates scarce / extreme condition data, feeding it back to the data transaction plane to enrich data resources. Simultaneously, it promotes the feedback of data optimized by strategies across each plane, forming a self-optimizing closed loop between data and models.

[0049] This embodiment provides a multi-plane data infrastructure system. The edge device plane collects multi-source heterogeneous sensing data and dynamically adjusts the collection parameters based on a self-optimization strategy, achieving targeted and controllable data sensing for GAI needs from the data source, ensuring the supply of multimodal data for GAI training and inference. The network orchestration plane plans the transmission links of the sensing data, allocates bandwidth, and schedules priorities, achieving efficient and deterministic transmission of multi-source data, providing stable network support for the flow of high-quality data required by GAI. The computing power processing plane relies on a virtualized and containerized computing power resource pool to complete the standardized preprocessing of the sensing data, transforming large-scale, multi-source heterogeneous raw data into standardized data that can be directly used by GAI models, significantly reducing the preprocessing cost of GAI models. The data transaction plane is based on blockchain and intelligent... The system enables the assetization and fair distribution of preprocessed data through contracts. Simultaneously, it accurately selects high-quality data suitable for the GAI model through value scoring, achieving precise supply and demand matching between data and GAI needs. This provides high-value, high-information-density core data for GAI model training. The generative intelligent plane utilizes high-quality data to complete the training and inference service output of the GAI model. Based on inference results and data quality characteristics, it reverse-engineers self-optimization strategies for each plane, driving improvements in the output of data collection, transmission, preprocessing, and filtering stages. Furthermore, the optimized, updated high-quality data flows back to the generative intelligent plane to feed back into the GAI model iteration. This self-optimizing closed loop allows the entire data infrastructure system to continuously adapt to the training and inference needs of the GAI model, achieving a two-way symbiotic improvement in data quality and GAI model performance. Ultimately, this multi-plane data infrastructure system becomes a dedicated data support system capable of accurately matching and dynamically adapting to the specific needs of GAI, and enabling end-to-end two-way collaboration and self-optimization iteration.

[0050] This application also provides an embodiment of a bidirectional optimization method. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0051] This embodiment provides a bidirectional optimization method, which includes a forward data flow process and a reverse optimization flow process.

[0052] It should be noted here that traditional data infrastructure only meets M t+1 =f(D t Its core is data-driven unidirectional model updates: the state of the GAI model at time t+1 is determined solely by the data set D provided by the data infrastructure at time t. tThe decision is that model optimization depends solely on the data supply, with no feedback mechanism from the model on the data side. This application proposes a bidirectional optimization method: Among them, M t+1 D represents the model state of the GAI model at time t+1. t For high-quality data available for use by the GAI model at time t; Q(D t V(D) represents the data quality score of high-quality data at time t; t ) represents the value score of high-quality data at time t; ΔP represents the model performance contribution score; Δcoverage represents the data coverage difference; and Δconfidence represents the change in inference confidence.

[0053] Forward data flow process: M t+1 =f(D t ,Q(D t ),V(D t At time t+1, the model state of the GAI model is no longer determined by a single data set D. t The decision is not made by the data set, but by the combination of data quality and data value.

[0054] Reverse optimization of the workflow: The optimized GAI model accurately feeds back into the data infrastructure system iteration: D t+1 =g(M t+1 The data set of the data infrastructure system at time t+1 is determined by the state of the newly optimized GAI model and three core quantitative indicators of model feedback (performance improvement contribution ΔP, data coverage difference Δcoverage, and change in inference confidence Δconfidence). Based on this feedback, the data infrastructure system will adjust the operating parameters of each plane (such as edge acquisition, computing power processing, and data transaction rules) to achieve the desired data set D. t The optimization and upgrades make subsequent data supply more aligned with the model's needs.

[0055] Figure 2 A flowchart of the forward data transfer process provided in the embodiments of this application is shown below. Figure 2 As shown, the process includes the following steps: Step S1: Collect sensing data from multi-source heterogeneous devices through the edge device plane and transmit it to the network orchestration plane.

[0056] Step S3: The network orchestration plane plans the transmission links and allocates resources for the sensing data, and then transmits it to the computing power processing plane.

[0057] Step S5: Select suitable computing units through the computing resource pool of the computing power processing plane, preprocess the sensing data, output standardized preprocessed data, and transmit it to the data transaction plane.

[0058] Step S7: The preprocessed data is value-scored through the data transaction plane, and high-quality data that is suitable for the GAI model is selected and transmitted to the generative intelligent plane.

[0059] Step S9: Train the GAI model using high-quality data through the generative intelligent plane, perform inference tasks based on the trained GAI model, generate self-optimization strategies adapted to each functional plane based on the inference results and the data quality characteristics of the high-quality data, and store the high-quality data in a cross-modal knowledge base after structuring.

[0060] This embodiment completes the initial acquisition of multi-source heterogeneous sensing data through the edge device plane, achieving comprehensive coverage of data from various scenarios and providing rich raw materials for subsequent GAI model training. The network orchestration plane integrates technologies such as Software-Defined Networking (SDN), Network Function Virtualization (NFV), Deterministic Networking (DetNet) / Time-Sensitive Networking (TSN), and Computing Power Network (CPN), along with access devices such as smart gateways and API gateways, forming an integrated intelligent network orchestration system. Each technology and device performs its specific function, ensuring efficient and stable data flow to the computing power processing plane through transmission link planning and resource allocation (bandwidth allocation and transmission priority scheduling). This solves the problem of disordered multi-source data transmission and ensures data timeliness and integrity.

[0061] Secondly, the computing power processing plane, through intelligent scheduling of the computing power resource pool, selects suitable computing power units to complete data preprocessing and output standardized data, eliminating data heterogeneity and quality defects, and providing a unified, high-quality data foundation for subsequent value assessment and GAI model use. The data transaction plane selects high-quality data suitable for the GAI model through value scoring, ensuring that the data delivered to the generative intelligence plane has high value and strong adaptability, and avoiding the waste of model training resources by low-quality data.

[0062] Finally, the generative intelligent plane not only utilizes high-quality data to complete GAI model training and inference tasks, but also generates self-optimizing strategies based on inference results and data quality characteristics. This enables precise guidance for preceding stages such as the edge device plane and network orchestration plane, while simultaneously structuring high-quality data and storing it in a cross-modal knowledge base to form data assets. This step allows the system to break free from the limitations of unidirectional data flow and construct a two-way collaborative closed loop of data application, feedback optimization, and data quality improvement.

[0063] Figure 3 The flowchart for step S5 provided in the embodiments of this application may include the following steps: Step S51: Clean and denoise the perceived data to remove redundant and abnormal data.

[0064] Specifically, to address the GAI model's requirements for semantic consistency and high information density in multimodal data, and to reduce data transmission costs between functional planes, the computing power processing plane integrates semantic perception and semantic communication mechanisms during the preprocessing of the perceived data. This is achieved through standardized processing at the semantic layer, transforming the raw data into data suitable for the GAI model. Furthermore, by reducing redundant transmission through cross-plane semantic feature sharing, the multi-source heterogeneous perceived data received from the network orchestration plane needs to be cleaned and denoised. Through pre-defined outlier identification algorithms and redundant data filtering rules, invalid redundant data and abnormal data collection are accurately removed from the perceived data. This filters out interference information that is of no value to GAI model training, improves the information purity of the raw data, and avoids the negative impact of low-quality data on subsequent semantic processing and model training.

[0065] Step S53: Use a multimodal encoder to perform semantic-level fusion of perceptual data from different modalities to generate a feature vector of a unified dimension.

[0066] Specifically, considering the characteristics of the cleaned and denoised data, which includes multiple modal types such as text, images, and time-series sensor data, a pre-trained multimodal encoder is used to perform semantic-level fusion processing on the perceptual data of different modalities. This multimodal encoder is built based on the semantic understanding requirements of the GAI model and can map perceptual data of different modalities and dimensions to a unified semantic space, generating feature vectors with consistent dimensions and semantic interoperability, solving the problem of heterogeneity of multimodal data, and enabling the processed data to have semantic features that can be directly recognized by the GAI model.

[0067] Step S55: Perform feature alignment on the feature vectors and output standardized preprocessed data.

[0068] Specifically, the unified dimensional feature vector generated by semantic-level fusion is subjected to feature alignment processing. Through operations such as temporal alignment, spatial alignment, and feature scale standardization, feature deviations caused by different acquisition devices and different transmission links are eliminated, and preprocessed data with unified format, consistent semantics, and standardized dimensions is output. This data can be directly connected to the data transaction plane for value scoring and asset circulation without the need for additional processing in downstream links, achieving accurate adaptation to the data input requirements of the GAI model.

[0069] The above semantic processing and transmission process can be modeled as a semantic mapping relationship: Where X represents the sensing data acquired by the edge device plane, Z represents the preprocessed data obtained after the above steps, and based on this semantic mapping, the transmission cost C of semantic communication is...semantic It can be represented as C semantic =α·|Z|, where α is a coefficient less than 1, representing the transmission loss coefficient of semantic features, and |Z| is the scale of the semantic feature representation compared to the transmission cost C of the original data. raw Satisfying C semantic ≤C raw This relationship can significantly reduce the communication bandwidth burden and data transmission cost between functional planes.

[0070] It should also be noted that while general-purpose large-scale models require massive amounts of data, redundant data significantly increases training computational and time costs; industry-specific small-scale models, on the other hand, emphasize small, high-quality data and have higher requirements for data scale adaptability. Therefore, the scale of the original data is compressed. To this end, the computing power processing plane performs real-time distillation on the perceptual data to achieve the optimal balance between minimizing the training data size of the GAI model and maximizing performance. Specifically, the data distillation process is as follows: Where D' represents the value obtained from dataset D during the data distillation process. t Multiple subsets of data to be verified are randomly selected from the data; P(M) t |D ' Let be the objective function of data distillation; the entire process of data distillation represents the process of extracting data from dataset D. t Select the smallest subset from the options. To maintain the model state M of the GAI model t .

[0071] This embodiment improves the information purity of input data by filtering redundant and abnormal data at the source through cleaning and denoising operations, avoiding low-quality data from interfering with the training accuracy of GAI models and providing a high-quality data foundation for multi-functional plane collaboration. Based on multi-modal semantic fusion, it generates unified-dimensional feature vectors, enabling data from different modalities to have interoperable semantic expressions, adapting to the multi-modal understanding and training needs of GAI models, and laying the foundation for cross-plane semantic sharing. By aligning features to output standardized data, it eliminates feature biases caused by devices and links, achieving seamless integration of preprocessed data with downstream data transaction plane value screening and generative intelligent plane GAI model training, significantly reducing data processing costs in downstream stages. The entire preprocessing process permeates the multi-functional plane collaboration logic, both transforming raw data into GAI model-adaptive data and providing a standardized semantic carrier for efficient data flow between various functional planes.

[0072] Figure 4 The flowchart for step S7 provided in the embodiments of this application may include the following steps: Step S71: Divide the preprocessed data into several data units.

[0073] Specifically, a data unit is the smallest unit of data that is priced and circulated in the data trading plane, such as a single labeled image, a piece of time-series sensor data, or a set of multimodal semantic alignment data. The splitting rules are formulated in combination with the training data requirements of the GAI model (such as the sample granularity of the model input and the task granularity of the industry scenario) and the pricing requirements of data transactions, to ensure that each data unit has independent value attributes (quality, frequency of use, and model contribution can be measured separately).

[0074] Step S73: Obtain the data usage frequency score, data quality score, and model performance contribution score for each data unit.

[0075] Specifically, the data usage frequency score U(d) is directly obtained from statistics by the data transaction plane. The data transaction plane records the historical usage records of each data unit (such as how many GAI models used it, the number of times it was called, and the coverage of usage scenarios). This data is then converted into a quantitative score through preset statistical rules (such as "the more times it is called and the more core the covered scenarios, the higher the score"), reflecting the market demand and universal adaptability of the data. The data quality score Q(d) is provided collaboratively by the computing power processing plane and the data transaction plane. The computing power processing plane outputs basic data quality information during the preprocessing stage (such as data integrity, anomaly rate, format standardization, and multimodal fusion consistency). The data transaction plane then adds its own quality checks (such as semantic accuracy and annotation accuracy), comprehensively converting this into a quantitative quality score to ensure the basic usability of the data. The model performance contribution score ΔP is obtained from feedback by the generative intelligent plane. The generative intelligent plane feeds each data unit into the GAI model for training, records the performance changes of the model before and after using the data (such as the improvement in accuracy, the reduction in inference error, and the improvement in generalization ability), and converts this performance improvement into a quantitative score.

[0076] Step S75: Determine the value score for each data unit based on the data usage frequency score, data quality score, and model performance contribution score.

[0077] Specifically, value rating: Here, d represents the smallest data unit that is priced and circulated in the data transaction plane; λ1, λ2, and λ3 are weight parameters, which are determined by the generative intelligent plane and then distributed to the data transaction plane after being determined in conjunction with the actual task scenario. Because the weights are directly related to "which values ​​of data are prioritized", and the core requirements of the task scenario (such as industrial fault detection and autonomous driving) are determined by the GAI model task of the generative intelligent plane (for example, industrial scenarios value data quality Q(d) more, while general large model training values ​​ΔP more), the weights are determined by the generative intelligent plane and then written into the smart contract for use by the data transaction plane.

[0078] Step S77: Select data units with value scores greater than the preset score threshold as high-quality data to adapt to the GAI model.

[0079] Specifically, the setting of the preset scoring threshold needs to be combined with the requirements of the GAI model of the generative intelligent plane (for example, if the general large model has a large data volume requirement, the threshold can be appropriately reduced; if the industry small model has extremely high data quality requirements, the threshold can be increased). After being determined by the generative intelligent plane, it is synchronized to the data transaction plane to filter out all data units with scores higher than the threshold, forming the final high-quality data adapted to the GAI model.

[0080] This embodiment achieves refined measurement of data value by splitting preprocessed data into independent data units, preventing high-value data from being diluted by low-value data. It leverages multi-functional plane collaboration to obtain multi-dimensional scores, and determines the value score through weighted aggregation, achieving unified quantification of multi-dimensional value and making data value comparable and rankable. Threshold-based filtering precisely eliminates low-value data, outputting high-quality data suitable for the GAI model. The entire process incorporates multi-functional plane collaborative logic, ensuring that the generative intelligent plane GAI model obtains highly adaptable training data.

[0081] Figure 5 The flowchart provided in this application embodiment for performing inference tasks based on a trained GAI model may include the following steps: Step S911: Train the GAI model using high-quality data to obtain the trained GAI model.

[0082] Specifically, high-quality data refers to cross-modal data units (text, images, videos, etc.) that have been filtered by the data transaction plane and are adapted to the requirements of the GAI model. The training process addresses the needs of different fields (such as industry and autonomous driving) by using high-quality data from the corresponding fields to complete model pre-training or fine-tuning, ensuring that the model possesses basic cross-modal understanding and generation capabilities.

[0083] Step S913: Generate the first semantic vector of each data unit in the cross-modal knowledge base and the second semantic vector of the input sample through the trained GAI model; wherein the dimensions of the first semantic vector and the second semantic vector are the same.

[0084] Specifically, the trained GAI model performs semantic encoding on each data unit in the cross-modal knowledge base and the current input sample, generating first and second semantic vectors with consistent dimensions. The cross-modal knowledge base is maintained by the generative intelligent plane and contains multiple types of data, including structured data and vectorized representations (all derived from high-quality data from the data transaction plane and model-synthesized data). The encoding process is executed by the multimodal encoder of the GAI model in the generative intelligent plane (e.g., a Transformer encoder for text and a convolutional neural network encoder for images). The core is to map heterogeneous modal data such as text, images, and videos to a unified semantic vector space through the multimodal semantic encoding function E(·).

[0085] Step S915: Determine the semantic similarity between the first semantic vector and the second semantic vector, and obtain the retrieval results based on the results of filtering similarity.

[0086] Specifically, semantic similarity: Where d is a data unit in the cross-modal knowledge base, consistent with the definition of the data unit in step S71; E(d) is the first semantic vector; E(x) is the second semantic vector; and sim(·) is the similarity function.

[0087] This involves selecting the top k data items with the highest semantic similarity to the input sample x from the cross-modal knowledge base K. This process involves calculating the similarity between all data items d in the cross-modal knowledge base and the input sample x, then sorting them according to the similarity scores, and finally selecting the retrieval results corresponding to the top k high-scoring items. This approach ensures the relevance of the retrieval results and can effectively improve the performance of subsequent tasks.

[0088] Step S917: The input sample and the retrieval result are concatenated and then input into the trained GAI model to output the inference result.

[0089] Specifically, the current input sample x is structurally concatenated with the filtered retrieval results R(K,x) (e.g., concatenating text samples with retrieved text knowledge, or image samples with retrieved related text / image features), and then input into the trained GAI model. The final output is a reasoning result that incorporates external knowledge. The core of this step is to allow the GAI model to reference accurate knowledge from a cross-modal knowledge base during the generation process, rather than relying solely on its own fixed knowledge from training. This improves the factual consistency and semantic integrity of the generated content and effectively reduces illusions.

[0090] This embodiment trains a GAI model using high-quality data filtered by the data transaction plane, achieving precise data collaboration between the data transaction plane and the generative intelligence plane, laying a high-quality model foundation for subsequent inference. The trained GAI model generates semantic vectors with consistent dimensions, relying on the multimodal coding capabilities of multi-plane collaboration to provide core support for cross-modal retrieval. Semantic similarity filtering accurately matches related knowledge, allowing the GAI model to call external resources from cross-modal knowledge bases, avoiding the limitations of relying solely on its own parameterized knowledge. Concatenation-enhanced inference effectively alleviates the GAI illusion problem and improves the generation quality.

[0091] Figure 6 A flowchart for generating a self-optimization strategy adapted to each functional plane, provided in an embodiment of this application, may include the following steps: Step S931: Obtain the prediction confidence of the input sample through the trained GAI model.

[0092] Specifically, prediction confidence is a quantitative indicator of the reliability of the inference results of the trained GAI model for the current input sample (such as the probability value of the classification task and the consistency score of the generation task), directly reflecting the certainty of the model's inference in that sample / scenario. For example, in an industrial fault detection scenario, if the model's classification confidence for a certain type of unknown fault image is only 0.6, it indicates that the model's recognition of that scenario has high uncertainty.

[0093] Step S933: Determine the change in inference confidence between the pre-set target confidence threshold and the predicted confidence.

[0094] Specifically, through the formula Δconfidence=conf target -conf current Calculate the change in inference confidence, where conf target This is the system's preset target confidence threshold (set according to the application scenario requirements of the GAI model, such as 0.95 for high-precision industrial inspection scenarios and 0.8 for general scenarios), conf current This represents the model's current prediction confidence. When Δconfidence > 0, it indicates that the current model's inference confidence is lower than the target requirement, and the model has high uncertainty in this type of sample / scenario, requiring reinforcement through optimized data collection; when Δconfidence ≤ 0, it indicates that the model's inference reliability for this scenario meets the standard, and no specific reinforcement is needed.

[0095] Step S935: Determine the data coverage gap based on the change in inference confidence.

[0096] Specifically, by combining the confidence change Δconfidence, we further analyze the root cause of model uncertainty and determine whether it is caused by "insufficient data coverage" (such as missing data for a certain type of work condition, a certain geographical area, or a certain modality, leading to the model's unfamiliarity with the scenario). Here, we need to link the cross-modal knowledge base of the generative intelligent plane (comparing the differences between scenarios with covered data and the current low-confidence scenario), and finally quantify the data coverage gap Δcoverage = 1-D using the data coverage difference Δcoverage. covered / D target D covered For the covered area, D target The larger the target coverage area, the more significant the data coverage gap.

[0097] Step S937: Based on the summarization of inference confidence changes and data coverage gaps, a self-optimization strategy is generated to guide edge device planes to adjust acquisition frequency and acquisition range, and to guide network orchestration planes to optimize transmission priorities.

[0098] Specifically, based on the summary analysis of Δconfidence (change in inference confidence) and Δcoverage (data coverage gap), the generated self-optimization strategy has the characteristics of dual-dimensional linkage optimization, and the core coverage adjusts the direction of two types of planes: First, it guides the edge device plane to adjust the acquisition parameters (such as increasing the acquisition frequency for the case where Δconfidence>0 and Δcoverage is large; expanding the acquisition range for the coverage gap in a certain geographical area; and adding corresponding sensor data acquisition for missing modes); Second, it guides the network orchestration plane to optimize the transmission priority (such as setting newly acquired high-value reinforcement data as high transmission priority to ensure rapid flow to the computing power processing plane).

[0099] The self-optimization strategy can be expressed as: Among them, S t S represents the current set of data acquisition strategies (including adjustable parameters such as acquisition frequency, mode selection, and spatial range). t+1 To guide the edge device plane and network orchestration plane in executing a specific set of data acquisition strategies in the next cycle, β1 and β2 are weight parameters that can be adjusted according to the actual scenario to adapt the strategy to different GAI application requirements.

[0100] This embodiment obtains prediction confidence through a trained GAI model, accurately capturing the model's uncertainty in the current input sample scenario, providing a core triggering basis for subsequent optimization. By calculating the change in inference confidence, the model uncertainty is quantified into a measurable optimization gap, clarifying the triggering criteria for strategy adjustment. Correlating the change in confidence to locate data coverage gaps, summarizing the two-dimensional gaps to generate a self-optimizing strategy, accurately guiding the edge device plane to adjust the collection frequency and range, and the network orchestration plane to optimize transmission priority. The entire process deeply links the generative intelligent plane, the edge device plane, and the network orchestration plane, upgrading data collection from blind coverage to precise reinforcement, and transmission scheduling from fixed configuration to dynamic adaptation. This ensures that the GAI model can continuously obtain high-quality data supplementation adapted to its needs to improve inference reliability, and strengthens the bidirectional collaborative evolution capability of model feedback, strategy execution, data optimization, and model improvement across multiple planes.

[0101] Figure 7 A flowchart illustrating another self-optimization strategy for generating adaptations to various functional planes, provided for embodiments of this application, may include the following steps: Step S932: Obtain the data quality score for each data unit in the high-quality data using the trained GAI model.

[0102] Specifically, the trained GAI model, based on its own training experience, evaluates the data quality of each data unit in the high-quality data output by the data transaction plane. This evaluation is not a repetition of the initial evaluation of the data transaction plane, but rather supplements the value dimensions (such as data semantic integrity, modal adaptability, and adaptability to scarce scenarios) from the perspective of actual users of the GAI model, resulting in a more accurate value quantification result.

[0103] Data quality scores can be expressed as: Where conf(·) is the prediction confidence function, which is consistent with the definition in step S931; consistency(·) is the data representative consistency detection, and the specific method can be determined according to the specific scenario and task requirements; ΔP is the model performance contribution score; γ1, γ2, and γ3 are weight parameters.

[0104] Step S934: Based on the data quality score, generate a self-optimizing strategy to guide the data trading plane in adjusting data pricing, sorting, and value screening rules.

[0105] Specifically, the self-optimization strategies generated based on data quality scores include, but are not limited to, the following: First, data pricing rules (such as increasing the price of data units that contribute significantly to the GAI model's value and have high labeling credibility to incentivize data providers to supply high-quality data); second, data ranking rules (prioritizing high-value, highly adaptable data to the trading market to ensure that the GAI model can quickly acquire core data and reduce training resource consumption); and third, value screening rules (such as adjusting screening thresholds and optimizing multi-dimensional value weights to make subsequently selected data more suitable for the current iteration needs of the GAI model). These strategies are directly deployed to the data trading platform to guide it in completing rule iterations.

[0106] This embodiment uses a trained GAI model to score the data quality of each data unit in high-quality data. Leveraging the strong semantic understanding and evaluation capabilities of the GAI model, it accurately quantifies data quality from the perspective of adapting to its own training needs, overcoming the limitations of traditional data quality evaluation that is detached from the actual needs of GAI. Based on this quality score, a self-optimizing strategy is generated to precisely guide the data trading platform in adjusting data pricing, ranking, and value screening rules, ensuring that high-quality data circulates first and is priced reasonably, and guaranteeing that the GAI model can efficiently acquire suitable data. The entire process constructs a two-way collaborative closed loop of generative intelligent platform quality assessment → data trading platform rule optimization → precise supply of high-quality data to GAI. This not only improves the dynamic adaptability of the data trading system to GAI needs but also enhances the performance of the GAI model through the supply of high-quality data, strengthening the collaborative evolution capability of the multi-plane system and helping to build a multi-plane data infrastructure system that accurately adapts to GAI needs.

[0107] Figure 8 The flowchart provided in this application embodiment shows the reverse optimization process, which may include the following steps: Step S2: The inference task is performed by the trained GAI model of the generative intelligent plane. The inference results and the data quality characteristics of high-quality data are combined to generate a self-optimizing strategy that is adapted to each functional plane.

[0108] In step S4, each functional plane adjusts its operating parameters based on its corresponding self-optimization strategy.

[0109] In step S6, after each functional plane adjusts its operating parameters, it re-executes the forward data flow process and outputs updated high-quality data back to the generative intelligent plane to form a self-optimizing closed loop.

[0110] This embodiment uses the inference results of the trained GAI model, combined with the quality characteristics of high-quality data, to accurately generate self-optimization strategies adapted to various functional planes such as edge devices, network orchestration, and computing power processing. This process deeply binds model requirements with the operational status of the entire data chain, ensuring that the strategy is targeted and can accurately solve problems existing in the preceding data flow (such as inaccurate data collection, low transmission efficiency, and insufficient preprocessing quality).

[0111] Secondly, each functional plane adjusts its operating parameters according to its corresponding self-optimization strategy. For example, the edge device plane optimizes the acquisition range and frequency, the network orchestration plane adjusts the transmission priority, and the computing power processing plane optimizes the allocation of computing power. This precise adjustment directly affects the core links of data flow, laying the foundation for improving data quality in subsequent iterations.

[0112] Finally, the adjusted functional planes re-execute the forward data flow process, and the updated, high-quality data output is fed back to the generative intelligent plane. This step allows the optimization effect to be verified and consolidated, providing more suitable training data for the GAI model to improve performance, and also providing new evidence for the generative intelligent plane to generate more accurate self-optimization strategies in the future, ultimately forming a self-optimization closed loop of strategy generation, parameter adjustment, data optimization, and strategy iteration.

[0113] Figure 9 This is a flowchart of the synthetic data generation and reflow process provided in the embodiments of this application. The process may include the following steps: Step S101: Obtain scarce seed samples through generative intelligent plane. The scarce seed samples are real scarce scene data obtained from edge device plane or auxiliary samples retrieved from cross-modal knowledge base.

[0114] Specifically, scarce seed samples can be obtained from real-world but difficult-to-collect scarce scene data (such as rare industrial fault images and extreme weather sensor data for intelligent driving) from the edge device plane, or auxiliary samples (such as multimodal data related to the semantics of scarce scenes) can be retrieved from the cross-modal knowledge base maintained by the generative intelligent plane.

[0115] Step S103: Using the trained GAI model, generate several synthetic data based on the features of scarce seed samples.

[0116] Specifically, the synthetic data generation function G(·) in the trained GAI model can be implemented by a diffusion model, generative adversarial network, or large-scale multimodal generation model to learn and expand the core features of scarce seed samples, generating large-scale, high-fidelity synthetic data (such as driving videos during extreme rainstorms or vibration time-series data of fatal industrial equipment failures). The generated data type is consistent with real scarce data and can cover multimodal forms such as text, images, videos, and time-series sensor data. The generation process is as follows: Where, d syn For synthetic data; M t The state of the GAI model; d seed This is a scarce seed sample.

[0117] Step S105: Fine-tune the trained GAI model using synthetic data and determine the changes in model performance indicators before and after fine-tuning.

[0118] Specifically, the generated synthetic data is used to fine-tune the trained GAI model, with a focus on monitoring the changes in key performance indicators ∆P(d) before and after fine-tuning. syn (e.g., inference accuracy in extreme scenarios, prediction confidence, generalization error, etc.), and calculate the performance change. The core of this step is to verify whether the actual value of synthetic data can truly improve the adaptability of the GAI model in scarce scenarios.

[0119] In step S107, if the change amount is greater than the preset change threshold, the synthesized data is determined to be valid data, and the valid data is returned to the data transaction plane.

[0120] Specifically, only when the change ∆P(d) syn Only when the value is greater than 0 is the synthetic data considered to have practical value and allowed to flow back to the data transaction plane to supplement the AIGC database, thereby avoiding the negative impact of low-quality or invalid generated data on the system.

[0121] This embodiment obtains scarce seed samples by combining real, scarce data from the edge device plane with auxiliary samples from a cross-modal knowledge base, providing real feature anchors and semantic support for synthetic data. Large-scale synthetic data is generated based on a pre-trained GAI model, overcoming the limitations of physical acquisition and accurately filling data gaps in scarce scenarios. The model is fine-tuned using synthetic data, and performance changes are calculated. The effectiveness of the synthetic data is verified by GAI model performance feedback, avoiding low-quality data contamination of the system. Effective synthetic data is then fed back to the data transaction plane. The entire process is upgraded from passive acquisition to proactive reconstruction and precise replenishment, providing high-quality training data across all scenarios to improve the generalization ability of the GAI model, and strengthening the bidirectional collaborative evolution of data-model-data across multiple planes.

[0122] Figure 10 A flowchart of the objective function optimization steps provided in this application embodiment is included, which may include the following steps: Step S201: Determine the current dataset state and the current model state through the generative intelligent plane at a preset period.

[0123] Specifically, the generative intelligent plane serves as the core coordinating unit, synchronously collecting and analyzing the current dataset status and the current model status at preset intervals (e.g., hours, days, dynamically adjustable according to business scenarios). The dataset status needs to be obtained in conjunction with the edge device plane (data volume and coverage), the computing power processing plane (preprocessed data quality and modal structure), and the data transaction plane (distribution and quality distribution of circulating data). The model status focuses on core indicators such as parameter configuration, knowledge scope, and inference ability of the current GAI model in the generative intelligent plane, ensuring the comprehensiveness and accuracy of status monitoring.

[0124] Step S203: Determine the data-side resource consumption cost corresponding to the current dataset state, and determine the model performance improvement benefit corresponding to the current model state.

[0125] Specifically, based on the acquired state data, two core indicators are quantified: one is the data-side resource consumption cost C(D). t The first aspect is the real resource overhead throughout the entire data lifecycle (bandwidth and energy costs in the acquisition phase, computing power consumption and latency costs in the preprocessing phase, database occupancy costs in the storage phase, incentive and pricing costs in the transaction phase, etc.); the second aspect is the model performance improvement benefit P(M). t The model employs quantifiable metrics such as generation quality (FID, IS, MOS), inference accuracy, task success rate, semantic consistency score, and output confidence stability to accurately reflect the positive contribution of data infrastructure to the improvement of GAI model capabilities.

[0126] Step S205: Using the difference between the model performance improvement benefit and the data side resource consumption cost as the objective function, compare the net benefits of multiple cycles and use an iterative optimization strategy to approximate the optimal solution.

[0127] Specifically, the improvement in model performance is P(M) t Data-side resource consumption cost C(D) tUsing the objective function as an example, by comparing the changes in net profit over multiple cycles, iterative optimization strategies (such as gradient descent, genetic algorithms, etc.) are employed to continuously adjust the dataset state (data volume, modal structure, coverage, etc.) and model state (parameter weights, knowledge scope, reasoning ability, etc.), gradually approaching the optimal solution that maximizes the objective function. During this process, each iteration optimizes the strategies for each plane based on the net profit feedback from the previous round (such as adjusting the collection range of edge devices, optimizing the allocation of computing power resources, and adjusting data transaction pricing rules, etc.).

[0128] Step S207: The target dataset state and target model state corresponding to the optimal solution are used as the steady-state benchmark of the system to guide the generation of the self-optimization strategy.

[0129] Specifically, the target dataset state (such as optimal data volume, suitable modal structure, and reasonable coverage) and the target model state (such as optimal model parameters and matching knowledge range) corresponding to the optimal solution obtained through iterative optimization are established as the steady-state benchmark of the system. The subsequent self-optimization strategies of the entire system (data acquisition, preprocessing, circulation, model training, etc.) are all generated with this benchmark as the guide, ensuring that the strategy adjustment of each plane always revolves around the optimal balance of benefit and cost.

[0130] This embodiment periodically coordinates data from various planes through generative intelligent planes, accurately grasping the current state of the dataset and the model. By quantifying the resource consumption costs and model performance improvement benefits throughout the entire data lifecycle, it addresses the pain point of traditional systems struggling to balance benefits and costs. Iterative optimization using the benefit-cost difference as the objective function achieves dynamic and precise resource allocation. A system steady-state benchmark is established to guide the generation of self-optimization strategies across the entire system. This ensures that the GAI model achieves performance improvements under optimal resource allocation while also enhancing multi-plane collaborative efficiency through global optimization.

[0131] The specific implementation of the present invention is described below. See also... Figure 11 This diagram illustrates the entire lifecycle of data flow in a multi-plane data infrastructure system, as provided in this application embodiment. The entire data lifecycle covers all stages from generation to utilization by a generative artificial intelligence (GAI) model, and ultimately, feedback from the GAI model to data generation. This process includes six key stages: data perception, data acquisition, data preprocessing, data circulation, data utilization, and data feedback. During this process, different planes (such as the edge device plane, network orchestration plane, computing power processing plane, data transaction plane, and generative intelligence plane) and multiple participants (including data providers, data infrastructure platforms, and data users) collaborate to form a closed-loop, self-optimizing, and self-evolving new data infrastructure. This infrastructure not only improves the efficiency of data utilization but also promotes the continuous iteration and optimization of the data ecosystem.

[0132] 1. Data perception phase (executed by the edge device plane) In the data perception phase, data providers collect real-time data, such as environmental parameters, business logs, video images, and text records, through various edge devices deployed in industrial sites, connected vehicle environments, or urban infrastructure. These devices not only perform preliminary local processing (such as simple filtering and compression) but also dynamically adjust their collection behavior based on the data infrastructure system's or generative intelligent plane's collection strategies (e.g., increasing nighttime image collection in specific areas or focusing on certain types of industrial fault phenomena). During this phase, the data infrastructure system monitors and manages the physical terminals to ensure device availability and the continuity of data collection. Finally, the edge data is transmitted and aggregated into the network for the next phase. Simultaneously, some data users can directly access the platform via mobile or wireless devices.

[0133] 2. Data Acquisition Phase (Executed by the Network Orchestration Plane) During the data acquisition phase, data generated in the perception phase is transmitted to the data infrastructure system via the network orchestration plane. The intelligent gateway is responsible for protocol conversion and aggregation of data from edge devices, while public cloud and web data are uniformly accessed through the API gateway. The SDN controller schedules data to appropriate computing nodes based on network status and computing power requirements; DetNet / TSN technology ensures deterministic latency and reliability of critical data during cross-domain transmission. At this stage, data users can also access the platform via wired or web-based connections to submit requests for data purchases or generative intelligent services.

[0134] 3. Data preprocessing stage (executed by the computing power processing plane) In the data preprocessing stage, the data infrastructure system schedules multiple computing nodes on the computing power processing plane to perform distributed preprocessing on the received multi-source data. This includes: data alignment, cleaning, and standardization based on data modality and task; compressing large-scale raw data into smaller datasets with sufficient information retention using data distillation methods, and ensuring consistency in statistical characteristics between the distilled data and the original data through techniques such as gradient matching and distribution matching; performing feature extraction and necessary semantic representation to transform the raw data into vector or structured forms suitable for input into generative models. During this stage, participating parties can also act as third-party computing power platforms, contributing computing resources to earn revenue, while the data infrastructure improves overall processing efficiency through a unified resource pool scheduling mechanism.

[0135] 4. Data circulation phase (executed by the data transaction plane) Preprocessed data is stored in the platform's global database. The data transaction plane maintains a multi-party data market through a blockchain system, recording data providers, data users, and their transaction activities. Smart contracts automatically execute data registration, access authorization, price settlement, and reward distribution. Data users, after paying the relevant fees, can download datasets of specific modalities, scenarios, and quality requirements as needed, directly obtaining the source data. Furthermore, this data can be used for subsequent stages in the form of model products and services. Data providers receive revenue based on metrics such as data quality, modality, and usage frequency, thereby incentivizing the continuous contribution of high-value data.

[0136] 5. Data utilization and data feedback phases (executed by the generative intelligent plane) During the data utilization phase, the generative intelligence plane receives high-quality data from the data trading plane, providing pre-training, fine-tuning, and inference services for different types of generative models. Data users can choose to host their purchased data on the data infrastructure platform for training, or they can directly call the inference interface provided by the optimized model to obtain various forms of generated results, such as images, text, policies, and scheduling schemes.

[0137] During the data feedback phase, both the data infrastructure system and the data provider can become "model users," generating new synthetic multimodal data (such as images of extreme industrial scenarios and complex task scheduling strategies) by calling the optimized model. This data is then re-entered into the AIGC database of the data transaction plane. In the data transaction flow, this data can be injected into the platform's global database to participate in further transactions and distribution, thus forming a closed loop of "using models to produce data and using data to optimize models." Furthermore, the optimized GAI model can also be used to generate self-optimizing strategies for various functional mechanisms, demonstrating a significant difference and improvement over traditional data infrastructure: the entire data lifecycle is not only self-driven and self-evolving, but also leverages the intelligent capabilities of generative artificial intelligence to achieve intrinsic intelligence in the architecture, directly driving the adaptive optimization of various functional mechanisms of the data infrastructure and possessing self-evolutionary capabilities.

[0138] In summary, the data infrastructure system spans all stages of the data lifecycle, undertaking platform-based functions such as equipment management, network scheduling, computing power orchestration, data market maintenance, model optimization, and model hosting, ensuring the continuity, controllability, and efficiency of the data throughout its entire lifecycle.

[0139] Accordingly, please refer to Figure 12 This is a block diagram of a bidirectional optimization device for the above-mentioned system provided in an embodiment of this application. The device includes: a forward data flow process component 101 and a reverse optimization flow process component 102.

[0140] In some optional implementations, the forward data flow process component 101 includes: Sensing data from multiple heterogeneous devices is collected via the edge device plane and transmitted to the network orchestration plane. The network orchestration plane plans the transmission links and allocates resources for the sensing data, which is then transmitted to the computing power processing plane. By selecting suitable computing units from the computing resource pool of the computing power processing plane, preprocessing the perceived data, outputting standardized preprocessed data, and transmitting it to the data transaction plane; The data transaction plane scores the value of the preprocessed data, filters out high-quality data that is compatible with the GAI model, and transmits it to the generative intelligence plane. Generative intelligent planes utilize high-quality data to train GAI models, perform inference tasks based on the trained GAI models, and generate self-optimization strategies adapted to each functional plane based on the inference results and data quality features of high-quality data. The high-quality data is then structured and stored in a cross-modal knowledge base.

[0141] In some optional implementations, the forward data flow process component 101 includes: Clean and denoise the perceived data to remove redundant and abnormal data; A multimodal encoder is used to perform semantic-level fusion of perceptual data from different modalities to generate feature vectors of a unified dimension. Feature alignment is performed on the feature vectors to output standardized preprocessed data.

[0142] In some optional implementations, the forward data flow process component 101 includes: The preprocessed data is split into several data units; Obtain the data usage frequency score, data quality score, and model performance contribution score for each data unit; The value score of each data unit is determined based on the data usage frequency score, data quality score, and model performance contribution score. Data units with value scores greater than a preset score threshold are selected as high-quality data for adaptation to the GAI model.

[0143] In some optional implementations, the forward data flow process component 101 includes: Train the GAI model using high-quality data to obtain a well-trained GAI model; The trained GAI model generates a first semantic vector for each data unit in the cross-modal knowledge base and a second semantic vector for the input sample; wherein the dimensions of the first semantic vector and the second semantic vector are the same. Determine the semantic similarity between the first semantic vector and the second semantic vector, and obtain the search results based on the similarity filtering results; The input sample and the retrieval result are concatenated and then fed into the trained GAI model to output the inference result.

[0144] In some optional implementations, the forward data flow process component 101 includes: The prediction confidence of the input sample is obtained by using the trained GAI model; Determine the change in inference confidence between the pre-set target confidence threshold and the predicted confidence level; Data coverage gaps are determined based on changes in inference confidence. Based on the summation of inference confidence changes and data coverage gaps, a self-optimizing strategy is generated to guide edge devices in adjusting their acquisition frequency and range, and to guide network orchestration in optimizing transmission priorities.

[0145] In some optional implementations, the forward data flow process component 101 includes: Obtain the data quality score for each data unit in high-quality data by using a trained GAI model; Based on data quality scores, a self-optimizing strategy is generated to guide the data trading platform in adjusting data pricing, sorting, and value screening rules.

[0146] In some alternative implementations, the reverse optimization workflow component 102 includes: By performing inference tasks using a pre-trained GAI model on the generative intelligent plane, and combining the inference results with the data quality characteristics of high-quality data, a self-optimizing strategy adapted to each functional plane is generated. Each functional plane adjusts its operating parameters based on its corresponding self-optimization strategy. After adjusting the operating parameters, each functional plane re-executes the forward data flow process, outputting updated high-quality data back to the generative intelligent plane to form a self-optimizing closed loop.

[0147] In some alternative implementations, the apparatus further includes a synthetic data generation and reflow process component, including: Scarce seed samples are obtained through generative intelligent plane. These scarce seed samples are real scarce scene data obtained from edge device plane or auxiliary samples retrieved from cross-modal knowledge base. Using a trained GAI model, several synthetic data are generated based on the features of scarce seed samples. The trained GAI model was fine-tuned using synthetic data, and the changes in the model's performance metrics before and after the fine-tuning were determined. If the change exceeds a preset change threshold, the synthesized data will be determined as valid data, and the valid data will be returned to the data transaction plane.

[0148] In some alternative implementations, the apparatus further includes an objective function optimization component, comprising: The current dataset state and the current model state are determined by the generative intelligent plane at preset cycles; Determine the data-side resource consumption cost corresponding to the current dataset state, and determine the model performance improvement benefit corresponding to the current model state; The objective function is the difference between the benefits of improved model performance and the cost of data-side resource consumption. By comparing the net benefits over multiple cycles, an iterative optimization strategy is adopted to approximate the optimal solution. The target dataset state and target model state corresponding to the optimal solution are used as the steady-state benchmark of the system to guide the generation of self-optimization strategies.

[0149] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0150] In this embodiment, the bidirectional optimization device of the above system is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0151] Please see Figure 13 , Figure 13 This application provides a schematic diagram of the structure of a computer device, as shown in the embodiment of the present application. Figure 13 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 13 Take a processor 10 as an example.

[0152] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0153] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0154] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0155] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0156] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0157] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods shown in the above embodiments are implemented.

[0158] The systems, devices, and units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0159] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0160] Those skilled in the art will understand that the embodiments of this application can be provided as methods or systems. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0161] This application is described with reference to flowchart illustrations and / or block diagrams of methods and systems according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0162] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0163] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0164] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0165] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0166] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

[0167] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A multi-plane data infrastructure system, characterized in that, The system comprises multiple functional planes that work in tandem, including an edge device plane, a network orchestration plane, a computing power processing plane, a data transaction plane, and a generative intelligence plane. Each functional plane forms a self-optimizing closed loop through bidirectional data flow. The edge device plane is used to collect sensing data from multi-source heterogeneous devices and adjust the collection parameters according to a self-optimization strategy. The network orchestration plane is used to plan transmission links, allocate bandwidth, and schedule transmission priorities for the sensed data, and output the transmitted sensed data to the computing power processing plane. The computing power processing plane includes a computing power resource pool encapsulated through virtualization and containerization technologies. It is used to select computing power units from the computing power resource pool based on business task requirements, preprocess the perceived data, and output standardized preprocessed data. The data transaction platform is used to realize the asset circulation and value distribution of the preprocessed data based on blockchain and smart contracts, and to filter high-quality data that is compatible with the GAI model through value scoring. The generative intelligent plane includes a cross-modal knowledge base and a GAI model, which are used to train the GAI model using the high-quality data, perform inference tasks based on the trained GAI model, and generate self-optimization strategies adapted to each functional plane based on the inference results and the data quality characteristics of the high-quality data. Each functional plane optimizes based on the corresponding self-optimization strategy and outputs updated high-quality data back to the generative intelligent plane to form a self-optimization closed loop.

2. A bidirectional optimization method based on the system of claim 1, characterized in that, The method includes a forward data flow process and a reverse optimization flow process.

3. The method according to claim 2, characterized in that, The forward data flow process includes: Sensing data from multi-source heterogeneous devices is collected through the edge device plane and transmitted to the network orchestration plane; The network orchestration plane performs transmission link planning and resource allocation on the sensed data, and then transmits it to the computing power processing plane. The computing power processing plane selects suitable computing power units from its computing power resource pool, preprocesses the perceived data, outputs standardized preprocessed data, and transmits it to the data transaction plane. The preprocessed data is value-scored through the data transaction plane, and high-quality data that is suitable for the GAI model is selected and transmitted to the generative intelligence plane. The generative intelligent plane uses the high-quality data to train a GAI model, performs inference tasks based on the trained GAI model, generates self-optimization strategies adapted to each functional plane based on the inference results and the data quality characteristics of the high-quality data, and stores the high-quality data in the cross-modal knowledge base after structuring.

4. The method according to claim 3, characterized in that, The preprocessing of the sensed data to output standardized preprocessed data includes: The sensed data is cleaned and denoised to remove redundant and abnormal data; A multimodal encoder is used to perform semantic-level fusion of perceptual data from different modalities to generate feature vectors of a unified dimension. The feature vectors are aligned to output standardized preprocessed data.

5. The method according to claim 3, characterized in that, The step of scoring the preprocessed data through the data transaction plane to filter out high-quality data suitable for the GAI model includes: The preprocessed data is divided into several data units; Obtain the data usage frequency score, data quality score, and model performance contribution score for each data unit; Based on the data usage frequency score, the data quality score, and the model performance contribution score, the value score of each data unit is determined. Data units with value scores greater than a preset score threshold are selected as high-quality data for the GAI model.

6. The method according to claim 3, characterized in that, The process of training a GAI model using the high-quality data and performing inference tasks based on the trained GAI model includes: The high-quality data is used to train the GAI model, resulting in a well-trained GAI model. The trained GAI model generates a first semantic vector for each data unit in the cross-modal knowledge base and a second semantic vector for the input sample; wherein the first semantic vector and the second semantic vector have the same dimension. Determine the semantic similarity between the first semantic vector and the second semantic vector, and obtain the retrieval results based on the similarity filtering results; The input sample and the retrieval result are concatenated and then input into the trained GAI model to output the inference result.

7. The method according to claim 3, characterized in that, The self-optimization strategy adapted to each functional plane is generated in reverse based on the inference results and the data quality characteristics of the high-quality data, including: The prediction confidence of the input sample is obtained through the trained GAI model; Determine the change in inference confidence between a pre-set target confidence threshold and the predicted confidence level; The data coverage gap is determined based on the change in the inference confidence level. Based on the sum of the inference confidence change and the data coverage gap, a self-optimizing strategy is generated to guide the edge device plane to adjust the acquisition frequency and acquisition range, and to guide the network orchestration plane to optimize transmission priority.

8. The method according to claim 3, characterized in that, The self-optimization strategy adapted to each functional plane is generated in reverse based on the inference results and the data quality characteristics of the high-quality data, including: The trained GAI model is used to obtain the data quality score of each data unit in the high-quality data. Based on the data quality score, a self-optimizing strategy is generated to guide the data trading platform in adjusting data pricing, sorting, and value screening rules.

9. The method according to claim 2, characterized in that, The reverse optimization process includes: The inference task is performed by the trained GAI model of the generative intelligent plane. The inference results and the data quality characteristics of the high-quality data are combined to generate a self-optimizing strategy that is adapted to each functional plane. Each functional plane adjusts its operating parameters based on its corresponding self-optimization strategy. After each functional plane adjusts its operating parameters, it re-executes the forward data flow process and outputs updated, high-quality data back to the generative intelligent plane to form a self-optimizing closed loop.

10. The method according to claim 2, characterized in that, The method also includes a synthetic data generation and reflow process, including: The generative intelligent plane obtains scarce seed samples, which are real scarce scene data obtained from the edge device plane or auxiliary samples retrieved from the cross-modal knowledge base. Using the trained GAI model, several synthetic data are generated based on the features of the scarce seed samples. The trained GAI model is fine-tuned using the synthetic data to determine the changes in model performance metrics before and after fine-tuning. If the amount of change is greater than a preset change threshold, the synthesized data is determined to be valid data, and the valid data is returned to the data transaction plane.

11. The method according to claim 2, characterized in that, The method further includes an objective function optimization step, including: The generative intelligent plane determines the current dataset state and the current model state at a preset period. Determine the data-side resource consumption cost corresponding to the current dataset state, and determine the model performance improvement benefit corresponding to the current model state; The objective function is the difference between the performance improvement benefit of the model and the resource consumption cost on the data side. The net benefits of multiple cycles are compared, and an iterative optimization strategy is used to approximate the optimal solution. The target dataset state and target model state corresponding to the optimal solution are used as the steady-state benchmark of the system to guide the generation of the self-optimization strategy.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the bidirectional optimization method according to any one of claims 2 to 11.