Data processing method, device, electronic device and storage medium
By parsing the data sample construction request and automatically starting the data engine and inference engine to extract data features, the problem of low efficiency in data sample construction is solved and efficient automatic construction of data samples is achieved.
Patent Information
- Application Number
- CN202510984402.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-17
AI Technical Summary
The efficiency of data sample construction in existing technologies is low, mainly because manual labeling consumes a lot of time.
By parsing the received data sample construction request, obtaining the dataset identifier and data feature identifier, determining the data sample construction plan based on the target dataset, data features and current computing resource information, and starting the data engine and inference engine to automatically extract the target data features.
It simplifies the data sample construction process, reduces construction time, improves data sample construction efficiency, achieves the effect of automatic construction of data samples, and avoids the inefficiency caused by manual calculation.
Smart Images

Figure CN120529134B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a data processing method, device, electronic device, storage medium, and computer program product. Background Art
[0002] With the development of computer technology, various AI (Artificial Intelligence) models have emerged. Corresponding data samples are needed when training various AI models.
[0003] In related technologies, current data sample construction methods generally involve manually collecting data sets (such as image data sets) and marking corresponding data features (such as motion features) on the data in the data sets. However, the demand for data samples involved in model training is large, and this manual labeling method consumes a lot of time, resulting in low data sample construction efficiency. Summary of the Invention
[0004] The present disclosure provides a data processing method, apparatus, electronic device, storage medium, and computer program product to at least address the problem of low efficiency in constructing data samples in related technologies. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a data processing method, including:
[0006] Parsing the received data sample construction request to obtain a data set identifier and a data feature identifier; the data set identifier is used to indicate the target data set for the data sample construction request; the data feature identifier is used to indicate the target data feature to be extracted from the target data set;
[0007] Determine a data sample construction plan corresponding to the data sample construction request based on the target data set, the target data characteristics, and the current computing resource information; the data sample construction plan includes at least a data processing link, deployment information of a data engine involved in the data processing link, and deployment information of an inference engine;
[0008] According to the data sample construction plan, the data engine and the inference engine are started to extract the target data features from the target data set, and a data sample construction result corresponding to the data sample construction request is obtained.
[0009] In an exemplary embodiment, determining a data sample construction scheme corresponding to the data sample construction request based on the target data set, the target data characteristics, and current computing resource information includes:
[0010] Acquire metadata of the target data set and metadata of the target data features;
[0011] Determine a data processing link, as well as deployment information of a data engine and an inference engine involved in the data processing link, based on the metadata of the target data set, the metadata of the target data features, and current computing resource information;
[0012] Based on the data processing link, and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link, a data sample construction scheme corresponding to the data sample construction request is determined.
[0013] In an exemplary embodiment, obtaining metadata of the target data set and metadata of the target data feature includes:
[0014] From the metadata management system, identify the dataset metadata management component and the data feature metadata management component;
[0015] Acquire metadata corresponding to the dataset identifier from the dataset metadata management component, and acquire metadata corresponding to the data feature identifier from the data feature metadata management component;
[0016] Based on the metadata corresponding to the dataset identifier, metadata of the target dataset is obtained, and based on the metadata corresponding to the data feature identifier, metadata of the target data feature is obtained.
[0017] In an exemplary embodiment, before starting the data engine and the inference engine to extract the target data features from the target data set according to the data sample construction plan and obtaining a data sample construction result corresponding to the data sample construction request, the method further includes:
[0018] Verifying the data sample construction scheme to obtain a verification result of the data sample construction scheme; the verification result is used to indicate whether the data sample construction scheme is correct and applicable;
[0019] The step of starting the data engine and the inference engine to extract the target data features from the target data set according to the data sample construction scheme to obtain a data sample construction result corresponding to the data sample construction request includes:
[0020] When the verification result indicates that the data sample construction scheme is correct and applicable, the data engine and the inference engine are started according to the data sample construction scheme to extract the target data features from the target data set to obtain a data sample construction result corresponding to the data sample construction request.
[0021] In an exemplary embodiment, verifying the data sample construction scheme to obtain a verification result of the data sample construction scheme includes:
[0022] Performing an initial verification process on the data sample construction scheme using preset verification rules to obtain an initial verification result; the preset verification rules are used to at least verify the resource constraints and link logic of the data sample construction scheme; the initial verification result is used to indicate whether the scheme information of the data sample construction scheme satisfies the preset verification rules;
[0023] In a case where the initial verification result indicates that the scheme information of the data sample construction scheme does not satisfy the preset verification rule, taking the initial verification result as the verification result;
[0024] When the initial verification result indicates that the scheme information of the data sample construction scheme meets the preset verification rules, the data sample construction scheme is re-verified through the preset verification instructions to obtain a re-verification result as the verification result of the data sample construction scheme; the preset verification instructions are at least used to verify the execution efficiency and feature extraction accuracy of the data sample construction scheme.
[0025] In an exemplary embodiment, when the verification result indicates that the data sample construction scheme is correct and applicable, starting the data engine and the inference engine to extract the target data features from the target data set according to the data sample construction scheme to obtain a data sample construction result corresponding to the data sample construction request includes:
[0026] If the verification result indicates that the data sample construction scheme is correct and applicable, starting the data engine to preprocess the data in the target data set according to the data sample construction scheme to obtain preprocessed data;
[0027] Starting the inference engine to load a target operator corresponding to the target data feature in an operator library, and extracting the target data feature of the preprocessed data through the target operator as the target data feature of the data in the target data set;
[0028] The data engine performs association processing on the data in the target data set and the target data features of the data in the target data set to obtain the data sample construction result.
[0029] In an exemplary embodiment, after constructing a solution according to the data sample and starting the data engine and the inference engine to extract the target data features from the target data set, the method further includes:
[0030] Real-time monitoring of the execution status of the data sample construction plan;
[0031] When the execution state does not satisfy the preset condition, updating the data sample construction scheme to obtain an updated data sample construction scheme;
[0032] According to the updated data sample construction scheme, the target data features are extracted from the unprocessed data in the target data set.
[0033] According to a second aspect of an embodiment of the present disclosure, there is provided a data processing apparatus, including:
[0034] a request parsing unit configured to parse the received data sample construction request to obtain a data set identifier and a data feature identifier; the data set identifier is used to indicate the target data set for which the data sample construction request is intended; and the data feature identifier is used to indicate the target data features to be extracted from the target data set;
[0035] A solution determination unit is configured to determine a data sample construction solution corresponding to the data sample construction request based on the target data set, the target data characteristics, and the current computing resource information; the data sample construction solution includes at least a data processing link, deployment information of a data engine, and deployment information of an inference engine;
[0036] The sample construction unit is configured to execute according to the data sample construction plan, start the data engine and the inference engine to extract the target data features from the target data set, and obtain a data sample construction result corresponding to the data sample construction request.
[0037] In an exemplary embodiment, the scheme determination unit is further configured to execute acquisition of the metadata of the target data set and the metadata of the target data features; determine the data processing link, and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link based on the metadata of the target data set, the metadata of the target data features and the current computing resource information; determine the data sample construction scheme corresponding to the data sample construction request based on the data processing link and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link.
[0038] In an exemplary embodiment, the scheme determination unit is further configured to identify a dataset metadata management component and a data feature metadata management component from a metadata management system; obtain metadata corresponding to the dataset identifier from the dataset metadata management component, and obtain metadata corresponding to the data feature identifier from the data feature metadata management component; obtain metadata of the target dataset based on the metadata corresponding to the dataset identifier, and obtain metadata of the target data feature based on the metadata corresponding to the data feature identifier.
[0039] In an exemplary embodiment, the apparatus further includes a solution verification unit configured to verify the data sample construction solution and obtain a verification result of the data sample construction solution; the verification result is used to indicate whether the data sample construction solution is correct and applicable;
[0040] The sample construction unit is further configured to execute, when the verification result indicates that the data sample construction scheme is correct and applicable, start the data engine and the inference engine according to the data sample construction scheme to extract the target data features from the target data set and obtain a data sample construction result corresponding to the data sample construction request.
[0041] In an exemplary embodiment, the scheme verification unit is further configured to perform initial verification processing on the data sample construction scheme through preset verification rules to obtain an initial verification result; the preset verification rules are at least used to verify the resource constraints and link logic of the data sample construction scheme; the initial verification result is used to indicate whether the scheme information of the data sample construction scheme satisfies the preset verification rules; if the initial verification result indicates that the scheme information of the data sample construction scheme does not satisfy the preset verification rules, the initial verification result is used as the verification result; if the initial verification result indicates that the scheme information of the data sample construction scheme satisfies the preset verification rules, the data sample construction scheme is re-verified through preset verification instructions to obtain a re-verification result as the verification result of the data sample construction scheme; the preset verification instructions are at least used to verify the execution efficiency and feature extraction accuracy of the data sample construction scheme.
[0042] In an exemplary embodiment, the sample construction unit is further configured to execute, when the verification result indicates that the data sample construction scheme is correct and applicable, starting the data engine to preprocess the data in the target data set according to the data sample construction scheme to obtain preprocessed data; starting the inference engine to load the target operator corresponding to the target data feature in the operator library, and extracting the target data feature of the preprocessed data through the target operator as the target data feature of the data in the target data set; and performing association processing on the data in the target data set and the target data feature of the data in the target data set through the data engine to obtain the data sample construction result.
[0043] In an exemplary embodiment, the sample construction unit is further configured to perform real-time monitoring of the execution status of the data sample construction scheme; when the execution status does not meet the preset conditions, the data sample construction scheme is updated to obtain an updated data sample construction scheme; and according to the updated data sample construction scheme, the target data features are extracted from the unprocessed data in the target data set.
[0044] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0045] processor;
[0046] a memory for storing instructions executable by the processor;
[0047] The processor is configured to execute the instructions to implement the data processing method as described in any one of the embodiments of the first aspect.
[0048] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the data processing method as described in any one of the embodiments of the first aspect.
[0049] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes instructions. When the instructions are executed by a processor of an electronic device, the electronic device is capable of executing the data processing method described in any one of the embodiments of the first aspect.
[0050] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:
[0051] By parsing the received data sample construction request, the data set identifier and data feature identifier are obtained; the data set identifier is used to represent the target data set targeted by the data sample construction request; the data feature identifier is used to represent the target data features to be extracted from the target data set; then, based on the target data set, target data features and current computing power resource information, the data sample construction plan corresponding to the data sample construction request is determined; the data sample construction plan includes at least the data processing link, the deployment information of the data engine involved in the data processing link and the deployment information of the inference engine; finally, according to the data sample construction plan, the data engine and the inference engine are started to extract the target data features from the target data set, and the data sample construction result corresponding to the data sample construction request is obtained. In this way, when constructing a data sample, a corresponding data sample construction plan is automatically generated through a data sample construction request containing a data set identifier and a data feature identifier, and through the data sample construction plan, the data engine and the inference engine are automatically started to extract target data features from the target data set. The entire process does not require manual collection of data sets and calculation of data features, thereby simplifying the data sample construction process, which is conducive to reducing data sample construction time and thus improving data sample construction efficiency; moreover, the purpose of automatically extracting target data features from the target data set based on the data sample construction request is achieved, achieving the effect of automatically constructing data samples and further improving data sample construction efficiency; at the same time, it avoids the defect of low data sample construction efficiency caused by manual calculation of data features on the data set.
[0052] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0054] Figure 1 The figure is a flow chart showing a data processing method according to an exemplary embodiment.
[0055] Figure 2 The flowchart shows the steps of determining a data sample construction solution corresponding to a data sample construction request according to an exemplary embodiment.
[0056] Figure 3 The flowchart shows the steps of obtaining a data sample construction result corresponding to a data sample construction request according to an exemplary embodiment.
[0057] Figure 4 The flowchart of another data processing method is shown according to an exemplary embodiment.
[0058] Figure 5 The figure is an overall architecture diagram of a native multimodal data sample construction system according to an exemplary embodiment.
[0059] Figure 6 It is a block diagram showing a metadata management system according to an exemplary embodiment.
[0060] Figure 7 The figure is a flowchart of a data sample construction method according to an exemplary embodiment.
[0061] Figure 8 It is a block diagram of a data processing device according to an exemplary embodiment.
[0062] Figure 9 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0063] In order to enable ordinary people in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0064] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0065] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0066] With the rapid development of generative AI technology, multimodal data (such as text, images, video, audio, point clouds, and motion patterns) has become a core component of the AI ecosystem. Understanding multimodal data directly determines the quality of AI models. Currently, the annual global generation of multimodal data has reached 175 zettabytes, with less than 15% being effectively utilized. This demonstrates that understanding the content of data in different modalities, mining cross-modal correlations, and natively integrating heterogeneous data sources are core data requirements in the AI era, presenting new challenges for the data engineering field. Currently, generative AI is developing rapidly, and companies are rapidly iterating their AI models. Building model training data is driven by the goal of rapidly delivering business needs, resulting in a patchwork approach. This leads to overlapping data requirements across various businesses, resulting in a siloed data chain and architecture. In the long term, this approach will hinder cross-functionality across data chains, increasing the need for reinventing the wheel, impacting R&D efficiency and operating costs. As demand complexity and data chains increase, engineering efficiency will gradually lag behind business growth. Based on this, the present invention proposes a data processing method, specifically a native multimodal data sample construction method, which can effectively improve the efficiency of data sample construction.
[0067] Figure 1 is a flow chart showing a data processing method according to an exemplary embodiment. Figure 1 As shown, the data processing method is used in a server; it is understood that the method can also be applied to a terminal, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, and tablet computers, and the server can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. In this exemplary embodiment, the method includes the following steps:
[0068] In step S110, the received data sample construction request is parsed to obtain a data set identifier and a data feature identifier; the data set identifier is used to indicate the target data set targeted by the data sample construction request; the data feature identifier is used to indicate the target data features to be extracted from the target data set.
[0069] Among them, a data sample construction request refers to a request to construct a corresponding data sample based on a selected target dataset (such as dataset A) and selected target data features (such as data feature B). For example, a data sample construction request is made for a short video comment text dataset and sentiment polarity features (such as positive and negative), a data sample construction request is made for an automobile engine cylinder X-ray image dataset and component defect features (such as crack length features and sand hole area features), a data sample construction request is made for a short video dataset and screen subject recognition features (such as people, pets, and products), and a data sample construction request is made for a song audio dataset and song style features (such as rock, folk, and pop).
[0070] Among them, the data sample construction request can be triggered by the user or the system; for example, Figure 7 To trigger a data sample creation request, users simply select the dataset for which they want to create a data sample (e.g., dataset ID: 123) and the data feature they want to extract from the data in this dataset (e.g., feature X). When selecting a data feature, users can also select a feature version (e.g., version Y). Feature versions refer to a versioning mechanism for managing and iterating data features, used to record, track, and manage changes to features at different stages.
[0071] Among them, the data sample construction request includes a dataset identifier and a data feature identifier, which are used to represent the data in the target dataset corresponding to the dataset identifier, and extract the target data features corresponding to the data feature identifier, such as extracting emotional polarity features from the comment text data in the short video comment text dataset, extracting component defect features from the X-ray image data in the automobile engine cylinder X-ray image dataset, extracting picture subject recognition features from the short video data in the short video dataset, and extracting song style features from the song audio data in the song audio dataset.
[0072] Among them, the dataset identifier is used to characterize the target dataset for the data sample construction request, such as a short video comment text dataset, a car engine cylinder X-ray image dataset, a short video dataset, a song audio dataset, etc., and can be specifically represented by an ID (such as 123). It should be noted that the target dataset involved in this disclosure refers to a multimodal dataset, that is, the target dataset involved in this disclosure can be a dataset of any modality, such as a text dataset, an image dataset, a video dataset, an audio dataset, a point cloud dataset, etc.; the target data features involved in this disclosure can be data features of any modality, such as text data features, image data features, video data features, audio data features, point cloud data features, etc., and this disclosure does not make any specific limitations.
[0073] Among them, the data feature identifier is used to represent the target data features that need to be extracted from the data in the target data set, such as emotional polarity features, component defect features, picture subject recognition features, song style features, etc., which can be specifically represented by symbols, such as using AGE_001 to represent user age, and PUR_FREQ_002 to represent purchase frequency, etc.
[0074] Exemplarily, the terminal generates a corresponding data sample construction request in response to the user's data sample construction operation, and sends the data sample construction request to the corresponding server; the server authenticates the data sample construction request, and if the data sample construction request is authenticated, the server parses the received data sample construction request through the request parsing instruction, obtains the data set identifier and data feature identifier contained in the data sample construction request, and then determines the target data set for the data sample construction request based on the data set identifier, and determines the target data feature to be extracted from the data in the target data set based on the data feature identifier; for example, the server queries the correspondence between the data set identifier and the data set based on the data set identifier, obtains the data set corresponding to the data set identifier, and uses it as the target data set; based on the data feature identifier, the server queries the correspondence between the data feature identifier and the data feature, obtains the data feature corresponding to the data feature identifier, and uses it as the target data feature.
[0075] For example, the user selects the ID of the short video comment text dataset (such as 456) and the symbol of the emotion polarity feature (such as Emotion_003) on the interface, and clicks the "Submit" option to trigger a data sample construction request for the short video comment text dataset and the emotion polarity feature, and sends the data sample construction request to the server through the terminal; the server parses the data sample construction request, determines that the target dataset is the "short video comment text dataset", and determines that the target data feature is the "emotion polarity feature".
[0076] In step S120, based on the target data set, target data characteristics and current computing resource information, a data sample construction plan corresponding to the data sample construction request is determined; the data sample construction plan includes at least the data processing link, the deployment information of the data engine involved in the data processing link and the deployment information of the inference engine.
[0077] Among them, current computing power resource information refers to the computing power resource information currently available on the server, specifically the hardware resources (such as computing power, storage resources, network resources, etc.) and software resources currently available on the server. For example, the currently available computing power resources in the machine learning platform (such as training cluster resources, inference service resources, feature storage resources, etc.) and the currently available computing power resources in the cloud native platform (such as container cluster resources, automatic scaling status, etc.). Training cluster resources refer to the number of idle nodes available for feature engineering and model training, and the remaining amount of GPU memory. Inference service resources refer to the number of deployed instances of the inference engine, the CPU / GPU utilization rate of each instance, and the concurrent processing capacity. Feature storage resources refer to the available space on the storage nodes of the feature library and the resource quota for feature calculation tasks.
[0078] Among them, the data sample construction plan refers to the plan for executing the data sample construction request, specifically refers to the execution plan for extracting target data features from the data in the target data set, including the data processing link, the deployment information of the data engine involved in the data processing link, and the deployment information of the inference engine.
[0079] The data processing chain refers to the processing chain for extracting target data features from the data in the target dataset. For example, the data in the target dataset is first preprocessed by the data engine to obtain preprocessed data. The target data features are then extracted from the preprocessed data by the inference engine, and used as the target data features of the data in the target dataset. Finally, the data engine associates the data in the target dataset with the target data features of the data in the target dataset to obtain the data sample construction results. For example, to extract the subject recognition features of short video data in a short video dataset, the corresponding data processing chain is as follows: the data engine (Spark) first extracts video frames from the short video data in the short video dataset → the inference engine (AI model service) performs subject recognition feature extraction processing on the extracted video frames → the data engine (Hive) stores the results of the short video data and the extracted subject recognition features.
[0080] Among them, the data engine refers to a tool or system that reads, cleans, converts and aggregates raw data to provide standardized data for subsequent processing (such as model inference), such as Spark (distributed data engine), Flink (real-time data engine), Pandas (lightweight data engine), etc.
[0081] Among them, the inference engine refers to a tool or system that loads a trained model, performs predictive calculations on the input standardized data, and outputs business results, while optimizing inference efficiency to support high concurrency or low latency requirements. Examples include TensorFlow Serving (supports online deployment and inference of TensorFlow models), Triton Inference Server (multi-framework GPU-accelerated inference), and ONNX Runtime (cross-platform lightweight inference).
[0082] The data engine deployment information refers to the number of deployed instances of the data engine and the concurrency of each deployed instance. The inference engine deployment information refers to the number of deployed instances of the data engine and the concurrency of each deployed instance. For example, for the inference engine, testing found that one instance could process 1,000 video slices in one hour. To complete 10,000 video slices in two hours, the number of instances was set to 5 (5 × 1,000 × 2 = 10,000), and the concurrency of each instance was set to 20 (i.e., processing 20 video slices simultaneously). For the data engine, to transmit data to five inference instances, the number of instances should be set to 2, and the concurrency of each instance should be set to 50 (i.e., outputting 50 video frames per second, matching the total concurrency requirement of 5 instances × 20).
[0083] Exemplarily, the server obtains current computing resource information and then, based on the target data set, target data features, and current computing resource information, queries the correspondence between the data set, data features, computing resource information, and the data sample construction plan, obtains the data sample construction plan corresponding to the target data set, target data features, and current computing resource information, and uses it as the data sample construction plan corresponding to the data sample construction request. Alternatively, the server determines the data processing link corresponding to the data sample construction request, as well as the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link, based on the target data set, target data features, and current computing resource information, and constructs the data sample construction plan corresponding to the data sample construction request based on the data processing link and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link.
[0084] For example, the server determines a data sample construction plan related to "extracting picture subject identification features from short video data in the short video dataset" based on the short video dataset, picture subject identification features and current computing resource information.
[0085] In step S130 , according to the data sample construction plan, the data engine and the inference engine are started to extract target data features from the target data set, and a data sample construction result corresponding to the data sample construction request is obtained.
[0086] Among them, extracting target data features from the target dataset refers to extracting target data features from the data in the target dataset, such as extracting emotional polarity features from the comment text data in the short video comment text dataset, extracting component defect features from the X-ray image data in the automobile engine cylinder X-ray image dataset, extracting picture subject recognition features from the short video data in the short video dataset, and extracting song style features from the song audio data in the song audio dataset.
[0087] The data sample construction result includes multiple data samples (such as multiple text data samples, multiple image data samples, multiple video data samples, multiple audio data samples, and multiple point cloud data samples). Each data sample includes one data in the target dataset (such as text data, image data, video data, audio data, and point cloud data), as well as the target data features corresponding to the data (such as text data features, image data features, video data features, audio data features, and point cloud data features). For example, for a short video comment text dataset and sentiment polarity features, each data sample includes comment text data and sentiment polarity features corresponding to the comment text data; for a car engine cylinder X-ray image dataset and component defect features, each data sample includes X-ray image data and component defect features corresponding to the X-ray image data; for a short video dataset and image subject recognition features, each data sample includes short video data and image subject recognition features corresponding to the short video data; for a song audio dataset and song style features, each data sample includes song audio data and song style features corresponding to the song audio data.
[0088] Exemplarily, the server starts the data engine and the inference engine according to the data sample construction plan; through the started data engine and the started inference engine, based on the data processing link, the corresponding target data features are extracted from each data in the target data set; then each data in the target data set and the target data features corresponding to each data are corresponded as each data sample, that is, each data sample includes a data in the target data set and the target data features corresponding to the data; finally, based on all data samples, a data sample construction result corresponding to the data sample construction request is obtained.
[0089] Furthermore, the server can train a corresponding AI model based on the data samples in the data sample construction result; wherein the data in the data sample is used as model input data, and the target data features corresponding to the data in the data sample are used as labels. For example, the server trains the review text processing model to be trained based on the comment text data and the emotional polarity features corresponding to the comment text data, so that the trained review text processing model can extract the corresponding emotional polarity features from the new comment text data. The server trains the X-ray image processing model to be trained based on the automobile engine cylinder X-ray image data and the component defect features corresponding to the automobile engine cylinder X-ray image data, so that the trained X-ray image processing model can extract the corresponding component defect features from the new automobile engine cylinder X-ray image data. The server trains the video data processing model to be trained based on the short video data and the picture subject recognition features corresponding to the short video data, so that the trained video data processing model can extract the corresponding picture subject recognition features from the new short video data. The server trains the audio data processing model to be trained based on the song audio data and the song style features corresponding to the song audio data, so that the trained audio data processing model can extract the corresponding song style features from the new song audio data.
[0090] In the above data processing method, the data set identifier and the data feature identifier are obtained by parsing the received data sample construction request; the data set identifier is used to represent the target data set targeted by the data sample construction request; the data feature identifier is used to represent the target data features to be extracted for the target data set; then, based on the target data set, the target data features and the current computing power resource information, the data sample construction plan corresponding to the data sample construction request is determined; the data sample construction plan includes at least the data processing link, the deployment information of the data engine involved in the data processing link and the deployment information of the inference engine; finally, according to the data sample construction plan, the data engine and the inference engine are started to extract the target data features of the target data set to obtain the data sample construction result corresponding to the data sample construction request. In this way, when constructing a data sample, a corresponding data sample construction plan is automatically generated through a data sample construction request containing a data set identifier and a data feature identifier, and through the data sample construction plan, the data engine and the inference engine are automatically started to extract target data features from the target data set. The entire process does not require manual collection of data sets and calculation of data features, thereby simplifying the data sample construction process, which is conducive to reducing data sample construction time and thus improving data sample construction efficiency; moreover, the purpose of automatically extracting target data features from the target data set based on the data sample construction request is achieved, achieving the effect of automatically constructing data samples and further improving data sample construction efficiency; at the same time, it avoids the defect of low data sample construction efficiency caused by manual calculation of data features on the data set.
[0091] In an exemplary embodiment, Figure 2 As shown, the above step S120 determines the data sample construction scheme corresponding to the data sample construction request based on the target data set, target data characteristics and current computing resource information, which can be specifically implemented by the following steps:
[0092] In step S210 , metadata of the target data set and metadata of target data features are acquired.
[0093] In step S220, based on the metadata of the target data set, the metadata of the target data features and the current computing resource information, the data processing link and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link are determined.
[0094] In step S230 , a data sample construction scheme corresponding to the data sample construction request is determined based on the data processing link and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link.
[0095] The target dataset's metadata describes the overall dataset's macro-attributes, enabling understanding of its overall context, source, and management information. Specifically, it includes basic information, source and ownership, scale and format, structural information, quality descriptions, and related information. Basic information includes the dataset name, unique identifier (e.g., ID), creation time, update time, and storage path (e.g., database table name, file path). Source and ownership include the data source (e.g., business system, log collection, third-party API), data owner, authorized use scope, and privacy agreement (e.g., whether sensitive information is included). Scale and format include the total data volume (e.g., 1 million records), data format (e.g., CSV, Parquet, JSON), and storage type (e.g., relational database, data lake, data warehouse). Structural information includes data dimensions (e.g., number of columns / fields), number of records, and data type distribution (e.g., number of numeric, string, and date fields). Quality descriptions include data completeness (e.g., percentage of missing values), consistency (e.g., compliance with business rules), and timeliness (e.g., time range covered by the data). Related information includes relationships with other datasets (e.g., the ID field of the "User Orders Dataset" linked to the "User Information Dataset").
[0096] The metadata of target data features describes the micro-attributes of the target data features, helping to understand their meaning, type, and processing rules. Specifically, it includes basic information, business significance, statistical attributes, quality information, processing history, and usage restrictions. Basic information includes the feature name, unique identifier, dataset, and data type (e.g., int, float, string, datetime). Business significance includes the feature's business definition (e.g., "user age" refers to the age calculated from the birth year entered during user registration) and business tags (e.g., "user attribute" and "behavioral characteristics"). Statistical attributes include numerical features (e.g., mean, median, maximum, minimum, standard deviation, and distribution type) and categorical features (e.g., enumerated values and their number, and the proportion of each value). Quality information includes the number and proportion of missing values, outlier flags (e.g., values outside a reasonable range), and duplicate values. Processing history includes the feature's generation method (e.g., "total amount" is calculated by aggregating the original order table), the upstream features or data sources it relies on, and the update frequency (e.g., real-time or daily). Usage restrictions include whether it is a sensitive feature (such as ID number, mobile phone number), desensitization rules (such as replacing some characters with *), and whether it is allowed to be used for model training or inference.
[0097] Exemplarily, the server obtains metadata of the target dataset, metadata of the target data features, and current computing resource information. Then, based on the metadata of the target dataset (such as magnitude), target data features (such as processing complexity), and current computing resource information (such as the current computing resource status of the machine learning platform and cloud native platform), it queries the correspondence between the metadata of the dataset, metadata of the data features, computing resource information, and the data sample construction plan, obtains the data sample construction plan corresponding to the metadata of the target dataset, metadata of the target data features, and current computing resource information, and uses it as the data sample construction plan corresponding to the data sample construction request. Alternatively, based on the metadata of the target dataset, metadata of the target data features, and current computing resource information, the server determines the data processing link corresponding to the data sample construction request, as well as the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link, and constructs the data sample construction plan corresponding to the data sample construction request based on the data processing link and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link.
[0098] The technical solution provided by the embodiment of the present disclosure comprehensively considers the metadata of the target data set, the metadata of the target data features and the current computing resource information when determining the data sample construction plan, which is conducive to improving the accuracy of determining the data sample construction plan; moreover, it also comprehensively considers the data processing link, as well as the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link, which is conducive to further improving the accuracy of determining the data sample construction plan.
[0099] In an exemplary embodiment, the above-mentioned step S210, obtaining metadata of the target data set and metadata of the target data feature, specifically includes the following contents: identifying the data set metadata management component and the data feature metadata management component from the metadata management system; obtaining metadata corresponding to the data set identifier from the data set metadata management component, and obtaining metadata corresponding to the data feature identifier from the data feature metadata management component; obtaining metadata of the target data set based on the metadata corresponding to the data set identifier, and obtaining metadata of the target data feature based on the metadata corresponding to the data feature identifier.
[0100] Among them, reference Figure 6 The metadata management system is used to manage different types of metadata, such as metadata of data features, metadata of data sets, metadata of requirements (i.e., data sample construction requests), and metadata of solutions. The metadata management system includes different types of metadata management components, such as data feature metadata management components (i.e., Figure 6 FeatureCataLog component shown), dataset metadata management component (i.e. Figure 6The DataSetCataLog component shown), the demand metadata management component (i.e. Figure 6 MissionCataLog component shown), solution metadata management component (i.e. Figure 6 ApproachCataLog component shown).
[0101] The dataset metadata management component is used to manage dataset metadata. The data feature metadata management component is used to manage data feature metadata. The target dataset metadata is the metadata corresponding to the dataset identifier. The target data feature metadata is the metadata corresponding to the data feature identifier.
[0102] Exemplarily, the server identifies the dataset metadata management component and the data feature metadata management component from the metadata management system through the dataset metadata management component identification instruction and the data feature metadata management component identification instruction; then, based on the dataset identifier, the server obtains the metadata corresponding to the dataset identifier from the dataset metadata management component, and based on the data feature identifier, the server obtains the metadata corresponding to the data feature identifier from the data feature metadata management component; finally, the metadata corresponding to the dataset identifier is used as the metadata of the target dataset, and the metadata corresponding to the data feature identifier is used as the metadata of the target data feature.
[0103] The technical solution provided by the embodiments of the present disclosure identifies a dataset metadata management component and a data feature metadata management component from a metadata management system, then obtains metadata corresponding to the dataset identifier from the dataset metadata management component as metadata of a target dataset, and obtains metadata corresponding to the data feature identifier from the data feature metadata management component as metadata of a target data feature, thereby achieving the purpose of accurately and efficiently obtaining metadata of a target dataset and metadata of a target data feature.
[0104] In an exemplary embodiment, the above step S130, before starting the data engine and the inference engine to extract target data features from the target data set according to the data sample construction plan and obtaining the data sample construction result corresponding to the data sample construction request, further includes the step of verifying the data sample construction plan, specifically including the following contents: verifying the data sample construction plan to obtain a verification result of the data sample construction plan; the verification result is used to indicate whether the data sample construction plan is correct and applicable. Then, the above step S130, which starts the data engine and the inference engine to extract target data features from the target data set according to the data sample construction plan and obtains the data sample construction result corresponding to the data sample construction request, specifically includes the following contents: if the verification result indicates that the data sample construction plan is correct and applicable, starting the data engine and the inference engine to extract target data features from the target data set according to the data sample construction plan and obtains the data sample construction result corresponding to the data sample construction request.
[0105] Among them, the verification of the data sample construction plan is mainly to verify whether the generated data sample construction plan is correct and usable.
[0106] Exemplarily, the server verifies the data sample construction scheme through a scheme verification instruction to determine whether the data sample construction scheme is correct and usable, thereby obtaining a verification result of the data sample construction scheme; for example, through artificial rules and some intelligent methods, to evaluate whether the generated data sample construction scheme is correct and usable; when the verification result indicates that the data sample construction scheme is correct and usable, the data engine and the inference engine are started according to the data sample construction scheme; through the started data engine and the started inference engine, based on the data processing link, the corresponding target data features are extracted from each data in the target data set; then each data in the target data set and the target data features corresponding to each data are corresponding to each data sample; finally, based on all data samples, the data sample construction result corresponding to the data sample construction request is obtained.
[0107] The technical solution provided by the embodiment of the present disclosure verifies the data sample construction scheme and obtains the verification result of the data sample construction scheme; when the verification result indicates that the data sample construction scheme is correct and applicable, the data engine and the inference engine are started according to the data sample construction scheme to extract the target data features of the target data set, and obtain the data sample construction result corresponding to the data sample construction request; in this way, the target data feature extraction operation is performed only after the data sample construction scheme is verified, thereby ensuring the accuracy of the obtained data sample construction result, thereby improving the data sample construction accuracy rate.
[0108] In an exemplary embodiment, the data sample construction scheme is verified to obtain a verification result of the data sample construction scheme, which specifically includes the following contents: the data sample construction scheme is initially verified through preset verification rules to obtain an initial verification result; the preset verification rules are at least used to verify the resource constraints and link logic of the data sample construction scheme; the initial verification result is used to indicate whether the scheme information of the data sample construction scheme satisfies the preset verification rules; when the initial verification result indicates that the scheme information of the data sample construction scheme does not satisfy the preset verification rules, the initial verification result is used as the verification result; when the initial verification result indicates that the scheme information of the data sample construction scheme satisfies the preset verification rules, the data sample construction scheme is re-verified through preset verification instructions to obtain a re-verification result as the verification result of the data sample construction scheme; the preset verification instructions are at least used to verify the execution efficiency and feature extraction accuracy of the data sample construction scheme.
[0109] Among them, the preset verification rules refer to the preset manual rules used to verify the resource constraints and link logic of the data sample construction plan, which at least include resource constraint verification rules and link logic verification rules. The resource constraint verification rules are used to verify the resource constraints of the data sample construction plan, and are specifically used to verify whether the resource allocation is reasonable, such as whether the GPU memory usage does not exceed 80%, and the computing node CPU load is within the threshold. The link logic verification rules are used to verify the link logic of the data sample construction plan, and are specifically used to verify whether the link logic is correctly closed, such as data extraction from Hive → Spark processing → inference engine recognition → result writing back to storage.
[0110] Among them, the initial verification result is used to indicate whether the scheme information of the data sample construction scheme (such as resource constraints and link logic) meets the preset verification rules. If not, the initial verification result is used as the verification result of the data sample construction scheme; if it meets, the data sample construction scheme is re-verified through the preset verification instructions.
[0111] Among them, the preset verification instructions refer to preset intelligent instructions used to verify the execution efficiency and feature extraction accuracy of the data sample construction plan, which at least include execution efficiency verification instructions and feature extraction accuracy verification instructions. The execution efficiency verification instructions are used to verify the execution efficiency of the data sample construction plan, such as selecting 100 video slices (covering different scenes and action types) for simulation calculation to verify the execution efficiency of the plan (such as whether the action recognition of 100 slices is completed within 1 hour). The feature extraction accuracy verification instructions are used to verify the feature extraction accuracy of the data sample construction plan, such as comparing the extracted recognition actions with the manually annotated action labels for 100 video slices to calculate the recognition accuracy. In actual scenarios, preset verification instructions refer to similarity matching instructions based on historical cases, resource load simulation and stress testing instructions, cost-benefit optimization analysis instructions, and machine learning prediction model evaluation instructions.
[0112] Exemplarily, the server obtains preset verification rules for verifying the resource constraints and link logic of the data sample construction plan, and performs initial verification processing on the data sample construction plan through the preset verification rules to determine whether the plan information of the data sample construction plan (such as resource constraints and link logic) meets the preset verification rules, thereby obtaining an initial verification result; when the initial verification result indicates that the plan information of the data sample construction plan does not meet the preset verification rules, it means that the data sample construction plan is not correctly available, and the initial verification result is used as the verification result of the data sample construction plan; when the initial verification result indicates that the plan information of the data sample construction plan meets the preset verification rules, the execution efficiency of the data sample construction plan is obtained. and preset verification instructions for feature extraction accuracy, and the data sample construction scheme is re-verified through the preset verification instructions to determine whether the execution efficiency and feature extraction accuracy of the data sample construction scheme meet the preset requirements (for example, whether the execution efficiency of the data sample construction scheme is greater than the preset execution efficiency, and whether the feature extraction accuracy is greater than the preset accuracy rate), thereby obtaining a re-verification result, and finally using the re-verification result as the verification result of the data sample construction scheme; for example, if the execution efficiency and feature extraction accuracy of the data sample construction scheme do not meet the preset requirements, then it means that the data sample construction scheme is not correct and available; if the execution efficiency and feature extraction accuracy of the data sample construction scheme meet the preset requirements, then it means that the data sample construction scheme is correct and available.
[0113] The technical solution provided by the embodiment of the present disclosure performs initial verification processing on the resource constraints and link logic of the data sample construction scheme through preset verification rules, and re-verifies the execution efficiency and feature extraction accuracy of the data sample construction scheme through preset verification instructions, thereby achieving the purpose of double verification of the data sample construction scheme, ensuring the accuracy of the verification results of the data sample construction scheme, and thus improving the accuracy of the determination of the verification results of the data sample construction scheme.
[0114] In an exemplary embodiment, Figure 3 As shown, when the verification result indicates that the data sample construction plan is correct and available, the data engine and the inference engine are started according to the data sample construction plan to extract the target data features from the target data set, and a data sample construction result corresponding to the data sample construction request is obtained. This can be specifically achieved through the following steps:
[0115] In step S310 , when the verification result indicates that the data sample construction scheme is correct and applicable, the data engine is started to preprocess the data in the target data set according to the data sample construction scheme to obtain preprocessed data.
[0116] In step S320, the inference engine is started to load the target operator corresponding to the target data feature in the operator library, and the target data feature of the preprocessed data is extracted through the target operator as the target data feature of the data in the target data set.
[0117] In step S330, the data engine performs association processing on the data in the target data set and the target data features of the data in the target data set to obtain a data sample construction result.
[0118] Preprocessing the data in the target dataset refers to performing preliminary processing on the data in the target dataset, such as extracting video frames from a video, extracting key text information from text information, etc.
[0119] Among them, the operator library includes a variety of operators, such as Figure 5 The following examples illustrate transcoding, frame extraction, MD5 deduplication, image description generation, optical flow, and variational autoencoders. Target operators are operators that extract target data features, such as frame extraction and transcoding.
[0120] Exemplarily, when the verification result indicates that the data sample construction scheme is correct and available, the server starts the corresponding data engine to preprocess the data in the target data set according to the data link in the data sample construction scheme and the deployment information of the data engine and the deployment information of the inference engine involved in the data link to obtain preprocessed data; then, based on the correspondence between the data features and the operators, the target operator corresponding to the target data features in the operator library is determined, and then the corresponding inference engine is started to load the target operator corresponding to the target data features in the operator library, and the target data features of the preprocessed data are extracted through the target operator, and the target data features of the preprocessed data are used as the target data features of the data in the target data set; finally, the data engine performs association processing on each data in the target data set and the target data features of each data in the target data set to obtain multiple data samples, that is, one data sample includes one data in the target data set and the target data features associated with the data, and based on the multiple data samples, the data sample construction result is obtained.
[0121] The technical solution provided by the embodiment of the present disclosure, when the verification result indicates that the data sample construction plan is correct and applicable, starts the data engine and the inference engine according to the data sample construction plan to extract the target data features from the target data set, and obtains the data sample construction result corresponding to the data sample construction request; in this way, the purpose of automatically constructing the corresponding data sample based on the data sample construction plan is achieved, and the entire process does not require manual calculation of data features, thereby simplifying the data sample construction process, which is conducive to saving data sample construction time and thus improving data sample construction efficiency.
[0122] In an exemplary embodiment, the above-mentioned step S130, after starting the data engine and the inference engine to extract target data features from the target data set according to the data sample construction plan, also includes the step of updating the data sample construction plan, which specifically includes the following contents: real-time monitoring of the execution status of the data sample construction plan; when the execution status does not meet the preset conditions, updating the data sample construction plan to obtain an updated data sample construction plan; and extracting target data features from the unprocessed data in the target data set according to the updated data sample construction plan.
[0123] The execution status of the monitoring data sample construction plan can refer to the Spark task progress (such as the number of processed video slices and the number remaining), inference engine latency (whether the inference time of a single video slice exceeds a threshold), resource utilization (such as GPU / CPU utilization), etc. The execution status not meeting the preset conditions can mean that the Spark task progress does not meet the preset progress, the inference engine latency exceeds the threshold, the resource utilization exceeds the threshold, etc.
[0124] Among them, updating the data sample construction plan refers to expanding or shrinking the number of instances of the data engine and the inference engine and adjusting the concurrency, such as adjusting the number of deployed instances of the data engine and the concurrency of each instance, and adjusting the number of deployed instances of the inference engine and the concurrency of each instance.
[0125] Exemplarily, the server monitors the execution status of the data sample construction plan in real time, and determines whether the execution status of the data sample construction plan meets the preset conditions, such as determining whether the execution progress of the data sample construction plan meets the preset progress; if not, it confirms that the execution status of the data sample construction plan does not meet the preset conditions, and based on the plan update instruction, updates the data sample construction plan to obtain an updated data sample construction plan; finally, according to the updated data sample construction plan, starts the data engine and the inference engine to extract target data features from the unprocessed data in the target data set, obtains the target data features corresponding to the unprocessed data in the target data set, and then combines the target data features corresponding to the processed data in the target data set to obtain the data sample construction result corresponding to the data sample construction request.
[0126] For example, if the monitoring detects that the inference engine instance only processes 800 video slices per hour due to complex actions in the video, the link orchestration (or Executor) will dynamically expand the number of instances to 6 (or increase the concurrency to 25) to ensure that the total task is completed on time; at the same time, the data engine concurrency will be adjusted to avoid data backlogs or insufficient supply.
[0127] The technical solution provided by the embodiments of the present disclosure monitors the execution status of the data sample construction scheme in real time; when the execution status does not meet the preset conditions, the data sample construction scheme is updated to obtain an updated data sample construction scheme; according to the updated data sample construction scheme, target data features are extracted from the unprocessed data in the target data set; in this way, by monitoring the execution status in real time and dynamically updating the data sample construction scheme, adaptive optimization of the target feature extraction of the unprocessed data of the target data set is achieved, thereby ensuring the effectiveness and stability of the feature extraction process.
[0128] Figure 4 is a flow chart showing another data processing method according to an exemplary embodiment. Figure 4 As shown, the data processing method is used in the server and can be implemented by the following steps:
[0129] In step S410, the received data sample construction request is parsed to obtain a data set identifier and a data feature identifier; the data set identifier is used to indicate the target data set targeted by the data sample construction request; the data feature identifier is used to indicate the target data features to be extracted from the target data set.
[0130] In step S420, a data set metadata management component and a data feature metadata management component are identified from the metadata management system; metadata corresponding to the data set identifier is obtained from the data set metadata management component, and metadata corresponding to the data feature identifier is obtained from the data feature metadata management component.
[0131] In step S430 , metadata of the target dataset is obtained based on the metadata corresponding to the dataset identifier, and metadata of the target data feature is obtained based on the metadata corresponding to the data feature identifier.
[0132] In step S440, based on the metadata of the target data set, the metadata of the target data features and the current computing resource information, the data processing link and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link are determined.
[0133] In step S450, a data sample construction scheme corresponding to the data sample construction request is determined based on the data processing link and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link.
[0134] In step S460 , the data sample construction scheme is verified to obtain a verification result of the data sample construction scheme; the verification result is used to indicate whether the data sample construction scheme is correct and applicable.
[0135] In step S470, when the verification result indicates that the data sample construction scheme is correct and applicable, the data engine is started to preprocess the data in the target data set according to the data sample construction scheme to obtain preprocessed data.
[0136] In step S480, the inference engine is started to load the target operator corresponding to the target data feature in the operator library, and the target data feature of the preprocessed data is extracted through the target operator as the target data feature of the data in the target data set.
[0137] In step S490, the data engine performs association processing on the data in the target data set and the target data features of the data in the target data set to obtain a data sample construction result corresponding to the data sample construction request.
[0138] In the above data processing method, when constructing a data sample, a corresponding data sample construction plan is automatically generated through a data sample construction request containing a data set identifier and a data feature identifier, and through the data sample construction plan, the data engine and the inference engine are automatically started to extract target data features from the target data set. The entire process does not require manual collection of data sets and calculation of data features, thereby simplifying the data sample construction process, which is conducive to reducing data sample construction time and thus improving data sample construction efficiency; moreover, the purpose of automatically extracting target data features from the target data set based on the data sample construction request is achieved, achieving the effect of automatically constructing data samples and further improving data sample construction efficiency; at the same time, it avoids the defect of low data sample construction efficiency caused by manual calculation of data features on the data set.
[0139] In order to more clearly illustrate the data processing method provided by the embodiment of the present disclosure, the data processing method is specifically described below using a specific embodiment. Figure 5 As shown, the present disclosure also provides a native multimodal data sample construction system, through which a native multimodal data sample construction method can be implemented. The core part of this method is mainly through the interaction of OneMetaLake (metadata management system), and the four core subsystems of Planner, Executor, data engine, and inference engine, to emerge new system capabilities. Through this method, the problem of increased TCO (Total Cost of Ownership, that is, the cost of the entire life cycle, including resource costs, operating costs, maintenance costs, etc.) caused by the chimney architecture can be reduced. Data from different links can be cross-empowered to avoid repeated storage, repeated development, and repeated calculations; self-service and automation capabilities can be significantly improved according to data needs, data operating costs can be reduced, and the efficiency and quality of data delivery can be continuously improved. Reference Figure 5 , the system can be divided into five levels:
[0140] 1. User Operation Layer: (1) Demand Management: Mainly responsible for the management of data requirements, such as calculating a certain feature for a certain data set; moreover, it is expected that 80% of the requirements can be completed through automation without manual intervention, and 20% require manual intervention after the requirements are reviewed; for example, refer to Figure 7, the user only needs to determine the data set constructed by the sample (for example, data set id: 123), as well as the feature list and version (for example, X feature, Y version) that needs to be refreshed for this set, without having to worry about the metadata related to the refreshed features and service startup information, thus avoiding the inefficiency and correctness problems caused by a large number of manual operations. (2) Dataset management: Generate a data set for the features to be calculated, such as registering 10,000 lines of video slice data with a resolution of 720P from a certain data source. For example, refer to Figure 7 For registered datasets, users can self-check the features and versions of various types and modalities that have been exposed in the data system, and then complete data feature refresh, retrieval, and data retrieval, achieving what you see is what you get. For example, when registering a dataset with the dataset ID 123, the user selects the data volume as 19 million, the data type as video, the pixel as 720_P, the data medium information as Hive table name, and other data descriptions as data ownership. (3) Retrieval and export: Supports retrieval and analysis of data in existing data systems, as well as asset ownership query and other functions.
[0141] 2. Link Plan Orchestration Layer: (1) Responsible for converting user data requirements into execution plans. Specifically, it orchestrates links and determines the number of deployed instances and concurrency of the data engine and inference engine based on the size of the dataset, the features to be acquired, and the current computing resources. (2) The plan evaluation layer mainly uses manual rule checking and some intelligent methods to evaluate whether the plan generated by the planner is correct and usable.
[0142] 3. Data engine and scheduling layer: (1) The data engineering engine consists of three parts: 1) Data engine: responsible for Hive (data warehouse tool), Spark (distributed computing framework), IDP (Integrated Data Platform) task scheduling, algorithm service control, intelligent parameter adjustment, etc. 2) Inference engine: responsible for the execution of inference tasks of algorithm services such as label / feature, covering correctness verification, service optimization, and inference optimization. 3) Operator library: contains various plug-in label algorithm services with unified interfaces, manual policy rules, etc. (2) Executor: According to the execution plan of the Planner, it starts the corresponding data engine and inference engine services, observes the service status, reports service information, and also provides the ability for manual intervention, such as expanding or shrinking the number of services and adjusting the number of concurrent users.
[0143] 4. Data infrastructure layer: This layer includes some common structured and unstructured storage media used to store multimodal data. In terms of computing, it mainly involves cloud-native platforms for service construction and deployment, as well as machine learning platforms for model training, deployment, and inference.
[0144] 5. OneMetaLake: Unifies the organization, storage, and access interfaces of multimodal data, and completes data modeling and metadata management for the four core objects of data, features, requirements, and solutions.
[0145] refer to Figure 6 The internal structure of OneMetaLake's system is mainly composed of an interface layer and an object model layer: (1) Interface layer: provides a unified API (Application Programming Interface) or RPC (Remote Procedure Call) calling method, and authenticates the access personnel to the data set, avoiding the maintenance costs and possible security issues caused by the different access methods of each link in the application layer. (2) Object model layer: It is divided into four layers. The four core objects of the bottom layer, Feature, training sample raw data, requirement list, and solution list, mainly obtain or associate real data metadata through the underlying storage. Then, in the middle layer, the four objects have their own Schema description, and complete the format conversion according to the corresponding modeling Schema. Then, they are managed by their respective CateLog (directory) components in the next upper layer. Finally, OneMetaLake at the top layer is responsible for the management of each CateLog component.
[0146] Figure 6 The unified asset management and governance module shown: OneMetaLake's core function is mainly responsible for the addition, deletion, modification and query of four core objects, as well as asset permission management, while completing the construction and tracking of blood relationships, which is the foundation for the construction of multimodal data samples.
[0147] The technical solutions provided by the embodiments of the present disclosure can achieve the following technical effects: (1) Delivery efficiency: manual operation time is reduced by 80%; overall delivery efficiency is increased by more than 30%. (2) Resource cost: through the cross-empowerment of feature operators, after standard algorithm engineering optimization, the throughput efficiency is increased by more than 50% on average. (3) System operation and maintenance and governance costs: manpower consumption is reduced by 50%. (4) Search capability upgrade: multimodal search capabilities are provided to facilitate the value mining of multimodal data. (5) Build a data processing system adapted to multimodal scenarios, encapsulate AI capabilities into data links, make the processing of multimodal data (such as video, pictures, audio, etc.) as convenient as structured ETL (extraction, transformation, loading), and use AI to drive the ETL paradigm shift. (6) Abstract the four main objects of requirements, solutions, data, and features, build a unified metadata management system for management, and realize the systematic management of data assets. By building this binding relationship, it has the ability to track data lineage, so that it can track the source and version changes of data assets, ensure the reliability and interpretability of data, and the correspondence between deliverables and requirements becomes clear and visible.
[0148] It should be understood that, although the various steps in the flowchart of the present disclosure are shown in sequence as indicated by the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0149] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can be referred to each other, and each embodiment focuses on the differences from other embodiments. For related parts, please refer to the description of other method embodiments.
[0150] Figure 8 FIG. 1 is a block diagram of a data processing device according to an exemplary embodiment. Figure 8 The device includes a request parsing unit 810, a solution determining unit 820 and a sample constructing unit 830.
[0151] The request parsing unit 810 is configured to parse the received data sample construction request to obtain a data set identifier and a data feature identifier; the data set identifier is used to indicate the target data set targeted by the data sample construction request; the data feature identifier is used to indicate the target data features to be extracted from the target data set.
[0152] The solution determination unit 820 is configured to determine the data sample construction solution corresponding to the data sample construction request based on the target data set, target data characteristics and current computing resource information; the data sample construction solution includes at least the data processing link, the deployment information of the data engine and the deployment information of the inference engine.
[0153] The sample construction unit 830 is configured to execute according to the data sample construction plan, start the data engine and the inference engine to extract target data features from the target data set, and obtain a data sample construction result corresponding to the data sample construction request.
[0154] In an exemplary embodiment, the scheme determination unit 820 is further configured to execute the acquisition of metadata of the target data set and metadata of the target data features; determine the data processing link, and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link based on the metadata of the target data set, the metadata of the target data features and the current computing resource information; determine the data sample construction scheme corresponding to the data sample construction request based on the data processing link, and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link.
[0155] In an exemplary embodiment, the scheme determination unit 820 is further configured to identify a data set metadata management component and a data feature metadata management component from the metadata management system; obtain metadata corresponding to the data set identifier from the data set metadata management component, and obtain metadata corresponding to the data feature identifier from the data feature metadata management component; obtain metadata of the target data set based on the metadata corresponding to the data set identifier, and obtain metadata of the target data feature based on the metadata corresponding to the data feature identifier.
[0156] In an exemplary embodiment, the data processing apparatus further includes a solution verification unit configured to verify the data sample construction solution and obtain a verification result of the data sample construction solution; the verification result is used to indicate whether the data sample construction solution is correct and applicable;
[0157] The sample construction unit 830 is also configured to execute, when the verification result indicates that the data sample construction scheme is correct and available, start the data engine and the inference engine according to the data sample construction scheme to extract target data features from the target data set and obtain a data sample construction result corresponding to the data sample construction request.
[0158] In an exemplary embodiment, the scheme verification unit is further configured to perform initial verification processing on the data sample construction scheme through preset verification rules to obtain an initial verification result; the preset verification rules are at least used to verify the resource constraints and link logic of the data sample construction scheme; the initial verification result is used to indicate whether the scheme information of the data sample construction scheme meets the preset verification rules; when the initial verification result indicates that the scheme information of the data sample construction scheme does not meet the preset verification rules, the initial verification result is used as the verification result; when the initial verification result indicates that the scheme information of the data sample construction scheme meets the preset verification rules, the data sample construction scheme is re-verified through preset verification instructions to obtain a re-verification result as the verification result of the data sample construction scheme; the preset verification instructions are at least used to verify the execution efficiency and feature extraction accuracy of the data sample construction scheme.
[0159] In an exemplary embodiment, the sample construction unit 830 is further configured to execute, when the verification result indicates that the data sample construction scheme is correct and available, start the data engine to preprocess the data in the target data set according to the data sample construction scheme to obtain preprocessed data; start the inference engine to load the target operator corresponding to the target data feature in the operator library, and extract the target data feature of the preprocessed data through the target operator as the target data feature of the data in the target data set; and perform association processing on the data in the target data set and the target data feature of the data in the target data set through the data engine to obtain the data sample construction result.
[0160] In an exemplary embodiment, the sample construction unit 830 is further configured to perform real-time monitoring of the execution status of the data sample construction plan; when the execution status does not meet the preset conditions, the data sample construction plan is updated to obtain an updated data sample construction plan; and according to the updated data sample construction plan, target data features are extracted from the unprocessed data in the target data set.
[0161] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0162] Figure 9 1 is a block diagram of an electronic device 900 for implementing a data processing method according to an exemplary embodiment. For example, the electronic device 900 may be a server. Figure 9The electronic device 900 includes a processing component 920, which further includes one or more processors, and a memory resource represented by a memory 922 for storing instructions executable by the processing component 920, such as an application. The application stored in the memory 922 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 920 is configured to execute the instructions to perform the above method.
[0163] The electronic device 900 may further include a power supply component 924 configured to perform power management of the electronic device 900, a wired or wireless network interface 926 configured to connect the electronic device 900 to a network, and an input / output (I / O) interface 928. The electronic device 900 may operate based on an operating system stored in the memory 922, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, or the like.
[0164] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as memory 922 including instructions. The instructions may be executed by a processor of the electronic device 900 to perform the above method. The storage medium may be a computer-readable storage medium, such as a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0165] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions, and the instructions can be executed by a processor of the electronic device 900 to implement the above method.
[0166] It should be noted that the above-mentioned devices, electronic devices, computer-readable storage media, computer program products, etc. can also include other implementation methods according to the description of the method embodiments. The specific implementation methods can refer to the description of the relevant method embodiments and will not be described one by one here.
[0167] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.
[0168] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A data processing method, characterized in that: include: Parse the received data sample construction request to obtain the data set identifier and data feature identifier; The data set identifier is used to indicate the target data set targeted by the data sample construction request; the data feature identifier is used to indicate the target data feature to be extracted from the target data set; Determining a data sample construction scheme corresponding to the data sample construction request based on the target data set, the target data characteristics, and current computing resource information; The data sample construction plan includes at least a data processing link, deployment information of a data engine involved in the data processing link, and deployment information of an inference engine; According to the data sample construction plan, starting the data engine and the inference engine to extract the target data features from the target data set, and obtaining a data sample construction result corresponding to the data sample construction request; When it is monitored that the execution status of the data sample construction scheme does not meet the preset conditions, the target data features are extracted from the unprocessed data in the target data set according to the updated data sample construction scheme.
2. The method according to claim 1, characterized in that The determining, based on the target data set, the target data characteristics, and the current computing resource information, a data sample construction scheme corresponding to the data sample construction request includes: Acquire metadata of the target data set and metadata of the target data features; Determine a data processing link, as well as deployment information of a data engine and an inference engine involved in the data processing link, based on the metadata of the target data set, the metadata of the target data features, and current computing resource information; Based on the data processing link, and the deployment information of the data engine and the deployment information of the inference engine involved in the data processing link, a data sample construction scheme corresponding to the data sample construction request is determined.
3. The method according to claim 2, characterized in that The acquiring of metadata of the target data set and metadata of the target data features includes: From the metadata management system, identify the dataset metadata management component and the data feature metadata management component; Acquire metadata corresponding to the dataset identifier from the dataset metadata management component, and acquire metadata corresponding to the data feature identifier from the data feature metadata management component; Based on the metadata corresponding to the dataset identifier, metadata of the target dataset is obtained, and based on the metadata corresponding to the data feature identifier, metadata of the target data feature is obtained.
4. The method according to claim 1, wherein Before starting the data engine and the inference engine to extract the target data features from the target data set according to the data sample construction plan and obtaining a data sample construction result corresponding to the data sample construction request, the method further includes: Verifying the data sample construction scheme to obtain a verification result of the data sample construction scheme; the verification result is used to indicate whether the data sample construction scheme is correct and applicable; The step of starting the data engine and the inference engine to extract the target data features from the target data set according to the data sample construction scheme to obtain a data sample construction result corresponding to the data sample construction request includes: When the verification result indicates that the data sample construction scheme is correct and applicable, the data engine and the inference engine are started according to the data sample construction scheme to extract the target data features from the target data set to obtain a data sample construction result corresponding to the data sample construction request.
5. The method according to claim 4, characterized in that The verifying the data sample construction scheme to obtain a verification result of the data sample construction scheme includes: Performing an initial verification process on the data sample construction scheme using preset verification rules to obtain an initial verification result; the preset verification rules are used to at least verify the resource constraints and link logic of the data sample construction scheme; the initial verification result is used to indicate whether the scheme information of the data sample construction scheme satisfies the preset verification rules; In a case where the initial verification result indicates that the scheme information of the data sample construction scheme does not satisfy the preset verification rule, taking the initial verification result as the verification result; When the initial verification result indicates that the scheme information of the data sample construction scheme meets the preset verification rules, the data sample construction scheme is re-verified through the preset verification instructions to obtain a re-verification result as the verification result of the data sample construction scheme; the preset verification instructions are at least used to verify the execution efficiency and feature extraction accuracy of the data sample construction scheme.
6. The method according to claim 4, characterized in that When the verification result indicates that the data sample construction scheme is correct and applicable, starting the data engine and the inference engine to extract the target data features from the target data set according to the data sample construction scheme to obtain a data sample construction result corresponding to the data sample construction request, including: If the verification result indicates that the data sample construction scheme is correct and applicable, starting the data engine to preprocess the data in the target data set according to the data sample construction scheme to obtain preprocessed data; Starting the inference engine to load a target operator corresponding to the target data feature in an operator library, and extracting the target data feature of the preprocessed data through the target operator as the target data feature of the data in the target data set; The data engine performs association processing on the data in the target data set and the target data features of the data in the target data set to obtain the data sample construction result.
7. The method according to any one of claims 1 to 6, characterized in that After constructing a solution according to the data sample and starting the data engine and the inference engine to extract the target data features from the target data set, the method further includes: Real-time monitoring of the execution status of the data sample construction plan; The step of extracting the target data features from the unprocessed data in the target data set according to an updated data sample construction scheme when it is monitored that the execution status of the data sample construction scheme does not satisfy a preset condition comprises: When the execution state does not satisfy a preset condition, updating the data sample construction scheme to obtain an updated data sample construction scheme; According to the updated data sample construction scheme, the target data features are extracted from the unprocessed data in the target data set.
8. A data processing device, characterized in that: include: A request parsing unit is configured to parse the received data sample construction request to obtain a data set identifier and a data feature identifier; The data set identifier is used to indicate the target data set targeted by the data sample construction request; the data feature identifier is used to indicate the target data feature to be extracted from the target data set; A solution determination unit is configured to determine a data sample construction solution corresponding to the data sample construction request based on the target data set, the target data characteristics, and the current computing resource information; the data sample construction solution includes at least a data processing link, deployment information of a data engine, and deployment information of an inference engine; A sample construction unit is configured to execute the data sample construction plan, start the data engine and the inference engine to extract the target data features from the target data set, and obtain a data sample construction result corresponding to the data sample construction request; The sample construction unit is further configured to extract the target data features from the unprocessed data in the target data set according to an updated data sample construction scheme of the data sample construction scheme when it is monitored that the execution status of the data sample construction scheme does not meet the preset conditions.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the data processing method according to any one of claims 1 to 7.
11. A computer program product comprising instructions, characterized in that: When the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute the data processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-model concurrent execution data interaction method and system
CN117555696A
Label data processing method and device, computer equipment and storage medium
CN119377720A