Intelligent data extraction method applied to data sharing platform
Through multiple rounds of cleaning strategies and deep learning network processing structured and semi-structured data, the problems of low data extraction efficiency and insufficient outlier processing in the prior art are solved, and high-quality data extraction and model training support are achieved.
Patent Information
- Application Number
- CN202510137591.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-07-04
AI Technical Summary
Existing data extraction methods are inefficient when processing structured and semi-structured data, lack universality and flexibility, and it is difficult to fully identify and process outliers, resulting in insufficient data cleaning and affecting the accuracy and efficiency of model training.
Multi-round cleaning strategies are adopted, including the combination of Canopy algorithm and K-means algorithm, to identify and process outliers, extract semi-structured data through a deep learning converter network, fill in missing data with timing correlation, and use ETL tools for data extraction and buffer storage.
It improves the accuracy and completeness of data, enhances the accuracy and efficiency of data extraction, and provides high-quality data support for subsequent model training.
Smart Images

Figure CN120256501A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a data intelligent extraction method applied to a data sharing platform. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, the application of data in various fields has become increasingly important. Especially when building large models (such as deep learning models, natural language processing models, etc.), high-quality data is the key to model training and optimization. Although existing data extraction methods meet the data requirements to a certain extent, there are still some deficiencies.
[0003] Existing data extraction methods mainly focus on the processing of structured data, such as table data in relational databases. These methods usually require manually writing complex SQL query statements, which are not only time-consuming and laborious but also error-prone. For semi-structured data (such as XML, JSON, etc.), existing extraction and parsing methods are more complex and often rely on specific parsing tools, lacking generality and flexibility. In addition, existing data cleaning methods mainly rely on statistical methods (such as mean, standard deviation, etc.) to identify and process outliers, but have limited ability to handle complex data features and temporal correlation degrees, resulting in insufficient data cleaning and affecting the accuracy and efficiency of subsequent model training.
[0004] In terms of data collection, existing methods usually require manually configuring connection parameters for different data sources, which not only increases the complexity of operations but also may lead to configuration errors. In addition, there is a lack of an effective buffer storage mechanism during the data extraction process, resulting in possible loss or delay problems during data transmission. In terms of data cleaning, existing methods mainly rely on a single statistical method and cannot comprehensively identify and process various types of outliers, especially when dealing with large-scale data, the cleaning effect is not good. In addition, for the handling of missing data, existing methods often adopt simple filling methods, such as mean filling, which may introduce biases and affect the accuracy and integrity of the data. Summary of the Invention
[0005] In view of the above-mentioned prior art, the present invention aims to provide a data intelligent extraction method applied to a data sharing platform, mainly solving the technical problems existing in the above background art.
[0006] To achieve the above object, the technical solution of the embodiment of the present invention is implemented as follows:
[0007] A data intelligent extraction method applied to a data sharing platform, the data intelligent extraction method includes the following steps:
[0008] S1. Establish connection channels with different data sources, and collect structured data and semi-structured data in the data sources;
[0009] S2. Clean the structured data and semi-structured data;
[0010] S3. Extract the structured data and the semi-structured data after data cleaning;
[0011] S4. Integrate and fill the data extraction results into a data model, and finally form model training metadata that can be used by large model components.
[0012] Optionally, step S1 specifically includes: adaptively matching corresponding ETL tools for different data sources, setting parameters for the ETL tools, extracting various types of data in different data sources through database interfaces, log file interfaces or stream data interfaces, and buffering and storing the extracted data based on HDFS or HBase.
[0013] Optionally, cleaning the structured data and semi-structured data specifically includes:
[0014] S201. Store the structured data in dataset D1, store the semi-structured data in dataset D2, and set distance T1 and distance T2, and distance T1 is greater than distance T2;
[0015] S202. Randomly select any corresponding data points in dataset D1 and dataset D2 as the data center points of the canopy, calculate the distances between the corresponding data points and other data points in dataset D1 / dataset D2. If the distance between the corresponding data points and other data points in dataset D1 / dataset D2 is less than T1, then put the data points with distances less than T1 into this canopy, regard the data points with distances less than T1 and less than T2 as abnormal data and delete them, and then put the data points with distances greater than T1 into a new canopy;
[0016] S203. Repeat the above process to finally achieve the first data cleaning of dataset D1 and dataset D2, and divide into K1 center points in dataset D1 and K2 center points in dataset D2.
[0017] Optionally, cleaning the structured data and semi-structured data further includes: respectively using the K-means algorithm based on the K1 center points in dataset D1 and the K2 center points in dataset D2 to perform the second data cleaning, and eliminating outliers or noise points in dataset D1 and dataset D2 again.
[0018] Optionally, for data cleaning of the structured data and semi-structured data, it further includes: for the abnormal data cleaned out from dataset D1 and dataset D2, identify the time series of the abnormal data, calculate the temporal correlation degree of the abnormal data, calculate the length of the time series with data missing based on the temporal correlation degree, and based on the calculation result, use the minimization of the missing function to fill in the missing part of the abnormal data.
[0019] Optionally, for the data after filling, perform length judgment, set the length outlier index. If there is a situation where the text length is greater than its length outlier, it indicates that the filled data is abnormal, and re-fill the missing data. The expression of the length outlier index is:
[0020]
[0021] In the formula, j1 represents the processing feature of the data to be processed, j2 represents the cleaning feature of the text length abnormal data, φ represents the length definition term coefficient, Υ represents the behavior vector of the data cleaning instruction, and Lδ represents the original parameter of the data to be processed.
[0022] Optionally, for the structured data and semi-structured data after data cleaning, perform data extraction. Specifically, use an ETL tool to extract data from the structured data, and use a deep learning-based transformer network to extract semi-structured data. The deep learning-based transformer network consists of an input layer, an embedding layer, an encoding layer, a bidirectional long short-term memory network layer, a decoding layer, and an output layer. Through the deep learning-based transformer network, accurately identify and extract the entities with specific meanings in the semi-structured data.
[0023] Optionally, use a third-party large model to establish a data model, which specifically includes: establishing an intent recognizer, the intent recognizer is connected to the third-party large model, obtain the dialogue instruction input by the user through the intent recognizer, the third-party large model forms a data standard based on the dialogue instruction, and generate a data model based on the data standard.
[0024] The beneficial effects of the present invention are as follows: A data intelligent extraction method applied to a data sharing platform disclosed herein adopts a multi-round cleaning strategy. During the initial cleaning of data, abnormal data points are deleted or redistributed to a new Canopy. Then, based on the divided center points, the K-means algorithm is used for the second data cleaning to further eliminate outliers or noise points in the data set. For the abnormal data cleaned out, the time series correlation degree is calculated to identify the length of the time series with missing data, and the missing part is filled using the minimum missing function. This filling method based on time series correlation can more accurately fill the missing data, avoiding the deviation caused by simple filling in traditional methods, and further improving the data quality. In summary, this multi-round cleaning strategy can more comprehensively identify and process various types of outliers, improving the accuracy and integrity of the data;
[0025] During the data extraction process, a deep learning-based transformer network is used for semi-structured data extraction. This transformer network consists of an input layer, an embedding layer, an encoding layer, a bidirectional long short-term memory network layer, a decoding layer, and an output layer, and can accurately identify and extract entities with specific meanings in semi-structured data. The application of this deep learning technology significantly improves the accuracy and efficiency of data extraction, providing high-quality data support for subsequent data analysis and model training. Brief Description of the Drawings
[0026] Figure 1 It is a schematic flowchart of the data intelligent extraction method applied to a large model in an embodiment of the present application. Detailed Embodiments
[0027] The technical solution of the present invention will be further elaborated in detail below in conjunction with the accompanying drawings of the specification and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In the following description, the expression "some embodiments" is described, which describes a subset of all possible embodiments. However, it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0028] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, some technical features well known to the public are not described to avoid confusion with the present invention.
[0029] It should be understood that the present invention can be implemented in different forms and should not be construed as limited to the embodiments presented herein. On the contrary, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. And the purpose of the terms used herein is only to describe specific embodiments and not to limit the present invention. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprising" and / or "including", when used in this specification, determine the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups. As used herein, the term "and / or" includes any and all combinations of the related listed items.
[0030] It should be further noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "inner", "outer", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only implementation.
[0031] To thoroughly understand the present invention, detailed structures will be presented in the following description to illustrate the technical solutions proposed by the present invention. The alternative embodiments of the present invention are described in detail below. However, in addition to these detailed descriptions, the present invention can also have other implementations.
[0032] Embodiment
[0033] Please refer to the attached Figure 1 , this application provides a data intelligent extraction method applied to a data sharing platform. The data intelligent extraction method includes the following steps:
[0034] S1. By establishing connection channels with different data sources, collect structured data and semi-structured data in the data sources; among them, structured data usually comes from data in relational databases (such as MySQL, Oracle), and semi-structured data is usually in JSON, XML format.
[0035] When collecting structured data and semi-structured data in the data source, adaptively match corresponding ETL tools for different data sources, set parameters for the ETL tools, extract various types of data in different data sources through database interfaces, log file interfaces or stream data interfaces, and buffer and store the extracted data based on HDFS or HBase.
[0036] The ETL tool is a data extraction, transformation, and loading tool that extracts multi-structured data from data sources and quickly and efficiently loads the original data into a big data container, enabling data conversion between spatial big data storage and traditional storage methods. According to different data types, it is divided into three tools, namely: a real-time data conversion tool that imports real-time data through web crawlers and Flume; a custom data conversion tool that uses the Sqoop big data access tool to improve storage efficiency, can customize the conversion tool according to specific business data types, and provides a file upload function; a spatial data conversion tool that converts spatial data in a spatial data format into a common format. In specific implementation, one of them can be selected for implementation, or other unlisted ETL tools can be selected for data extraction from the data source.
[0037] S2. Clean the structured data and semi-structured data;
[0038] S3. Extract the structured data and the semi-structured data after data cleaning;
[0039] S4. Integrate and fill the data extraction results into a data model, and finally form model training metadata that can be used by big model components.
[0040] Some components of its data model have the following format:
[0041]
[0042]
[0043] In a possible implementation, cleaning the structured data and semi-structured data specifically includes:
[0044] S201. Store the structured data in dataset D1, store the semi-structured data in dataset D2, and set distance T1 and distance T2, and distance T1 is greater than distance T2;
[0045] S202. Randomly select any corresponding data point in dataset D1 and dataset D2 as the data center point of the canopy, calculate the distance between the corresponding data point and other data points in dataset D1 / dataset D2. If the distance between the corresponding data point and other data points in dataset D1 / dataset D2 is less than T1, then put the data points with a distance less than T1 into the canopy, regard the data points with a distance less than T1 and less than T2 as abnormal data and delete them, and put the data points with a distance greater than T1 into a new canopy;
[0046] S203. Repeat the above process to finally achieve the first data cleaning of dataset D1 and dataset D2, and divide them into K1 center points in dataset D1 and K2 center points in dataset D2.
[0047] In this embodiment, the structured data and semi-structured data are respectively subjected to data cleaning. That is, first, the structured data is stored in dataset D1, and the semi-structured data is stored in dataset D2. Then, a feature selection method based on Logsf is adopted for dataset D1 and dataset D2. The feature selection method based on Logsf is a technology for big data dimensionality reduction and feature extraction. By processing the training sample set, this method extracts the main features of the data, thereby achieving data dimensionality reduction. The specific steps of this method are as follows:
[0048] Define the training sample set: Set the training sample set R = {M, N} = {m i , n i}, where m i is the i-th training sample, n i is the label of the i-th sample, and each sample m i contains a d-dimensional vector;
[0049] Define the loss function: The loss function of sample m i is defined as:
[0050] L(β, m i ) = log(1 + exp(-β T * F i ))
[0051]
[0052] where m i ’ is the data point closest to m i and with a different label, n i ’ is the data point closest to m i and with the same label, and β is the feature weight.
[0053] Define the evaluation function:
[0054]
[0055] where k represents the number of sample features, and e(β) is the sum of the loss functions of all feature samples.
[0056] Adopt the gradient descent method to obtain the optimal weight, and based on the feature sample similarity, the optimal feature weight β can be obtained, thereby achieving feature selection and dimensionality reduction of big data. This method has high accuracy and data processing speed when dealing with high-dimensional data.
[0057] After extracting the main features of the data, randomly select any corresponding data points in dataset D1 and dataset D2 respectively as the data center points of the canopy. Calculate the distances between the corresponding data points and other data points in dataset D1 / dataset D2. Based on the rule that the greater the distance difference, the stronger the abnormality, perform the division of abnormal data. For example, if the distance between the corresponding data point and other data points in dataset D1 / dataset D2 is less than T1, then put all the data points with a distance less than T1 into this canopy, consider the data points with a distance less than T1 and less than T2 as abnormal data and delete them, and put the data points with a distance greater than T1 into a new canopy, finally realizing the first cleaning of abnormal data.
[0058] Furthermore, after performing the first data cleaning on the structured data and semi-structured data, continue to perform the second data cleaning on the structured data and semi-structured data, which includes: based on K1 center points in dataset D1 and K2 center points in dataset D2, respectively use the K-means algorithm to perform the second data cleaning, and eliminate the outliers or noise points in dataset D1 and dataset D2 again.
[0059] Specifically, after completing the initial data cleaning and determining the center points in the dataset, the K-means algorithm can be used for the second data cleaning to further eliminate the outliers or noise points in the dataset. The following are the specific steps:
[0060] Use the center points determined in the initial data cleaning process as the initial center points of the K-means algorithm. These center points have been verified as reasonable clustering centers in the initial cleaning;
[0061] For each data point in the dataset, calculate the distance between it and each center point. Common distance measurement methods include Euclidean distance, Manhattan distance, etc., and assign each data point to the cluster to which the center point with the closest distance belongs. Specifically, for each data point x i , find the center point μ i with the closest distance, and assign x i to the cluster C i ;
[0062] For each cluster C i , calculate the average value of all data points in this cluster, and use it as the new center point μ j , and its specific calculation formula is:
[0063]
[0064] Repeat the process of redistributing data points and updating the center points until the clustering result converges, that is, the center points no longer change or change very little. The convergence condition can be that the change in the center points is less than a certain threshold, or the maximum number of iterations is reached;
[0065] For each data point, calculate the distance between it and the center point of the cluster it belongs to. Data points with a distance exceeding the threshold T2 are regarded as outliers, and the outlier data is removed from the data set D1 or the data set D2. The remaining data in the data set D1 or the data set D2 is the normal data.
[0066] In a possible implementation, after the second data cleaning, the abnormal data screened out includes missing cases. Therefore, for the abnormal data cleaned out from the data set D1 and the data set D2, identify the time series of the abnormal data, calculate the temporal correlation degree of the abnormal data, calculate the length of the time series with data missing based on the temporal correlation degree, and based on the calculation result, use the minimization of the missing function to fill in the missing part of the abnormal data.
[0067] Specifically, calculate the temporal correlation degree of the outlier data. If the temporal correlation degree of the outlier is significantly lower than that of the normal data points, it may be missing data. The temporal correlation degree is calculated by the following formula:
[0068]
[0069] H = {h1, h1, h1,..., h j}
[0070]
[0071] where, represents the mean of the time series of data features, H represents the time series of data features, L represents the mean of the time series, h represents the ordered efficacy function, V j represents the effectiveness degree of data j, V max represents the maximum ordered degree, V min represents the minimum ordered degree, represents the temporal correlation degree;
[0072] When the temporal correlation degree of the outlier is significantly lower than that of the normal data points, judge that the data is missing data. For the missing data, first calculate the length of the time series with data missing: where θ’ represents the data density, and then fill in the missing data by minimizing the missing function. Its expression is:
[0073] L = ||s - Y|| 2
[0074] Among them, L represents the missing content filled for abnormal data, and Y represents the time series length of normal data with the same temporal correlation degree as the abnormal data.
[0075] Further, after filling in the missing data, the length of the filled data is judged, and a length outlier index is set. If there is a situation where its text length is greater than its length outlier value, it indicates that the filled data is abnormal, and the missing data filling is performed again. The expression of the length outlier index is:
[0076]
[0077] In the formula, j1 represents the processing feature of the data to be processed, j2 represents the cleaning feature of the text length abnormal data, φ represents the length definition term coefficient, Υ represents the behavior vector of the data cleaning instruction, and Lδ represents the original parameter of the data to be processed.
[0078] In a possible implementation manner, data extraction is performed on the structured data and the semi-structured data after data cleaning. Specifically, an ETL tool is used to extract data from the structured data. The ETL (Extract, Transform, Load) tool is a commonly used data extraction tool that can efficiently extract data from various data sources. In addition, for semi-structured data, a deep learning-based transformer network is used for data extraction. The deep learning-based transformer network consists of an input layer, an embedding layer, an encoding layer, a bidirectional long short-term memory network layer, a decoding layer, and an output layer. Specific entities with specific meanings in the semi-structured data are accurately identified and extracted through the deep learning-based transformer network.
[0079] In the embedding layer, in addition to word embeddings, positional encoding is added to capture word order information. The positional encoding can be a sine / cosine function or a learnable positional vector. The encoding layer consists of multiple self-attention mechanisms and feed-forward neural networks (FFNNs). The self-attention mechanism is used to capture the dependencies between words. The calculation process of the self-attention mechanism is divided into two steps. To mitigate the impact of random initial values on self-attention, the dot product of the weight matrix and the word vector is used as the sub-vector. Based on the input vectors of each encoder, sub-vector queries Q, keys K, and values V are established, and the softmax function is introduced to reduce the difficulty of weight calculation and solve the attention scores. The FFNN is used to perform non-linear transformations on the features of each word. The number of encoding layers can be adjusted according to the task complexity. After the encoding layer, by adding a bidirectional long short-term memory network layer (BiLSTM), context information is further captured. The BiLSTM can capture the dependencies of the context and enhance the semantic understanding ability of the model. The decoding layer is responsible for converting the encoded features into the final output. The decoding layer can also include self-attention mechanisms and feed-forward neural networks to generate the final entity labels or relationship labels.
[0080] In a possible implementation, a third-party large model is used to establish a data model, which specifically includes: establishing an intent recognizer connected to the third-party large model. The intent recognizer obtains the dialogue instructions input by the user, and the third-party large model forms a data standard based on the dialogue instructions and generates a data model based on the data standard.
[0081] Specifically, the user inputs dialogue instructions. The intent recognizer clarifies the user's intent based on the corresponding dialogue instructions and determines the data standard that the user wants. The third-party large model outputs the data standard, where the data standard refers to the normative constraints that ensure the consistency and accuracy of the internal and external use and exchange of data. Exemplarily, the data standard can include multiple aspects such as data layering, data domain partitioning, field library, word segmentation library, time dimension, space dimension, other dimensions, data model, asset theme, professional category, data type, etc. The user can perform a series of management operations on the data standard through the data sharing platform, including but not limited to adding, modifying, deleting, adding a new version, deleting a version, copying a version, importing / exporting the data standard, enabling / disabling a version, etc., and finally form a corresponding data model based on the data standard. The data model usually includes multiple tables or documents, and each table or document contains multiple fields. The designed data model describes the static characteristics, dynamic behaviors, and constraint conditions of the system at an abstract level, providing an abstract framework for the information representation and operation of the database system. The framework of the formed data model is as follows:
[0082]
[0083]
[0084] Finally, the extracted content is filled into the above data model. During the filling process, the extracted data fields are mapped to the fields in the data model. Ensure that each extracted data field can correspond to a field in the data model, and then insert the cleaned and standardized data into the data model. According to the design of the data model, batch insertion or individual insertion can be selected. Finally, after filling the data model, model training metadata that can be used by the large model component is formed, and the formed model training metadata is used in the training of the large model component, where the large model component is one of the components of the data sharing platform, and the data sharing platform is used to achieve unified data collection, unified storage, unified processing, unified sharing, and unified governance, acting as a data processing factory internally and providing a unified data consumption entry externally.
[0085] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A data intelligent extraction method applied to a data sharing platform, characterized in that, The described data intelligent extraction method includes the following steps: S1. Establish connection channels with different data sources to collect structured data and semi-structured data in the data sources; S2. Clean the structured data and semi-structured data; S3. Extract the structured data and the semi-structured data after data cleaning; S4. Integrate and fill the data extraction results into a data model, and finally form model training metadata that can be used by large model components.
2. The data intelligent extraction method applied to a data sharing platform according to claim 1, wherein Step S1 specifically includes: adaptively matching corresponding ETL tools for different data sources, setting parameters for the ETL tools, extracting various types of data in different data sources through database interfaces, log file interfaces or stream data interfaces, and buffering and storing the extracted data based on HDFS or HBase.
3. A data intelligent extraction method applied to a data sharing platform according to claim 1, characterized in that, Cleaning the structured data and semi-structured data specifically includes: S201. Store the structured data in dataset D1, store the semi-structured data in dataset D2, and set distance T1 and distance T2, and distance T1 is greater than distance T2; S202. Randomly select any corresponding data points in dataset D1 and dataset D2 as the data center points of the canopy, calculate the distances between the corresponding data points and other data points in dataset D1 / dataset D2. If the distance between the corresponding data point and other data points in dataset D1 / dataset D2 is less than T1, then put the data points with distances less than T1 into the canopy, regard the data points with distances less than T1 and less than T2 as abnormal data and delete them, and put the data points with distances greater than T1 into a new canopy; S203. Repeat the above process to finally achieve the first data cleaning of dataset D1 and dataset D2, and divide into K1 center points in dataset D1 and divide into K2 center points in dataset D2.
4. A data intelligent extraction method applied to a data sharing platform according to claim 3, characterized in that, Cleaning the structured data and semi-structured data also includes: Based on the K1 center points in dataset D1 and the K2 center points in dataset D2, respectively use the K-means algorithm to perform the second data cleaning to eliminate outliers or noise points in dataset D1 and dataset D2 again.
5. A data intelligent extraction method applied to a data sharing platform according to claim 4, characterized in that, Cleaning the structured data and semi-structured data also includes: For the abnormal data cleaned out in dataset D1 and dataset D2, identify the time series of the abnormal data, calculate the temporal correlation degree of the abnormal data, calculate the length of the time series with missing data based on the temporal correlation degree, and based on the calculation result, use the minimization of the missing function to fill in the missing part of the abnormal data.
6. The data intelligent extraction method applied to a data sharing platform according to claim 5, characterized in that, Judge the length of the filled data, set the length outlier index. If there is a situation where its text length is greater than its length outlier, it means that the filled data is abnormal and the missing data filling is redone. The expression of the length outlier index is: Where j1 represents the processing feature of the data to be processed, j2 represents the cleaning feature of the text length abnormal data, φ represents the length definition term coefficient, Υ represents the behavior vector of the data cleaning instruction, and L δ represents the original parameter of the data to be processed.
7. A data intelligent extraction method applied to a data sharing platform according to claim 6, characterized in that, Perform data extraction on the structured data and the semi-structured data after data cleaning. Specifically, use an ETL tool to perform data extraction on the structured data, and use a deep learning-based transformer network for semi-structured data extraction. The deep learning-based transformer network consists of an input layer, an embedding layer, an encoding layer, a bidirectional long short-term memory network layer, a decoding layer, and an output layer. Accurately identify and extract entities with specific meanings in the semi-structured data through the deep learning-based transformer network.
8. A data intelligent extraction method applied to a data sharing platform according to claim 7, characterized in that, Build a data model using a third-party large model, which specifically includes: building an intent recognizer, connecting the intent recognizer to the third-party large model, obtaining the dialogue instructions input by the user through the intent recognizer, and the third-party large model forming data standards based on the dialogue instructions and generating a data model based on the data standards.