Method and system for constructing high-quality data set for large model in urban rail field
Through data collection, cleaning, labeling and maintenance methods for the urban rail field, the problem of difficult to ensure the quality of urban rail data is solved, the construction of high-quality data sets is realized, the systematization and security of data processing are improved, and the learning ability of the model is enhanced.
Patent Information
- Application Number
- CN202510425437.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-25
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology lacks preprocessing technology specifically targeting the characteristics of urban rail data, which makes it difficult to ensure the quality of the data set. General data annotation tools and platforms are difficult to meet the specific scenarios and feature requirements in the urban rail field. Data security and privacy protection are insufficient, and reliable high-quality urban rail large-scale data sets cannot be formed.
It provides a method for building high-quality data sets for large models in the urban rail field, including data collection, cleaning, labeling and maintenance, desensitization of data, missing value processing, outlier value monitoring, etc. through data preprocessing, uses K nearest neighbor algorithm and support vector machine to judge data quality, uses tools such as LabelBox, BRAT, WebAnno and other tools to add labels, and data security services throughout the entire process.
It improves the systemization and standardization of the data processing process, improves data quality and model training effect, ensures that data is accurately marked, enhances the model's learning ability and generalization ability, and ensures data security.
Smart Images

Figure CN120277328A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of urban rail transit data processing, and more specifically, relates to a method and system for constructing a high-quality data set for a large model in the urban rail transit field. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the excellent performance of large models, all industries are actively exploring how to use artificial intelligence to promote the intelligent transformation of business. Focusing on the field of urban rail transit, large models also have great potential, which can not only improve operation and management efficiency and optimize the passenger experience, but also further empower the digital transformation of enterprises. However, to fully utilize the advantages of large models in the field of urban rail transit, constructing a high-quality data set is a key step.
[0003] Currently, one of the main challenges in constructing a large model in the urban rail transit field is how to effectively collect, clean, annotate, and manage large-scale urban rail transit data, which usually has the characteristics of wide sources, diverse formats, and different storage methods. The data information includes equipment monitoring information, train operation status, passenger flow, and other information.
[0004] Currently, there is a lack of preprocessing technology specifically for the characteristics of urban rail transit data on the market, resulting in difficulty in guaranteeing the quality of the data set. General data annotation tools and platforms are difficult to meet the requirements of specific scenarios and features in the urban rail transit field. The existing data set maintenance methods are not perfect, and the lack of data security and privacy protection cannot form a reliable high-quality urban rail transit large model data set. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for constructing a high-quality data set for a large model in the urban rail transit field, aiming to solve the technical problem in the existing technology that there is a lack of preprocessing technology specifically for the characteristics of urban rail transit data on the market, resulting in difficulty in guaranteeing the quality of the data set.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is: to provide a method for constructing a high-quality data set for a large model in the urban rail transit field, including the following steps:
[0007] Collect and summarize the data that meets the business requirements as the input data source;
[0008] Clean the data through data preprocessing to improve the data quality;
[0009] Add labels to the preprocessed data so that the key features of the data can be better understood and learned;
[0010] Conduct quality review and dynamic management on the annotated data through data maintenance.
[0011] Preferably, the data cleaning is performed on the data through data preprocessing to improve the data quality, including: one or more of data desensitization processing, missing value processing, low-quality filtering processing, redundancy removal processing, and outlier monitoring processing.
[0012] Preferably, the data desensitization processing includes:
[0013] Protecting the sensitive information of the data to be processed by using noise addition or hash desensitization strategy; wherein, the noise includes one or more of Gaussian noise and Laplace noise.
[0014] Preferably, the outlier monitoring processing includes performing data anomaly detection using the K-nearest neighbor algorithm, and the performing data anomaly detection using the K-nearest neighbor algorithm includes the following steps:
[0015] Selecting the Euclidean distance as the metric:
[0016]
[0017] where x i and x j represent two data sample points, k represents the k-th eigenvalue, n represents the number of data point features, x ik represents the value of the i-th sample on the k-th feature, x jk represents the value of the j-th sample on the k-th feature, d(x i , x j ) represents the distance between two sample points x i and x j ;
[0018] For each data point x i , find the nearest K neighbors, where x j represents a neighbor of x i , representing one of the K nearest neighbors of x i , x k represents any data point, used to compare the distance from x i to the K nearest neighbors to determine the set of K nearest neighbors of x i , and Neighbors(x i ) represents the set of K nearest neighbors of the data point x i :
[0019] Neighbors(x i ) = {x j |d(x i , x j ) < d(x i , x k ), k > j}
[0020] Use the average distance of the K nearest neighbors to measure the local density. Density(x i ) represents the local density of x i , that is, calculate the average distance of the K nearest neighbors of x i , reflecting the data distribution in the area around the data point:
[0021]
[0022] Calculate the outlier score index OutlierScore(x i ) for each data point to measure the degree to which the data point deviates from the normal pattern. The larger the value, the more likely the data point is an outlier:
[0023]
[0024] Set a threshold according to the outlier score, and mark the data points with scores higher than the threshold as abnormal data.
[0025] Preferably, the data is cleaned through data preprocessing to improve the data quality, including judging the content quality. The judging of the content quality includes the following steps:
[0026] Find the optimal hyperplane in the data distribution feature space to classify the sample quality.
[0027] The hyperplane formula is: ω T x + b = 0, where ω is the normal vector, x is the point on the hyperplane, and b is the constant term.
[0028] The distance from a point to the plane is The classification interval is Minimize the objective function while satisfying the constraint conditions where x i is the feature vector of the i-th data, and y i is the label of the i-th data. The constraint conditions ensure that the data quality is correctly judged.
[0029] Preferably, adding labels to the preprocessed data so that the key features of the data can be better understood and learned, including using tools to add labels to the preprocessed data. The tools include one or more of LabelBox, BRAT, WebAnno, Audacity.
[0030] Preferably, adding labels to the preprocessed data so that the key features of the data can be better understood and learned, including using one or more of NLP text data annotation methods such as machine translation, named entity recognition, question-answer matching, association extraction, taxonomy, text classification, file summarization.
[0031] Preferably, tags are added to the pre - processed data so that the key features of the data can be better understood and learned, and the annotation objects are one or more of text, speech, and images.
[0032] Preferably, it further includes a data security service that runs through all the processes of dataset construction.
[0033] The present invention also provides a system for constructing a high - quality dataset for large models in the urban rail transit field, which is used to execute the steps included in the method for constructing a high - quality dataset for large models in the urban rail transit field as described in any one of the above, and includes:
[0034] A data collection module, which is used to collect and summarize data that meets business requirements as an input data source;
[0035] A pre - processing module, which is used to clean the data through data pre - processing to improve the quality of the data;
[0036] An annotation module, which is used to add tags to the pre - processed data so that the key features of the data can be better understood and learned;
[0037] A maintenance module, which is used to conduct quality review and dynamic management of the annotated data through data maintenance.
[0038] The beneficial effects of the method and system for constructing a high - quality dataset for large models in the urban rail transit field provided by the present invention are as follows: Compared with the prior art, the method and system for constructing a high - quality dataset for large models in the urban rail transit field of the present invention cover multiple aspects such as data collection, pre - processing, annotation, maintenance, and security. It improves the systematization and standardization of the entire data - processing process, which is beneficial to improving the data quality and the effect of model training. The two core service modules of data pre - processing and data annotation clarify the cleaning strategy and annotation process of data in the urban rail transit field. Ensure that the data is accurately annotated, enhancing the learning ability and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] Figure 1 It is an architecture diagram of a method for constructing a high - quality dataset for large models in the urban rail transit field provided by an embodiment of the present invention;
[0041] Figure 2It is a flowchart of the data annotation process link in a method for constructing a high-quality dataset for large models in the urban rail transit field provided by an embodiment of the present invention;
[0042] Figure 3 It is a block diagram of the structure of a system for constructing a high-quality dataset for large models in the urban rail transit field provided by an embodiment of the present invention;
[0043] Figure 4 It is a block diagram of the structure of another system for constructing a high-quality dataset for large models in the urban rail transit field provided by an embodiment of the present invention.
[0044] In the figure: 1. Data acquisition module; 2. Preprocessing module; 3. Annotation module; 4. Maintenance module; 5. Security service module. Detailed implementation manners
[0045] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0046] Please refer to Figures 1 to 4 together, and now a method for constructing a high-quality dataset for large models in the urban rail transit field provided by the present invention will be described. The method for constructing a high-quality dataset for large models in the urban rail transit field includes the following steps:
[0047] S1. Collect and summarize the data that meets the business requirements as the input data source;
[0048] In this step, the data is specifically data in the urban planning field. And different acquisition methods need to be adopted for data from different sources in the urban rail transit field data. Through the data acquisition service, the data that meets the business requirements is collected and summarized as the input data source of the data preprocessing service.
[0049] The sources of the data include big data platform data, data of each professional business system, industry public data, and data transmitted by sensors in real time. For urban rail transit data from different sources, different acquisition methods are used for data acquisition. Among them, the acquisition methods include one or more of API interface access, SDK method, data download, and data subscription service. Preferably, the acquisition methods are API interface access, SDK method, data download, and data subscription service.
[0050] The API interface supports multiple API standards such as RESTful API and SOAP API, can communicate using protocols such as HTTP and HTTPS, and obtain data in formats such as JSON and XML. It provides an SDK method to encapsulate the API call logic, simplify the data acquisition process, interact with the data acquisition platform more conveniently, and achieve efficient data extraction. With the permission of the authority, users can directly download and export, which is supported in file formats such as TXT, CSV, JSON, Excel, XML, etc. The export function is set in various forms such as periodic (e.g., daily, weekly), manually triggered, or regularly generated, and it supports regularly synchronizing data to the user's local storage or cloud storage service through a dedicated tool or service; the data subscription method supports users to set subscription rules, and according to the preset update frequency or event trigger mechanism, push data in real time through message queue technologies such as WebSocket, MQTT, and Kafka, and supports configuring data subscription topics and filtering conditions.
[0051] S2. Clean the data through data preprocessing to improve the data quality;
[0052] Data preprocessing strictly screens, corrects, and improves the original data provided by the data acquisition service. It is a key link to form high-quality data available for the urban rail large model.
[0053] The implementation process of this step includes:
[0054] S2.1. Perform data desensitization on the data to protect sensitive data information. Among them, the methods of cleaning include one or more of data desensitization, missing value processing, low-quality filtering processing, redundancy removal processing, and outlier monitoring processing.
[0055] Data desensitization is to transform sensitive data through desensitization rules without reducing data security, effectively reducing the exposure of sensitive data in the acquisition, transmission, use and other links.
[0056] According to the differences in the data sensitivity, usage scenario requirements, and security levels of different scenarios in urban rail, adopt noise addition or / and hash desensitization strategies to protect sensitive data information for the sensitive data to be processed.
[0057] Specifically, use Laplace noise to cover the real urban rail data for data desensitization.
[0058] y ′ = y + Laplace(Δf / ε)
[0059] where y is the original data value, y ′The data after adding noise, Δf is the sensitivity of the dataset, that is, the maximum impact of individual data changes on the function output, and ε represents the differential privacy parameter.
[0060] Gaussian noise is used to retain the statistical characteristics of urban rail data for data desensitization;
[0061] y ′ = y + N(0, σ 2 )
[0062] where y is the original data, and y ′ is the data after adding Gaussian noise, N(0, σ 2 ) is a Gaussian distribution with a mean of 0 and a variance of σ 2 .
[0063] Protect the privacy of urban rail data by modifying and replacing specific fields, and use hash desensitization processing:
[0064] h = H(x)
[0065] where x is the original data, h is the hashed hash value, and H is the hash function, also called the hashing function, such as SHA-256, which is the established rule for data mapping.
[0066] S2.2. Process missing values
[0067] Identify records or fields with missing values in the dataset, including null values, special values, inconsistent formats, and outlier identification, and select appropriate processing strategies according to the data characteristics of the urban rail field, such as deletion, interpolation based on relevant variables, etc.
[0068] S2.2. Detect outliers in the data. Specifically, the K-nearest neighbor algorithm is used for data outlier detection. More specifically, the Euclidean distance is selected as the metric:
[0069]
[0070] where x i and x j represent two data sample points, k represents the kth eigenvalue, n represents the number of data point features, x ik represents the value of the ith sample on the kth feature, and x jk represents the value of the jth sample on the kth feature, and d(x i , x j ) represents the distance between two sample points x i and x j ;
[0071] For each data point x i , find the nearest K neighbors, where x jDenote x i A neighbor of the data point, representing x i One of the K nearest neighbors of x k Represents any data point, used for comparison with x i The distance to the K nearest neighbors to determine the K nearest neighbor set of x i The set of K nearest neighbors of x, Neighbors(x i ) represents the set of K nearest neighbors of the data point x i The set of K nearest neighbors of x:
[0072] Neighbors(x i ) = {x j | d(x i , x j ) < d(x i , x k ), k > j}
[0073] Use the average distance of the K nearest neighbors to measure the local density, Density(x i ) represents the local density of x i , that is, calculate the average distance of the K nearest neighbors of x i , reflecting the data distribution in the area around the data point:
[0074]
[0075] Calculate the outlier score index OutlierScore(x i ) of each data point, measuring the degree to which the data point deviates from the normal pattern. The larger the value, the more likely the data point is an outlier:
[0076]
[0077] Set a threshold according to the outlier score, and mark the data points with scores higher than the threshold as abnormal data
[0078] Judge the content quality, and judge the quality through the support vector machine method.
[0079] Aiming at the problem of uneven data quality collected, it is necessary to judge the content quality, and judge the quality through the support vector machine method.
[0080] Find the optimal hyperplane in the data distribution feature space to classify the sample quality.
[0081] The hyperplane formula is: ω T x + b = 0, where ω is the normal vector, x is the point on the hyperplane, and b is the constant term.
[0082] The distance from a point to the plane is The classification interval is Minimize the objective function While satisfying the constraint conditions Where x i Is the feature vector of the i-th data, and y i Is the label of the i-th data. By ensuring the constraint conditions, the data quality can be correctly judged.
[0083] Identify outliers in the dataset based on the business knowledge scenario, and use statistical methods (such as box plots, clustering, etc.) or set thresholds based on business knowledge to identify outliers in the dataset. Outliers can be corrected by setting deletion, replacement, or smoothing techniques to ensure the stability of the dataset.
[0084] The data preprocessing service completes data cleaning through multiple technical processes, further ensuring the high quality of the data, and at the same time entering the data annotation link as the data source.
[0085] S3. Add labels to the preprocessed data so that the key features of the data can be better understood and learned; among them, the methods used include one or more of NLP text-based data annotation methods such as machine translation, named entity recognition, question-answer matching, association extraction, taxonomy, text classification, and file summarization.
[0086] The collected data has various formats
[0087] (1) From the perspective of file types, the annotation objects are divided into text, speech, and image annotations.
[0088] Text annotation uses NLP text-based data annotation methods such as machine translation, named entity recognition, question-answer matching, association extraction, taxonomy, text classification, and file summarization;
[0089] Speech annotation uses annotation methods such as segmented speech automatic recognition, automatic speech recognition, intent classification, signal quality detection, sound event detection, semantic segmentation, and speech transcription;
[0090] Image annotation selects appropriate annotation shapes (such as linear, rectangular, polygon, etc.) according to the requirements of the annotation task for adding labels to the selected area and other annotation methods.
[0091] (2) From the perspective of annotation forms, the forms are divided into structured, semi-structured, and unstructured annotations.
[0092] Structured annotation matches the annotation object with the structured label, and the corresponding rules are reasonable and clear;
[0093] Unstructured annotation formulates binding annotation standards and guidelines for data annotation differentiation or label matching for data lacking a fixed format;
[0094] Semi-structured annotation targets data with certain structural characteristics, associates structured tag values with semi-structured data, and performs annotation by means of document structure recognition, content classification, etc.
[0095] (3) From the perspective of annotation methods, annotation methods are divided into two categories: manual annotation and machine annotation.
[0096] For the conventional annotation tasks of common sense understanding in the urban rail transit field, annotators need to use annotation tools to perform operations such as tagging and bounding box selection on the data. It is easy to get started, simple to learn, and has strong repeatability. For the complex annotations with high professionalism in the urban rail transit field, annotators are required to have rich professional capabilities. Manual annotation is difficult and task-heavy. Through automated annotation with annotation tools, different annotation tools are selected according to the annotation objects and annotation requirements. The data annotation tools are shown in the following table, which can improve the annotation efficiency and reduce costs at the same time.
[0097]
[0098] Data annotation services convert data into information recognizable by large models through reprocessing for large model training.
[0099] Data annotation is usually realized by predefined rules or machine learning methods, which is similar to the concept of the data quality control link, based on predefined rules or machine learning methods.
[0100] Methods based on predefined rules include hidden Markov models, conditional random fields, etc. Machine learning methods include deep learning, naive Bayes, support vector machines, etc.
[0101] Select different methods for application according to different business requirements. When applying these methods, there is no need to re-derive the formulas, which depends on the specific task requirements and available resources.
[0102] The tasks and differences of annotation are specifically described in the text.
[0103] S4. Conduct quality review and dynamic management of the annotated data through data maintenance.
[0104] At this step, the implementation process of quality review includes:
[0105] The administrator decomposes the annotation tasks; the annotator submits the standard results to the reviewer for review; the task administrator decomposes the tasks; the reviewer can view the content annotated by the annotator through the assigned review tasks, and supports operations such as marking the annotated content as wrong and rejecting it for re-annotation. Ensure the rigor and accuracy of the baseline of the model training data and be responsible for the final model training effect from the source.
[0106] The implementation process of dynamic management includes: creating new data versions for each data update or major change, recording the change logs of each version (such as the amount of newly added data, data sources, label modification situations, etc.) to facilitate tracking the change history of the dataset. To manage and maintain the latest state of the dataset, set a fixed update cycle according to business requirements, determine the scope of incremental or full-scale updates, and set monitoring metrics to track data quality and update progress.
[0107] S5, Data security services throughout the entire process of dataset construction
[0108] Following the architecture design of the internal production network, management network, and external service network in the urban rail transit field, and following the principle of security isolation, perform user access control permissions and log monitoring.
[0109] Allocate data access permissions according to the principle of least privilege to ensure that only authorized personnel can access specific data, protecting the security of enterprise rights and interests and enterprise data privacy. Combine log monitoring metrics to obtain data operation logs and avoid data leakage.
[0110] A method and system for constructing a high-quality dataset for a large model in the urban rail transit field provided by the present invention, compared with the prior art, cover multiple aspects such as data collection, preprocessing, annotation, maintenance, and security. It improves the systematization and standardization of the entire data processing process, which is conducive to improving data quality and the effect of model training. The two core service modules of data preprocessing and data annotation clarify the data cleaning strategy and annotation process in the urban rail transit field. Ensure that the data is accurately annotated, enhancing the learning ability and generalization ability of the model. By defining a clear data management mechanism and task decomposition process in the urban rail transit field, the standardization and efficiency of data processing are ensured. Improve the overall efficiency of data processing and ensure the accuracy and timeliness of the dataset.
[0111] The present invention also provides a system for constructing a high-quality dataset for a large model in the urban rail transit field, characterized in that it is used to execute the steps included in the method for constructing a high-quality dataset for a large model in the urban rail transit field as described in any one of the above, including: a data collection module 1, a preprocessing module 2, an annotation module 3, and a maintenance module 4. The data collection module 1 is used to collect and summarize data that meets business requirements as an input data source; the preprocessing module 2 is used to clean the data through data preprocessing to improve data quality; the annotation module 3 is used to add labels to the preprocessed data so that the key features of the data can be better understood and learned; the maintenance module 4 is used to perform quality review and dynamic management on the annotated data through data maintenance.
[0112] In some feasible embodiments, the system for constructing a high-quality dataset for the large model in the urban rail transit field further includes a security service module 5, which is used to run through all the process links of dataset construction and perform security monitoring on each link respectively.
[0113] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a high-quality dataset for large models in the urban rail transit field, characterized in that, Including the following steps: Collect and summarize the data that meets the business requirements as the input data source; Clean the data through data preprocessing to improve the data quality; Add labels to the preprocessed data so that the key features of the data can be better understood and learned; Conduct quality review and dynamic management of the labeled data through data maintenance.
2. The method for constructing a high-quality data set for a large model in the urban rail field according to claim 1, wherein The step of cleaning the data through data preprocessing to improve the data quality includes one or more of: data desensitization processing, missing value processing, low-quality filtering processing, redundancy removal processing, and outlier monitoring processing.
3. The method for constructing a high-quality data set for a large model in the urban rail field according to claim 2, wherein The data desensitization processing includes: Protecting the sensitive information of the data that needs to be processed by using noise addition or hash desensitization strategy; wherein, the noise includes one or more of Gaussian noise and Laplace noise.
4. The method for constructing a high-quality data set for a large model in the urban rail field according to claim 2, wherein The outlier monitoring processing includes using the K-nearest neighbor algorithm for data anomaly detection, and the step of using the K-nearest neighbor algorithm for data anomaly detection includes: Selecting the Euclidean distance as the metric: where x i and x j represent two data sample points, k represents the k-th eigenvalue, n represents the number of data point features, x ik represents the value of the i-th sample on the k-th feature, x jk represents the value of the j-th sample on the k-th feature, d(x i , x j ) represents the distance between two sample points x i and x j ; For each data point x i , find the K nearest neighbors, where x j represents a neighbor of x i and is one of the K nearest neighbors of x i . x k represents any data point used to compare the distances from x i to its K nearest neighbors in order to determine the set of K nearest neighbors of x i . Neighbors(x i ) represents the set of K nearest neighbors of the data point x i : Neighbors(x i ) = {x j | d(x i , x j ) < d(x i , x k ), k > j} Use the average distance of the K nearest neighbors to measure the local density. Density(x i ) represents the local density of x i , that is, calculate the average distance of the K nearest neighbors of x i , which reflects the data distribution in the area around the data point: Calculate the outlier score metric OutlierScore(x i ), which measures the degree to which a data point deviates from the normal pattern. The larger the value, the more likely the data point is an outlier: Setting a threshold according to the anomaly score and marking the data points with a score higher than the threshold as abnormal data.
5. A method for constructing a high-quality data set for large models in the urban rail transit field according to any one of claims 1-4, characterized in that The step of cleaning the data through data preprocessing to improve the data quality includes judging the content quality, and the step of judging the content quality includes: Finding the optimal hyperplane in the data distribution feature space to classify the sample quality; The hyperplane formula is: ω T · x + b = 0, where ω is the normal vector, x is a point on the hyperplane, and b is the constant term; The distance from a point to a plane is The classification margin is Minimize the objective function While satisfying the constraint conditions Where x i Is the feature vector of the i-th data, y i Is the label of the i-th data, and the data quality is correctly judged by the constraint conditions.
6. The method for constructing a high-quality data set for a large model in the urban rail transit field according to claim 1, wherein The step of adding labels to the preprocessed data so that the key features of the data can be better understood and learned includes adding labels to the preprocessed data by using tools, and the tools include one or more of LabelBox, BRAT, WebAnno, and Audacity.
7. A method for constructing a high-quality data set for a large model in the urban rail transit field as claimed in claim 1, characterized in that The step of adding labels to the preprocessed data so that the key features of the data can be better understood and learned includes using one or more of NLP text data annotation methods such as machine translation, named entity recognition, question-answer matching, association extraction, taxonomy, text classification, and file summarization.
8. A method for constructing a high-quality data set for a large model in the urban rail transit field according to claim 1, characterized in that, In the step of adding labels to the preprocessed data so that the key features of the data can be better understood and learned, the annotation objects are one or more of text, speech, and image.
9. A method for constructing a high-quality data set for a large model in the urban rail transit field according to any one of claims 1 to 8, characterized in that It also includes a data security service throughout the entire process of dataset construction.
10. A system for constructing a high-quality data set for large models in the urban rail transit field, characterized in that, Used to execute the steps included in the method for constructing a high-quality dataset for the large model in the urban rail transit field as described in any one of claims 1 to 9, including: A data collection module (1) for collecting and summarizing the data that meets the business requirements as the input data source; A preprocessing module (2) for cleaning the data through data preprocessing to improve the data quality; A labeling module (3) for adding labels to the preprocessed data so that the key features of the data can be better understood and learned; A maintenance module (4) for conducting quality review and dynamic management of the labeled data through data maintenance.
Citation Information
Patent Citations
Load identification method and system based on user cooperation, electronic equipment and medium
CN116089820A
Method for improving data quality of heterogeneous system
CN117312290A
Corpus construction method and system for subway field and storage medium
CN117875304A
Big data medical information sharing system and method for medical treatment
CN118299016A
Data docking detection method for multi-service system
CN118708961A