Large language model high-quality text data set construction method and system
By building source tracking models, data preprocessing, label definition and text enhancement steps, the noise problem in the construction of high-quality industry text data sets is solved, and the collection and enhancement of high-quality text data is achieved, and the quality and quantity of data sets are improved.
Patent Information
- Application Number
- CN202510812339.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult to construct high-quality industry text data sets in the prior art, and traditional data enhancement methods are prone to generate noisy data.
By building source tracking models, data preprocessing, labeling and text enhancement steps, collect, clean, annotate and enhance industry text data, use industry dictionary to update large language models, perform unsupervised training and data testing.
Improves the quality and quantity of text data, reduces noise data, and provides high-quality industry text analysis and model training support.
Smart Images

Figure CN120336527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and specifically to a method and system for constructing a high-quality text dataset of a large language model. Background Art
[0002] Constructing a high-quality text dataset is of great significance for efficiently utilizing historical data and empirical knowledge in a vast amount of scientific materials. As an emerging information extraction technology, text mining can parse character sequences through constructing deep learning algorithms to obtain logical information. However, the effect of text mining highly depends on high-quality data. Limited by the complexity and high heterogeneity of industry-specific texts, traditional dataset construction methods are difficult to be directly applicable to the construction of industry high-quality text datasets, which makes the dataset construction method a key issue.
[0003] Currently, there are two major problems in constructing a high-quality industry text dataset: one is how to improve the quality of text data, and the other is how to increase the quantity of text data. In terms of improving data quality, traditional dataset construction methods are difficult to be directly applicable to the construction of specific industry high-quality text datasets. In terms of increasing data quantity, data augmentation technology (Data Augmentation, DA) can be adopted to generate more new samples by learning the prior knowledge of training instances to expand the original dataset. Currently, the simple data augmentation (Easy Data Augmentation, EDA) method is insensitive to labels and is prone to generating noisy samples when processing supervised text corpora. If industry knowledge can be integrated to annotate and augment text data, it will help improve data quality and increase the quantity of available data for modeling, thereby obtaining a high-quality industry text mining dataset.
[0004] How to construct a high-quality text dataset and reduce noisy data is a technical problem that needs to be solved. Summary of the Invention
[0005] The technical task of the present invention is to, in view of the above deficiencies, provide a method and system for constructing a high-quality text dataset of a large language model to solve the technical problem of how to construct a high-quality text dataset and reduce noisy data.
[0006] In a first aspect, a method for constructing a high-quality text dataset of a large language model according to the present invention includes the following steps: Data collection: Collect relevant documents from industry materials to obtain a text corpus; Data processing: Perform data preprocessing on the collected text corpus, and through data preprocessing, perform data cleaning, format conversion, text cutting, and parsing and marking operations to obtain a pre-annotated text corpus, and output the pre-annotated text corpus in a predetermined format; Label Definition: Build a label system based on industry-specific terms and metrics. Annotate the entities and the relationships between entities in the pre-annotated text corpus based on the label system. The obtained entity labels and the relationships between entities are used as label information. Text Enhancement: Build an industry dictionary based on the collected industry data. Update the original vocabulary of the pre-trained large language model based on the industry dictionary. Use the collected text corpus samples as input to perform unsupervised training on the pre-trained large language model to obtain a fine-tuned large language model. For the text corpus to be enhanced, randomly mask some words in the text corpus. Use the masked text corpus and the corresponding label information as input attributes. Through the unsupervised trained large model, perform upper and lower semantic analysis on the input attributes, predict the masked words in the text corpus, output the enhanced text data, and conduct data testing on the text data. Through data testing, perform data quality evaluation and application quality evaluation on the enhanced text data.
[0007] Preferably, data collection includes the following steps: Build a source tracking model. The source tracking model is a triple structure including an object, metadata, and a data source. The object represents various objects in the industry that need to be source-tracked. The metadata is used to describe the metadata information involved in the acquisition, sharing, and application of the object. The data source is a tracking operation mechanism built based on the metadata source. Divide the metadata information in the source tracking model into basic metadata and derivative metadata. The basic metadata is used to describe the basic attribute information of industry-related literature and materials, including the literature title, literature author, literature abstract, keywords, literature content, literature publication time, journal impact factor, and literature storage path. The derivative metadata is used to record the derivative information generated during the data processing process and related to industry literature analysis and application, including the key corpus extracted from the original industry literature, the supplementary data generated by simulating different industry working conditions, the input parameters and output optimization results involved in various industry process optimization models, and the industry process knowledge base formed during the practice and research process. According to the characteristics of the source website of the target literature, select an appropriate web crawler framework, and parse the web page content through the web crawler framework to obtain the target literature. For literature in the form of paper materials, scan the paper materials based on scanning and optical character recognition technology to obtain relevant text corpus. For literature in the form of network electronic resources, call the corresponding database interface or obtain relevant text corpus through web crawler technology. For the collected text corpus, digital fingerprint identification is performed on the text corpus based on the source tracking model to establish the association relationship between the literature, the text corpus, and the derivative data, and the entire process of literature analysis and application tasks is recorded through the derivative data, and an ordered association between the derivative data and the corresponding associated processing operations is constructed, where the digital fingerprint identification includes the original source website, the release time, the author information, and the collection time.
[0008] Preferably, the data processing includes the following steps: Perform duplicate checking on the text corpus based on the hash algorithm; Match HTML tags in the text corpus based on regular expressions, replace all matched HTML tags with an empty string, retain the pure content part in the text corpus, identify special characters in the text corpus based on regular expressions, replace the special characters with an empty string, and locate and delete the advertising content in the text corpus based on the constructed advertising keyword library to obtain the denoised text corpus; Perform spelling and grammar checks on the denoised text corpus to obtain the checked text corpus; Convert the checked text corpora from different sources into a unified format of plain text form, and store the basic metadata corresponding to each text corpus in the source database; Segment the text corpus, and each paragraph obtained by segmentation constitutes an independent text unit. For each text unit of a paragraph, disassemble the paragraph into individual sentences based on natural language processing tools; Parse and perform part-of-speech tagging on each sentence to obtain the pre-annotated text corpus; Convert the pre-annotated text corpus into a predetermined format for output. Among them, for the text classification task, convert the pre-annotated text corpus into JSON format for output, and store the text corpus and the corresponding label information in the form of key-value pairs. For the sentiment analysis task, convert the pre-annotated text corpus into XML format for output, and store the pre-annotated text corpus through the XML tree structure; Among them, storing the pre-annotated text corpus through the XML tree structure includes the following operations: Create an XML root element; Create corresponding sub-elements and grandchild elements according to the text content and structure of the text corpus, fill the text content into the corresponding element nodes, and add necessary attributes. The necessary attributes include the language type and the text category to construct a complete XML tree structure.
[0009] Preferably, the label definition includes the following operations: Construct an entity label system: Based on the analysis results of industry characteristics, determine the scope of entity types, set detailed attributes for each entity type, and construct a multi-level entity label structure according to the level of detail of entities and business semantics to ensure that entity labels can precisely describe various objects in the data; Construct an entity relationship label system: Analyze the interaction methods between entities in industry business processes and typical scenarios, identify the types of entity relationships, determine the participants and participation methods of entity relationships according to business logic and data flow, set corresponding attributes for each entity relationship, and standardize the directionality of entity relationships to clarify the main entity and the guest entity; Use the annotation format of "Entity 1 - Relationship Type - Entity 2" to annotate the potential relationships between two entities, where the relationship type follows the predefined entity relationship types.
[0010] Preferably, text enhancement includes the following steps: Collect industry-specific terms, construct an industry-specific industry dictionary based on industry-specific data, and filter the professional words in the industry dictionary; Update the original vocabulary of the large language model based on the industry dictionary. For professional data not covered in the original vocabulary, add the professional data to the original vocabulary and add the embedding vectors corresponding to the professional terms to the embedding matrix of the corresponding original vocabulary in the large language model; Obtain text corpus as text corpus samples, perform unsupervised training on the large language model based on the text corpus samples, and fine-tune the parameters of the large language model to obtain a fine-tuned large language model; For the text corpus to be enhanced, randomly mask some words in the text corpus and record the corresponding label information of the masked words. Use the masked text corpus and the corresponding label information as input attributes. The fine-tuned large language model establishes the association between words and label information, generates semantically rich embedding vectors, represents the information of each word through the embedding vectors, predicts the masked words according to the embedding vectors, generates a candidate enhanced vocabulary list, and selects words that match the original semantics and label information from the candidate enhanced vocabulary list as the enhanced text data; Perform data tests on the enhanced text data, including data quality assessment and application quality assessment. When performing data quality assessment, conduct integrity assessment by checking whether the data lacks necessary fields and whether there are blank or invalid values, verify the data correctness by comparing the enhanced text data with the original corpus text and authoritative data in the field, and conduct consistency assessment by evaluating whether the data format and encoding are unified and whether the logic of the data is coherent in different scenarios. When performing application quality assessment, train an industrial large model using the enhanced text data, and by comparing the accuracy, recall rate, and F1 value of the large models trained using the original data and the enhanced text data on the validation set and the test set, examine the improvement effect of the enhanced text data on the model generalization ability, and conduct manual evaluation in combination with the prediction results of domain experts on the large language model, and let the experts judge the rationality and accuracy of the output of the large language model.
[0011] In a second aspect, a high-quality text data set construction system for a large language model of the present invention is used to construct a high-quality text data set by using a high-quality text data set construction method according to any one of the first aspects. The system includes a data collection module, a data processing module, a label definition module, and a text enhancement module; The data collection module is used to perform the following: collect relevant documents from industry materials to obtain text corpora; The data processing module is used to perform the following: perform data preprocessing on the collected text corpora, perform data cleaning, format conversion, text cutting, and parsing and marking operations through data preprocessing to obtain pre-annotated text corpora, and output the pre-annotated text corpora in a predetermined format; The label definition module is used to perform the following: construct a label system based on industry professional terms and indicators, and label the entities and the relationships between entities in the pre-annotated text corpora based on the label system, and use the obtained entity labels and the relationships between entities as label information; The text enhancement module is used to perform the following: construct an industry dictionary based on the collected industry data, update the original vocabulary of the pre-trained large language model based on the industry dictionary, use the collected text corpus samples as inputs to perform unsupervised training on the pre-trained large language model to obtain a fine-tuned large language model. For the text corpus to be enhanced, randomly mask some words in the text corpus, use the masked text corpus and the corresponding label information as input attributes, perform upper and lower semantic analysis on the input attributes through the unsupervised trained large model, predict the masked words in the text corpus, output the enhanced text data, and perform data tests on the text data, and perform data quality assessment and application quality assessment on the enhanced text data through the data tests.
[0012] Preferably, the data collection module is used to perform the following: Build a source tracking model, which is a triple structure including an object, metadata, and a data source. The object represents various objects in the industry that need to be tracked for their sources. The metadata is used to describe the metadata information involved in the acquisition, sharing, and application of the object. The data source is a tracking operation mechanism built based on the metadata source; Divide the metadata information in the source tracking model into basic metadata and derivative metadata. The basic metadata is used to describe the basic attribute information of industry-related literature and materials, including the literature title, literature author, literature abstract, keywords, literature content, literature publication time, journal impact factor, and literature storage path. The derivative metadata is used to record the derivative information generated during the data processing process and related to industry literature analysis and application, including the key corpus extracted from the original industry literature, the supplementary data generated by simulating different industry working conditions, the input parameters and output optimization results involved in various industry process optimization models, and the industry process knowledge base formed during the practice and research process; According to the characteristics of the source website of the target literature, select an appropriate web crawler framework, and obtain the target literature by parsing the web page content through the web crawler framework; For literature in the form of paper materials, scan the paper materials based on scanning and optical character recognition technology to obtain relevant text corpus. For literature in the form of network electronic resources, call the corresponding database interface or obtain relevant text corpus through web crawler technology; For the collected text corpus, perform digital fingerprint identification on the text corpus based on the source tracking model to establish the association relationship between the literature, the text corpus, and the derivative data, and record the entire process of the literature analysis and application task through the derivative data, and build an ordered association between the derivative data and the corresponding associated processing operations. Among them, the digital fingerprint identification includes the original source website, publication time, author information, and collection time.
[0013] Preferably, the data processing module is used to perform the following: Check for duplicate text corpus based on the hash algorithm; Match HTML tags in the text corpus based on regular expressions, replace all matched HTML tags with an empty string, retain the pure content part in the text corpus, identify special characters in the text corpus based on regular expressions, replace the special characters with an empty string, and locate and delete the advertising content in the text corpus based on the constructed advertising keyword library to obtain the denoised text corpus; Check the spelling and grammar of the denoised text corpus to obtain the checked text corpus; Convert the checked text corpus from different sources into a text corpus in a unified format of plain text, and store the basic metadata corresponding to each text corpus in the source database; Segment the text corpus. Each paragraph obtained by segmentation constitutes an independent text unit. For each text unit of a paragraph, disassemble the paragraph into individual sentences based on natural language processing tools; Parse and perform part-of-speech tagging on each sentence to obtain a pre-annotated text corpus; Convert the pre-annotated text corpus into a predetermined format for output. Among them, for text classification tasks, convert the pre-annotated text corpus into JSON format for output, and store the text corpus and the corresponding label information in the form of key-value pairs. For sentiment analysis tasks, convert the pre-annotated text corpus into XML format for output, and store the pre-annotated text corpus through an XML tree structure; Among them, when storing the pre-annotated text corpus through an XML tree structure, the data processing module is used to perform the following operations: Create an XML root element; Create corresponding child elements and grandchild elements according to the text content and structure of the text corpus, fill the text content into the corresponding element nodes, and add necessary attributes. The necessary attributes include language type and text category to build a complete XML tree structure.
[0014] Preferably, the label definition module is used to perform the following operations: Construct an entity label system: Based on the analysis results of industry characteristics, determine the scope of entity types, set detailed attributes for each entity type, and construct a multi-level entity label structure according to the detailed degree of entities and business semantics to ensure that entity labels can finely describe various objects in the data; Construct an entity relationship label system: Analyze the interaction methods between entities in the industry business process and typical scenarios, identify the relationship types between entities, determine the participants and participation methods of entity relationships according to business logic and data flow, set corresponding attributes for each entity relationship, and standardize the directionality of entity relationships to clarify the main entity and the guest entity; Use the annotation format of "Entity 1 - Relationship Type - Entity 2" to annotate the potential relationships between two entities, where the relationship type follows the pre-defined entity relationship types.
[0015] Preferably, the text enhancement module is used to perform the following operations: Collect industry-specific terms, build an industry-specific industry dictionary based on industry professional data, and filter the professional vocabulary in the industry dictionary; Update the original vocabulary of the large language model based on the industry dictionary. For the professional data not covered in the original vocabulary, add the embedding vectors corresponding to the professional terms to the original vocabulary of the large language model; Obtain text corpus as a text corpus sample, perform unsupervised training on the large language model based on the text corpus sample, and fine-tune the parameters of the large language model to obtain a fine-tuned large language model; For the text corpus to be enhanced, randomly mask some words in the text corpus, and record the corresponding label information of the masked words. Take the masked text corpus and the corresponding label information as input attributes. The fine-tuned large language model establishes the association between vocabulary and label information, generates semantic-rich embedding vectors, represents the information of each word through the embedding vectors, and predicts the masked words based on the embedding vectors to generate a candidate enhanced vocabulary list, and selects words that match the original semantics and label information from the candidate enhanced vocabulary list as the enhanced text data; Conduct data testing on the enhanced text data, including data quality assessment and application quality assessment. When performing data quality assessment, conduct integrity assessment by checking whether the data lacks necessary fields and whether there are blanks or invalid values, verify the data correctness by comparing the enhanced text data with the original corpus text and authoritative data in the field, and conduct consistency assessment by evaluating whether the data format and encoding are unified and whether the logic of the data is coherent in different scenarios. When performing application quality assessment, use the enhanced text data to train an industrial large model, and by comparing the accuracy, recall rate, and F1 value of the large models trained using the original data and the enhanced text data on the validation set and the test set, test the improvement effect of the enhanced text data on the model generalization ability, and combine the prediction results of domain experts on the large language model for manual evaluation, and let the experts judge the rationality and accuracy of the output of the large language model.
[0016] The method and system for constructing a high-quality text data set of the large language model of the present invention have the following advantages: improving the quality of text corpus in the industry field through data collection, data processing, and label definition operations. The text enhancement operation incorporates domain knowledge into pre-training large models to perform supervised text data expansion, improving the industry text corpus. The expanded data fully relies on industry semantic knowledge, which can effectively reduce the generation of noise data and provide high-quality data support for industry text analysis and model training. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] The present invention will be further described below with reference to the drawings.
[0019] Figure 1 It is a flowchart of a method for constructing a high-quality text dataset of a large language model in Embodiment 1. Detailed implementation manners
[0020] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the exemplified embodiments are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0021] The embodiments of the present invention provide a method and system for constructing a high-quality text dataset of a large language model, which are used to solve the technical problems of how to construct a high-quality text dataset and reduce noise data.
[0022] Embodiment 1: A method for constructing a high-quality text dataset of a large language model according to the present invention includes four steps: data collection, data processing, label definition, and text enhancement.
[0023] Step S100 Data collection: Collect relevant literatures from industry materials to obtain text corpora.
[0024] The data collection includes the following steps: (1) Construct a source tracking model, which is a triple structure including an object, metadata, and a data source. The object represents various objects in the industry that need to be tracked for their sources. The metadata is used to describe the metadata information involved in the acquisition, sharing, and application links of the object. The data source is a tracking operation mechanism constructed based on the metadata source; (2) Divide the metadata information in the source tracking model into basic metadata and derivative metadata. The basic metadata is used to describe the basic attribute information of industry-related literatures and materials, including literature titles, literature authors, literature abstracts, keywords, literature contents, literature publication times, journal impact factors, and literature storage paths. The derivative metadata is used to record the derivative information generated during the data processing process and related to industry literature analysis and application, including key corpora extracted from original industry literatures, supplementary data generated by simulating different industry working conditions, input parameters involved in various industry process optimization models and the optimized results output, as well as industry process knowledge bases accumulated during the practice and research processes; (3) According to the characteristics of the source website of the target literature, select an appropriate web crawler framework, and parse the web page content through the web crawler framework to obtain the target literature; (4) For literatures in the form of paper materials, scan the paper materials based on scanning and optical character recognition technologies to obtain relevant text corpora. For literatures in the form of network electronic resources, call the corresponding database interface or obtain relevant text corpora through web crawler technologies; (5) For the collected text corpus, digital fingerprint identification is performed on the text corpus based on the source tracking model to establish the association relationship between the literature, the text corpus, and the derivative data, and the entire process of literature analysis and application tasks is recorded through the derivative data, and an ordered association between the derivative data and the corresponding associated processing operations is constructed. Among them, the digital fingerprint identification includes the original source website, the release time, the author information, and the collection time.
[0025] The data collection step accurately extracts relevant literature from a vast amount of industry materials by establishing a source tracking model, ensuring that the origin of each piece of data is clearly traceable. As the specific implementation of data processing, it includes the following five operations.
[0026] Operation 1: Construct a triple structure <object, metadata, source> of the Source Tracking Model (STM). Among them, the object represents various objects that need to be source-tracked within the injection molding industry; the metadata is used to describe in detail the metadata information involved in the acquisition, analysis, and application of the object; and the source is the tracking operation mechanism constructed based on the metadata source.
[0027] Operation 2: Classify the metadata information in the source tracking model into two categories: basic metadata and derivative metadata. The basic metadata is mainly used to describe the basic attribute information of injection molding-related literature and materials, including the literature title, such as "Research on the Influence of Injection Molding Process Parameter Optimization on Product Performance"; the literature author; the literature abstract, which refines and summarizes the core content of the literature; keywords, which accurately reflect the research focus of the literature, such as "injection molding process", "product performance", "parameter optimization", etc.; the specific text content of the literature; the name of the professional journal in which the literature is published; the publication time of the literature; the impact factor of the journal; the DOI number of the literature; the specific storage path of the literature in the storage system, etc. The derivative metadata records the derivative information generated during the data processing process and closely related to the analysis and application of injection molding literature, including the key corpus obtained from the original injection molding literature through processing operations such as screening and refining, such as the core expression regarding the relationship between injection molding temperature and product forming quality; the supplementary data generated by simulating different injection molding conditions; the input parameters and the optimized output results involved in various injection molding process optimization models; and the injection molding process knowledge base accumulated during the long-term practice and research process.
[0028] Operation 3: Select a web crawler framework: Based on the characteristics of the source website of the target literature, select a suitable web crawler framework. For example, Scrapy can be used for structured crawling, and Beautiful Soup can be used for parsing web content. Configure the crawling rules: Set the starting crawling URL of the crawler, the target link selection rules, the crawling depth, the crawling interval, etc. Deploy a crawler cluster: Deploy multiple crawler instances in a distributed environment to ensure their coordinated work, improve the crawling efficiency, and avoid causing excessive access pressure on the target website.
[0029] Operation 4: Data collection and verification: Professional text corpora in the injection molding field widely exist in various document carriers, such as academic papers, technical reports, patent documents, equipment operation manuals, injection molding process parameter manuals, process flowcharts, industry standards, quality management system documents, etc.
[0030] For the relevant paper materials collected by enterprises, industry associations, and standard organizations, use scanning combined with optical character recognition (OCR) technology. Use a professional high-resolution scanner to carefully scan the paper documents to ensure that information such as text and charts is digitized and collected completely without omission. Then, with the help of advanced OCR recognition software, accurately convert the scanned image files into editable electronic documents, which greatly facilitates subsequent corpus processing and analysis work and improves work efficiency and accuracy. Second, for electronic resources on the Internet, you can use the application programming interfaces provided by relevant databases for data retrieval and acquisition, or use web crawler technology. Use a pre-written crawler program, such as Crawl4AI, to automatically access and accurately capture the relevant content of the injection molding industry on the web page to obtain the required corpus.
[0031] Operation 5: Attach a digital fingerprint identifier containing key elements such as the original source URL, publication time, author information, and collection time to each piece of collected literature, thereby establishing the correlation between the literature, text data, and derivative data. Use derivative metadata to record the entire process of literature analysis and application tasks, construct an orderly correlation between derivative data and corresponding processing operations, and ensure the traceability of the text data processing process.
[0032] Step S200: Data processing: Perform data preprocessing on the collected text corpus. Through data preprocessing, perform data cleaning, format conversion, text cutting, and parsing and marking operations to obtain pre-annotated text corpus and output the pre-annotated text corpus in a predetermined format.
[0033] In this embodiment, the data processing includes the following steps: (1) Check for duplicate texts based on the hash algorithm; (2)Match HTML tags in the text corpus based on regular expressions, replace all matched HTML tags with empty strings, retain the pure content part in the text corpus, identify special characters in the text corpus based on regular expressions, replace special characters with empty strings, and locate and delete the advertising content in the text corpus based on the constructed advertising keyword library to obtain the denoised text corpus; (3)Perform spelling and grammar checks on the denoised text corpus to obtain the text corpus after checking; (4)Convert the text corpus after checking from different sources into a unified format of plain text, and store the corresponding basic metadata of each text corpus in the source database; (5)Segment the text corpus, and each paragraph obtained by segmentation constitutes an independent text unit. For each text unit of a paragraph, disassemble the paragraph into individual sentences based on natural language processing tools; (6)Parse and perform part-of-speech tagging on each sentence to obtain the pre-annotated text corpus; (7)Convert the pre-annotated text corpus into a predetermined format for output. Among them, for the text classification task, convert the pre-annotated text corpus into JSON format for output, and store the text corpus and the corresponding label information in the form of key-value pairs. For the sentiment analysis task, convert the pre-annotated text corpus into XML format for output, and store the pre-annotated text corpus through the XML tree structure.
[0034] Among them, when storing the pre-annotated text corpus through the XML tree structure, create an XML root element, create corresponding child elements and grandchild elements according to the text content and structure of the text corpus, fill the text content into the corresponding element nodes, and add necessary attributes. The necessary attributes include language type and text category to construct a complete XML tree structure.
[0035] The data processing steps in this embodiment perform preprocessing such as data cleaning and format conversion and post-processing such as text segmentation on the collected literature based on subsequent specific task requirements.
[0036] As a specific implementation of data processing, this step includes seven operations: data deduplication, noise information removal, spelling and grammar checking, format conversion, text segmentation, parsing and annotation, and data output.
[0037] Data deduplication: Use the hash algorithm to check for duplicates in the obtained corpus to ensure the uniqueness of each piece of literature data.
[0038] Noise information removal: Use the regular expression "<.*?>" to match HTML tags in the text. This regular expression accurately identifies HTML tags by looking for strings that start with "<" and end with ">". During text processing, all matched HTML tags are replaced with an empty string, thus retaining the pure content part of the text and removing things like " ”" Implement preliminary purification of text by using various HTML tags such as "etc."; Use the regular expression "[^\a-zA-Z0-9\s.,!?-]" to identify special characters in the text other than letters, numbers, and common punctuation marks (including spaces, periods, commas, exclamation marks, question marks, hyphens, etc.). These special characters often interfere with subsequent text analysis tasks. Through string replacement operations, replace them all with empty strings and remove them from the text, only retaining the meaningful text content. For example, remove special symbols such as "@", "#", "$" in the text; Build an advertising keyword library containing common advertising words and phrases, such as "Get for free", "Limited-time discount", "Click to get a gift", etc. During text processing, use a string matching algorithm to search for these keywords one by one in the text. Once advertising content is found in the text, locate and delete it to ensure that the processed text does not contain any advertising-related information, thereby improving the quality and relevance of the text.
[0039] Spelling and grammar checking: For possible spelling mistakes of professional terms or errors caused by inaccurate OCR recognition in the denoised corpus, use natural language processing tools such as NLTK (Natural Language Toolkit) or StanfordCoreNLP for spelling and grammar checking.
[0040] Format conversion: Uniformly convert the corpus from different sources to UTF-8 encoding format and convert it to TXT plain text data. Store basic metadata such as the title, author, keywords, DOI number, etc. of each piece of preprocessed literature corpus in the metadata database to ensure the traceability of the source of the literature corpus.
[0041] Text segmentation: Considering that different paragraphs in the literature in the injection molding field describe different links of the injection molding process, different dimensions of mold design, etc., first segment the abstract and text content in the preprocessed literature corpus according to paragraphs, so that each paragraph forms an independent text unit. Then, for each segmented paragraph, use natural language processing tools such as the Punkt sentence splitter to further break it down into individual sentences. This way can accurately label specific semantic units during the annotation process, thereby improving the accuracy and consistency of the annotation.
[0042] For each segmented paragraph, design sentence boundary recognition models in different languages. According to the grammar rules and punctuation habits of the language, accurately identify the start and end positions of sentences, and use natural language processing tools such as the sentence splitters in NLTK or SpaCy or the Punkt sentence splitter to further break it down into individual sentences; Taking NLTK as an example, after loading the English sentence splitter, call its tokenize.sent_tokenize() function and pass in the text to be processed. This function will split the text into individual sentences based on punctuation marks such as full stops (.), question marks (?), exclamation marks (!) in English, and return a list containing all the sentences. For Chinese text, corresponding Chinese sentence splitting models or rules can be used for processing, such as splitting sentences based on Chinese full stops (。), question marks (?), exclamation marks (!), etc.
[0043] Parse and annotate: Parse and perform part-of-speech tagging on each segmented sentence to obtain a fine-grained pre-annotated text corpus.
[0044] Data output: Text format conversion: Determine the target text format according to the specific requirements of subsequent tasks.
[0045] For text classification tasks, select the JSON format and store information such as the text and its corresponding classification labels in the form of key-value pairs. Complete the conversion with the help of the json library in Python. Integrate the plain text and its related metadata (such as document unique identifier, source information, preprocessing timestamp, etc.) after the above processing into a dictionary object, and then use the dumps method of the json library to convert the dictionary into a JSON-formatted string to achieve the JSON format encapsulation of the text; For sentiment analysis tasks, select the XML format and use the tag feature of XML to annotate the sentiment tendency of the text (such as positive, negative, neutral, etc.). Use the xml.etree.ElementTree library for conversion. First, create an XML root element, then create corresponding child elements and grandchild elements according to the text content and structure, fill the text content into the corresponding element nodes, and add necessary attributes (such as language type, text category, etc.) to build a complete XML tree structure, thus obtaining the XML format text that meets the requirements.
[0046] Step S300 tag definition: Build a tag system based on industry-specific terms and metrics, and annotate the entities and the relationships between entities in the pre-annotated text corpus. The obtained entity tags and the relationships between entities are used as tag information.
[0047] As a specific implementation of tag definition, this step includes the following operations: (1) Build an entity tag system: Based on the analysis results of industry characteristics, determine the scope of entity types, set detailed attributes for each entity type, and build a multi-level entity tag structure according to the level of detail of the entities and business semantics to ensure that the entity tags can precisely describe various objects in the data; (2)Construct an entity relationship label system: Analyze the interaction methods between entities in the industry business processes and typical scenarios, identify the types of relationships between entities, determine the participants and participation methods of entity relationships based on business logic and data flow directions, set corresponding attributes for each entity relationship, and standardize the directionality of entity relationships to clarify the main entity and the guest entity; (3)Use the annotation format of "Entity 1 - Relationship Type - Entity 2" to annotate the potential relationships between two entities, where the relationship type follows the predefined entity relationship types.
[0048] In this embodiment, the label definition steps formulate a dedicated label system based on industry-specific terms, key indicators, etc., accurately annotate the data, and accurately mark the entity information and their mutual relationships. Specifically, it includes two operations: constructing an entity label system and constructing an entity relationship label system.
[0049] Construct an entity label system: Based on the results of industry characteristics analysis, determine the scope of entity types, and set detailed attributes for each entity type. Construct a multi-level entity label structure according to the detail level and business semantics of the entities to ensure that the entity labels can precisely describe various objects in the data.
[0050] Construct an entity relationship label system: Analyze the interaction methods between entities in the industry business processes and typical scenarios, identify common relationship types, determine the participants and participation methods of the relationships based on business logic and data flow directions. Set corresponding attributes for each entity relationship, and standardize the directionality of the relationships to clarify the main entity and the guest entity.
[0051] Based on the four core elements of material - equipment - process - product, eight entity label categories suitable for the entire injection molding field system are constructed as shown in Table 1.
[0052] Table 1. Entity Label Category Table
[0053] Perform entity annotation on the preprocessed fine-grained pre-annotated text corpus. The label format is <Entity Label Type>Entity Content< / Entity Label Type>. For example, for the material "polypropylene (PP)", it should be annotated as <Material>polypropylene (PP)< / Material>; for the equipment "Haitian MA2000 injection molding machine", it is annotated as <Equipment>Haitian MA2000 injection molding machine< / Equipment>.
[0054] Abstract and summarize the relationships between every two of the eight types of entities, and define eight relationship categories to represent the potential associations between entities, as shown in Table 2.
[0055] Table 2. Relationship Table between Entities
[0056] Use the annotation format of "Entity 1 - Relationship Type - Entity 2" to annotate the potential relationship between two entities, where the relationship type follows the pre - defined entity relationship types for the injection molding field.
[0057] Step S400 Text Enhancement: Build an industry dictionary based on the collected industry data, update the original vocabulary of the pre - trained large - language model based on the industry dictionary, use the collected text corpus samples as input to perform unsupervised training on the pre - trained large - language model to obtain a fine - tuned large - language model. For the text corpus to be enhanced, randomly mask some words in the text corpus, use the masked text corpus and the corresponding label information as input attributes, and through the unsupervised - trained large model, perform upper - and - lower semantic analysis on the input attributes, predict the masked words in the text corpus, output the enhanced text data, and conduct data testing on the text data, and evaluate the data quality and application quality of the enhanced text data through the data testing.
[0058] As a specific implementation of text enhancement, it includes the following steps: (1) Collect industry - specific terms, build an industry - exclusive industry dictionary based on industry - specific data, and filter the professional vocabulary in the industry dictionary; (2) Update the original vocabulary of the large - language model based on the industry dictionary. For the professional data not covered in the original vocabulary, add the professional data to the original vocabulary, and add the embedding vectors corresponding to the professional terms to the embedding matrix of the corresponding original vocabulary of the large - language model; (3) Obtain the text corpus as a text corpus sample, perform unsupervised training on the large - language model based on the text corpus sample, and fine - tune the parameters of the large - language model to obtain a fine - tuned large - language model; (4) For the text corpus to be enhanced, randomly mask some words in the text corpus and record the corresponding label information of the masked words. Use the masked text corpus and the corresponding label information as input attributes. The fine - tuned large - language model establishes the association between the vocabulary and the label information, generates semantic - rich embedding vectors, represents the information of each word through the embedding vectors, predicts the masked words according to the embedding vectors, generates a candidate enhanced vocabulary list, and selects words that match the original semantics and label information from the candidate enhanced vocabulary list as the enhanced text data; (5) Conduct data testing on the enhanced text data, including data quality assessment and application quality assessment. When performing data quality assessment, integrity assessment is carried out by checking whether the data lacks necessary fields and whether there are blank or invalid values. Data correctness is verified by comparing the enhanced text data with the original corpus text and authoritative data in the field. Consistency assessment is conducted by evaluating whether the data format and encoding are unified and whether the logic of the data is coherent in different scenarios. When performing application quality assessment, use the enhanced text data to train the industrial large model. By comparing the accuracy, recall rate, and F1 value of the large models trained using the original data and the enhanced text data on the validation set and the test set, examine the improvement effect of the enhanced text data on the model generalization ability, and conduct manual evaluation in combination with the prediction results of domain experts on the large language model. Let the experts judge the rationality and accuracy of the output of the large language model.
[0059] The text enhancement step of this embodiment uses a pre-trained language model processed by knowledge distillation. With the rich experience and professional knowledge accumulated in the industry, it expands and optimizes the existing data to generate more valuable data samples. As a specific implementation of text enhancement, this step includes the following operations: (1) Widely collect industry-specific terms, build an exclusive industry dictionary, and filter the professional terms in the industry dictionary. Add the professional terms not covered in the original vocabulary of the pre-trained language model into the vocabulary using the "add_tokens" method. Subsequently, use the "resize_token_embeddings" method to add the embedding vectors corresponding to the new words into the embedding matrix, thereby constructing a tokenizer strengthened by industry knowledge; (2) Use the text corpus as input to perform unsupervised training on the pre-trained language model and complete parameter fine-tuning. During this process, the strengthened tokenizer can effectively avoid excessive segmentation of special words, completely retain the original semantics of the words, and greatly improve the fine-tuning efficiency; (3) Use the fine-tuned model to capture the context semantics of the input statement and achieve supervised injection text data augmentation. Set the text data to be enhanced and its corresponding label information as input attributes, randomly mask some words in the statement, and record the corresponding label information. By establishing the association between the words and the labels, the model generates semantic-rich embedding vectors to represent the information of each word. Based on this, predict the masked positions according to the embedding vectors to generate a candidate enhanced vocabulary list, and select the word that best matches the original semantics and label information from it as the enhanced text data; (4) Conduct comprehensive data testing work, mainly including data quality assessment and application quality assessment.
[0060] The data quality assessment is carried out from dimensions such as data integrity, accuracy, and consistency.
[0061] The integrity assessment is conducted by checking whether the data lacks necessary fields and whether there are excessive blanks or invalid values; the accuracy assessment compares the enhanced data with the original data and authoritative data sources in the field to verify the correctness of the data; the consistency assessment focuses on whether the data formats, encodings, etc. are unified, and whether the logic of the data is coherent in different scenarios.
[0062] The application quality assessment mainly examines the impact of data enhancement on the model performance. On the one hand, the enhanced data is used to train the industrial large model. By comparing the accuracy, recall rate, F1 value and other indicators of the models trained with the original data and the enhanced data on the validation set and the test set, the improvement effect of the enhanced data on the model generalization ability is tested. On the other hand, the model prediction results are manually evaluated in combination with domain experts. Representative industry cases are selected, and experts judge the rationality and accuracy of the model output, proving that the enhanced data has good application value in the actual scenario.
[0063] The method of this embodiment provides a practical solution strategy for the high-quality construction of the industry text dataset with a coherent process system. The augmented data obtained based on this method fully relies on professional semantic knowledge, can effectively reduce the generation of noise data, and provides high-quality data support for industry text analysis and model training.
[0064] Embodiment 2: A system for constructing a high-quality text dataset of a large language model according to the present invention includes a data acquisition module, a data processing module, a label definition module, and a text enhancement module.
[0065] The data acquisition module is used to perform the following: collect relevant documents from industry materials to obtain text corpora.
[0066] Among them, the data acquisition module is used to perform the following operations: (1) Construct a source tracking model, the source tracking model is a triple structure including an object, metadata, and a data source. The object represents various objects in the industry that need to be tracked for their sources. The metadata is used to describe the metadata information involved in the acquisition, sharing, and application links of the object. The data source is a tracking operation mechanism constructed based on the metadata source; (2)Divide the metadata information in the source tracking model into basic metadata and derivative metadata. The basic metadata is used to characterize the basic attribute information of industry-related literature and materials, including literature title, literature author, literature abstract, keywords, literature content, literature publication time, journal impact factor, and literature storage path. The derivative metadata is used to record the derivative information generated during the data processing process and related to industry literature analysis and application, including key corpus extracted from original industry literature, supplementary data generated by simulating different industry working conditions, input parameters involved in various industry process optimization models and the optimized results output, as well as the industry process knowledge base formed during the practice and research process. (3)According to the characteristics of the source website of the target literature, select an appropriate web crawler framework, and obtain the target literature by parsing the web page content through the web crawler framework. (4)For literature in the form of paper materials, scan the paper materials based on scanning and optical character recognition technology to obtain relevant text corpus. For literature in the form of network electronic resources, call the corresponding database interface or obtain relevant text corpus through web crawler technology. (5)For the collected text corpus, perform digital fingerprint identification on the text corpus based on the source tracking model to establish the association relationship between literature, text corpus, and derivative data, and record the entire process of literature analysis and application tasks through derivative data, and construct an ordered association between derivative data and corresponding associated processing operations. Among them, the digital fingerprint identification includes the original source website, publication time, author information, and collection time.
[0067] The data collection module accurately captures relevant literature from a vast amount of industry materials by establishing a source tracking model, ensuring that the origin of each piece of data is clearly traceable. As the specific implementation of the data processing module, the following five operations are provided.
[0068] Operation 1: Construct a triple structure <object, metadata, source> of the Source Tracking Model (STM). Among them, the object represents various objects that need to be source-tracked within the injection molding industry; the metadata is used to describe in detail the metadata information involved in the acquisition, analysis, and application links of the object; and the source is the tracking operation mechanism constructed based on the metadata source.
[0069] Operation 2: Divide the metadata information in the source tracking model into two major categories: basic metadata and derivative metadata. Basic metadata is mainly used to describe the basic attribute information of injection molding-related literature and materials, including the literature title, such as "Research on the Influence of Injection Molding Process Parameter Optimization on Product Performance"; the literature author; the literature abstract, which is a refined summary of the core content of the literature; keywords, which accurately reflect the research focus of the literature, such as "injection molding process", "product performance", "parameter optimization", etc.; the specific text content of the literature; the name of the professional journal in which the literature is published; the publication time of the literature; the impact factor of the journal; the DOI number of the literature; the specific storage path of the literature in the storage system, etc. Derivative metadata records the derivative information generated during the data processing process and closely related to the analysis and application of injection molding literature, including the key corpus obtained through processing operations such as screening and refining from the original injection molding literature, such as the core expression regarding the relationship between injection temperature and product forming quality; the supplementary data generated by simulating different injection molding working conditions; the input parameters involved in various injection molding process optimization models and the optimized results output; and the injection molding process knowledge base formed through long-term practice and research.
[0070] Operation 3 Select a crawler framework: According to the characteristics of the source website of the target literature, select a suitable web crawler framework, such as Scrapy for structured crawling and Beautiful Soup for parsing web content. Configure crawling rules: Set the starting crawling URL, target link selection rules, crawling depth, crawling interval, etc. of the crawler. Deploy a crawler cluster: Deploy multiple crawler instances in a distributed environment to ensure that they work together to improve the crawling efficiency and avoid causing excessive access pressure on the target website.
[0071] Operation 4 Data collection and verification: Professional text corpora in the field of injection molding widely exist in various document carriers, such as academic papers, technical reports, patent documents, equipment operation manuals, injection molding process parameter manuals, process flowcharts, industry standards, quality management system documents, etc.
[0072] For the relevant paper materials collected within the enterprise, industry associations, and standard organizations, the scanning combined with Optical Character Recognition (OCR) technology is adopted. A professional high-resolution scanner is used to carefully scan the paper documents to ensure that information such as text and charts is completely and accurately digitized. Then, with the help of advanced OCR recognition software, the scanned image files are precisely converted into editable electronic documents, thus greatly facilitating subsequent corpus processing and analysis work and improving work efficiency and accuracy. Second, for electronic resources on the Internet, the application programming interfaces provided by relevant databases can be used for data retrieval and acquisition, or web crawler technology can be used. A pre-written crawler program, such as Crawl4AI, is used to automatically access and accurately capture the relevant content of the injection molding industry on the web pages to obtain the required corpus.
[0073] Operation Five: Attach a digital fingerprint identifier containing key elements such as the original source website, release time, author information, and collection time to each collected document, thereby establishing the association relationship between the document, text data, and derivative data. Use derivative metadata to record the entire process of document analysis and application tasks, construct an ordered association between derivative data and corresponding processing operations, and ensure the traceability of the text data processing process.
[0074] The data processing module is used to perform the following: preprocess the collected text corpus, perform data cleaning, format conversion, text segmentation, and parsing and marking operations through data preprocessing to obtain pre-annotated text corpus, and output the pre-annotated text corpus in a predetermined format.
[0075] In this embodiment, the data processing module is used to perform the following operations: (1) Check the duplication of the text corpus based on the hash algorithm; (2) Match HTML tags in the text corpus based on regular expressions, replace all matched HTML tags with empty strings, retain the pure content part in the text corpus, identify special characters in the text corpus based on regular expressions, replace special characters with empty strings, and locate and delete the advertising content in the text corpus based on the constructed advertising keyword library to obtain the denoised text corpus; (3) Check the spelling and grammar of the denoised text corpus to obtain the checked text corpus; (4) Convert the format of the checked text corpus from different sources to obtain a text corpus in a unified format of plain text, and store the basic metadata corresponding to each text corpus in the source database; (5) Segment the text corpus. Each paragraph obtained by segmentation constitutes an independent text unit. For each paragraph text unit, disassemble the paragraph into individual sentences based on natural language processing tools; (6) Parse and perform part-of-speech tagging on each sentence to obtain a pre-annotated text corpus; (7) Convert the pre-annotated text corpus into a predetermined format for output. Specifically, for text classification tasks, convert the pre-annotated text corpus into JSON format for output, and store the text corpus and its corresponding label information in the form of key-value pairs. For sentiment analysis tasks, convert the pre-annotated text corpus into XML format for output, and store the pre-annotated text corpus through an XML tree structure.
[0076] Among them, when storing the pre-annotated text corpus through an XML tree structure, create an XML root element, create corresponding child elements and grandchild elements according to the text content and structure of the text corpus, fill the text content into the corresponding element nodes, and add necessary attributes. The necessary attributes include language type and text category to construct a complete XML tree structure.
[0077] The data processing module in this embodiment performs preprocessing such as data cleaning and format conversion on the collected documents based on subsequent specific task requirements, as well as post-processing such as text segmentation.
[0078] As a specific implementation of the data processing module, seven operations including data deduplication, noise information removal, spelling and grammar checking, format conversion, text segmentation, parsing and annotation, and data output are provided.
[0079] Data deduplication: Use the hash algorithm to check for duplicates in the obtained corpus to ensure the uniqueness of each piece of literature data.
[0080] Noise information removal: Use the regular expression "<.*?>" to match HTML tags in the text. This regular expression accurately identifies HTML tags by searching for strings starting with "<" and ending with ">". During text processing, replace all matched HTML tags with an empty string, thereby retaining the pure content part of the text and removing, for example, ”" Implement preliminary purification of the text by using various HTML tags such as "etc."; Use the regular expression "[^\a-zA-Z0-9\s.,!?-]" to identify special characters in the text other than letters, numbers, and common punctuation marks (including spaces, periods, commas, exclamation marks, question marks, hyphens, etc.). These special characters often interfere with subsequent text analysis tasks. Through string replacement operations, they are all replaced with empty strings and removed from the text, leaving only meaningful text content. For example, remove special symbols such as "@", "#", "$" from the text; Construct an advertising keyword library containing common advertising words and phrases, such as "Get for free", "Limited-time discount", "Click to get a gift", etc. During text processing, use a string matching algorithm to search for these keywords one by one in the text. Once advertising content is found in the text, it is located and deleted to ensure that the processed text does not contain any advertising-related information, thereby improving the quality and relevance of the text.
[0081] Spelling and grammar checking: For possible spelling mistakes of professional terms or errors caused by inaccurate OCR recognition in the denoised corpus, use natural language processing tools such as NLTK (Natural Language Toolkit) or StanfordCoreNLP to perform spelling and grammar checking.
[0082] Format conversion: Uniformly convert the corpus from different sources to the UTF-8 encoding format and convert it into TXT plain text data. Store the basic metadata such as the title, author, keywords, DOI number, etc. of each preprocessed literature corpus in the metadata database to ensure the traceability of the source of the literature corpus.
[0083] Text segmentation: Considering that different paragraphs in the literature in the injection molding field elaborate on different links of the injection molding process, different dimensions of mold design, etc., first segment the abstract and text content in the preprocessed literature corpus according to paragraphs, so that each paragraph forms an independent text unit. Then, for each divided paragraph, use natural language processing tools, such as the Punkt sentence splitter, to further disassemble it into individual sentences. This method can accurately label specific semantic units during the annotation process, thereby improving the accuracy and consistency of the annotation.
[0084] For each divided paragraph, design sentence boundary recognition models in different languages. According to the grammar rules and punctuation habits of the language, accurately identify the start and end positions of sentences, and use natural language processing tools, such as the sentence splitters in NLTK or SpaCy or the Punkt sentence splitter, to further disassemble it into individual sentences; Taking NLTK as an example, after loading the English sentence splitter, call its tokenize.sent_tokenize() function and pass in the text to be processed. This function will split the text into individual sentences based on punctuation marks such as the period (.), question mark (?), exclamation point (!) in English, and return a list containing all the sentences. For Chinese text, corresponding Chinese sentence segmentation models or rules can be used for processing, such as splitting sentences based on Chinese punctuation marks like the period (。), question mark (?), exclamation point (!).
[0085] Parse and annotate: Parse and perform part-of-speech tagging on each segmented sentence to obtain a fine-grained pre-annotated text corpus.
[0086] Data output: Text format conversion: Determine the target text format according to the specific requirements of subsequent tasks.
[0087] For text classification tasks, select the JSON format and store information such as the text and its corresponding classification labels in the form of key-value pairs. Complete the conversion with the help of the json library in Python. Integrate the processed plain text and its related metadata (such as document unique identifier, source information, preprocessing timestamp, etc.) into a dictionary object, and then use the dumps method of the json library to convert the dictionary into a JSON-formatted string to achieve the JSON format encapsulation of the text; For sentiment analysis tasks, select the XML format and use the tag features of XML to annotate the sentiment tendency of the text (such as positive, negative, neutral, etc.). Use the xml.etree.ElementTree library for conversion. First, create an XML root element, then create corresponding child elements and grandchild elements according to the text content and structure, fill the text content into the corresponding element nodes, and add necessary attributes (such as language type, text category, etc.) to build a complete XML tree structure, thus obtaining the XML format text that meets the requirements.
[0088] The label definition module is used to perform the following: Build a label system based on industry-specific terms and metrics, and annotate the entities and the relationships between entities in the pre-annotated text corpus. The obtained entity labels and the relationships between entities are used as label information.
[0089] As a specific implementation of the label definition module, this module is used to provide the following operations: (1) Build an entity label system: Based on the analysis results of industry characteristics, determine the scope of entity types, set detailed attributes for each entity type, and build a multi-level entity label structure according to the level of detail of the entities and business semantics to ensure that the entity labels can accurately describe various objects in the data; (2) Construct an entity relationship label system: Analyze the interaction methods between entities in the industry business process and typical scenarios, identify the types of relationships between entities, determine the participants and participation methods of entity relationships based on business logic and data flow, set corresponding attributes for each entity relationship, standardize the directionality of entity relationships, and clarify the main entity and the guest entity; (3) Use the annotation format of "Entity 1 - Relationship Type - Entity 2" to annotate the potential relationships between two entities, where the relationship type follows the predefined entity relationship types.
[0090] In this embodiment, the label definition module formulates a special label system based on industry-specific terms, key indicators, etc., accurately annotates the data, and accurately marks the entity information and their mutual relationships. It specifically includes two operations: constructing an entity label system and constructing an entity relationship label system.
[0091] Construct an entity label system: Based on the analysis results of industry characteristics, determine the scope of entity types, and set detailed attributes for each entity type. According to the level of detail of the entity and business semantics, construct a multi-level entity label structure to ensure that entity labels can precisely describe various objects in the data.
[0092] Construct an entity relationship label system: Analyze the interaction methods between entities in the industry business process and typical scenarios, identify common relationship types, determine the participants and participation methods of the relationships based on business logic and data flow. Set corresponding attributes for each entity relationship, standardize the directionality of the relationships, and clarify the main entity and the guest entity.
[0093] The text enhancement module is used to perform the following: Construct an industry dictionary based on the collected industry data, update the original vocabulary of the pre-trained large language model based on the industry dictionary, use the collected text corpus samples as input to perform unsupervised training on the pre-trained large language model to obtain a fine-tuned large language model. For the text corpus to be enhanced, randomly mask some words in the text corpus, use the masked text corpus and the corresponding label information as input attributes, perform upper and lower semantic analysis on the input attributes through the unsupervised trained large model to predict the masked words in the text corpus, output the enhanced text data, and perform data testing on the text data. Through data testing, perform data quality assessment and application quality assessment on the enhanced text data.
[0094] As a specific implementation of the text enhancement module, the following operations are provided: (1) Collect industry-specific terms, construct an industry-specific industry dictionary based on industry professional data, and filter the professional vocabulary in the industry dictionary; (2) Update the original vocabulary of the large language model based on the industry dictionary. For professional data not covered in the original vocabulary, add the professional data to the original vocabulary and add the embedding vectors corresponding to the professional terms to the embedding matrix of the large language model corresponding to the original vocabulary; (3) Obtain text corpus as text corpus samples, perform unsupervised training on the large language model based on the text corpus samples, and fine-tune the parameters of the large language model to obtain the fine-tuned large language model; (4) For the text corpus to be enhanced, randomly mask some words in the text corpus and record the corresponding label information of the masked words. Use the masked text corpus and the corresponding label information as input attributes. The fine-tuned large language model establishes the association between vocabulary and label information, generates rich semantic embedding vectors, represents the information of each word through the embedding vectors, and predicts the masked words based on the embedding vectors to generate a candidate enhanced vocabulary, and selects words that match the original semantics and label information from the candidate enhanced vocabulary as the enhanced text data; (5) Conduct data testing on the enhanced text data, including data quality assessment and application quality assessment. When performing data quality assessment, conduct integrity assessment by checking whether the data lacks necessary fields and whether there are blanks or invalid values, verify the data correctness by comparing the enhanced text data with the original corpus text and authoritative data in the field, and conduct consistency assessment by evaluating whether the data format and encoding are unified and whether the logic of the data is coherent in different scenarios. When performing application quality assessment, use the enhanced text data to train an industrial large model, and by comparing the accuracy, recall rate, and F1 value of the large models trained using the original data and the enhanced text data on the validation set and the test set, test the improvement effect of the enhanced text data on the model generalization ability, and combine the prediction results of domain experts on the large language model for manual evaluation, and let the experts judge the rationality and accuracy of the output of the large language model.
[0095] The text enhancement module in this embodiment uses a pre-trained language model processed by knowledge distillation, and with the help of the rich experience and professional knowledge accumulated in the industry, expands and optimizes the existing data to generate more valuable data samples. As a specific implementation of text enhancement, this module provides the following operations: (1)Collect industry-specific terms extensively, build an exclusive industry dictionary, and filter the professional terms in the industry dictionary. Add the professional terms not covered in the original vocabulary of the pre-trained language model into the vocabulary by means of the "add_tokens" method. Subsequently, use the "resize_token_embeddings" method to add the embedding vectors corresponding to the new words into the embedding matrix, thus constructing a tokenizer enhanced by industry knowledge; (2)Use the text corpus as input to conduct unsupervised training on the pre-trained language model and complete parameter fine-tuning. During this process, the enhanced tokenizer can effectively avoid over-segmenting special words, completely retain the original semantics of the words, and greatly improve the fine-tuning efficiency; (3)Use the fine-tuned model to capture the context semantics of the input statement and achieve supervised injection text data augmentation. Set the text data to be enhanced and its corresponding label information as input attributes, randomly mask some words in the statement, and record the corresponding label information. By establishing the association between the words and the labels, the model generates semantic-rich embedding vectors to represent the information of each word. Based on this, predict the masked positions according to the embedding vectors to generate a candidate enhanced vocabulary list, and select the word that best matches the original semantics and label information from it as the enhanced text data; (4)Conduct comprehensive data testing work, mainly including data quality assessment and application quality assessment.
[0096] The data quality assessment starts from dimensions such as data integrity, accuracy, and consistency.
[0097] The integrity assessment is carried out by checking whether the data lacks necessary fields and whether there are too many blanks or invalid values; the accuracy assessment compares the enhanced data with the original data and authoritative data sources in the field to verify the correctness of the data; the consistency assessment focuses on whether the data formats, encodings, etc. are unified, and whether the logic of the data is coherent in different scenarios.
[0098] The application quality assessment mainly examines the impact of data augmentation on the model performance. On the one hand, use the enhanced data to train the industrial large model, and test the improvement effect of the enhanced data on the model generalization ability by comparing the accuracy, recall rate, F1 value and other indicators of the models trained with the original data and the enhanced data on the validation set and the test set. On the other hand, combine domain experts to conduct manual evaluation on the model prediction results, select representative industry cases, and let the experts judge the rationality and accuracy of the model output, proving that the enhanced data has good application value in the actual scenario.
[0099] The system of this embodiment can execute the method disclosed in Embodiment 1 to construct a high-quality text data set.
[0100] The above has introduced in detail the method and system for constructing a high-quality text dataset of large language models. In this article, specific examples are used to elaborate on the principle and implementation mode of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for constructing a high-quality text dataset of large language models, characterized in that It includes the following steps: Data collection: Collect relevant literature from industry materials to obtain text corpora; Data processing: Perform data preprocessing on the collected text corpora. Through data preprocessing, operations such as data cleaning, format conversion, text segmentation, and parsing marking are carried out to obtain pre-annotated text corpora, and the pre-annotated text corpora are output in a predetermined format; Label definition: Construct a label system based on industry-specific terms and indicators, and label the entities and the relationships between entities in the pre-annotated text corpora based on the label system. The obtained entity labels and the relationships between entities are used as label information; Text enhancement: Construct an industry dictionary based on the collected industry data, update the original vocabulary of the pre-trained large language model based on the industry dictionary. Use the collected text corpus samples as input to perform unsupervised training on the pre-trained large language model to obtain a fine-tuned large language model. For the text corpora to be enhanced, randomly mask some words in the text corpora. Use the masked text corpora and the corresponding label information as input attributes. Through the unsupervised trained large model, perform up and down semantic analysis on the input attributes, predict the masked words in the text corpora, output enhanced text data, and perform data testing on the text data. Through data testing, perform data quality assessment and application quality assessment on the enhanced text data; 2. The method for constructing a high-quality text dataset of a large language model according to claim 1, characterized in that Data collection includes the following steps: Construct a source tracking model. The source tracking model is a triple structure including an object, metadata, and a data source. The object represents various objects in the industry that need to be tracked for their sources. Metadata is used to describe the metadata information involved in the acquisition, sharing, and application of the object. The data source is a tracking operation mechanism constructed based on the metadata source; Divide the metadata information in the source tracking model into basic metadata and derivative metadata. The basic metadata is used to describe the basic attribute information of industry-related literature and materials, including literature titles, literature authors, literature abstracts, keywords, literature content, literature publication times, journal impact factors, and literature storage paths. The derivative metadata is used to record the derivative information generated during the data processing process and related to industry literature analysis and application, including key corpora extracted from original industry literature, supplementary data generated by simulating different industry working conditions, input parameters involved in various industry process optimization models and the optimized results output, as well as the industry process knowledge base formed during the practice and research process; According to the characteristics of the source website of the target literature, select an appropriate web crawler framework, and parse the web page content through the web crawler framework to obtain the target literature; For literature in the form of paper materials, scan the paper materials based on scanning and optical character recognition technology to obtain relevant text corpora. For literature in the form of network electronic resources, call the corresponding database interface or obtain relevant text corpora through web crawler technology; For the collected text corpus, digital fingerprint identification is performed on the text corpus based on the source tracking model to establish the association relationship among documents, text corpus, and derivative data, and the entire process of document analysis and application tasks is recorded through derivative data, and an ordered association between derivative data and corresponding associated processing operations is constructed. Among them, the digital fingerprint identification includes the original source website, release time, author information, and collection time.
3. The method for constructing a high-quality text dataset of a large language model according to claim 1, wherein Data processing includes the following steps: Perform duplicate checking on the text corpus based on the hash algorithm; Match HTML tags in the text corpus based on regular expressions, replace all matched HTML tags with empty strings, retain the pure content part in the text corpus, identify special characters in the text corpus based on regular expressions, replace special characters with empty strings, and locate and delete advertising content in the text corpus based on the constructed advertising keyword library to obtain the denoised text corpus; Perform spelling and grammar checks on the denoised text corpus to obtain the checked text corpus; Convert the checked text corpus from different sources into a unified format of plain text, and store the basic metadata corresponding to each text corpus in the source database; Segment the text corpus, and each paragraph obtained by segmentation constitutes an independent text unit. For each paragraph text unit, disassemble the paragraph into individual sentences based on natural language processing tools; Parse and perform part-of-speech tagging on each sentence to obtain the pre-annotated text corpus; Convert the pre-annotated text corpus into a predetermined format for output. Among them, for text classification tasks, convert the pre-annotated text corpus into JSON format for output, and store the text corpus and corresponding label information in the form of key-value pairs. For sentiment analysis tasks, convert the pre-annotated text corpus into XML format for output, and store the pre-annotated text corpus through the XML tree structure; Among them, storing the pre-annotated text corpus through the XML tree structure includes the following operations: Create an XML root element; Create corresponding sub-elements and grandchild elements according to the text content and structure of the text corpus, fill the text content into the corresponding element nodes, and add necessary attributes. The necessary attributes include language type and text category to construct a complete XML tree structure.
4. The method for constructing a high-quality text dataset of a large language model according to claim 1, characterized in that, Label definition includes the following operations: Construct an entity label system: Based on the analysis results of industry characteristics, determine the scope of entity types, set detailed attributes for each entity type, and construct a multi-level entity label structure according to the detailed degree of entities and business semantics to ensure that entity labels can finely describe various objects in the data; Construct an entity relationship label system: Analyze the interaction methods between entities in industry business processes and typical scenarios, identify the relationship types between entities, determine the participants and participation methods of entity relationships based on business logic and data flow, set corresponding attributes for each entity relationship, and standardize the directionality of entity relationships to clarify the main entity and the guest entity; Adopt the annotation format of "entity 1 - relationship type - entity 2" to annotate the potential relationship between two entities, where the relationship type follows the predefined entity relationship type.
5. The method for constructing a high-quality text dataset of a large language model according to claim 1, wherein Text enhancement includes the following steps: Collect industry-specific terms, build an industry-exclusive dictionary based on industry-specific data, and filter the professional terms in the industry dictionary; Update the original vocabulary of the large language model based on the industry dictionary. For the professional data not covered in the original vocabulary, add the professional data to the original vocabulary, and add the embedding vectors corresponding to the professional terms to the embedding matrix of the corresponding original vocabulary of the large language model; Obtain text corpus as text corpus samples, perform unsupervised training on the large language model based on the text corpus samples, and fine-tune the parameters of the large language model to obtain a fine-tuned large language model; For the text corpus to be enhanced, randomly mask some words in the text corpus, and record the corresponding label information of the masked words. Use the masked text corpus and the corresponding label information as input attributes. The fine-tuned large language model establishes the association between words and label information, generates rich semantic embedding vectors, represents the information of each word through the embedding vectors, and predicts the masked words according to the embedding vectors to generate a candidate enhanced vocabulary list, and selects words that match the original semantics and label information from the candidate enhanced vocabulary list as the enhanced text data; Conduct data testing on the enhanced text data, including data quality assessment and application quality assessment. When performing data quality assessment, conduct integrity assessment by checking whether the data lacks necessary fields and whether there are blanks or invalid values, verify the data correctness by comparing the enhanced text data with the original corpus text and authoritative data in the field, and conduct consistency assessment by evaluating whether the data format and encoding are unified and whether the logic of the data is coherent in different scenarios. When performing application quality assessment, use the enhanced text data to train an industrial large model, and by comparing the accuracy, recall rate, and F1 value of the large models trained using the original data and the enhanced text data on the validation set and the test set, examine the improvement effect of the enhanced text data on the model generalization ability, and combine the prediction results of the domain experts on the large language model for manual evaluation, and let the experts judge the rationality and accuracy of the output of the large language model.
6. A high-quality text dataset construction system for large language models, characterized in that, For constructing a high-quality text data set through the method for constructing a high-quality text data set of a large language model described in any one of claims 1-5, the system includes a data acquisition module, a data processing module, a label definition module, and a text enhancement module; The data acquisition module is used to perform the following: collect relevant documents from industry materials to obtain text corpus; The data processing module is used to perform the following: perform data preprocessing on the collected text corpus, perform data cleaning, format conversion, text cutting, and parsing and marking operations through data preprocessing to obtain pre-annotated text corpus, and output the pre-annotated text corpus in a predetermined format; The label definition module is used to perform the following: build a label system based on industry-specific terms and metrics, and label the entities and the relationships between entities in the pre-annotated text corpus based on the label system, and use the obtained entity labels and the relationships between entities as label information; The text enhancement module is used to perform the following: construct an industry dictionary based on the collected industry data, update the original vocabulary of the pre-trained large language model based on the industry dictionary, use the collected text corpus samples as input to perform unsupervised training on the pre-trained large language model to obtain a fine-tuned large language model. For the text corpus to be enhanced, randomly mask some words in the text corpus, use the masked text corpus and the corresponding label information as input attributes, perform upper and lower semantic analysis on the input attributes through the large model after unsupervised training, predict the masked words in the text corpus, output the enhanced text data, and perform data testing on the text data, and perform data quality evaluation and application quality evaluation on the enhanced text data through data testing.
7. The large language model high-quality text dataset construction system according to claim 6, wherein The data collection module is used to perform the following: Construct a source tracking model, where the source tracking model is a triple structure including an object, metadata, and a data source. The object represents various objects in the industry that need to be tracked for their sources. The metadata is used to describe the metadata information involved in the acquisition, sharing, and application of the object. The data source is a tracking operation mechanism constructed based on the metadata source; Divide the metadata information in the source tracking model into basic metadata and derivative metadata. The basic metadata is used to describe the basic attribute information of industry-related literature and materials, including the literature title, literature author, literature abstract, keywords, literature content, literature publication time, journal impact factor, and literature storage path. The derivative metadata is used to record the derivative information generated during the processing of data and related to industry literature analysis and application, including key corpus extracted from the original industry literature, supplementary data generated by simulating different industry working conditions, input parameters and output optimization results involved in various industry process optimization models, and the industry process knowledge base accumulated during the practice and research process; According to the characteristics of the source website of the target literature, select an appropriate web crawler framework, and parse the web page content through the web crawler framework to obtain the target literature; For literature in the form of paper materials, scan the paper materials based on scanning and optical character recognition technology to obtain relevant text corpus. For literature in the form of network electronic resources, call the corresponding database interface or obtain relevant text corpus through web crawler technology; For the collected text corpus, perform digital fingerprint identification on the text corpus based on the source tracking model to establish the association relationship between the literature, the text corpus, and the derivative data, and record the entire process of literature analysis and application tasks through the derivative data, and construct an ordered association between the derivative data and the corresponding associated processing operations. Among them, the digital fingerprint identification includes the original source website, publication time, author information, and collection time.
8. The large language model high-quality text dataset construction system according to claim 6, characterized in that, The data processing module is used to perform the following: Perform duplicate checking on the text corpus based on the hash algorithm; Match HTML tags in the text corpus based on regular expressions, replace all matched HTML tags with empty strings, retain the pure content part in the text corpus, identify special characters in the text corpus based on regular expressions, replace the special characters with empty strings, and locate and delete the advertising content in the text corpus based on the constructed advertising keyword library to obtain the denoised text corpus; Perform spelling and grammar checks on the denoised text corpus to obtain the text corpus after checking; Convert the text corpora after checking from different sources into a text corpus in a unified format of plain text, and store the corresponding basic metadata of each text corpus in the source database; Segment the text corpus, and each paragraph obtained by segmentation constitutes an independent text unit. For each text unit of a paragraph, disassemble the paragraph into individual sentences based on natural language processing tools; Parse and perform part-of-speech tagging on each sentence to obtain the pre-annotated text corpus; Convert the pre-annotated text corpus into a predetermined format for output. Among them, for text classification tasks, convert the pre-annotated text corpus into JSON format for output, and store the text corpus and the corresponding tag information in the form of key-value pairs. For sentiment analysis tasks, convert the pre-annotated text corpus into XML format for output, and store the pre-annotated text corpus through the XML tree structure; Among them, when storing the pre-annotated text corpus through the XML tree structure, the data processing module is used to perform the following operations: Create an XML root element; Create corresponding sub-elements and grandchild elements according to the text content and structure of the text corpus, fill the text content into the corresponding element nodes, and add necessary attributes. The necessary attributes include language type and text category to build a complete XML tree structure.
9. The large language model high-quality text dataset construction system according to claim 6, characterized in that, The tag definition module is used to perform the following operations: Construct an entity tag system: Based on the analysis results of industry characteristics, determine the scope of entity types, set detailed attributes for each entity type, and build a multi-level entity tag structure according to the detailed degree of entities and business semantics to ensure that the entity tags can precisely describe various objects in the data; Construct an entity relationship tag system: Analyze the interaction methods between entities in the industry business process and typical scenarios, identify the relationship types between entities, determine the participants and participation methods of entity relationships according to business logic and data flow, set corresponding attributes for each entity relationship, and standardize the directionality of entity relationships to clarify the main entity and the guest entity; Use the annotation format of "entity 1 - relationship type - entity 2" to annotate the potential relationship between two entities, where the relationship type follows the pre-defined entity relationship type.
10. The large language model high-quality text dataset construction system according to claim 6, characterized in that The text enhancement module is used to perform the following operations: Collect industry-specific terms, build an industry-specific industry dictionary based on industry-specific data, and filter the professional vocabulary in the industry dictionary; Update the original vocabulary of the large language model based on the industry dictionary. For the professional data not covered in the original vocabulary, add the embedding vector corresponding to the professional term to the original vocabulary of the large language model; Obtain a text corpus as a text corpus sample, perform unsupervised training on a large language model based on the text corpus sample, and fine-tune the parameters of the large language model to obtain a fine-tuned large language model; For the text corpus to be enhanced, randomly mask some words in the text corpus and record the corresponding label information of the masked words. Use the masked text corpus and the corresponding label information as input attributes. The fine-tuned large language model establishes the association between words and label information, generates semantically rich embedding vectors, represents the information of each word through the embedding vectors, predicts the masked words based on the embedding vectors, generates a candidate enhanced vocabulary list, and selects words that match the original semantics and label information from the candidate enhanced vocabulary list as the enhanced text data; Conduct data testing on the enhanced text data, including data quality assessment and application quality assessment. When performing data quality assessment, conduct integrity assessment by checking whether the data lacks necessary fields and whether there are blanks or invalid values, verify the data correctness by comparing the enhanced text data with the original corpus text and authoritative data in the field, and conduct consistency assessment by evaluating whether the data format and encoding are unified and whether the logic of the data is coherent in different scenarios. When performing application quality assessment, use the enhanced text data to train an industrial large model. By comparing the accuracy, recall rate, and F1 value of the large models trained using the original data and the enhanced text data on the validation set and the test set, examine the improvement effect of the enhanced text data on the model generalization ability, and combine the prediction results of domain experts on the large language model for manual evaluation. Let the experts judge the rationality and accuracy of the output of the large language model.
Citation Information
Patent Citations
Multi-label inspection work order problem traceability identification method and device
CN113868422A
Address named entity recognition tuning method based on deep learning model
CN114169332A
Entity tag attribute identification method, apparatus and device, and storage medium
CN116069932A
Electromagnetic space domain entity relation joint extraction method with additional time information
CN116911296A
NLP-based industry data analysis method and system
CN118070812A
Cited By
Large-scale high-speed text training comparison data set production device
CN121278393A
Data enhancement and generalization method and system for vertical large model in insurance field
CN121786191A
High-quality metal material process data set construction method based on large language model
CN121789818A