Method and device for converting clipboard data into label data based on specific standard
Through the automated method of converting clipboard data into label data, the problems of low data conversion efficiency and error-prone in the prior art are solved, efficient and accurate data structured processing is achieved, and label data that meets the standards is generated.
Patent Information
- Application Number
- CN202510165753.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is inefficient and error-prone when converting unstructured data into structured data, and requires manual copy-paste or verbatim input, which cannot meet the needs of efficient automation.
By obtaining the source data of clipboard data, determining the target files of international or industry standards, extracting key features and topic information, using semantic analysis and vector similarity to calculate matching scene rules files, automatically parsing and converting them into label data that complies with the standards, and generating structured XML documents.
It significantly improves data conversion efficiency and accuracy, reduces manual intervention, ensures data consistency and compliance with predetermined standards, and improves the intelligence and automation level of data processing.
Smart Images

Figure CN120257942A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and specifically relates to a method and device for converting clipboard data into tag data based on specific standards. Background Art
[0002] On the one hand, with the popularization of information systems, the management mode of paper technical materials has been gradually phased out, and the structuring of technical materials has become a trend. On the other hand, with the growth of market demand, the structured data management software market is also constantly maturing, providing various functions, including data migration, data quality management, data governance, and compliance management, etc., to help enterprises better utilize their structured data assets.
[0003] In the process of managing structured data, a large amount of data conversion work is faced. The whole lines, paragraphs, and pages of data in the existing files are input into the target file to achieve data structuring. Such operations usually require defining the corresponding tags in the target file first, and then copying the data out one by one from the source file and pasting it under the corresponding tags, or inputting it word by word in the target file. This traditional management mode requires a large amount of repetitive operations, with a large workload and is extremely error-prone. Therefore, a method is needed to improve the efficiency of data conversion. Summary of the Invention
[0004] This application provides a method and device for converting clipboard data into tag data based on specific standards, which can improve the efficiency of data conversion.
[0005] In the first aspect of this application, a method for converting clipboard data into tag data based on specific standards is provided. The method includes:
[0006] Obtain the source data of the clipboard data;
[0007] Determine a target file based on international standards or industry standards;
[0008] Determine the data standard according to the target file;
[0009] Analyze the source data and convert the source data according to the data standard to obtain tag data.
[0010] Based on the above technical solutions, preferably, the determination of the target file based on international standards or industry standards specifically includes:
[0011] Extract the key features and theme information of the source data;
[0012] Determine the scenario rule file corresponding to the source data according to the key features and the subject information, where there is a second key description in the source data that meets the preset similarity condition with the first key description, and the scenario rule file includes the first key description;
[0013] Perform semantic analysis on the source data to check whether the source data contains a target data module, where the target data module is any one of the multiple data modules included in the target data standard, and the target data standard is any one of the multiple data standards included in the scenario rule file;
[0014] If it is determined that the source data contains the target data module, determine the rule file corresponding to the target data standard as the target file.
[0015] On the basis of the above technical solutions, preferably, the determining the scenario rule file corresponding to the source data according to the key features and the subject information specifically includes:
[0016] Determine the first vector of the first key description;
[0017] Determine the second vectors of the key descriptions included in each training data during the model training phase;
[0018] Calculate the vector similarity between the first vector and the second vectors respectively;
[0019] Determine the target similarity that meets the preset similarity condition among the multiple vector similarities, and determine the target training data corresponding to the target similarity;
[0020] Determine the rule file corresponding to the target training data as the scenario rule file.
[0021] On the basis of the above technical solutions, preferably, analyze the source data and convert the source data according to the data standard to obtain labeled data, specifically including:
[0022] Parse the source data item by item, and convert each content of the source data into a target label, where each target label is attached with a unique attribute;
[0023] Determine the hierarchical structure and label relationship according to the rules of the data standard, including determining the hierarchical structure and order among the multiple target labels, and each attribute follows the generation standard and is assigned by a random number;
[0024] Generate the labeled data that meets the data standard and output it in the XML document format with a structured data format.
[0025] Based on the above technical solutions, preferably, after analyzing the source data and converting the source data according to the data standard to obtain labeled data, the method further includes:
[0026] Obtain an XSD file that defines the structure and data type of the XML document;
[0027] Load the XSD file through an XSD validator and read the XML document;
[0028] Check one by one whether the tags, attributes, data types, and hierarchical structures in the XML file conform to the XSD rules included in the XSD file;
[0029] If multiple checks for the XML file conform to the XSD rules, the verification for the labeled data passes.
[0030] Based on the above technical solutions, preferably, before analyzing the source data and converting the source data according to the data standard to obtain labeled data, the method further includes:
[0031] Obtain clipboard data in different formats and contents;
[0032] Obtain pre-labeled pre-stored labeled data corresponding to the clipboard data;
[0033] Identify the correspondence between the clipboard data and the pre-stored labeled data, and adjust the model parameters based on the evaluation result for the correspondence;
[0034] Use independent validation sets and test sets to evaluate the recognition accuracy and generalization ability, and perform cross-validation, where the validation set includes clipboard data and pre-stored labeled data, and the test set includes clipboard data and pre-stored labeled data.
[0035] Based on the above technical solutions, preferably, determining the target similarity that satisfies the preset similarity condition among the multiple vector similarities specifically includes:
[0036] Judge whether the target similarity is greater than or equal to the preset threshold. If it is determined that the target similarity is greater than or equal to the preset threshold, it is determined that the target similarity satisfies the preset similarity condition; or,
[0037] Sort the multiple vector similarities in ascending order of numerical values, and based on the sorting result, determine the vector similarity with the largest numerical value among the multiple vector similarities to obtain the target similarity.
[0038] In the second aspect of the present application, a device for converting clipboard data into tag data based on specific standards is provided. The device includes an acquisition module, a processing module, and an output module, where:
[0039] The acquisition module is configured to acquire the source data of the clipboard data;
[0040] The processing module is configured to determine a target file based on international standards or industry standards;
[0041] The processing module is configured to determine a data standard according to the target file;
[0042] The output module is configured to analyze the source data and convert the source data according to the data standard to obtain tag data.
[0043] Based on the above technical solutions, preferably, the processing module is configured to extract the key features and theme information of the source data;
[0044] The processing module is configured to determine a scenario rule file corresponding to the source data according to the key features and the theme information, where there is a second key description in the source data that satisfies a preset similarity condition with a first key description, and the scenario rule file includes the first key description;
[0045] The processing module is configured to perform semantic analysis on the source data to check whether the source data includes a target data module, where the target data module is any one of a plurality of data modules included in a target data standard, and the target data standard is any one of a plurality of data standards included in the scenario rule file;
[0046] The processing module is configured to, if it is determined that the source data includes the target data module, determine the rule file corresponding to the target data standard as the target file.
[0047] Based on the above technical solutions, preferably, the processing module is configured to determine a first vector of the first key description;
[0048] The processing module is configured to determine second vectors of key descriptions included in each training data during the model training phase;
[0049] The processing module is configured to calculate the vector similarity between the first vector and the second vectors respectively;
[0050] The processing module is configured to determine a target similarity that satisfies a preset similarity condition among the multiple vector similarities and determine the target training data corresponding to the target similarity;
[0051] The processing module is used to determine that the rule file corresponding to the target training data is the scenario rule file.
[0052] Based on the above technical solutions, preferably, the processing module is used to parse the source data item by item, convert each content of the source data into a target label, and each of the target labels is attached with a unique attribute.
[0053] The processing module is used to determine the hierarchical structure and label relationship according to the rules of the data standard, including determining the hierarchical structure and order among multiple target labels, and each of the attributes follows a generation standard and is assigned a value by a random number.
[0054] The processing module is used to generate the label data that conforms to the data standard and output an XML document with a structured data format.
[0055] Based on the above technical solutions, preferably, the acquisition module is used to acquire an XSD file that defines the structure and data type of the XML document.
[0056] The acquisition module is used to load the XSD file through an XSD validator and read the XML document.
[0057] The processing module is used to check item by item whether the labels, attributes, data types, and hierarchical structures in the XML file conform to the XSD rules included in the XSD file.
[0058] The processing module is used to determine that the verification of the label data passes if multiple checks on the XML file conform to the XSD rules.
[0059] Based on the above technical solutions, preferably, the acquisition module is used to acquire clipboard data in different formats and contents.
[0060] The acquisition module is used to acquire pre-annotated pre-stored label data corresponding to the clipboard data.
[0061] The processing module is used to identify the corresponding relationship between the clipboard data and the pre-stored label data, and adjust the model parameters based on the evaluation result of the corresponding relationship.
[0062] The processing module is used to use independent validation sets and test sets to evaluate the recognition accuracy and generalization ability, and through cross-validation, where the validation set includes clipboard data and pre-stored label data, and the test set includes clipboard data and pre-stored label data.
[0063] Based on the above technical solutions, preferably, the processing module is configured to determine whether the target similarity is greater than or equal to a preset threshold. If it is determined that the target similarity is greater than or equal to the preset threshold, it is determined that the target similarity meets the preset similarity condition; or,
[0064] The processing module is configured to sort the multiple vector similarities in ascending order of numerical value, and determine the vector similarity with the largest numerical value among the multiple vector similarities according to the sorting result, so as to obtain the target similarity.
[0065] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is configured to execute the instructions stored in the memory, so that the electronic device executes the method described in any one of the above.
[0066] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed.
[0067] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0068] 1. In the present application, by automatically converting clipboard data into tag data that conforms to specific international standards or industry standards, the data conversion efficiency is significantly improved. By automatically parsing and converting the source data into standardized tag data, manual processing and repetitive labor are avoided. At the same time, the data structure is automatically matched and verified according to the preset target file and data standard, ensuring accuracy and consistency. This automated processing flow not only reduces manual intervention but also speeds up the data processing speed and improves the efficiency of large-scale data conversion tasks.
[0069] 2. Automatically extract the key features and theme information of the source data, and combine similarity analysis and semantic understanding to accurately determine the matching relationship between the source data and the scenario rule file, and then automatically identify and map the target data module. In this way, the applicable data standard and rule file can be automatically selected, reducing manual intervention and errors, improving the accuracy and efficiency of data conversion, and ensuring that the generated data conforms to the predetermined standard and structure, thereby greatly improving the intelligence and automation level of data processing.
[0070] 3. Convert the key descriptions of the source data into vector representations, and use vector similarity calculation to match with the descriptions in the training data, realizing the automatic recognition and selection of the most suitable scenario rule file. This method based on vector space and similarity analysis can improve the accuracy and efficiency of the matching between the source data and the rule file.
[0071] 4. Parse the source data item by item and convert it into target tags that conform to the data standard, automatically generate structured tag data that conforms to the specification, and determine the hierarchical structure and order of the tags according to the standard rules. Each tag is attached with unique attributes to ensure the consistency and accuracy of the data. Significantly improve the automation degree and accuracy of data conversion, ensure that the generated tag data conforms to the predetermined standard, and output it in a structured XML format for convenient subsequent data processing and use. Brief Description of the Drawings
[0072] Figure 1 is a schematic flowchart of a method for converting clipboard data into tag data based on a specific standard disclosed in an embodiment of the present application;
[0073] Figure 2 is a schematic diagram of tag data generated from source data disclosed in an embodiment of the present application;
[0074] Figure 3 is a schematic diagram of modules of a device for converting clipboard data into tag data based on a specific standard disclosed in an embodiment of the present application;
[0075] Figure 4 is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application.
[0076] Description of the Reference Numerals: 301, acquisition module; 302, processing module; 303, output module; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. Detailed Embodiments
[0077] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0078] In the description of the embodiments of the present application, words such as "for example" or "for instance" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "for example" or "for instance" is intended to present relevant concepts in a specific manner.
[0079] In the description of the embodiments of the present application, the term "plurality" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0080] With the popularization of information systems, the traditional paper-based technical data management mode has gradually been replaced by structured data management. The growth of market demand has also promoted the maturity of related software, providing functions such as data migration, data quality management, and data governance. However, in the process of managing structured data, a large amount of data conversion work is still faced, and unstructured content needs to be gradually sorted into target files. The traditional method relies on manual copying and pasting or word-by-word input, which is inefficient and error-prone. Therefore, there is an urgent need for a more efficient method to improve the accuracy and automation level of data conversion.
[0081] This embodiment discloses a method for converting clipboard data into tag data based on specific standards. Refer to Figure 1 , and it includes the following steps S110 - S140:
[0082] S110, obtain the source data of the clipboard data.
[0083] The method for converting clipboard data into tag data based on specific standards disclosed in the embodiments of the present application is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a background server running a method for converting clipboard data into tag data based on specific standards. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0084] The source data format is data that supports copying to the clipboard, including but not limited to data in word, data in pdf, data in txt documents, data in html, and data of other types. Copy part, the whole line, the whole paragraph, or the whole page of data from the source file to the clipboard.
[0085] S120, determine a target file based on international standards or industry standards.
[0086] In a possible implementation, determining the target file based on international standards or industry standards specifically includes: extracting the key features and theme information of the source data; determining the scenario rule file corresponding to the source data according to the key features and theme information, where there is a second key description in the source data that satisfies the preset similarity condition with the first key description, and the scenario rule file includes the first key description; performing semantic analysis on the source data to check whether the source data contains a target data module, where the target data module is any one of the multiple data modules included in the target data standard, and the target data standard is any one of the multiple data standards included in the scenario rule file; if it is determined that the source data contains the target data module, then determining the rule file corresponding to the target data standard as the target file.
[0087] Specifically, first, preprocess the source data obtained from the clipboard, including operations such as text cleaning, word segmentation, and stop word filtering. Subsequently, use natural language processing techniques (such as TF-IDF, word vectors, or deep learning models) to extract the key features and theme information in the text. These features can be keywords, phrases, and semantic vector representations that describe the main content of the data. Through this information, the business background and application scenario of the data can be initially understood, providing a basis for the subsequent matching of the rule file.
[0088] The extracted key features and theme information are compared with a pre-stored rule library. The rule library stores rule files corresponding to various scenarios, and each rule file contains a set of specific descriptions (the first key description). If there is a second key description in the source data that is similar to the preset first key description (meeting the preset similarity condition), then it is considered that the source data meets the requirements of a specific scenario, and at this time, the corresponding scenario rule file is determined. This process mainly relies on similarity calculation algorithms (such as cosine similarity, Euclidean distance, etc.) to ensure the accuracy of the matching.
[0089] In a possible implementation, determining the scenario rule file corresponding to the source data according to the key features and theme information specifically includes: determining the first vector of the first key description; determining the second vectors of the key descriptions included in each training data during the model training stage; calculating the vector similarities between the first vector and the second vectors respectively; determining the target similarity that meets the preset similarity condition among the multiple vector similarities, and determining the target training data corresponding to the target similarity; determining the rule file corresponding to the target training data as the scenario rule file.
[0090] Specifically, extract a preset first key description from the source data obtained from the clipboard. This description usually reflects the core content or scenario characteristics of the data. Then, use a pre-selected word vector model (such as Word2Vec, GloVe, or BERT, etc.) to convert this text description into a vector representation of a fixed dimension, that is, the first vector. This vector can capture the implicit semantic information in the text and provide basic features for subsequent similarity calculation.
[0091] In the model training stage, each piece of training data contains a pre-annotated key description to reflect the scenario or rule it belongs to. Use the same word vector model as in the first step to vectorize the key description in each piece of training data, thereby generating the corresponding second vector. In this way, the key descriptions of all training data are converted into vectors in a unified semantic space, facilitating comparative analysis with the source data vector.
[0092] Use a vector similarity calculation method (such as cosine similarity, Euclidean distance, or Manhattan distance, etc.) to compare the first vector generated from the source data with the second vector generated from each piece of training data. For each piece of training data, calculate the similarity score between its key description vector and the source data vector. This score reflects the degree of semantic proximity between the two and provides a quantitative basis for subsequent judgment of whether it matches a specific scenario rule.
[0093] In a possible implementation manner, determining the target similarity that satisfies the preset similarity condition among multiple vector similarities specifically includes: judging whether the target similarity is greater than or equal to the preset threshold. If it is determined that the target similarity is greater than or equal to the preset threshold, then determine that the target similarity satisfies the preset similarity condition; or, sort the multiple vector similarities in ascending or descending order of numerical values, and according to the sorting result, determine the vector similarity with the largest numerical value among the multiple vector similarities to obtain the target similarity.
[0094] Specifically, after calculating the vector similarities between the source data and each piece of training data, the system pre-sets a similarity threshold, which represents the minimum requirement for matching. For each calculated vector similarity, the system judges whether its value is greater than or equal to this preset threshold. If a certain vector similarity meets this condition, it is considered that this similarity has reached the preset matching standard, and thus determine that this similarity is the target similarity that satisfies the preset similarity condition. This step ensures that only when the similarity is high enough, it is considered that there is sufficient semantic matching between the training data and the source data, thereby providing a basis for subsequent determination of the rule file.
[0095] When multiple vector similarity results exist simultaneously, the system sorts all the calculated similarities according to their numerical values. After the sorting is completed, the vector similarity with the largest value is selected as the target similarity because this value represents the strongest matching degree between the source data and a certain training data. This method is particularly useful when there is no explicitly preset threshold or all similarities do not reach the threshold. By selecting the highest similarity, it ensures that the determined target training data has the optimal matching effect, thus providing the most powerful reference for the determination of the subsequent scenario rule file.
[0096] According to the target training data determined in the previous step, search for and read the rule file information bound by this training data during the training phase. This rule file contains the rules and structure definitions that need to be followed for data conversion and label generation in a specific scenario. Finally, this rule file is determined as the scenario rule file for the current source data, thus providing clear criteria and guiding basis for subsequent data conversion and label generation.
[0097] Use deep semantic analysis technology to deeply interpret the source data and detect whether there are target data modules. The target data module refers to each component that constitutes a complete data structure as defined in the target data standard (international standard or industry standard). The system will obtain several target data standards from the scenario rule file and compare each standard with multiple data modules included in it in turn to determine whether the source data contains the feature information or semantic identifiers of these modules. If the feature of any one target data module is detected, it is considered that the source data meets a certain part of the requirements in the corresponding data standard, providing a basis for the determination of the subsequent rule file.
[0098] When the semantic analysis in the previous step confirms that the source data indeed contains the target data module, the system determines the applicable target data standard according to the matching result. Subsequently, the rule file corresponding to the target data standard (i.e., the rule associated with the target data standard in the scenario rule file) is selected as the final target file. This rule file details the generation requirements, structural order, label names, and attribute settings of the label data, ensuring that during the conversion process, the source data can be correctly mapped and converted into compliant label data according to international or industry standards.
[0099] S130, determine the data standard according to the target file.
[0100] The data standard can be either an international standard or an industry standard, such as: S1000D, Airbus FlightOpsXML Exchange, Boeing FTID, ATA iSpec 2300, etc. When operating on a target file, the data standard applicable to the file (described by XSD or DTD) is usually defined in the header of the target file, and the program will automatically obtain the corresponding XSD or DTD description of the target file. XSD or DTD defines XML, determines whether the XML content is legal, and also stipulates where the content in the clipboard will be inserted into the target file and in what hierarchical structure it will be presented.
[0101] S140, analyze the source data, and convert the source data according to the data standard to obtain labeled data.
[0102] In a possible implementation, before analyzing the source data and converting the source data according to the data standard to obtain labeled data, the method further includes: obtaining clipboard data in different formats and contents; obtaining pre-labeled pre-stored label data corresponding to the clipboard data; identifying the correspondence between the clipboard data and the pre-stored label data, and adjusting the model parameters based on the evaluation result of the correspondence; using independent validation sets and test sets to evaluate the accuracy and generalization ability of the identification, and through cross-validation, where the validation set includes the clipboard data and the pre-stored label data, and the test set includes the clipboard data and the pre-stored label data.
[0103] Specifically, obtaining various formats of data from the user's clipboard may include, but is not limited to, text, tables, HTML content, PDF files, etc. To process these different formats, a data acquisition module can be written to automatically extract data from the clipboard. The acquisition module can utilize corresponding format parsers, such as: for text data, directly extract the plain text in the clipboard. For table data, use, for example, the pandas library to parse Excel or CSV format data in the clipboard. For HTML or PDF data, specialized libraries (such as BeautifulSoup to parse HTML and PyPDF2 to parse PDF) can be used to extract the content.
[0104] Then obtain the pre-labeled label data corresponding to the clipboard data. These label data are usually pre-labeled manually or obtained from other data sources. The pre-stored label data is structured label information that corresponds one-to-one with the content in the clipboard and conforms to a certain standard (such as XML tags, JSON tags, etc.). These pre-stored label data contain the key information in the clipboard data, such as the category to which a text paragraph belongs, the attributes of each data item, etc. These label data provide labeled data for model training, and learn the mapping relationship between the clipboard data and the target labels through these pre-stored labels.
[0105] Analyze the correspondence between the clipboard data and the pre-stored label data. Generally, this relationship can be identified through natural language processing (NLP) techniques, pattern matching, or deep learning models. By comparing the clipboard data with the pre-stored labels, the label information for each piece of data can be identified. For example, semantic similarity calculations (such as cosine similarity, Jaccard similarity, etc.) can be used to determine the degree of association between the data and the labels. Or machine learning algorithms (such as classifiers, sequence labeling models, etc.) can be used to learn the mapping relationship between the clipboard data and the pre-stored labels. For the relationship between labels and data, the training model can output corresponding labels by inputting features (such as text content, context information, etc.).
[0106] After identifying the correspondence, the accuracy of the model's recognition will be evaluated, and the model's parameters will be adjusted based on the evaluation results (such as precision, recall, etc.). For example, the gradient descent method can be used to adjust the model parameters to optimize the recognition accuracy and ensure that the model can better match the clipboard data with the target labels.
[0107] To evaluate the performance of the model in actual applications, independent validation sets and test sets need to be used for evaluation. Cross-validation is a common evaluation method. It divides the dataset into multiple subsets (such as 5-fold cross-validation), trains and tests the model multiple times to ensure the accuracy and generalization ability of the model. The clipboard data and the pre-stored label data are divided into training sets, validation sets, and test sets. The training set is used for model training, the validation set is used to tune the model hyperparameters, and the test set is used to finally evaluate the performance of the model. By dividing the dataset into multiple subsets and using different combinations of training sets and validation sets to train the model and evaluate the results of each training. Cross-validation can effectively avoid overfitting and ensure that the model can show good generalization ability on different datasets. During the validation and testing process, evaluate the accuracy, recall, F1-score and other metrics of the model to ensure its high recognition accuracy and good adaptability to different data patterns.
[0108] In a possible implementation, the source data is analyzed and the source data is transformed according to the data standard to obtain label data, specifically including: parsing the source data item by item, converting each content item of the source data into a target label, where each target label is attached with a unique attribute; according to the rules of the data standard, determining the hierarchical structure and label relationship, including determining the hierarchical structure and order between multiple target labels, and each attribute follows the generation standard and is assigned by a random number; generating label data that conforms to the data standard and outputting the data in the format of a structured XML document.
[0109] Specifically, it is first necessary to parse the source data in detail. The source data can be text data, table data, or other formats of data, for example:
[0110] "One can be inoperative when the following conditions are met:
[0111] 1. The related air conditioning pack switch is in the OFF position;
[0112] 2. The aircraft is maintained at 31000 ft and below:
[0113] 3. Confirm that the emergency ram air system is operating normally;
[0114] 4. If the ground outside air temperature (OAT) is greater than or equal to ISA + 30°C, do not use APU bleed air to supply air to the air conditioning pack on the ground. Monitor the operating status of the other air conditioning pack on the AIR schematic page;
[0115] 6. Passenger operation;
[0116] 7. Complete the repair within 15 flight hours."
[0117] If the source data is text, it can be split into several items of content through sentences, paragraphs, or other identifiers (such as keywords). Each item of content represents an independent data unit. According to the predefined data standard, each item of content needs to be converted into the target tag. For example, if the source data describes the status of a certain device, the converted target tag may be <status>, where the content is the status description of the device. Each tag (such as <status>)All need to be accompanied by a unique attribute, such as the id attribute. The value of this attribute is assigned according to a generation standard, usually using a random number generation method to ensure the uniqueness of each tag. For example, id = "status-12345".
[0118] Based on the target data standard (such as XSD or other industry standards), we need to determine the hierarchy and the relationships between tags for each target tag. According to the structural rules defined by the target standard, determine the position of each tag. For example, if some tags have a parent-child relationship (such as <listitem>Containing multiple <para>If there are tags, the relationship between the parent tag and the child tags needs to be determined according to the hierarchy. Determine the order of each tag. For example, certain child tags need to appear first under a certain parent tag. Ensure that the tags are in the specified order in the XML document. According to the requirements of the target standard, each target tag needs to be attached with an attribute (such as id). The values of these attributes are generated by random numbers. For example, id = "para-67890". This attribute ensures that each tag has a unique identifier throughout the document. After the hierarchy is determined, the tags at each level can be verified according to the rules to ensure that the order, parent-child relationship, and attributes of each tag meet the target data standard.
[0119] After completing the tag parsing and hierarchy definition, the next step is to generate tag data that meets the target data standard and output it as a structured XML document. Use a programming language (such as Python, Java, or others) to construct the XML document, referring to Figure 3 , and the figure shows a schematic diagram of the generated tag data. For example, use the lxml or xml.etree.ElementTree library in Python to generate an XML structure that meets the target data standard. Each tag and sub-tag are nested and arranged according to the target standard. Save the generated XML data in the form of a file, such as output.xml, to ensure that the data meets the requirements of the target standard.
[0120] In a possible implementation, after analyzing the source data and converting the source data according to the data standard to obtain tag data, the method further includes: obtaining an XSD file that defines the structure and data type of the XML document; loading the XSD file through an XSD validator and reading the XML document; checking one by one whether the tags, attributes, data types, and hierarchy in the XML file meet the XSD rules included in the XSD file; if multiple checks for the XML file meet the XSD rules, the verification for the corresponding tag data passes.
[0121] Specifically, first, an XSD file for validating the XML file needs to be obtained. The XSD (XML Schema Definition) file defines the structure, elements, attributes, and their data types of the XML document. Usually, the XSD file is prepared in advance and corresponds to the target data standard (for example, dispatchItem.xsd), or is customized according to specific industry standards. This file contains the rules that the XML document must follow, such as which elements are required, the order of the elements, the hierarchical relationship between the elements, and the data types of the elements and attributes. The system needs to correctly load this XSD file to provide a rule basis for subsequent verification steps.
[0122] Then use the XSD validator to load the XSD file and XML document into memory for subsequent verification. Usually, the programming language provides corresponding libraries to load and process XSD files and XML files. The XSD validator will check the elements in the XML file one by one according to the rules defined in the XSD file, verify whether the elements in the XML file are consistent with the elements in the XSD definition, and whether there are tags that do not comply with the rules. Ensure that the attributes of each element in the XML file are consistent with the attributes defined in the XSD file, including the attribute name, attribute type, and whether it is required. Check whether the content data type of each tag in the XML is consistent with the data type specified in the XSD file. For example, a tag may require an integer, but the one given in the XML file is a string or a floating point number. This inconsistency will cause the verification to fail. According to the XSD rules, verify whether the parent-child relationship and order of the elements in the XML file conform to the specified hierarchical structure. That is, some elements may require to be nested under a specific parent element. If this hierarchical relationship is wrong, the verification will also fail.
[0123] If the XML file passes all the rule checks defined in the XSD file (tags, attributes, data types, hierarchical structures, etc.) in the above checks, the XML file is considered valid, meets the data standard, and passes the verification. In this process, the XSD validator will return the result of the verification and output relevant information, indicating that the XML file meets the predetermined structure and data type requirements.
[0124] If any of the items do not meet the requirements, the XSD validator will throw an error or warning, indicating the specific location of the non-compliance (for example, the id attribute type of an element is wrong, or the order of an element is wrong). When validation fails, the system usually outputs an error log, detailing which parts do not meet the XSD rules and need to be modified.
[0125] Finally, if all checks pass, the validation result indicates that the generated XML data fully complies with the expected standards and can continue to be used for storage, transmission or other applications.
[0126] This embodiment also discloses a device for converting clipboard data into label data based on a specific standard, referring to Figure 3 The device includes an acquisition module 301, a processing module 302 and an output module 303, wherein:
[0127] The acquisition module 301 is used to acquire source data of the clipboard data.
[0128] The processing module 302 is used to determine a target file based on an international standard or an industry standard.
[0129] The processing module 302 is used to determine the data standard according to the target file.
[0130] The output module 303 is used to analyze the source data and convert the source data according to the data standard to obtain labeled data.
[0131] In a possible implementation, the processing module 302 is used to extract the key features and theme information of the source data.
[0132] The processing module 302 is used to determine the scenario rule file corresponding to the source data according to the key features and theme information, where there is a second key description in the source data that satisfies a preset similarity condition with the first key description, and the scenario rule file includes the first key description.
[0133] The processing module 302 is used to perform semantic analysis on the source data to check whether the source data contains a target data module, where the target data module is any one of the multiple data modules included in the target data standard, and the target data standard is any one of the multiple data standards included in the scenario rule file.
[0134] The processing module 302 is used to determine the rule file corresponding to the target data standard as the target file if it is determined that the source data contains the target data module.
[0135] In a possible implementation, the processing module 302 is used to determine the first vector of the first key description.
[0136] The processing module 302 is used to determine the second vectors of the key descriptions included in each training data during the model training phase.
[0137] The processing module 302 is used to calculate the vector similarity between the first vector and the second vectors respectively.
[0138] The processing module 302 is used to determine the target similarity that satisfies the preset similarity condition among the multiple vector similarities and determine the target training data corresponding to the target similarity.
[0139] The processing module 302 is used to determine the rule file corresponding to the target training data as the scenario rule file.
[0140] In a possible implementation, the processing module 302 is used to parse the source data item by item and convert each content item of the source data into a target label, where each target label is attached with a unique attribute.
[0141] The processing module 302 is used to determine the hierarchical structure and label relationship according to the rules of the data standard, including determining the hierarchical structure and order among multiple target labels, and each attribute follows the generation standard and is assigned a value by a random number.
[0142] The processing module 302 is used to generate tag data that conforms to the data standard and output an XML document with a structured data format.
[0143] In a possible implementation, the acquisition module 301 is used to acquire an XSD file that defines the structure and data type of the XML document.
[0144] The acquisition module 301 is used to load the XSD file through an XSD validator and read the XML document.
[0145] The processing module 302 is used to check one by one whether the tags, attributes, data types, and hierarchies in the XML file conform to the XSD rules included in the XSD file.
[0146] The processing module 302 is used to determine that the verification of the tag data passes if multiple checks on the XML file conform to the XSD rules.
[0147] In a possible implementation, the acquisition module 301 is used to acquire clipboard data in different formats and contents.
[0148] The acquisition module 301 is used to acquire pre-annotated pre-stored tag data corresponding to the clipboard data.
[0149] The processing module 302 is used to identify the correspondence between the clipboard data and the pre-stored tag data and adjust the model parameters based on the evaluation result of the correspondence.
[0150] The processing module 302 is used to use independent validation sets and test sets to evaluate the recognition accuracy and generalization ability, and through cross-validation, where the validation set includes clipboard data and pre-stored tag data, and the test set includes clipboard data and pre-stored tag data.
[0151] In a possible implementation, the processing module 302 is used to determine whether the target similarity is greater than or equal to a preset threshold. If it is determined that the target similarity is greater than or equal to the preset threshold, it is determined that the target similarity meets the preset similarity condition. Or,
[0152] The processing module 302 is used to sort multiple vector similarities in ascending order of numerical values, and based on the sorting result, determine the vector similarity with the largest numerical value among the multiple vector similarities to obtain the target similarity.
[0153] It should be noted that: when the device provided in the above embodiment realizes its functions, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0154] This embodiment also discloses an electronic device. Referring to Figure 4 , the electronic device may include: at least one processor 401, at least one communication bus 402, a user interface 403, a network interface 404, and at least one memory 405.
[0155] Among them, the communication bus 402 is used to realize the connection and communication between these components.
[0156] Among them, the user interface 403 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 403 may further include a standard wired interface and a wireless interface.
[0157] Among them, the network interface 404 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0158] Among them, the processor 401 may include one or more processing cores. The processor 401 connects various parts within the entire server through various interfaces and lines, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 405, as well as calling data stored in the memory 405, it executes various functions of the server and processes data. Optionally, the processor 401 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 401 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 401 and may be implemented separately by a single chip.
[0159] Among them, the memory 405 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory 405 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 405 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above method embodiments, etc.; the data storage area can store the data involved in the above method embodiments. The memory 405 is optionally also at least one storage device located far from the aforementioned processor 401. The memory 405, as a computer storage medium, may include an operating system, a network communication module, a user interface 403 module, and an application program for a method of converting clipboard data into tag data based on a specific standard.
[0160] In Figure 4 In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and obtain the data input by the user; while the processor 401 can be used to call the application program for a method of converting clipboard data into tag data based on a specific standard stored in the memory 405. When executed by one or more processors 401, the electronic device is caused to execute the method of one or more of the above embodiments.
[0161] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0162] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0163] In several embodiments provided in this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of devices or units can be in electrical or other forms.
[0164] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0166] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 405 and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. And the aforementioned memory 405 includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0167] This application also discloses a computer-readable storage medium storing instructions. When executed by one or more processors 401, it causes the electronic device to execute the method as described in one or more of the above embodiments.
[0168] The foregoing are only exemplary embodiments of the present disclosure, and thus cannot limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Other embodiments of the present disclosure will be readily contemplated by those skilled in the art after considering the specification and the practice of the disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.< / para> < / listitem> < / status> < / status>
Claims
1. A method for converting clipboard data into label data based on specific criteria, characterized in that, The method includes: Obtaining the source data of the clipboard data; Determining a target file based on international standards or industry standards; Determining a data standard according to the target file; Analyzing the source data and converting the source data according to the data standard to obtain labeled data.
2. The method for converting clipboard data into tag data based on specific criteria according to claim 1, wherein The determining of the target file based on international standards or industry standards specifically includes: Extracting the key features and theme information of the source data; Determining a scenario rule file corresponding to the source data according to the key features and the theme information, wherein there is a second key description in the source data that satisfies a preset similarity condition with a first key description, and the scenario rule file includes the first key description; Performing semantic analysis on the source data to check whether the source data contains a target data module, where the target data module is any one of multiple data modules included in a target data standard, and the target data standard is any one of multiple data standards included in the scenario rule file; If it is determined that the source data contains the target data module, then determining the rule file corresponding to the target data standard as the target file.
3. A method for converting clipboard data into tag data based on specific criteria according to claim 2, characterized in that, The determining of the scenario rule file corresponding to the source data according to the key features and the theme information specifically includes: Determining a first vector of the first key description; Determining second vectors of the key descriptions included in each training data during the model training phase; Calculating the vector similarity between the first vector and each of the second vectors respectively; Determining a target similarity that satisfies a preset similarity condition among the multiple vector similarities, and determining the target training data corresponding to the target similarity; Determining the rule file corresponding to the target training data as the scenario rule file.
4. A method for converting clipboard data into tag data based on a specific standard according to claim 1, characterized in that, Analyzing the source data and converting the source data according to the data standard to obtain labeled data specifically includes: Parsing the source data item by item and converting each content of the source data into a target label, where each target label is attached with a unique attribute; Determining the hierarchical structure and label relationship according to the rules of the data standard, including determining the hierarchical structure and order among the multiple target labels, and each attribute follows a generation standard and is assigned a value by a random number; Generating the labeled data that conforms to the data standard and outputting an XML document with a structured data format.
5. A method for converting clipboard data into tag data based on specific criteria according to claim 4, characterized in that, After analyzing the source data and converting the source data according to the data standard to obtain labeled data, the method further includes: Obtaining an XSD file that defines the structure and data type of the XML document; Loading the XSD file through an XSD validator and reading the XML document; Checking one by one whether the labels, attributes, data types, and hierarchical structures in the XML file conform to the XSD rules included in the XSD file; If multiple checks on the XML file conform to the XSD rules, then the verification of the labeled data is passed.
6. A method for converting clipboard data into label data based on specific criteria according to claim 2, characterized in that, Before analyzing the source data and converting the source data according to the data standard to obtain labeled data, the method further includes: Obtain clipboard data in different formats and with different contents; Obtain pre-annotated pre-stored label data corresponding to the clipboard data; Identify the correspondence between the clipboard data and the pre-stored label data, and adjust the model parameters based on the evaluation result for the correspondence; Use independent validation sets and test sets to evaluate the recognition accuracy and generalization ability, and perform cross-validation, where the validation set includes clipboard data and pre-stored label data, and the test set includes clipboard data and pre-stored label data.
7. A method for converting clipboard data into tag data based on a specific standard according to claim 3, characterized in that, Determining the target similarity that satisfies the preset similarity condition among the multiple vector similarities specifically includes: Judge whether the target similarity is greater than or equal to a preset threshold. If it is determined that the target similarity is greater than or equal to the preset threshold, then determine that the target similarity satisfies the preset similarity condition; or, Sort the multiple vector similarities in ascending order of numerical values, and according to the sorting result, determine the vector similarity with the largest numerical value among the multiple vector similarities to obtain the target similarity.
8. An apparatus for converting clipboard data into label data based on specific criteria, characterized in that, The device includes an acquisition module (301), a processing module (302), and an output module (303), where: The acquisition module (301) is configured to acquire the source data of the clipboard data; The processing module (302) is configured to determine a target file based on international standards or industry standards; The processing module (302) is configured to determine a data standard according to the target file; The output module (303) is configured to analyze the source data and convert the source data according to the data standard to obtain label data.
9. An electronic device, characterized in that, It includes a processor (401), a communication bus (402), a user interface (403), a network interface (404), and a memory (405). The memory (405) is used to store instructions. The user interface (403) and the network interface (404) are both used to communicate with other devices. The communication bus (402) is used to realize the connection and communication between components in the electronic device. The processor (401) is used to execute the instructions stored in the memory (405) so that the electronic device executes the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1-7 is executed.