A method and system for improving the quality of industrial corpus collected at the gateway end

CN122527338APending Publication Date: 2026-08-07SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610698445.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]当前工业语料处理方法普遍存在两大制约:一是质量管控与脱敏合规环节相互割裂,难以在确保全流程数据安全的同时完整保留语料的预训练价值;二是网关端仅承担基础采集与简单预处理,语料处理的核心参数均为人工预设的固定值,无法根据云服务器端大模型的实时训练反馈进行反向动态优化,导致产出语料与预训练需求持续脱节,适配性与应用效果受限

Benefits of technology

本发明能够根据云服务器端预训练大模型的实际训练效果,自动反向调整网关端语料质量评估体系的参数权重,使网关端产出的标准化语料对在持续迭代中更加贴合大模型的训练需求,实现语料供给与模型优化的自适应协同进化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122527338A_ABST
    Figure CN122527338A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and in particular to a method and system for improving the quality of industrial corpus collected at the gateway end. The method comprises: collecting industrial data and converting it into initial corpus, and extracting sensitive entity features; defining a corpus quality evaluation system based on a knowledge graph; using the corpus quality evaluation system to score the quality of the initial corpus; dividing the initial corpus into high-quality corpus, good corpus and invalid corpus according to the scoring results, performing semantic enhancement on the good corpus, and taking the high-quality corpus and the good corpus after semantic enhancement as optimized corpus; according to the sensitive entity features, dividing the corresponding optimized corpus into different secret levels, and executing corresponding desensitization strategies for the optimized corpus of different secret levels; performing domain standardization annotation on the desensitized optimized corpus to obtain standardized corpus pairs; sending the standardized corpus pairs to the cloud server end, receiving the training results returned by the cloud server end, and updating the parameter weights in the corpus quality evaluation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a method and system for improving the quality of industrial corpora collected at gateways. Background Technology

[0002] In the industrial sector, the key to supporting the depth and effectiveness of pre-trained large-scale models lies in providing high-quality, compliant corpora specific to the industrial field for the entire pre-training process. Industrial corpora are highly heterogeneous in origin, have strong scenario-specific attributes, and carry core enterprise process secrets and production operation data. Their semantic integrity, temporal relevance, and security compliance requirements are far higher than those of general corpora.

[0003] Current industrial corpus processing methods generally suffer from two major constraints: First, quality control and desensitization compliance are disconnected, making it difficult to fully preserve the pre-training value of the corpus while ensuring data security throughout the entire process; second, the gateway only undertakes basic collection and simple preprocessing, and the core parameters of corpus processing are all fixed values ​​preset by humans. It is impossible to perform reverse dynamic optimization based on the real-time training feedback of the large model on the cloud server, resulting in a continuous disconnect between the output corpus and the pre-training requirements, limiting adaptability and application effectiveness. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method and system for improving the quality of industrial corpora collected at gateways; To achieve the above objectives, the present invention adopts the following technical solution: Firstly, a method for improving the quality of industrial corpora collected at the gateway is provided, including: Collect industrial data and convert it into initial corpus, then extract sensitive entity features from the initial corpus; Based on knowledge graphs, a corpus quality assessment system is defined; the initial corpus is scored using the corpus quality assessment system; based on the scoring results, the initial corpus is divided into high-quality corpus, good corpus, and invalid corpus; semantic enhancement is performed on the good corpus; and the high-quality corpus and the semantically enhanced good corpus are used as optimized corpus. Based on the characteristics of sensitive entities, the corresponding optimized corpus is divided into different security levels, and corresponding desensitization strategies are implemented for the optimized corpus of different security levels; the desensitized optimized corpus is then labeled with domain standardization to obtain standardized corpus pairs. The standardized corpus pairs are sent to the cloud server, the training results returned by the cloud server are received, and the parameter weights in the corpus quality evaluation system are updated.

[0005] Secondly, a device for improving the quality of industrial corpora collected at a gateway is provided, applied at the gateway, and the device includes: The feature extraction module is configured to: collect industrial data, convert the industrial data into initial corpus, and extract sensitive entity features from the initial corpus; The corpus quality assessment module is configured to: define a corpus quality assessment system based on a knowledge graph; use the corpus quality assessment system to score the quality of the initial corpus; divide the initial corpus into high-quality corpus, good corpus, and invalid corpus according to the scoring results; perform semantic enhancement on the good corpus; and use the high-quality corpus and the semantically enhanced good corpus as optimized corpus. The desensitization module is configured to: divide the corresponding optimized corpus into different security levels based on the characteristics of sensitive entities; execute corresponding desensitization strategies for the optimized corpus at different security levels; and perform domain-standardized annotation on the desensitized optimized corpus to obtain standardized corpus pairs. The transmission and iterative update module is configured to: send the standardized corpus pairs to the cloud server, receive the training results returned by the cloud server, and update the parameter weights in the corpus quality evaluation system.

[0006] Thirdly, a system for improving the quality of industrial corpus collection at the gateway end is provided, including a gateway end and a cloud server end; the gateway end is used to connect to machine equipment, and the gateway end and the cloud server end are connected via a network; The gateway is used to collect industrial data and convert it into initial corpus, extracting sensitive entity features from the initial corpus; defining a corpus quality assessment system based on a knowledge graph; scoring the initial corpus using the corpus quality assessment system; classifying the initial corpus into high-quality, good, and invalid corpus according to the scoring results; semantically enhancing the good corpus; and using the high-quality corpus and the semantically enhanced good corpus as optimized corpus; classifying the corresponding optimized corpus into different security levels according to the sensitive entity features; implementing corresponding desensitization strategies for optimized corpus of different security levels; performing domain standardization annotation on the desensitized optimized corpus to obtain standardized corpus pairs; sending the standardized corpus pairs to the cloud server; receiving the training results returned by the cloud server; and updating the parameter weights in the corpus quality assessment system. The cloud server is used to receive standardized corpus pairs, fine-tune the pre-trained large model based on the standardized corpus pairs, and send the training results to the gateway.

[0007] Fourthly, a gateway device is also provided, including: Memory, used for non-transitory storage of computer-readable instructions; and Processor, for executing the computer-readable instructions, When the computer-readable instructions are executed by the processor, they perform the method described in the first aspect above.

[0008] Fifthly, a computer-readable storage medium is provided having a program stored thereon that, when executed by a processor, implements the method described in the first aspect above.

[0009] The above technical solution has the following advantages or beneficial effects: This invention can automatically adjust the parameter weights of the gateway-side corpus quality evaluation system based on the actual training effect of the pre-trained large model on the cloud server side, so that the standardized corpus produced by the gateway side can better meet the training needs of the large model in continuous iteration, and realize the adaptive co-evolution of corpus supply and model optimization.

[0010] This invention solves the problem that existing industrial data acquisition gateways cannot convert heterogeneous data of various formats from the industrial site into standardized corpora that can be directly learned by large models on the local side, resulting in a large amount of valuable industrial data being unusable and extremely low data utilization.

[0011] This invention solves the problem that existing industrial corpus processing solutions cannot adapt to the low-computing-power edge environment of industrial gateways. Either they can only transmit all the raw data containing sensitive information to the cloud for processing, which poses a very high risk of data leakage; or they cannot run complex algorithms on the gateway side, and cannot produce corpora that meet the requirements of pre-trained large models. Attached Figure Description

[0012] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0013] Figure 1 This is a flowchart illustrating a method for improving the quality of industrial corpus collected by a gateway in a specific embodiment of the present invention. Detailed Implementation

[0014] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0015] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the invention. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] In this embodiment of the invention, "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of this invention, "multiple" refers to two or more.

[0017] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0018] All data acquisition in this embodiment is carried out in accordance with laws and regulations and with user consent, and the data is used legally.

[0019] Example 1 This embodiment provides a method for improving the quality of industrial corpora collected at the gateway, including: S1: Collect industrial data and convert it into initial corpus, then extract sensitive entity features from the initial corpus; S2: Based on knowledge graphs, define a corpus quality assessment system; use the corpus quality assessment system to score the quality of the initial corpus; divide the initial corpus into high-quality corpus, good corpus, and invalid corpus according to the scoring results; perform semantic enhancement on the good corpus; and use the high-quality corpus and the good corpus after semantic enhancement as the optimized corpus. S3: Based on the sensitive entity characteristics, the corresponding optimized corpus is divided into different security levels, and the corresponding desensitization strategy is executed for the optimized corpus of different security levels; the desensitized optimized corpus is labeled with domain standardization to obtain standardized corpus pairs; S4: Send the standardized corpus pairs to the cloud server, receive the training results returned by the cloud server, and update the parameter weights in the corpus quality evaluation system.

[0020] Step S1 specifically includes: S1.1: Complete the collection of industrial data from the industrial site at the gateway.

[0021] Industrial data comes from multiple sources, specifically including: 1) Time-series data: equipment operating parameters and process control parameters such as temperature, pressure, flow rate, and speed; 2) Protocol data: register data and control command data in PLC and DCS systems; 3) Text data: equipment fault alarm logs, equipment maintenance work orders, and repair records; 4) Document data: process specification documents, equipment manuals, industry standard documents, etc. Protocol data and text data are further combined into unstructured text.

[0022] S1.2: Transform the multi-source data into a unified semantic space, complete semantic normalization processing, and form the initial corpus. The specific steps are as follows: By integrating data from the industrial sector, entity, relation, and attribute triples are extracted from the industrial data, and triples are constructed and corresponding knowledge graphs are built.

[0023] Multi-source data undergoes unified transformation. For time-series data, knowledge graphs are used to perform semantic mapping of numerical values, equipment, processes, and scenarios, converting discrete numerical values ​​into textual semantic descriptions with time-series associations and process attributes. For unstructured text and document data, the jieba word segmentation tool is used to segment the data into sentences, words, and remove stop words, extracting core semantic paragraphs. The textual semantic descriptions of time-series data and the unstructured text and document data after extracting core semantic paragraphs are used as the initial corpus.

[0024] S1.3: Extract sensitive entity features from the initial corpus.

[0025] Extracting sensitive entity features from the initial corpus: Identifying sensitive entities in the corpus using a sensitive word list and regular expressions.

[0026] In this embodiment, step S1.3 further includes: extracting three other types of features from the initial industrial corpus: 1) Semantic structure features: using the spaCy library to perform syntactic parsing on the initial industrial corpus, extracting sentence features such as subject-verb-object attributive-adverb-complement structures and the position of core predicates; extracting entity relationships in the corpus using the BERT model; and constructing templates based on causal conjunctions to extract causal logic features from the corpus. 2) Temporal correlation features: using the TSFresh library to extract trend features such as rising, falling, stable, volatile, periodic peaks, and troughs in the temporal corpus of the initial industrial corpus; using the Apriori association rule algorithm to mine the association rules between temporal data and operating conditions; comparing similar temporal data, calculating cosine similarity, and identifying inconsistent abnormal segments. 3) Industrial entity features: Using an industrial terminology dictionary and regular expression book, entities are extracted from the corpus and categorized; using an industrial thesaurus, synonyms in the corpus are uniformly converted, such as normalizing "reaction vessel 101" and "reaction vessel No. 101" to "reaction vessel-101", to avoid feature dimension confusion.

[0027] Step S2 specifically includes: S2.1: Based on the knowledge graph built in S1.2, a corpus quality assessment system is defined, consisting of five dimensions: semantic completeness, domain adaptability, temporal relevance, information density, and noise ratio. The assessment formula for this system is: ; Wherein, Q(C) is the overall quality score of the corpus C to be processed, ranging from 0 to 100; S1(C): Semantic integrity dimension score (0-100 points), based on knowledge graph, evaluates the completeness of entities, relationships, and causal logic in the corpus; S2(C): Domain adaptability dimension score (0-100 points), based on knowledge graph, evaluates the matching degree between the corpus and industrial scenarios and industrial large model pre-training tasks; S3(C): Temporal relevance dimension score (0-100 points), for industrial time-series data, evaluates the consistency of temporal features and working condition change relationship logic in the corpus; S4(C): Information density dimension score (0-100 points), evaluates the proportion of effective industrial information in the corpus; S5(C): Noise proportion dimension score (0-100 points), evaluates the proportion of outliers, erroneous information, and irrelevant content in the corpus. The higher the noise proportion, the lower the corresponding dimension score. The adaptive weighting coefficients for the corresponding scores must satisfy the following constraints. The initial weights should be preset according to the industrial scenario.

[0028] S2.2: Calculate a quality assessment score for each piece of text using the formula in the above evaluation system, and classify the text based on the score: High-quality corpus: Quality assessment score ≥ first threshold, the corpus quality fully meets the pre-training requirements, and it directly proceeds to the subsequent de-identification stage; Good corpus: Second threshold < quality assessment score < first threshold, indicating that the corpus has problems such as missing core information and incomplete semantic logic, requiring semantic enhancement; Invalid corpus: Quality assessment score ≤ second threshold, the corpus has a high noise ratio and lacks effective industrial information, and is directly discarded. In this embodiment, the first threshold is set to 90 and the second threshold is set to 70, which can be set according to the needs of those skilled in the art.

[0029] In this embodiment, the specific steps for semantic enhancement of good corpora are as follows: the knowledge graph built in S1.2 is used as an external knowledge base for the semantic enhancement model. Under the guidance of the knowledge graph, the semantic enhancement model is used to perform enhancement processing such as knowledge expansion and logical error correction on the corpora. The enhanced good corpora are then re-evaluated for quality until the score is ≥90 points, at which point semantic enhancement stops.

[0030] In this embodiment, the semantic enhancement model used is a general-purpose basic open-source model, which is mainly used to expand the knowledge base. The semantic enhancement model includes, but is not limited to: Tongyi Qianwen Qwen-7B, Qwen-14B, Qwen-72B; Zhipu ChatGLM3-6B, GLM-4-9B; Baichuan2-7B, Baichuan2-13B; Llama 2-7B, Llama 2-13B, Llama 2-70B.

[0031] S2.3: Use high-quality corpora and semantically enhanced good corpora as optimized corpora.

[0032] Step S3 specifically includes: S3.1: Based on sensitive entity characteristics, the optimized corpus is divided into three levels: Top Secret, Confidential, and General. The three levels have the following characteristics: Top Secret: Core process formulas, core production and operation data, core trade secrets; leakage would cause significant losses to the enterprise; Confidential: Equipment-specific operating parameters, non-public process control thresholds, non-public equipment maintenance data; leakage would cause some losses to the enterprise; General: Public industry standards, public equipment parameters, routine operating condition data; no sensitive information.

[0033] S3.2: Adopt corresponding desensitization strategies for optimized corpora of different security levels.

[0034] S3.2.1: For top-secret corpora: Irreversible desensitization is performed using industrial semantic conformal one-way encryption technology. Specifically: First, semantic embedding mapping is completed. Based on the semantic structure features, temporal correlation features, and industrial domain entity features extracted from the initial industrial corpora in step S1.3, the top-secret corpora are mapped into a fixed-dimensional dense semantic vector through the pre-deployed embedding model at the gateway. This vector completely preserves the core semantics, process causal logic, and temporal correlation of the corpora, and possesses the semantic similarity conservation property. Then, layered one-way encryption is performed. The national cryptographic SM3 hash algorithm is used to perform layered irreversible encryption processing on the semantic vector. First, the semantic vector is divided into blocks according to the semantic dimension. A corresponding SM3 hash value is generated for each block. Then, the block hash values ​​are fused and hashed again to finally generate a fixed-length encrypted semantic representation. Finally, the representation is formatted. The encrypted semantic representation is bound and encapsulated with the domain label, temporal label, and working condition label of the corpora to generate desensitized corpora that can only be used for large-scale model semantic learning and cannot be reversed to restore the original plaintext information, ensuring that the original core confidential information is completely unrecoverable. In this embodiment, the embedding model is m3e-base or bge-m3, but those skilled in the art can also use other models with vectorization capabilities.

[0035] S3.2.2: For confidential corpora: Semantic differential privacy technology is used for desensitization. By superimposing semantically preserving substitution functions with controllable noise, the original sensitive information is irreversibly hidden without changing the core semantic logic, temporal correlation and pre-training value of the corpus.

[0036] S3.2.3: For general-level corpora: the original semantic information is fully preserved without any desensitization processing.

[0037] In some embodiments, during the desensitization process for top-secret and confidential corpora, a semantic fidelity hard constraint mechanism is also introduced to perform real-time semantic fidelity verification and quantitatively control the quality of the corpora before and after desensitization. The semantic fidelity quantitative constraint formula is as follows: ; in, The quality score of the corpus after anonymization is τ, which is a preset quality tolerance threshold used to control the degree of semantic distortion.

[0038] Anonymizing top-secret and confidential corpora may result in the loss of important information, leading to issues such as semantic incoherence, lack of domain identification, missing temporal information, and the loss of key information, rendering the data noisy. Therefore, it is necessary to score the anonymization results using evaluation formulas from a corpus quality assessment system. The system substitutes the scores before and after desensitization into the semantic fidelity quantification constraint formula for judgment. If the constraint is met, the current desensitization result is retained; if it exceeds the threshold, it is judged as semantic distortion exceeding the standard, and the corpus is used to perform semantic enhancement optimization using a large model. After enhancement, the above steps are repeated until the constraint conditions are met. The semantic fidelity hard constraint mechanism can ensure that the desensitization operation does not damage the quality of the corpus.

[0039] The S3 step of this invention first classifies the data according to its sensitivity level, then matches the corresponding desensitization method, and at the same time adopts a semantic fidelity hard constraint mechanism. This ensures that the core sensitive information is absolutely irreversible and not leaked, and that the desensitized corpus fully complies with compliance requirements. It also does not destroy the core semantics in the corpus that are useful for training large models. This solves the problem that existing technologies cannot accurately identify the specific sensitive information in industrial scenarios, either missing sensitive content and causing compliance risks, or removing useful industrial information as sensitive content, resulting in a significant decrease in the value of industrial corpus.

[0040] S3.3: Perform domain-standardized annotation on the de-identified optimized corpus to obtain standardized corpus pairs. The specific steps are as follows: S3.3.1: Perform double verification on the anonymized optimized corpus; the double verification includes semantic integrity verification and quality score review. Semantic integrity ensures the integrity of the core semantics and temporal logic of the corpus, while the quality score is reviewed again using the pre-evaluation system to ensure the consistency of the corpus quality level before and after anonymization.

[0041] Semantic integrity verification specifically involves comparing the entities and logic of the corpora before and after desensitization based on the knowledge graph constructed in S1.2 to ensure that the core semantics are not lost or tampered with. Quality score review specifically involves re-evaluating the quality of the desensitized corpora through a quality assessment system to ensure that the score is ≥90 points and still meets the standards for high-quality corpora.

[0042] S3.3.2: Perform domain-standardized annotation on the corpus that passes double verification, and generate... <prompt-completion>Standardized corpus pairs in the specified format. Here, `prompt` represents the task instruction / question for the domain scenario, and `completion` represents the corresponding standard answer.

[0043] In some embodiments, the standardized corpus is encrypted using the national cryptographic SM4 algorithm and then stored in the local database at the gateway.

[0044] This invention directly converts various scattered data collected on-site, such as equipment parameters, process data, and operation and maintenance logs, into standardized corpus pairs that meet the pre-training requirements of large industrial models at the gateway end. The cloud server end can use these corpus pairs directly, which greatly reduces the processing workload of the cloud server end and turns massive amounts of industrial data that were originally unusable into training materials for large models.

[0045] The specific steps of S4 are as follows: S4.1: After encrypting the standardized corpus using the national cryptographic SM4 algorithm, upload it to the cloud server. The cloud server includes a pre-trained large model.

[0046] The pre-trained large model refers to the large model used to train the high-quality corpus generated by this invention, and does not participate in the content implementation of this invention. The role of the pre-trained large model is to input the standardized corpus pairs generated in steps S1-S3 into the pre-trained large model for training, and then optimize the relevant parameters in steps S1-S3 based on the training results.

[0047] In this embodiment, the pre-trained large models used include, but are not limited to: Tongyi Qianwen Qwen-7B, Qwen-14B, Qwen-72B; Zhipu ChatGLM3-6B, GLM-4-9B; Baichuan2-7B, Baichuan2-13B; Llama 2-7B, Llama 2-13B, Llama 2-70B.

[0048] S4.2: The cloud server receives standardized corpus pairs from the gateway, fine-tunes the pre-trained large model, and then statistically analyzes data such as training speed, inference accuracy, and the number of hallucination errors to calculate the comprehensive loss function. After encryption, the comprehensive loss function is sent back to the gateway.

[0049] The formula for the comprehensive loss function is as follows: ; in, It is the model's overall loss value, V t This is the normalized training speed value, Acc is the inference accuracy (0~1), and H... r λ1 represents the proportion of hallucination errors (0~1), and λ2 and λ3 are the weight coefficients corresponding to training speed, inference accuracy, and hallucination errors, respectively.

[0050] S4.3: Comprehensive Loss Function Based on Cloud Server Backhaul The weight coefficients in the gateway-side corpus quality evaluation system are dynamically updated through an iterative gain formula, enabling the linkage and iterative adaptation between the gateway-side corpus processing parameters and the training effect of the pre-trained large model on the cloud server. The iterative gain formula is as follows: ; in: Let be the weighting coefficient of the quality assessment system formula in the t-th iteration; η represents the weight coefficients after iterative updates; η is the edge iterative learning rate, which is 0.01 in this embodiment. The partial derivative of the comprehensive loss function with respect to the weight coefficient of the i-th dimension represents the degree of influence of the change of the weight of this dimension on the training effect of the pre-trained large model on the cloud server, and provides a quantitative direction and magnitude reference for adjusting the weight coefficient.

[0051] The gateway extracts the comprehensive loss function returned from the cloud server and assigns weights to the coefficients of each dimension. Solve for partial derivatives to clarify the adjustment direction and quantification range of each weight; then calculate the weight coefficients for the current round. Substituting the results of the edge iterative learning rate η and partial derivatives into the iterative gain formula, the updated weight system is obtained by calculating it dimension by dimension. The new weighting coefficients are then normalized to ensure that they meet the requirements. The constraints; the gateway end based on the updated The weights of the corpus quality assessment system will be updated to serve as the calculation weights for the quality score of the next batch of industrial corpora, ensuring that the dimensions of corpus quality assessment align more closely with the training needs of the pre-trained large model. Based on the new weight coefficients, steps S1-S5 will be repeated to ensure that the standardized corpus output from the gateway continuously adapts to the pre-training requirements of the pre-trained large model on the cloud server, gradually reducing the overall loss function of the model. This improves the training effect of the model.

[0052] Example 2 This embodiment provides a device for improving the quality of industrial corpora collected at a gateway. The device is applied at the gateway and includes: The feature extraction module is configured to: collect industrial data, convert the industrial data into initial corpus, and extract sensitive entity features from the initial corpus; The corpus quality assessment module is configured to: define a corpus quality assessment system based on a knowledge graph; use the corpus quality assessment system to score the quality of the initial corpus; divide the initial corpus into high-quality corpus, good corpus, and invalid corpus according to the scoring results; perform semantic enhancement on the good corpus; and use the high-quality corpus and the semantically enhanced good corpus as optimized corpus. The desensitization module is configured to: divide the corresponding optimized corpus into different security levels based on the characteristics of sensitive entities; execute corresponding desensitization strategies for the optimized corpus at different security levels; and perform domain-standardized annotation on the desensitized optimized corpus to obtain standardized corpus pairs. The transmission and iterative update module is configured to: send the standardized corpus pairs to the cloud server, receive the training results returned by the cloud server, and update the parameter weights in the corpus quality evaluation system.

[0053] It should be noted that the feature extraction module, corpus quality assessment module, desensitization processing module, and transmission and iterative update module mentioned above correspond to steps S1 to S4 in Embodiment 1. The examples and application scenarios implemented by these modules and their corresponding steps are the same, but they are not limited to the content disclosed in Embodiment 1. It should be noted that these modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0054] Example 3 This embodiment provides a system for improving the quality of industrial corpus collected by a gateway, including a gateway and a cloud server; the gateway is used to connect to machine equipment, and the gateway and the cloud server are connected via a network; The gateway is used to collect industrial data and convert it into initial corpus, extracting sensitive entity features from the initial corpus; defining a corpus quality assessment system based on a knowledge graph; scoring the initial corpus using the corpus quality assessment system; classifying the initial corpus into high-quality, good, and invalid corpus according to the scoring results; semantically enhancing the good corpus; and using the high-quality corpus and the semantically enhanced good corpus as optimized corpus; classifying the corresponding optimized corpus into different security levels according to the sensitive entity features; implementing corresponding desensitization strategies for optimized corpus of different security levels; performing domain standardization annotation on the desensitized optimized corpus to obtain standardized corpus pairs; sending the standardized corpus pairs to the cloud server; receiving the training results returned by the cloud server; and updating the parameter weights in the corpus quality assessment system. The cloud server is used to receive standardized corpus pairs, fine-tune the pre-trained large model based on the standardized corpus pairs, and send the training results to the gateway.

[0055] The descriptions of each embodiment in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0056] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and the division of modules described above is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.

[0057] Example 4 This embodiment provides a gateway device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory to cause the electronic device to perform the method described in Embodiment 1 above.

[0058] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0059] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.

[0060] In the implementation process, each step of the above method can be completed by the integrated logic circuits in the processor hardware or by software instructions.

[0061] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.

[0062] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.

[0063] Example 5 Embodiment 5 of the present invention provides a computer-readable storage medium.

[0064] A computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps of the method as described in Embodiment 1 of the present invention.

[0065] The detailed steps are the same as those provided in Example 1, and will not be repeated here.

[0066] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for improving the quality of industrial corpora collected at the gateway, characterized in that, include: Collect industrial data and convert it into initial corpus, then extract sensitive entity features from the initial corpus; Based on knowledge graphs, a corpus quality assessment system is defined. The initial corpus was scored using the aforementioned corpus quality assessment system. Based on the scoring results, the initial corpus is divided into high-quality corpus, good corpus, and invalid corpus. The good corpus is semantically enhanced, and the high-quality corpus and the semantically enhanced good corpus are used as the optimized corpus. Based on the characteristics of sensitive entities, the corresponding optimized corpus is divided into different security levels, and corresponding desensitization strategies are implemented for the optimized corpus of different security levels. The anonymized and optimized corpus is labeled with domain-standardized annotations to obtain standardized corpus pairs. The standardized corpus pairs are sent to the cloud server, the training results returned by the cloud server are received, and the parameter weights in the corpus quality evaluation system are updated.

2. The method for improving the quality of industrial corpus collected at the gateway as described in claim 1, characterized in that, While extracting sensitive entity features from the initial corpus, semantic structure features, temporal correlation features, and industrial domain entity features are also extracted.

3. The method for improving the quality of industrial corpus collected at the gateway as described in claim 1, characterized in that, The knowledge graph-based corpus quality assessment system defines five evaluation dimensions: semantic completeness, domain adaptability, temporal relevance, information density, and noise ratio. The specific evaluation formula is as follows: ; Where Q(C) is the overall quality score of the corpus C to be processed, and S1(C) is the score for the semantic integrity dimension. S2(C) is the domain suitability score, S3(C) is the temporal relevance score, S4(C) is the information density score, and S5(C) is the noise ratio score. This refers to the adaptive weighting coefficients for the corresponding scores.

4. The method for improving the quality of industrial corpus collected at the gateway as described in claim 1, characterized in that, The process involves dividing the corresponding optimized corpus into different security levels based on the characteristics of sensitive entities, and then implementing corresponding desensitization strategies for the optimized corpus at different security levels. Specifically: The optimized corpus is divided into three levels: top secret, confidential, and ordinary. For top-secret corpora, irreversible desensitization processing is performed using industrial semantic conformal one-way encryption technology; For classified corpora, semantic differential privacy technology is used for anonymization. No anonymization processing is performed on ordinary-level corpora.

5. The method for improving the quality of industrial corpus collected at the gateway as described in claim 4, characterized in that, In the process of desensitizing top-secret and confidential corpora, a semantic fidelity hard constraint mechanism is introduced.

6. The method for improving the quality of industrial corpus collected at the gateway as described in claim 1, characterized in that, The process of receiving the training results returned by the cloud server and updating the parameter weights in the corpus quality evaluation system is as follows: Receive the comprehensive loss function returned by the cloud server and update the weight coefficients in the corpus quality assessment system based on the iterative gain formula; The iterative gain formula is: ; In the formula, Let be the weighting coefficient of the quality assessment system formula in the t-th iteration; The weights are the updated weights; η is the edge learning rate. Let be the partial derivative of the overall loss function with respect to the weight coefficients of the i-th dimension.

7. A device for improving the quality of industrial corpus collected at a gateway, characterized in that, The device, applied to a gateway, includes: The feature extraction module is configured to: collect industrial data, convert the industrial data into initial corpus, and extract sensitive entity features from the initial corpus; The corpus quality assessment module is configured to: define a corpus quality assessment system based on a knowledge graph; use the corpus quality assessment system to score the quality of the initial corpus; divide the initial corpus into high-quality corpus, good corpus, and invalid corpus according to the scoring results; perform semantic enhancement on the good corpus; and use the high-quality corpus and the semantically enhanced good corpus as optimized corpus. The desensitization module is configured to: divide the corresponding optimized corpus into different security levels based on the characteristics of sensitive entities; execute corresponding desensitization strategies for the optimized corpus at different security levels; and perform domain-standardized annotation on the desensitized optimized corpus to obtain standardized corpus pairs. The transmission and iterative update module is configured to: send the standardized corpus pairs to the cloud server, receive the training results returned by the cloud server, and update the parameter weights in the corpus quality evaluation system.

8. A system for improving the quality of industrial corpora collected at gateways, characterized in that, This includes both the gateway and the cloud server. The gateway is used to connect to the machine equipment, and the gateway is connected to the cloud server via a network. The gateway is used to collect industrial data, convert the industrial data into initial corpus, and extract sensitive entity features from the initial corpus. Based on knowledge graphs, a corpus quality assessment system is defined; the initial corpus is then scored using the corpus quality assessment system. Based on the scoring results, the initial corpus is divided into high-quality corpus, good corpus, and invalid corpus. The good corpus is semantically enhanced, and the high-quality corpus and the semantically enhanced good corpus are used as the optimized corpus. Based on the characteristics of sensitive entities, the corresponding optimized corpus is divided into different security levels, and corresponding desensitization strategies are implemented for the optimized corpus of different security levels. The desensitized optimized corpus is labeled with domain-standardized annotations to obtain standardized corpus pairs; the standardized corpus pairs are sent to the cloud server, the training results returned by the cloud server are received, and the parameter weights in the corpus quality evaluation system are updated. The cloud server is used to receive standardized corpus pairs, fine-tune the pre-trained large model based on the standardized corpus pairs, and send the training results to the gateway.

9. A gateway device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the method for improving the quality of industrial corpus acquisition at a gateway as described in any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the method for improving the quality of industrial corpus acquisition at the gateway as described in any one of claims 1-6.