An industrial health data purification method and system based on cloud-edge collaboration

By using continuous read interval analysis based on register addresses and dynamic packet processing, combined with asynchronous event-driven scheduling and edge cache queues, the high-frequency concurrency problem of industrial data processing under centralized architecture is solved, and efficient, real-time industrial health data purification and consistency analysis are achieved.

CN122640469APending Publication Date: 2026-08-25ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610742914.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In existing technologies, industrial data processing methods adopt a centralized architecture, which leads to high data link pressure, poor real-time transmission, and high cloud storage resource consumption in high-frequency concurrency scenarios. Furthermore, the lack of an efficient event-driven scheduling mechanism results in insufficient data integrity and accuracy.

Method used

By collecting the operating parameters of the target industrial equipment based on the preset register address, and performing continuous reading interval analysis and dynamic packet processing in combination with the register address distribution, asynchronous event-driven scheduling is realized. Batch uploading is performed in combination with the edge cache queue, and consistency analysis is performed using standard process parameters to obtain a healthy industrial dataset.

Benefits of technology

It improves the efficiency and real-time performance of industrial data acquisition, reduces communication redundancy and polling overhead, achieves low-latency processing and cloud transmission pressure sharing in high-concurrency scenarios, and ensures the accuracy and consistency of industrial health data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122640469A_ABST
    Figure CN122640469A_ABST
Patent Text Reader

Abstract

The application provides a cloud-edge collaboration-based industrial health data purification method and system, and relates to the technical field of data processing. The method comprises the following steps: collecting operation parameters of a target industrial equipment; sorting and spacing analyzing a target register to obtain a continuous reading interval; performing dynamic packet processing on the continuous reading interval to obtain data reading messages corresponding to each packet interval; polling and reading the target industrial equipment based on the data reading messages to obtain industrial time series data; performing asynchronous event-driven scheduling on the industrial time series data and writing the industrial time series data after the asynchronous event-driven scheduling into an edge cache queue; batch-asynchronously transmitting the industrial time series data in the edge cache queue to a cloud server according to a preset uploading strategy; comparing the industrial time series data to obtain a process deviation result; and performing group consistency analysis on the same batch of industrial equipment according to the process deviation result to obtain a healthy industrial data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for purifying industrial health data based on cloud-edge collaboration. Background Technology

[0002] With the continuous improvement of the level of intelligence in the textile industry, the twisting machines, spindles and supporting transmission equipment in the twisting workshop are gradually being digitally connected, and a large number of operating parameters can be collected in real time through industrial communication protocols. Especially in the twisting production process, each spindle will continuously generate high-frequency time-series data such as rotation speed, tension, temperature and process status, and the data scale exhibits the characteristics of high concurrency, high density and continuous streaming growth.

[0003] Due to the complex environment in the twisting workshop, the operating status of equipment is easily affected by factors such as mechanical vibration, transient fluctuations, communication jitter, and process switching. This results in the presence of noise data, abnormal drift data, and invalid contamination data in the original industrial time-series data. Therefore, it is necessary to establish an industrial health data purification mechanism based on cloud-edge collaboration. This mechanism uses the edge computing layer to asynchronously schedule and cache high-frequency industrial data for short periods. Combined with standard process parameters issued by the central cloud, it performs group consistency analysis and health status screening on equipment in the same batch, thereby obtaining a stable and reliable healthy industrial dataset.

[0004] However, most existing methods for industrial data processing employ a centralized architecture, where edge devices directly upload raw data to a central server for unified analysis. This approach suffers from high data link pressure, poor real-time transmission, and high cloud storage resource consumption when facing high-frequency concurrent industrial scenarios. Furthermore, existing solutions lack efficient event-driven scheduling mechanisms for the large number of asynchronous requests and responses generated during industrial communication, easily leading to blocking, message loss, and low data parsing efficiency, thus affecting the integrity and accuracy of industrial time-series data. Summary of the Invention

[0005] To address the issue that most existing methods for industrial data processing employ a centralized architecture, where edge devices directly upload raw data to a central server for unified analysis, this approach suffers from technical problems such as high data link pressure, poor real-time transmission performance, and high cloud storage resource consumption when facing high-frequency concurrent industrial scenarios.

[0006] The technical solution provided by this invention is as follows: S1: Collect operating parameters of the target industrial equipment based on the preset register address; S2: Based on the register address distribution corresponding to the running parameters, sort and analyze the spacing of the target registers to obtain the continuous reading interval; S3: Combined with the preset register number threshold, perform dynamic packet processing on the continuous read interval to obtain the data read message corresponding to each packet interval; S4: Based on the data reading message, poll and read the target industrial equipment to obtain industrial time-series data in a unified format; S5: Perform asynchronous event-driven scheduling on industrial time-series data and write the industrial time-series data after asynchronous event-driven scheduling into the edge cache queue; S6: According to the preset upload strategy, the industrial time series data in the edge cache queue is transmitted asynchronously in batches to the cloud server; S7: Based on standard process parameters, compare the operating status of industrial time-series data to obtain process deviation results; S8: Based on the process deviation results, perform a population consistency analysis on the same batch of industrial equipment to obtain a healthy industrial dataset.

[0007] A second aspect of this invention provides an industrial health data purification system based on cloud-edge collaboration, comprising: a processor and a memory; The memory stores programs or instructions that can run on the processor, and when the programs or instructions are executed by the processor, they implement the steps of the cloud-edge collaborative industrial health data purification method of the first aspect.

[0008] The beneficial effects of the technical solution provided by this invention include: By collecting operating parameters of target industrial equipment based on preset register addresses and performing continuous interval analysis and dynamic packet processing in conjunction with register address distribution, optimized organization of industrial communication messages is achieved. This reduces communication redundancy and polling overhead, improving data acquisition efficiency in high-concurrency industrial scenarios. Asynchronous event-driven scheduling of industrial time-series data, combined with short-term buffering and batch asynchronous uploading via edge cache queues, enables low-latency processing of high-frequency data at the edge and reduces cloud transmission pressure, thereby improving the real-time performance and system stability of industrial data transmission. Simultaneously, by comparing operating status based on standard process parameters and analyzing the group consistency of industrial equipment in the same batch, baseline offset anomalies, local mutation anomalies, and high-frequency oscillation anomalies in the original industrial time-series data are identified and removed. This effectively purifies industrial health data, providing a high-quality, highly consistent healthy industrial dataset for subsequent anomaly detection model training and online fault diagnosis. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating a cloud-edge collaborative method for purifying industrial health data, provided as an embodiment of the present invention.

[0010] Figure 2This is a schematic diagram of the structure of an industrial health data purification system based on cloud-edge collaboration, provided as an embodiment of the present invention. Detailed Implementation

[0011] S1: Collect the operating parameters of the target industrial equipment based on the preset register address.

[0012] In one possible implementation, the operating parameters include: rotational speed, tension, temperature, and process status data.

[0013] The preset register address refers to the data storage address pre-configured according to the target industrial equipment's communication protocol, equipment manual, or controller register mapping table. It identifies the corresponding location of different operating parameters within the equipment's internal storage space. The target industrial equipment refers to industrial field equipment requiring operational status monitoring and health analysis, such as a doubling twister, spindle drive unit, motor controller, or tension control module. Operating parameters refer to data variables that reflect the real-time operating status of industrial equipment.

[0014] It should be noted that by selectively collecting operating parameters of target industrial equipment based on preset register addresses, edge computing nodes can directly obtain core data related to the operating status from within the industrial equipment. This avoids data redundancy issues associated with traditional manual configuration or random polling methods, improving the targeting and accuracy of industrial data collection. Furthermore, by pre-establishing a mapping relationship between register addresses and operating parameters, unified data access and standardized management across different industrial devices can be achieved, thereby enhancing the data compatibility of heterogeneous devices in the industrial field.

[0015] S2: Based on the register address distribution corresponding to the running parameters, sort and analyze the spacing of the target registers to obtain the continuous reading interval.

[0016] Register address distribution refers to the arrangement and dispersion of register addresses corresponding to various operating parameters in the storage space of industrial equipment. Target registers refer to the set of registers used to store target operating parameter data. Continuous read intervals refer to dividing multiple adjacent registers or registers meeting preset spacing conditions into the same batch read range based on the spacing relationship between register addresses, so that multiple register data can be read in a single communication message.

[0017] It should be noted that by sorting and analyzing the spacing of register addresses corresponding to operating parameters, the originally discrete target registers are reorganized into multiple continuous read intervals. This allows edge computing nodes to read multiple operating parameters in a batch manner, thereby reducing the number of message transmissions and redundant connection overhead during industrial communication and improving the data acquisition efficiency of industrial equipment. Furthermore, analyzing register spacing can prevent the problem of a large number of invalid registers being read due to excessively large address spans, thus reducing communication bandwidth usage and equipment response load.

[0018] S3: Combined with the preset register number threshold, perform dynamic packet processing on the continuous read interval to obtain the data read message corresponding to each packet interval.

[0019] The preset register count threshold refers to the maximum number of registers allowed to be read in a single communication message, used to limit the length of a single data read message to avoid excessive communication load. Dynamic packet splitting refers to adaptively splitting the read interval based on the number of registers, address span, and communication constraints in the current continuous read interval to generate multiple data read sub-intervals that meet communication requirements. The data read message refers to a data request message generated based on an industrial communication protocol, used to send register read instructions to the target industrial equipment and obtain corresponding operating parameter data.

[0020] It should be noted that by combining a preset register count threshold with dynamic packet segmentation for continuous read intervals, industrial communication messages can be adaptively adjusted according to the actual register distribution. This avoids communication timeouts, slow device response, or message parsing failures caused by excessive register reads in a single session, thus improving the stability and reliability of industrial communication. Simultaneously, by jointly judging the idle interval and estimated message length, redundant reads of invalid register addresses can be effectively reduced, lowering communication bandwidth usage and device resource consumption, thereby improving industrial data acquisition efficiency.

[0021] In one possible implementation, S3 specifically includes: S301: Obtain the target register address sequence corresponding to the continuous read interval.

[0022] S302: Sort the target register address sequence in ascending order to obtain the target register address list.

[0023] S303: Use the first register address in the target register address list as the starting address of the current packet segment, and initialize the number of registers corresponding to the current packet segment.

[0024] The number of registers refers to the number of valid registers contained in the current packet interval.

[0025] S304: Traverse the list of target register addresses.

[0026] S305: Calculate the unloaded gap between the current target register address and the end address of the current packet interval.

[0027] The unoccupied interval refers to the address span between the current target register address and the end address of the current packet interval that is not occupied by a valid register.

[0028] S306: Determine the estimated message length by combining the idle interval and the number of registers after initialization.

[0029] The estimated message length refers to the expected size of the communication message calculated based on the number of registers and address span corresponding to the current packet segment.

[0030] S307: Determine whether the estimated message length is less than a preset register quantity threshold and whether the idle interval is less than a preset interval threshold. If so, merge the current target register address into the current packet segment and update the register quantity corresponding to the current packet segment. Otherwise, determine the current packet segment as the target data reading segment and generate data reading messages corresponding to each packet segment based on the target data reading segment.

[0031] The preset register number threshold refers to the maximum number of registers allowed in a single data read message. The preset spacing threshold refers to the maximum unused address span between adjacent registers. The target data read interval refers to the final data read range determined after meeting the communication constraints.

[0032] It should be noted that by sorting and traversing the target register address sequence in ascending order, and combining this with dynamic packet processing based on idle intervals, the number of registers, and the estimated message length, the system can adaptively generate multiple reasonably structured data read messages according to the actual distribution of industrial equipment registers. This avoids the problem of reading a large number of invalid addresses in traditional fixed-length read methods, improving the utilization rate of industrial communication resources. Furthermore, by jointly judging the estimated message length and a preset threshold, it can effectively prevent equipment response delays, communication blockages, and message timeouts caused by excessively long single communication messages, improving the stability and real-time performance of the industrial data acquisition process. In addition, by dynamically dividing the target data read interval, the communication load of edge computing nodes in high-concurrency polling scenarios can be reduced, providing a more granular and efficient data read foundation for subsequent asynchronous event-driven scheduling and batch data processing, thereby improving the throughput and operational reliability of the entire cloud-edge collaborative industrial health data purification system.

[0033] S4: Based on the data reading message, poll and read the target industrial equipment to obtain industrial time-series data in a unified format.

[0034] Polling read refers to a data acquisition method in which edge computing nodes periodically send data read messages to target industrial equipment according to a preset order and receive corresponding response messages. Unified format industrial time-series data refers to industrial operation data that has been parsed and standardized, possessing unified data fields, timestamp formats, and parameter structures for subsequent storage and analysis.

[0035] In one possible implementation, S4 specifically includes: S401: Create an asynchronous Future object corresponding to the data read message and register the asynchronous Future object to the mailbox pool.

[0036] In this context, an asynchronous Future object is a non-blocking object used to represent the result of an asynchronous task, and it is used to associate data read requests with corresponding response results. A mailbox pool is a mapping container used to store and manage multiple asynchronous Future objects, enabling fast matching between requests and responses.

[0037] S402: Write the data read message to the send queue and switch the business coroutine to the non-blocking waiting state corresponding to the asynchronous Future object to wait for the return of the corresponding response message.

[0038] Here, the sending queue refers to the data queue used to buffer data read messages to be sent. A service coroutine is a lightweight asynchronous execution unit implemented based on a coroutine mechanism, used to schedule industrial communication tasks in non-blocking mode. The non-blocking waiting state refers to the asynchronous waiting state in which the service coroutine does not occupy main thread resources while waiting for a response message. The response message refers to the data response message returned by the target industrial equipment in response to the data read message.

[0039] S403: Listen to the response messages returned by the target industrial equipment and extract the message prefix information from the response messages.

[0040] Among them, message prefix information refers to the information field located in the response message header that is used to identify the business type or the source of the request.

[0041] S404: Determine the service type corresponding to the response message based on the message prefix information.

[0042] The service type refers to the data processing category to which the current communication task belongs.

[0043] S405: If an asynchronous Future object corresponding to the business type is found in the mailbox pool, the response message is delivered to the corresponding asynchronous Future object.

[0044] In one possible implementation, if no asynchronous Future object corresponding to the business type is found in the mailbox pool, the response message is determined to be an ownerless response, and the ownerless response is logged or discarded.

[0045] S406: Based on the delivery result, the response message is parsed through a business coroutine to obtain industrial time-series data in a unified format.

[0046] It should be noted that by performing asynchronous polling reads on target industrial equipment based on data read messages, and utilizing asynchronous Future objects, mailbox pools, and business coroutines to construct an event-driven data interaction mechanism, the industrial communication process can complete request sending and response processing in a non-blocking mode. This avoids the thread resource waste caused by blocking and waiting in traditional synchronous communication methods, improving data acquisition efficiency in high-concurrency industrial scenarios. Simultaneously, by quickly matching and targeting the business type of response messages, precise association between requests and responses can be achieved, reducing the probability of message mismatches, packet loss, and parsing anomalies, thereby improving the stability and reliability of the industrial communication process. Furthermore, by performing unified format data parsing on response messages, standardized processing between different industrial devices and different protocol data can be achieved, providing a structurally consistent industrial time-series data foundation for subsequent asynchronous event scheduling, edge caching, and cloud health data purification, thereby improving the real-time performance, scalability, and compatibility of the entire cloud-edge collaborative industrial data processing system.

[0047] S5: Perform asynchronous event-driven scheduling on industrial time-series data and write the industrial time-series data after asynchronous event-driven scheduling into the edge cache queue.

[0048] Asynchronous event-driven scheduling refers to a non-blocking scheduling and execution method for industrial data processing tasks based on an event-triggered mechanism. Different business tasks can run asynchronously according to their corresponding events without waiting for other tasks to complete. Edge cache queues refer to data buffer storage structures deployed in edge computing nodes, used to temporarily cache industrial time-series data to be uploaded.

[0049] In one possible implementation, S5 specifically includes: S501: Establish a template dictionary corresponding to the business type. The template dictionary includes function address, business parameters, request message configuration, and context key-value pairs.

[0050] Here, the template dictionary refers to a pre-established business template mapping structure used to store configuration parameters and processing rules corresponding to different business types. The function address refers to the call path used to locate specific business processing functions or modules. Business parameters refer to the parameter configurations required to execute the corresponding business logic. The request message configuration refers to the communication format, register address, and protocol field configurations required to generate industrial communication request messages. Context key-value pairs refer to information fields used to describe the business operating environment, device identification, or task status.

[0051] S502: In response to a business task trigger command, load the target business template corresponding to the business type from the template dictionary.

[0052] Among them, the business task trigger instruction refers to the control signal used to trigger the corresponding industrial data processing task. The target business template refers to the template configuration loaded from the template dictionary that corresponds to the current business type.

[0053] S503: Combine the function addresses in the target business template and load the business execution module corresponding to the business type through a dynamic import mechanism.

[0054] The dynamic import mechanism refers to the method of loading corresponding business execution modules on demand during program execution. A business execution module is a software functional module used to complete data parsing, calculation, or business logic processing.

[0055] S504: Generate a data read message corresponding to the target industrial equipment based on the request message configuration in the target business template.

[0056] S505: Based on the business parameters in the target business template, construct the business execution parameters for the business execution module to call.

[0057] S506: Pass the business execution parameters to the business execution module to perform data parsing on the data read message and obtain the data parsing result.

[0058] Among them, the data parsing result refers to the data result obtained after decoding, converting and structuring the industrial response message.

[0059] S507: Based on the data parsing results, generate asynchronous event-driven scheduled industrial time-series data and write the scheduled industrial time-series data into the edge cache queue.

[0060] It should be noted that by establishing a template dictionary corresponding to business types and combining it with an asynchronous event-driven scheduling mechanism to dynamically process industrial time-series data, the system can flexibly load corresponding business execution modules according to different industrial business needs. This avoids the problems of high coupling and poor scalability in traditional fixed-process architectures, improving the modularity and scalability of the industrial data processing system. Simultaneously, by generating corresponding data reading messages and business execution parameters through a dynamic import mechanism and business template configuration, rapid adaptation between different industrial equipment and different business logics can be achieved, thereby improving data processing compatibility in heterogeneous industrial scenarios. Furthermore, by writing the scheduled industrial time-series data into an edge cache queue, short-term buffering and peak-shaving processing of high-frequency industrial data can be achieved, reducing the impact of instantaneous cloud transmission pressure and network fluctuations on system stability. This improves the real-time performance of data processing, task scheduling efficiency, and system reliability of the entire cloud-edge collaborative industrial health data purification system.

[0061] S6: Based on the preset upload strategy, the industrial time-series data in the edge cache queue is transmitted asynchronously in batches to the cloud server.

[0062] The preset upload strategy refers to the pre-defined rules for uploading industrial data, which control the conditions under which industrial time-series data in the edge cache queue is uploaded. These conditions include time period, cache capacity, data priority, network status, or abnormal triggering conditions. The cloud server refers to the data processing platform deployed in the central cloud layer, used for storing, analyzing, refining, and detecting anomalies in the uploaded industrial time-series data.

[0063] It should be noted that by performing batch asynchronous transmission of industrial time-series data in the edge cache queue according to a preset upload strategy, edge computing nodes can dynamically adjust the data upload rhythm based on the network status and data load of the industrial site. This avoids the network congestion and cloud storage pressure problems caused by continuous high-frequency data uploads in traditional real-time direct transmission modes, thereby improving the stability and resource utilization of industrial data transmission. Simultaneously, batch uploading of industrial time-series data reduces the number of communication connections and protocol interaction overhead, thus improving data transmission efficiency and reducing network bandwidth consumption.

[0064] In one possible implementation, the process after S6 and before S7 includes: The industrial time-series data, after being transmitted asynchronously in batches, are stored in relational databases and non-relational time-series databases respectively.

[0065] S7: Based on standard process parameters, compare the operating status of industrial time-series data to obtain process deviation results.

[0066] Standard process parameters refer to a set of reference parameters determined by the central cloud layer based on normal production processes, historical healthy operation data, or equipment operating standards. These parameters describe the target operating range of industrial equipment under healthy conditions and include standard speed, standard tension, standard temperature, and process state thresholds. Process deviation results refer to the deviation analysis results between the current operating state of industrial equipment and the standard process parameters. They are used to characterize whether the equipment has operational abnormalities, process fluctuations, or state drift.

[0067] It should be noted that by comparing the operational status of industrial time-series data with standard process parameters, the system can quickly identify data changes deviating from standard process conditions during the real-time operation of industrial equipment. This enables early detection of abnormal operating trends, process drift, and potential fault states, improving the accuracy and real-time performance of industrial equipment operational status monitoring. Simultaneously, comparing industrial time-series data with unified standard process parameters eliminates data interference caused by differences in processes, operating loads, or environmental fluctuations between different equipment. This improves data consistency during subsequent health data purification and anomaly analysis. Furthermore, generating process deviation results provides crucial input for subsequent group consistency analysis, enabling the system to further identify abnormal equipment and unhealthy data. This enhances the reliability and purity of the healthy industrial dataset and provides a more stable data foundation for subsequent anomaly detection model training and online fault diagnosis.

[0068] S8: Based on the process deviation results, perform a population consistency analysis on the same batch of industrial equipment to obtain a healthy industrial dataset.

[0069] Among them, "industrial equipment in the same batch" refers to a set of industrial equipment that performs production tasks under the same production process, the same production time period, and the same operating conditions. "Healthy industrial dataset" refers to the final set of high-quality healthy operation data obtained through screening.

[0070] In one possible implementation, S8 specifically includes: S801: Obtain the rotational speed deviation sequence of each spindle in the same batch of industrial equipment.

[0071] Here, a spindle refers to a single-node execution unit in a twisting machine, used to complete the yarn twisting operation. The speed deviation sequence refers to the sequence of deviation data formed over time between the actual operating speed of each spindle and the standard process speed.

[0072] S802: Constructing a group reference trajectory for single-node spindles based on rotational speed deviation sequence:

[0073] in, C ( t) indicates the group baseline reference trajectory at time step t Reference value, The standard deviation is expressed as The Gaussian kernel function, median() represents the median operation. d s ( t ) indicates the first s Each spindle in time step t The speed deviation value, S This indicates the total number of spindles.

[0074] Among them, the group benchmark reference trajectory refers to the group health operation reference curve generated by median statistics and Gaussian smoothing based on the operating status of the same batch of spindles.

[0075] S803: Calculate the first-round robust Modified Z-score by combining the single-spindle time-series mean, the median population mean, and the initial absolute deviation.

[0076]

[0077] in, Indicates the first s The first round robust Modified Z-score of each spindle. Indicates the first s The time-series mean of each spindle, This represents the median of the time series mean for all spindles. Indicates the first s The absolute deviation of the median of the first round of spindles.

[0078] Among them, the robust Modified Z-score is a robust anomaly scoring index constructed based on the absolute deviation between the median and the median, used to measure the degree to which the operating state of a single spindle deviates from the center of the group.

[0079] S804: By using a preset lenient threshold, the first round of robust Modified Z-scores is used for initial screening to obtain the first surviving dataset.

[0080] The first surviving dataset refers to the set of data that survives the first round of anomaly screening.

[0081] S805: Calculate the second round robust Modified Z-score based on the first surviving dataset.

[0082] S806: Combine the preset dynamic threshold to filter the second round of robust Modified Z-score scores and identify abnormal baseline offset data.

[0083] Among them, baseline offset anomaly data refers to data whose overall operating trend deviates from the group's baseline reference trajectory over a long period of time.

[0084] S807: By using a one-dimensional Gaussian smoothing filter to remove the low-frequency trend from the speed deviation sequence, a high-frequency noise residual sequence is obtained.

[0085] in, Indicates the first s Each spindle in time step t The high-frequency noise residual value.

[0086] Among them, the high-frequency noise residual sequence refers to the high-frequency fluctuation data retained after stripping away the low-frequency trend.

[0087] S808: Calculate the first-order forward difference peak value and local rolling standard deviation based on the high-frequency noise residual sequence.

[0088] Among them, the first-order forward difference peak value refers to the peak value of the data change between adjacent time steps, which is used to characterize the degree of abrupt change. The local rolling standard deviation is a data fluctuation intensity index calculated within a local time window.

[0089] S809: Combining the first-order forward difference peak value and the local rolling standard deviation, local abrupt change anomaly data and high-frequency oscillation anomaly data are identified.

[0090] Local mutation anomaly data refers to data that undergoes drastic changes within a short period of time. High-frequency oscillation anomaly data refers to data that exhibits continuous high-frequency fluctuations.

[0091] S810: Remove baseline offset anomalies, local mutation anomalies, and high-frequency oscillation anomalies to obtain the dataset to be scored.

[0092] The dataset to be scored refers to the set of candidate health data that remains after removing outlier data.

[0093] S811: Calculate the comprehensive health score of the dataset to be scored by combining the group baseline reference trajectory.

[0094] in, H s Indicates the first s The overall health score of each spindle This represents the weight adjustment coefficient for the RMSE dimension. This represents the weight adjustment coefficient for the Osc dimension. This represents the weight adjustment coefficient for the Jerk dimension. Indicates the first s Individual spindle and group baseline reference trajectory C The root mean square error between them This represents the median of the RMSE index for the entire sample. Indicates the first s High-frequency oscillation indicators of individual spindles This represents the median of the Osc index for all samples. Indicates the first s Transient change index of each spindle, This represents the median of the Jerk index for all samples.

[0095] Among them, the comprehensive health score refers to the health evaluation index calculated based on the root mean square error, high-frequency oscillation index and transient change index.

[0096] S812: Based on the comprehensive health score and dynamic cutoff threshold, the dataset to be scored is sorted in ascending order and layered to obtain the core health layer, the sub-health isolation layer, and the abnormal layer.

[0097] The dynamic cutoff threshold refers to the stratified threshold that is dynamically adjusted based on the current distribution of group data. The core health layer refers to the data set with the best overall health score and the most stable operational status. The sub-health isolation layer refers to the data set with minor abnormalities but not reaching the level of serious malfunction. The abnormal layer refers to the data set that significantly deviates from the group's health status.

[0098] S813: Determine the data corresponding to the health core layer as the health industry dataset for output.

[0099] It should be noted that by performing group consistency analysis on industrial equipment of the same batch based on process deviation results, and combining robust Modified Z-score analysis, Gaussian smoothing filtering, and a multi-dimensional health scoring mechanism, baseline offset anomalies, local mutation anomalies, and high-frequency oscillation anomalies in industrial time-series data are screened and eliminated layer by layer. This enables the system to accurately identify truly stable and representative healthy operation data from a large amount of high-frequency industrial data, thereby avoiding the problems of high false positive rate and weak noise resistance in traditional single-threshold screening methods, and improving the accuracy and robustness of industrial health data purification. At the same time, by constructing a group baseline reference trajectory and adopting a group consensus mechanism for health status assessment, the impact of random fluctuations of single equipment, local process interference, and sensor noise on the analysis results can be effectively reduced, thereby improving the consistency and credibility of the health dataset. In addition, by hierarchically managing the dataset to be scored, a clear and reliable data foundation can be provided for subsequent anomaly detection model training, online fault diagnosis, and process optimization analysis, thereby improving the modeling accuracy, generalization ability, and real-time diagnostic performance of the industrial intelligent detection system.

[0100] In one possible implementation, the process after S8 includes: S9: Determine the anomaly detection results based on the health industry dataset.

[0101] Among them, the abnormal detection results refer to the analysis results of abnormal operating states identified based on the health industry dataset, which are used to characterize whether industrial equipment has a failure trend, process deviation, local abnormal fluctuation or potential failure risk.

[0102] It should be noted that by determining anomaly detection results based on a healthy industrial dataset, the subsequent anomaly identification process is built upon purified, high-quality health data. This avoids the high false alarm rate and low detection accuracy problems of traditional anomaly detection methods, which suffer from high noise, drift, and contamination in the raw data. This improves the accuracy and reliability of industrial anomaly detection results. Furthermore, by constructing a stable health operation reference model using the healthy industrial dataset, the anomaly detection system's ability to identify subtle anomaly trends, latent faults, and complex process fluctuations can be enhanced, thereby improving the sensitivity and real-time performance of industrial equipment fault warnings.

[0103] Specifically, in traditional centralized architectures, all device data is directly reported to a single relational database in the cloud. However, when faced with high-concurrency writes from twisting machine clusters, traditional row-level locking and strict transaction verification mechanisms are prone to database connection pool exhaustion and deadlocks, causing the system to be unable to meet the performance requirements of high-throughput writes and multi-table join queries. To address this, this system designs a cloud-edge collaborative hybrid storage architecture of "edge buffering + cloud persistence".

[0104] (1) High-frequency concurrent buffer pool based on MongoDB At the edge computing layer, the system uses the NoSQL database MongoDB as a high-concurrency buffer. Because the high-frequency data from the twisting machine is characterized by large data volume and high write frequency but no strong transactional requirements, the system encapsulates the parsed single-spindle data into BSON format and stores it directly into MongoDB. This approach eliminates the need for pre-defining a strict two-dimensional table structure, thus avoiding table locking overhead.

[0105] Furthermore, due to the limited disk capacity of edge computing nodes, the system utilizes MongoDB's built-in TTL (Time-To-Live) automatic index expiration mechanism, setting the data retention period to 7 days. This mechanism automatically cleans up outdated data in the form of a sliding window, ensuring that the single-node real-time anomaly inference (H-1CLA) at the edge computing layer has sufficient recent data while simultaneously cleaning up edge storage resources.

[0106] (2) Data persistence and alignment based on "relational + non-relational" At the central cloud layer, the system adopts a hybrid storage strategy of "MySQL + MongoDB" as shown below: 1) MySQL focuses on managing low-frequency business data with high consistency requirements, such as twisting machine ledgers, batch process parameter tables, equipment topology relationships, and operation logs, and supports business join table retrieval.

[0107] 2) MongoDB is specifically responsible for receiving historical data uploaded by edge computing nodes and performing long-term persistent cold backup archiving.

[0108] The system developed a lightweight hybrid data scheduling middleware at the central cloud layer. When the central cloud layer performs offline model training or initiates a yarn quality traceability request, the middleware performs a two-stage cross-database alignment query: First, the middleware accesses the MySQL database and uses relational indexes to lock the target production batch, the corresponding machine space topology matrix, and the precise time window boundary; then, using this time boundary and machine ID as a composite index condition, it drives the middleware to send concurrent aggregation retrieval requests to the cloud MongoDB to quickly pull the corresponding high-frequency time-series speed slices.

[0109] The cloud-edge collaborative hybrid storage strategy designed in this section can alleviate the concurrent write pressure of the central cloud layer. In terms of logical architecture, it achieves the separation of transaction consistency and high throughput, and provides data support for subsequent anomaly detection models (H-1CLA and MCT) to extract data and schedule features.

[0110] The beneficial effects of the technical solution provided by this invention include: By collecting material parameters, processing environment temperature, and machine tool operating status in real time using multi-source sensors, and combining finite element simulation and dimensionality reduction modeling techniques, a low-dimensional parametric expression model of the residual stress field is constructed. Furthermore, temperature field information is integrated to establish a thermo-mechanical coupled deformation model, enabling efficient prediction of structural deformation of alloy plates under complex processing conditions. Simultaneously, by constructing a rapid prediction model of the machine tool temperature field and calculating its thermal displacement error, this model is fused with workpiece deformation under a unified processing coordinate system to form a comprehensive deformation error field. This achieves system-level unified modeling and accurate characterization of processing error sources. Based on this, by introducing processing accuracy constraints to optimize the control variables, the optimal compensation parameters are obtained. This enables feedforward compensation and dynamic correction of errors during processing, effectively improving processing accuracy and stability, reducing trial-and-error costs, and enhancing the model's real-time performance and engineering applicability.

[0111] This invention provides a readable storage medium that stores a program or instructions on the medium. When the program or instructions are executed by a processor, they implement the steps of the above-described cloud-edge collaborative industrial health data purification method and achieve the same technical effect. To avoid repetition, this invention will not elaborate further.

[0112] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for refining industrial health data based on cloud-edge collaboration, characterized in that, include: S1: Collect operating parameters of the target industrial equipment based on the preset register address; S2: Based on the register address distribution corresponding to the running parameters, sort and analyze the spacing of the target registers to obtain the continuous reading interval; S3: Combine the preset register number threshold to perform dynamic packet processing on the continuous reading interval to obtain the data reading message corresponding to each packet interval; S4: Based on the data reading message, poll the target industrial equipment to obtain industrial time-series data in a unified format; S5: Perform asynchronous event-driven scheduling on the industrial time-series data, and write the industrial time-series data after asynchronous event-driven scheduling into the edge cache queue; S6: According to the preset upload strategy, the industrial time-series data in the edge cache queue is transmitted asynchronously in batches to the cloud server; S7: Based on standard process parameters, the operating status of the industrial time series data is compared to obtain the process deviation results; S8: Based on the process deviation results, perform a population consistency analysis on the same batch of industrial equipment to obtain a healthy industrial dataset.

2. The method for purifying industrial health data based on cloud-edge collaboration according to claim 1, characterized in that, The operating parameters include: rotational speed, tension, temperature, and process status data.

3. The method for purifying industrial health data based on cloud-edge collaboration according to claim 1, characterized in that, S3 specifically includes: S301: Obtain the target register address sequence corresponding to the continuous read interval; S302: Sort the target register address sequence in ascending order to obtain a target register address list; S303: Using the first register address in the target register address list as the starting address of the current packet interval, initialize the number of registers corresponding to the current packet interval; S304: Traverse the target register address list; S305: Calculate the unused interval between the current target register address and the end address of the current packet interval; S306: Determine the estimated message length by combining the idle spacing and the number of registers after initialization; S307: Determine whether the estimated message length is less than a preset register number threshold and whether the idle interval is less than a preset interval threshold; if so, merge the current target register address into the current packet interval and update the register number corresponding to the current packet interval; otherwise, determine the current packet interval as the target data reading interval and generate data reading messages corresponding to each packet interval based on the target data reading interval.

4. The method for purifying industrial health data based on cloud-edge collaboration according to claim 1, characterized in that, S4 specifically includes: S401: Create an asynchronous Future object corresponding to the data read message, and register the asynchronous Future object to the mailbox pool; S402: Write the data read message into the sending queue and switch the business coroutine to the non-blocking waiting state corresponding to the asynchronous Future object to wait for the return of the corresponding response message; S403: Listen to the response message returned by the target industrial equipment and extract the message prefix information in the response message; S404: Determine the service type corresponding to the response message based on the message prefix information; S405: If an asynchronous Future object corresponding to the business type is found in the mailbox pool, the response message is delivered to the corresponding asynchronous Future object; S406: Based on the delivery result, the response message is parsed using a business coroutine to obtain industrial time-series data in a unified format.

5. The method for purifying industrial health data based on cloud-edge collaboration according to claim 4, characterized in that, If no asynchronous Future object corresponding to the business type is found in the mailbox pool, the response message is determined to be an ownerless response, and the ownerless response is logged or discarded.

6. The method for purifying industrial health data based on cloud-edge collaboration according to claim 4, characterized in that, S5 specifically includes: S501: Establish a template dictionary corresponding to the business type, wherein the template dictionary includes function address, business parameters, request message configuration and context key-value pairs; S502: In response to the business task triggering instruction, load the target business template corresponding to the business type from the template dictionary; S503: Combining the function addresses in the target business template, load the business execution module corresponding to the business type through a dynamic import mechanism; S504: Generate a data read message corresponding to the target industrial equipment according to the request message configuration in the target service template; S505: Based on the business parameters in the target business template, construct business execution parameters for the business execution module to call; S506: Pass the service execution parameters to the service execution module to perform data parsing on the data read message and obtain the data parsing result; S507: Based on the data parsing results, generate asynchronous event-driven scheduled industrial time-series data and write the scheduled industrial time-series data into the edge cache queue.

7. The method for purifying industrial health data based on cloud-edge collaboration according to claim 1, characterized in that, The process includes the following steps after S6 and before S7: The industrial time-series data, after being transmitted asynchronously in batches, are stored in relational databases and non-relational time-series databases respectively.

8. The method for purifying industrial health data based on cloud-edge collaboration according to claim 1, characterized in that, S8 specifically includes: S801: Obtain the rotational speed deviation sequence of each spindle in the same batch of industrial equipment; S802: Based on the aforementioned rotational speed deviation sequence, construct a group reference trajectory for the single-node spindle: ; in, C ( t ) indicates the group baseline reference trajectory at time step t Reference value, The standard deviation is expressed as The Gaussian kernel function, median() represents the median operation. d s ( t ) indicates the first s Each spindle in time step t The speed deviation value, S Indicates the total number of spindles; S803: Calculate the first-round robust Modified Z-score by combining the single-spindle time-series mean, the population mean median, and the initial absolute deviation. ; ; in, Indicates the first s The first round robust Modified Z-score of each spindle Indicates the first s The time-series mean of each spindle, This represents the median of the time series mean for all spindles. Indicates the first s The absolute deviation of the median in the first round of the spindles; S804: By using a preset lenient threshold, the robust Modified Z-score of the first round is initially screened to obtain the first surviving dataset; S805: Based on the first surviving dataset, calculate the second round robust Modified Z-score; S806: Combine the preset dynamic threshold to filter the second round of robust Modified Z-score scores and identify abnormal baseline offset data; S807: The low-frequency trend is stripped from the speed deviation sequence using a one-dimensional Gaussian smoothing filter to obtain a high-frequency noise residual sequence. ; in, Indicates the first s Each spindle in time step t The high-frequency noise residual value; S808: Calculate the first-order forward difference peak value and the local rolling standard deviation based on the high-frequency noise residual sequence; S809: Combine the first-order forward differential peak value and the local rolling standard deviation to determine the local abrupt change anomaly data and the high-frequency oscillation anomaly data; S810: Remove the baseline offset anomaly data, the local mutation anomaly data, and the high-frequency oscillation anomaly data to obtain the dataset to be scored; S811: Calculate the comprehensive health score of the dataset to be scored by combining the aforementioned group baseline reference trajectory. ; in, H s Indicates the first s The overall health score of each spindle This represents the weight adjustment coefficient for the RMSE dimension. This represents the weight adjustment coefficient for the Osc dimension. This represents the weight adjustment coefficient for the Jerk dimension. Indicates the first s Individual spindle and group baseline reference trajectory C The root mean square error between them This represents the median of the RMSE index for the entire sample. Indicates the first s High-frequency oscillation indicators of individual spindles This represents the median of the Osc index for all samples. Indicates the first s Transient change index of each spindle, This represents the median of the Jerk index for all samples. S812: Based on the comprehensive health score and the dynamic cutoff threshold, the dataset to be scored is sorted in ascending order and layered to obtain the health core layer, the sub-health isolation layer and the abnormal layer. S813: The data corresponding to the health core layer is determined as the health industrial dataset and output.

9. The method for purifying industrial health data based on cloud-edge collaboration according to claim 1, characterized in that, Following S8, the following is also included: S9: Based on the aforementioned health industry dataset, determine the anomaly detection results.

10. A cloud-edge collaborative industrial health data purification system, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement a cloud-edge collaborative industrial health data purification method as described in any one of claims 1 to 9.