Full life cycle quality trustworthy data collection method and related devices
By updating data density and deploying nodes through the ant colony algorithm, a trusted data path model for DCS equipment is constructed, which solves the problems of low data collection efficiency and delay throughout the life cycle of DCS equipment and achieves more efficient trusted data collection.
Patent Information
- Application Number
- CN202410709148.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-03
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-06-03
AI Technical Summary
The existing technology has low data collection efficiency throughout the life cycle of DCS equipment, the error elimination cycle has a great impact on data transmission, the amount of reliable data collected is small and the delay is serious.
The data density is updated through the ant colony algorithm, the trusted data deployment nodes are determined, and a trusted data path model is constructed to screen the target trusted data.
The data acquisition stability is improved, the data acquisition delay is reduced, and the impact of the error elimination cycle on data transmission is reduced.
Smart Images

Figure CN119484598B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present application relate to the technical field, and in particular to a method for collecting reliable quality data throughout the entire life cycle and related equipment. Background Art
[0002] A distributed control system (DCS) is a microprocessor-based system characterized by decentralized risk control and centralized operation and management. It integrates advanced computer, communication, CRT, and control technologies, known as the "4C" technology. With the rapid development of modern computer and communication network technologies, DCS is moving toward diversification, networking, openness, and integrated management. This allows different DCS models to interconnect and exchange data, and connect DCS systems to factory management networks via Ethernet, enabling real-time data access to the internet. This has become the mainstream of process industry automation control.
[0003] Generally speaking, the service life of DCS equipment is 20 years, which can be extended through upgrades and repairs. Therefore, it is necessary to obtain data on the entire life cycle of DCS equipment to develop a maintenance plan for DCS equipment.
[0004] Related technologies collect display value data from DCS devices and then filter out reliable data from this display value data using methods such as optimizing objective functions. However, this method requires collecting all quality data throughout the entire lifecycle before performing the screening, and ignores the dynamic nature of data density. Related technologies suffer from low data collection efficiency, a significant impact of the error elimination cycle on data transmission, a small amount of reliable data collected, and data collection delays. Summary of the Invention
[0005] In view of this, the purpose of one or more embodiments of the present application is to propose a full life cycle quality trusted data collection method and related equipment to solve the problems pointed out in the background technology.
[0006] Based on the above objectives, one or more embodiments of the present application provide a method for collecting reliable quality data throughout the entire life cycle, which is applied to a distributed control system, including:
[0007] Determine updated data density based on current full lifecycle quality data, wherein the data density is used to represent the relative closeness between trusted data;
[0008] Determining a trusted data deployment node based on the updated data density;
[0009] A trusted data path model is constructed based on the structure of the trusted data deployment node, and target trusted data in the trusted data path model is screened.
[0010] Optionally, the step of calculating the updated data density includes:
[0011] Based on the current full life cycle quality data, obtaining the possibility that any full life cycle quality data becomes credible data;
[0012] Based on the likelihood, an updated data density is obtained.
[0013] Optionally, obtaining the possibility of any full life cycle quality data becoming credible data based on the current full life cycle quality data includes:
[0014] Based on the current full life cycle quality data, the possibility of any full life cycle quality data becoming credible data is obtained through the following formula;
[0015]
[0016] Wherein, i represents any of the life cycle quality data, j represents another life cycle quality data, α represents the effect of pheromone concentration on path selection during ant crawling, and β represents the importance of selecting life cycle quality data under ideal conditions. represents the data density at time t, represents the ideal crawling length of the ant from the full life cycle quality data i to the full life cycle quality data j, κ represents the actual crawling position of the ant colony, G represents the allowed crawling range of the ant colony, It represents the pheromone obtained by the ant colony of the sth sample at time t when crawling.
[0017] Optionally, obtaining an updated data density based on the possibility includes: obtaining an updated data density based on the possibility using the following formula;
[0018] ρ i ' j (t) = p ij (t)+Δρ ij ;
[0019] Where Δρ ij Indicates the pheromone concentration generated by the ant colony when the expected data is selected.
[0020] Optionally, the trusted data deployment node includes a fixed trusted data deployment node and a dynamic trusted data deployment node;
[0021] The step of determining the trusted data deployment node includes:
[0022] Determining the collection similarity of the trusted data based on a preset fixed trusted data deployment node and the updated data density;
[0023] Based on the collection similarity, the dynamic trusted data deployment node is determined.
[0024] Optionally, determining the dynamic trusted data deployment node based on the collection similarity includes:
[0025] Determining the distance between the trusted data deployment nodes based on the collection similarity;
[0026] Based on the distance, the dynamic trusted data deployment node is determined.
[0027] Optionally, the spacing is obtained by the following formula:
[0028]
[0029] in, Represents the total coverage of trusted data collection, It represents the unit distance of the trusted data deployment node, e represents the number of data collection times of the trusted data node, and c represents the deviation range of the trusted data collection.
[0030] Optionally, determining the dynamic trusted data deployment node based on the collection similarity includes:
[0031] Determining attribute classification of credible data based on the collected similarity;
[0032] Based on the collection similarity and target attribute classification, the spacing between the trusted data deployment nodes is determined.
[0033] Optionally, the acquisition similarity is obtained by the following formula:
[0034]
[0035] Among them, η represents the data length, ε represents the interval frequency of data collection, V represents the foot feeling deviation, and ψ represents the collection ratio of reliable data.
[0036] Optionally, the trusted data path model includes a plurality of trusted data deployment nodes, each of which corresponds to at least one cluster head; and screening the target trusted data in the trusted data path model includes:
[0037] Determining a data transmission rate between the trusted data deployment node and the corresponding cluster head;
[0038] In response to determining that the data transmission rate is less than or equal to a preset threshold, determining the energy consumption of transmitting data of the trusted data deployment node using a free space model; or,
[0039] In response to determining that the data transmission rate is greater than a preset threshold, determining the energy consumption of transmitting data of the trusted data deployment node using a multipath fading model;
[0040] Obtaining the total energy consumption of the data collection process based on the energy consumption of transmitting data of the trusted data deployment node;
[0041] The data of the cluster head corresponding to the path with the lowest total energy consumption in the trusted data path model is determined as the updated trusted data.
[0042] Based on the same inventive concept, one or more embodiments of the present application further provide a full lifecycle quality trusted data acquisition device, including:
[0043] A first calculation module is configured to determine an updated data density based on current full life cycle quality data, wherein the data density is used to represent a relative closeness between trusted data;
[0044] a node deployment module configured to determine a trusted data deployment node based on the updated data density;
[0045] The second computing module is configured to construct a trusted data path model based on the structure of the trusted data deployment node and filter target trusted data in the trusted data path model.
[0046] Optionally, the first calculation module is specifically used to: obtain the possibility of any full life cycle quality data becoming credible data based on the current full life cycle quality data; and obtain updated data density based on the possibility.
[0047] Optionally, the node deployment module is specifically configured to: determine the collection similarity of the trusted data based on preset fixed trusted data deployment nodes and the updated data density; and determine the dynamic trusted data deployment node based on the collection similarity.
[0048] Optionally, the second calculation module is specifically used to: determine the data transmission rate between the trusted data deployment node and the corresponding cluster head; in response to determining that the data transmission rate is less than or equal to a preset threshold, determine the transmission data energy consumption of the trusted data deployment node using a free space model; or, in response to determining that the data transmission rate is greater than a preset threshold, determine the transmission data energy consumption of the trusted data deployment node using a multipath fading model; obtain the total energy consumption of the data acquisition process based on the transmission data energy consumption of the trusted data deployment node; and determine that the data of the cluster head corresponding to the path with the lowest total energy consumption in the trusted data path model is the updated trusted data.
[0049] Based on the same inventive concept, one or more embodiments of the present application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the full life cycle quality trusted data collection method as described in any one of the above items.
[0050] Based on the same inventive concept, one or more embodiments of the present application also provide a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute any of the above-mentioned full life cycle quality trusted data collection methods.
[0051] From the above description, it can be seen that the method for collecting trusted data on quality throughout the life cycle provided by one or more embodiments of the present application determines the updated data density based on the current quality data throughout the life cycle; the data density is used to represent the relative closeness between trusted data; based on the updated data density, the trusted data deployment node is determined; based on the structure of the trusted data deployment node, a trusted data path model is constructed, and the target trusted data in the trusted data path model is screened.
[0052] One or more embodiments of this application provide a method for collecting trusted data for lifecycle quality assurance. Based on updated data density, the method identifies trusted data deployment nodes with a higher proportion of trusted data within the DCS device. The trusted data path model constructed by the trusted data deployment nodes is then used to further filter the trusted data. This technical solution can effectively improve data collection stability, reduce data collection latency, and mitigate the impact of error elimination cycles on data transmission.
[0053] The full life cycle quality trusted data collection device, electronic device and computer-readable storage medium provided in this application can all implement the steps of the above-mentioned full life cycle quality trusted data collection method, and therefore also have the beneficial effects of the above-mentioned full life cycle quality trusted data collection method. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate one or more embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only one or more embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0055] Figure 1 A flowchart of a method for collecting reliable quality data throughout the entire life cycle according to one or more embodiments of the present application;
[0056] Figure 2 This is a schematic diagram of the structure of a full life cycle quality trustworthy data acquisition device according to one or more embodiments of the present application;
[0057] Figure 3 This is a graph showing experimental results of the effect of the error elimination period on data transmission according to one or more embodiments of the present application;
[0058] Figure 4 This is a graph showing the experimental results of data acquisition stability testing of one or more embodiments of the present application;
[0059] Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to one or more embodiments of the present application. DETAILED DESCRIPTION
[0060] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0061] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in one or more embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0062] As mentioned in the technical background section, DCS equipment is the primary means for power plants and industrial enterprises like the oil industry to collect important internal data. It provides timely feedback on the data required by the enterprise and has both decentralized and centralized control capabilities. This significantly improves industry production capacity and management, enhancing enterprise flexibility and adaptability. The typical service life of DCS equipment is 20 years, but this lifespan can be extended through upgrades and repairs. Specifically, accurate data on the entire lifecycle of DCS equipment is required. Based on this accurate data and the current lifecycle of the DCS equipment, developing a personalized maintenance plan can effectively extend the lifecycle of the DCS equipment.
[0063] The above solution can be regarded as two stages: data collection and solution customization. For the data collection stage, the solutions of related technologies include two categories. The first type of solution first collects all the quality data of the entire life cycle of the DCS equipment in an iterative round-robin manner, and then filters out the reliable data through the preset data optimization objective function; the second type of solution also collects all the quality data of the entire life cycle of the DCS equipment in an iterative round-robin manner, and then uses the preset data feature statement to search the data and obtain the reliable data. There are other solutions in related technologies to achieve data collection.
[0064] However, all of the aforementioned solutions rely on processing the entire lifecycle quality data from DCS devices to generate reliable data. This data collection often contains a significant amount of useless information, resulting in significant computational and time-consuming collection and processing. This leads to low efficiency and poorly validated results. Furthermore, the error elimination cycle significantly impacts data transmission, the amount of reliable data collected is low, and data collection is often delayed.
[0065] Therefore, the full life cycle quality trustworthy data collection method of one or more embodiments of the present application can effectively improve the stability of data collection, reduce data collection delay, and reduce the impact of the error elimination cycle on data transmission through the technical solution of the present application.
[0066] refer to Figure 1 The method for collecting reliable quality data throughout the life cycle of one or more embodiments of the present application includes the following steps:
[0067] Step S101: Based on the current full life cycle quality data, determine the updated data density, wherein the above data density is used to represent the relative closeness between trusted data.
[0068] The above-mentioned lifecycle quality data represents data that can characterize the lifecycle of a DCS device. This lifecycle quality data may include one or more of the following: display data, operating data, fault data, and maintenance data of the DCS device. This lifecycle quality data may also be other DCS data. This application does not limit the type or quantity of lifecycle quality data.
[0069] Optionally, the above-mentioned full life cycle quality data can be obtained by collecting equipment display data, or from workers' maintenance records, or from monitoring equipment of DCS equipment. This application does not limit the method of obtaining full life cycle quality data.
[0070] The updated data density is determined based on the current full life cycle quality data, which means that the updated data density is obtained according to the distribution of credible data in the full life cycle quality data.
[0071] Alternatively, the updated data density can be derived based on an ant colony algorithm (ACA). The ACA is an optimization algorithm that simulates the foraging behavior of ants and is used to solve graph path optimization problems. The ACA relies on the fact that ants secrete pheromones when searching for food and returning to their nests, marking paths and guiding the behavior of other ants. These pheromones evaporate over time, and paths used by more ants become more concentrated, leading more ants to choose those paths.
[0072] In this application, we consider a path from one data point to another by several ants simultaneously accessing it. This requires updating the data attributes along that path, thereby determining the pheromone concentration. This pheromone concentration represents the quantitative relationship between different factors, including their interconnections and interactions, and their impact on the system state. This path can pass through multiple data points. A single database or data set can contain multiple paths.
[0073] Based on the above, in some embodiments of the present application, the step of determining the updated data density includes: based on the current full life cycle quality data, obtaining the possibility of any full life cycle quality data becoming credible data by the following formula; Wherein, i represents any of the life cycle quality data, j represents another life cycle quality data, α represents the effect of pheromone concentration on path selection during ant crawling, and β represents the importance of selecting life cycle quality data under ideal conditions. represents the data density at time t, represents the ideal crawling length of the ant from the full life cycle quality data i to the full life cycle quality data j, κ represents the actual crawling position of the ant colony, G represents the allowed crawling range of the ant colony, represents the pheromone obtained by the ant colony of the sth sample at time t when crawling; based on the above possibility, the updated data density is obtained by the following formula; ρ i ' j (t) = p ij (t)+Δρ ij ; where Δρ ij Indicates the pheromone concentration generated by the ant colony when the expected data is selected.
[0074] It's important to note that data selection is crucial, as it directly impacts the algorithm's performance and results. Ideally, data selection should fully reflect the characteristics and requirements of the problem, enabling ants to efficiently discover and utilize useful information during the search process. The importance of quality data selection throughout the entire lifecycle is automatically accomplished by the ant colony algorithm based on specific rules and mechanisms. In the ant colony algorithm, each ant selects the next data point to visit based on the current pheromone distribution and heuristic information (such as distance and quality). As the algorithm iterates, the ants continuously update their pheromones, gradually discovering better paths or data combinations.
[0075] The timing for calculating the updated data density is when the ants complete their path selection and reach the terminal. At this point, the ant colony algorithm is considered to have completed a cycle and the data density can be updated.
[0076] The significance of obtaining the updated data density through this step is to effectively utilize the pheromone in the ant colony algorithm, reduce the search space, improve the convergence speed of the algorithm of this application, and reduce the data processing burden.
[0077] Specifically, high data density means a high concentration of data points within a certain area. Combined with the ant colony algorithm, this high-density data point distribution leads to a greater accumulation of pheromones. This high-density data distribution helps ants more quickly locate high-concentration pheromones within the area and make decisions based on them, reducing search time. High data density reduces the potential paths or solution space the algorithm must search. Ants can traverse these points more quickly when searching for the optimal path, reducing unnecessary search steps. Furthermore, as iterations proceed, pheromones gradually accumulate on the optimal paths, forming a positive feedback mechanism. This positive feedback effect is more pronounced when data density is high, as ants can perceive high pheromone concentrations in a shorter period of time. This helps the algorithm converge to the optimal solution more quickly, reducing overall convergence time. High data density also means that the algorithm can more efficiently utilize memory and computing resources when processing data. Because the correlation and similarity between data points are higher, the algorithm may require less computation when aggregating the data.
[0078] Because the lifecycle of DCS equipment changes in real time, data density updates enable timely acquisition of the latest reliable data on the quality of DCS equipment throughout its lifecycle, ensuring that the collected data is current and accurate. This allows for more accurate calculation of node deployment spacing and reduces misjudgments or incorrect decisions caused by outdated or inaccurate data.
[0079] Step S102: Based on the updated data density, a trusted data deployment node is determined.
[0080] Based on the updated data density obtained in step S101, it is necessary to further determine the trusted data deployment node for collecting trusted data.
[0081] In the technical solution of the present application, in order to ensure the efficiency of automated collection of trusted data, it is necessary to further extract and classify the data features to obtain specific data trust features in order to obtain a suitable node deployment solution.
[0082] Optionally, this application can use collection similarity to represent the characteristics between trusted data and determine the trusted data deployment node based on the above collection similarity. Among them, collection similarity refers to comparing the similarity or consistency between data collected at different time points or different detection nodes during the continuous data collection process. This similarity calculation helps to identify outliers, noise or potential deviations in the data, thereby ensuring the reliability and accuracy of the collected data.
[0083] Specifically, the calculation formula for the above acquisition similarity can be expressed as: Where η represents the data length, ε represents the interval frequency of data collection, V represents the foot feel deviation, and ψ represents the collection ratio of reliable data. The collection ratio refers to the proportional relationship between the collection frequency, data volume or data type between different detection nodes or different data sources during the data collection process. The collection ratio may involve how to balance the data contributions of different data sources to ensure the accuracy and completeness of the overall data. For example, in some cases, certain detection nodes may be more critical than other nodes and therefore require a higher collection frequency or a larger data volume. The setting of the collection ratio should be based on actual needs and the importance of the data to ensure the overall quality of the data. Foot feel deviation is used to describe the deviation related to data collection or data processing.
[0084] In some practical application scenarios, the time required to collect a single piece of data is 0.5 minutes, and the data conversion ratio is 1.2. Unidirectional data refers to the data collected within the collection space formed by fixed detection nodes installed in the DCS equipment. Unidirectional data refers to the data set transmitted independently and completely from each node to the data processing center. These data sets together form a trusted data foundation for the entire equipment lifecycle, providing critical support for subsequent data analysis and equipment status monitoring. The data conversion ratio refers to the conversion relationship between the raw data and the actual data required during the data collection process. This conversion may involve data amplification, reduction, unit conversion, or other mathematical transformations. For example, if a sensor measures data in one unit, but the actual analysis requires data in a different unit or ratio, a conversion ratio is required to adjust the data. Specifically, the conversion ratio of 1.2 here means that each acquired data set must be multiplied by 1.2 to obtain the actual analysis value. The above formula is based on the above application scenario.
[0085] Preferably, attribute features of the trusted data may also be acquired based on the above-mentioned collection similarity, and a targeted data packet may be constructed based on the above-mentioned attribute features to meet specific application scenarios or business requirements.
[0086] Specifically, similarity calculations can help discover patterns, trends, and outliers in data. When two or more data points show a high degree of similarity in a similarity calculation, they may share certain attribute characteristics. These characteristics may include the distribution, range, and changing trends of the data. By analyzing the common characteristics between these similar data points, some attribute characteristics of the entire data set can be inferred. In some embodiments, the step of determining the attribute characteristics obtained by the similarity calculation includes: selecting a suitable similarity calculation method, such as Euclidean distance, cosine similarity, Pearson correlation coefficient, etc., based on the characteristics and requirements of the data; calculating a similarity matrix, using the selected similarity calculation method to calculate the similarity between all data points in the data set to obtain a similarity matrix; analyzing the similarity matrix, by observing and analyzing the similarity matrix, identifying clusters of data points with high similarity. These clusters may correspond to certain common attribute characteristics; extracting attribute characteristics, based on the analysis results of the similarity matrix, extracting the attribute characteristics of the data. These features can be numerical (such as mean, standard deviation, etc.) or descriptive (such as data distribution pattern, change trend, etc.); verification and evaluation, by comparing and verifying with actual data or domain knowledge, evaluate the accuracy and effectiveness of the extracted attribute features.
[0087] In some embodiments, after obtaining the aforementioned attribute characteristics, a unique conversion format can be constructed to classify data attributes, enabling subsequent data collection to form targeted data packets. Conversion refers to converting raw collected data or preliminarily processed data into a standardized, unified format suitable for subsequent analysis. This conversion ensures data consistency and comparability, facilitating more efficient extraction of data attribute characteristics, classification of data attributes, and ultimately the construction of targeted data packets. A targeted data packet is constructed based on specific business needs or data analysis objectives. It contains specific data that has been converted and classified to meet specific application scenarios or business requirements. The construction of a targeted data packet is based on the classification and conversion format of data attributes, ensuring that the data in the packet is closely aligned with business requirements, thereby improving the efficiency of data collection, transmission, and processing. This conversion format may include the following steps: data analysis and preprocessing. First, an in-depth analysis of the raw data is performed to understand its data structure, data type, and data range. Data cleaning and preprocessing are used to remove outliers and missing values to ensure data accuracy and completeness. Conversion objectives are determined, and the target format for conversion is determined based on business needs and data characteristics. This may include data representation (e.g., numeric, text), storage structure (e.g., relational database, NoSQL database), and data granularity (e.g., seconds, minutes). Select or develop transformation tools: Based on the transformation goals, select appropriate transformation tools or algorithms, or develop customized transformation programs based on specific needs. Implement and validate the transformation, using transformation tools or algorithms to transform raw or pre-processed data. Verify and test the transformed data to ensure it meets business needs and data quality requirements.
[0088] After obtaining the aforementioned collection similarity, trusted data deployment nodes can be deployed. By deploying trusted deployment data nodes, the collected data links can be tightly connected, forming cyclical and synchronous data collection structures to improve device management capabilities. A cyclical data collection structure forms a closed loop in the data collection process, with data continuously circulating through the collection, processing, analysis, and feedback stages. This structure ensures the real-time and continuity of data, facilitating the timely identification and resolution of issues. A synchronous data collection structure involves multiple data collection points performing data collection simultaneously, ensuring that the data collected at each point is consistent in time. This structure eliminates data inconsistencies caused by time differences and improves data accuracy and comparability.
[0089] In some embodiments, the trusted data deployment nodes include fixed trusted data deployment nodes and dynamic trusted data deployment nodes. Fixed trusted data deployment nodes can be pre-deployed. The method for pre-deploying nodes for DCS equipment lifecycle data primarily involves pre-planning and setting data collection points at key stages of DCS equipment design, production, operation, and maintenance.
[0090] Dynamic trusted data deployment nodes can be dynamically arranged based on DCS equipment. Specifically, they can adjust the frequency and accuracy of data collection in real time based on changes in network traffic, differences in equipment performance, and the importance of the data, ensuring timely and accurate collection of critical data. Dynamic trusted data deployment nodes can also intelligently select the type of data and information to be collected based on equipment operating status and fault warnings, further improving data relevance and usability. Dynamic layout enables data collection in any environment, thereby achieving regional trusted data collection.
[0091] The distance between trusted data deployment nodes can be obtained by the following formula: in, Represents the total coverage of trusted data collection, It represents the unit distance of the trusted data deployment node, e represents the number of data collection times of the trusted data node, and c represents the deviation range of the trusted data collection.
[0092] This step results in trusted data deployment nodes based on updated data density. The placement of trusted data deployment nodes determines the scope and coverage of data collection. Dynamically deployed trusted data deployment nodes can cover the data areas that require collection in real time, ensuring comprehensiveness and accuracy, thereby facilitating faster automated trusted data collection for DCS devices.
[0093] Step S103: constructing a trusted data path model based on the structure of the trusted data deployment node, and screening target trusted data in the trusted data path model.
[0094] In this step, a trusted data path model is first established. In the above trusted data path model, the mark S k Represents any trusted data node, C k Indicates the cluster head corresponding to the above node.
[0095] The above-mentioned screening of target trusted data in the above-mentioned trusted data path model includes: determining the data transmission rate between the above-mentioned trusted data deployment node and the corresponding cluster head; in response to determining that the above-mentioned data transmission rate is less than or equal to a preset threshold, using a free space model to determine the transmission data energy consumption of the above-mentioned trusted data deployment node; or, in response to determining that the above-mentioned data transmission rate is greater than a preset threshold, using a multipath fading model to determine the transmission data energy consumption of the above-mentioned trusted data deployment node; obtaining the total energy consumption of the data acquisition process based on the transmission data energy consumption of the above-mentioned trusted data deployment node; and determining that the data of the cluster head corresponding to the path with the lowest total energy consumption in the above-mentioned trusted data path model is the above-mentioned updated trusted data.
[0096] Specifically, the data transmission rate between the above-mentioned trusted data deployment node and the corresponding cluster head can be calculated by the following steps: determining the line-of-sight link probability between the above-mentioned trusted data deployment node and the corresponding cluster head Where θ represents the angle between the above-mentioned trusted data deployment node and the corresponding cluster head; based on the above-mentioned line-of-sight link probability, the average path loss between the above-mentioned trusted data deployment node and the corresponding cluster head is calculated in, represents the average path loss of the above-mentioned trusted data deployment node line-of-sight link, P NLoS represents the signal path loss of the non-line-of-sight link of the above-mentioned trusted data deployment node, represents the average path loss of the non-line-of-sight link of the trusted data deployment node.
[0097] Optionally, path loss and The calculation formulas are: Among them, f c represents the data carrier frequency, H represents the size of the acquisition device, ξ NLoS and ξ LoS represent the additional path loss in non-line-of-sight and line-of-sight conditions, respectively.
[0098] Therefore, the data transmission rate between the above trusted data deployment node and the cluster head is expressed as: Among them, R Ck represents the data transmission rate in the node, B represents the available bandwidth, N0 represents the power spectrum density of the noise in the DCS equipment, Indicates the power of the trusted data deployment node.
[0099] In the technical solution of this application, the relationship between the data transmission rate and the energy consumption of the transmission data of the trusted data deployment node is expressed as follows: Among them, l represents the trusted data of DCS equipment, represents the energy consumption of node data transmission, E elec Indicates the power consumed by the node during transmission and reception. represents the length between the cluster head and the node, d0 represents the length threshold, ε mp represents the energy parameter of the free space model, ε fs Represents the energy parameter of the multipath fading model.
[0100] Based on the above content, the energy consumption calculation formula for the trusted data deployment node receiving data is: If the data collection amount corresponding to all trusted data deployment nodes in the DCS equipment is the same, the energy consumed by the cluster head during data transmission is calculated as follows: Among them, |Q k | represents the number of trusted data deployment nodes, Represents the time required to collect reliable data from DCS devices. Thus, during one cycle of data collection, the nodes in the device transmit data to the corresponding cluster head, which then transmits it to the collection system. The total energy consumption of data collection is: K represents the number of data clusters.
[0101] As mentioned above, each trusted data deployment node corresponds to a cluster head. The K data clusters correspond to K paths, and the starting points of the K paths correspond to the same trusted data deployment node.
[0102] The path with the lowest total energy consumption for data collection is selected, and the trusted data in it is transmitted to the collection system according to the corresponding cluster head, thus completing the automated collection of trusted data. Specifically, according to the deployed DCS equipment data transmission nodes, there are K clusters in the data collection process, and the trusted data in the clusters needs to be collected according to the optimal path. Assuming that the initial node position of data collection is S0, the data collection path is Q = {S0, Q1, ..., Q K}. Thus, automatic data collection is achieved, and its expression is:
[0103]
[0104] To verify the technical effects of one or more embodiments of this application, the applicant conducted a series of test experiments. The experimental parameters were set as follows: Hollysys' sixth-generation DCS-TS series DCS equipment, model TC-1000, and a PT100 sensor. The sensor sampling frequency was set to 1 per second, and the total experimental time was 60 minutes. This means that a total of 3,600 data points were collected during the experiment (60 minutes * 60 seconds / minute).
[0105] During the experiment, a local storage device is used to store the collected temperature data. This local storage device can be an SD card. Pre-configured data processing software is used to process and analyze the collected temperature data. This software can plot temperature curves, calculate indicators such as average temperature, maximum and minimum values, and analyze how the data changes under different settings. During the experiment, the collected temperature data is transmitted to a computer in real time via an Ethernet connection. Ensure the stability of the Ethernet connection and the data transmission rate to facilitate real-time temperature monitoring.
[0106] The experiment focuses on three key aspects: the impact of the error elimination cycle on data transmission, data acquisition stability testing, and data acquisition latency. The final experimental results of the three methods are compared under these three indicators, and the data acquisition capabilities of the proposed method are demonstrated in detail based on the experimental results. The following experiments were completed using the ant colony algorithm with a pheromone heuristic factor of 0.8, a pheromone evaporation rate of 0.3, an ant colony size of 50, a data interval of 5 minutes, a data sample size of 100, and a data range of 10-100.
[0107] Regarding the impact of the error elimination cycle on data transmission, the premise of data collection is to ensure data integrity, but some acquisition errors are inevitable during the acquisition process, and the data transmission volume under different error elimination cycles is also different. In order to verify the ability of the proposed method to automatically collect data, different error elimination cycles are selected and the average data transmission volume under each error elimination cycle is calculated at different data densities. The calculation formula for the average data transmission volume during the data transmission process is: Represents the number of non-zero elements, N T represents the average number of data transmissions, υ represents the data transmission error elimination period, N P Represents the number of step load changes. The experimental results are as follows Figure 3 shown.
[0108] according to Figure 3 The experimental results show that under each error elimination cycle, the average data transmission volume corresponding to data acquisition under 50% data density is larger than the transmission data volume under the other two data densities, which proves that data density can be used for data screening very well and has an absolute data transmission advantage. Moreover, under different error elimination cycles, the average data transmission volume of high data density changes very little, and when the error elimination cycle reaches a certain value, the average data transmission volume of the proposed method no longer fluctuates, which proves that the error elimination cycle has little effect on the data transmission volume, and the errors can be eliminated at a larger interval period.
[0109] In terms of data collection stability testing, energy consumption is inevitable during the trusted data collection process, and the amount of data collected will increase over time. Although increased energy consumption will enhance data collection capabilities, without reasonable energy consumption planning, it will have the opposite effect. When a large amount of node energy is used during data collection to ensure the collection of trusted data, most nodes will die due to energy depletion after a period of time, resulting in a reduction in data collection nodes. At this time, only farther nodes or alternative nodes can be selected for data collection. In this case, not only will the social ability of the node be reduced, but the burden on other nodes will also be increased.
[0110] In order to explore the data collection situation of the three data credibility, the data collection volume under the three data credibility in different time periods is described. If the data volume increases proportionally, it proves that the energy consumption used for data collection is reasonable. If the data collection volume shows a downward trend after a period of time, it means that the initial energy consumption is high, and some nodes are exhausted later, which proves that the energy consumption used by the collection method is not reasonable. The experimental results are as follows. Figure 4 shown.
[0111] By comparing the experimental data, it was found that in the data science environment, the amount of data increased steadily and orderly during the collection of trusted data, with no downward trend. However, in the other data environments, one method showed a sudden decrease after 1500 seconds, while the other method showed large fluctuations. This directly shows that the two data collection environments other than the proposed method did not establish a reasonable data trust environment during the collection process, resulting in the inability to collect trusted data stably and orderly during the data collection process, causing the data collection results to be very unstable and incomplete, which further verifies the effectiveness of the proposed method.
[0112] Regarding data collection delay, the process of collecting trusted data requires ensuring not only data integrity and stability but also data collection speed. Data shows that many data collection methods experience long collection delays in the later stages. To more intuitively demonstrate the capabilities of the proposed method, we consider data collection delay as a key metric. We randomly selected different databases, labeled Database 1 through Database 10, and collected trusted data from each of these three methods. We assumed that the data integrity and packet loss rates obtained by the three methods were the same, and compared the delays under the three different data environments. The experimental results are shown in Table 1.
[0113] Table 1
[0114]
[0115]
[0116] According to the experimental results, delays are inevitable during data collection. However, in each experimental environment, only the delay under credible data is always lower than 3.53, which proves that the data collection capability of the proposed method is better than that of the other two methods.
[0117] The method for collecting trusted data on quality throughout the entire life cycle provided by this application first updates the data density of the entire device's life cycle, and then deploys the channels required for data collection. By judging the credibility of the data, it realizes automated collection of trusted data, solving the problems of the error elimination cycle having a large impact on data transmission, the small amount of data collected, and severe data collection delays.
[0118] It can be understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.
[0119] It should be noted that the method of one or more embodiments of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of one or more embodiments of the present application, and the multiple devices will interact with each other to complete the described method.
[0120] It should be noted that the above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0121] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a full life cycle quality trusted data acquisition device.
[0122] like Figure 2 As shown, the full life cycle quality trusted data acquisition device includes:
[0123] A first calculation module 11 is configured to determine an updated data density based on the current full life cycle quality data, wherein the data density is used to represent the relative closeness between the trusted data;
[0124] The node deployment module 12 is configured to determine a trusted data deployment node based on the updated data density;
[0125] The second calculation module 13 is configured to construct a trusted data path model based on the structure of the trusted data deployment node, and filter the target trusted data in the trusted data path model.
[0126] Optionally, the first calculation module 11 is specifically configured to: obtain the possibility of any full life cycle quality data becoming credible data based on the above current full life cycle quality data; and obtain updated data density based on the above possibility.
[0127] Optionally, the node deployment module 12 is specifically configured to: determine the collection similarity of the trusted data based on preset fixed trusted data deployment nodes and the updated data density; and determine the dynamic trusted data deployment node based on the collection similarity.
[0128] Optionally, the above-mentioned second calculation module 13 is specifically used to: determine the data transmission rate between the trusted data deployment node and the corresponding cluster head; in response to determining that the data transmission rate is less than or equal to a preset threshold, determine the transmission data energy consumption of the trusted data deployment node using a free space model; or, in response to determining that the data transmission rate is greater than a preset threshold, determine the transmission data energy consumption of the trusted data deployment node using a multipath fading model; obtain the total energy consumption of the data acquisition process based on the transmission data energy consumption of the trusted data deployment node; and determine that the data of the cluster head corresponding to the path with the lowest total energy consumption in the trusted data path model is the updated trusted data.
[0129] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing one or more embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0130] The apparatus of the above embodiment is used to implement the corresponding method in the above embodiment and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0131] Figure 5 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0132] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0133] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of the present application are implemented through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0134] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0135] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0136] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0137] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of the present application, and does not necessarily include all the components shown in the figure.
[0138] The electronic devices of the above embodiments are used to implement the corresponding methods in the above embodiments and have the beneficial effects of the corresponding method embodiments, which will not be described in detail here.
[0139] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0140] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0141] In addition, to simplify the description and discussion, and in order not to make one or more embodiments of the present application difficult to understand, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making one or more embodiments of the present application difficult to understand, and this also takes into account the following fact, that is, the details of the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of the present application will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that one or more embodiments of the present application can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0142] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0143] The one or more embodiments of the present application are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the one or more embodiments of the present application should be included within the scope of protection of this disclosure.
Claims
1. A method for collecting reliable quality data throughout the entire life cycle, characterized in that: Applied to distributed control systems, including: Determine updated data density based on current full lifecycle quality data, wherein the data density is used to represent the relative closeness between trusted data; Determining a trusted data deployment node based on the updated data density; Building a trusted data path model based on the structure of the trusted data deployment node, and screening target trusted data in the trusted data path model; The step of calculating the updated data density includes: Based on the current full life cycle quality data, obtaining the possibility that any full life cycle quality data becomes credible data; Based on the likelihood, obtaining an updated data density; The possibility of obtaining any full life cycle quality data as credible data based on the current full life cycle quality data includes: Based on the current full life cycle quality data, the possibility of any full life cycle quality data becoming credible data is obtained through the following formula; ; in, represents any of the life cycle quality data, Indicates another life cycle quality data, represents the effect of pheromone concentration on the path selection of ants during crawling. Indicates the importance of selecting quality data throughout the entire life cycle under ideal circumstances, express The data density at each moment, In an ideal situation, ants can obtain quality data from their entire life cycle. Crawling to full life cycle quality data length, represents the actual crawling position of the ant colony, Indicates the allowed crawling range of the ant colony, Indicates the time t The pheromones obtained by the sample ant colony while crawling; The obtaining of updated data density based on the possibility includes: Based on the above possibilities, the updated data density is obtained by the following formula: ; in, represents the pheromone concentration generated by the ant colony when the expected data is selected; The trusted data deployment nodes include fixed trusted data deployment nodes and dynamic trusted data deployment nodes; The step of determining the trusted data deployment node includes: Determining the collection similarity of the trusted data based on a preset fixed trusted data deployment node and the updated data density; Based on the collection similarity, the dynamic trusted data deployment node is determined.
2. The method according to claim 1, characterized in that The determining the dynamic trusted data deployment node based on the collection similarity includes: Determining the distance between the trusted data deployment nodes based on the collection similarity; Based on the distance, the dynamic trusted data deployment node is determined.
3. The method according to claim 2, characterized in that The spacing is obtained by the following formula: ; in, Indicates the acquisition similarity, Represents the total coverage of trusted data collection, represents the unit distance of the trusted data deployment node, Indicates the number of data collection times of the trusted data node, Indicates the deviation range of reliable data collection.
4. The method according to claim 1, wherein The determining the dynamic trusted data deployment node based on the collection similarity includes: Determining attribute classification of credible data based on the collected similarity; Based on the collection similarity and target attribute classification, the spacing between the trusted data deployment nodes is determined.
5. The method according to claim 3 or 4, characterized in that The acquisition similarity is obtained by the following formula: ; in, Indicates the data length, Indicates the interval frequency of data collection, Indicates foot feel deviation, Indicates the collection ratio of trusted data.
6. The method according to claim 5, characterized in that The trusted data path model includes a plurality of trusted data deployment nodes, each of the trusted data deployment nodes corresponds to at least one cluster head; The screening of target trusted data in the trusted data path model includes: Determining a data transmission rate between the trusted data deployment node and the corresponding cluster head; In response to determining that the data transmission rate is less than or equal to a preset threshold, determining the data transmission energy consumption of the trusted data deployment node using a free space model; or, in response to determining that the data transmission rate is greater than a preset threshold, determining the data transmission energy consumption of the trusted data deployment node using a multipath fading model; Obtaining the total energy consumption of the data collection process based on the energy consumption of transmitting data of the trusted data deployment node; The data of the cluster head corresponding to the path with the lowest total energy consumption in the trusted data path model is determined as the updated trusted data.
7. A full life cycle quality trustworthy data acquisition device, characterized in that: include: A first calculation module is configured to determine an updated data density based on current full life cycle quality data, wherein the data density is used to represent a relative closeness between trusted data; a node deployment module configured to determine a trusted data deployment node based on the updated data density; a second computing module configured to construct a trusted data path model based on the structure of the trusted data deployment node, and filter target trusted data in the trusted data path model; The first computing module is specifically configured to: Based on the current full life cycle quality data, obtaining the possibility that any full life cycle quality data becomes credible data; Based on the likelihood, obtaining an updated data density; The first computing module is further specifically configured to: Based on the current full life cycle quality data, the possibility of any full life cycle quality data becoming credible data is obtained through the following formula; ; in, represents any of the life cycle quality data, Indicates another life cycle quality data, represents the effect of pheromone concentration on the path selection of ants during crawling. Indicates the importance of selecting quality data throughout the entire life cycle under ideal circumstances, express The data density at each moment, In an ideal situation, ants can obtain quality data from their entire life cycle. Crawling to full life cycle quality data length, represents the actual crawling position of the ant colony, Indicates the allowed crawling range of the ant colony, Indicates the time t The pheromones obtained by the sample ant colony while crawling; The first computing module is further specifically configured to: Based on the above possibilities, the updated data density is obtained by the following formula: ; in, represents the pheromone concentration generated by the ant colony when the expected data is selected; The trusted data deployment nodes include fixed trusted data deployment nodes and dynamic trusted data deployment nodes; The node deployment module is specifically configured as follows: Determining the collection similarity of the trusted data based on a preset fixed trusted data deployment node and the updated data density; Based on the collection similarity, the dynamic trusted data deployment node is determined.
8. The device according to claim 7, characterized in that The second calculation module is specifically configured to: determine a data transmission rate between the trusted data deployment node and the corresponding cluster head; in response to determining that the data transmission rate is less than or equal to a preset threshold, determine the energy consumption of transmitting data of the trusted data deployment node using a free space model; Alternatively, in response to determining that the data transmission rate is greater than a preset threshold, determining the energy consumption of transmitting data of the trusted data deployment node using a multipath fading model; Obtaining the total energy consumption of the data collection process based on the energy consumption of transmitting data of the trusted data deployment node; The data of the cluster head corresponding to the path with the lowest total energy consumption in the trusted data path model is determined as the updated trusted data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.