A fast data acquisition method based on the integration of intelligent data processing and industrial connection technology

By integrating intelligent data processing with industrial connection technology, the automated parsing and configuration of industrial control protocols is realized, and a standard collection model is generated. This solves the problems of low manual operation efficiency and high error rate in existing technologies, and achieves efficient and accurate data collection.

CN120358288BActive Publication Date: 2025-09-09JIANGXI TONGRUI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510851293.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-09
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing industrial data collection methods rely on manual operations, resulting in low protocol adaptation efficiency and high configuration error rate, making it difficult to achieve rapid deployment and flexible adjustment, seriously affecting the accuracy and timeliness of data collection.

Method used

By integrating intelligent data and industrial connection technologies, and using AI process orchestration and semantic understanding technologies, it can realize the automated parsing and configuration of industrial control protocols, generate standard acquisition models, and automatically identify and reuse device models through multi-dimensional fingerprint data. Combined with the precise delivery of edge gateways, it ensures configuration consistency.

Benefits of technology

The efficiency and accuracy of data collection have been significantly improved. The modeling and configuration time for a single device has been shortened to 5 minutes, the overall efficiency has been increased by 90%, the model reuse rate has reached 98%, and the configuration consistency accuracy has reached 100%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358288B_ABST
    Figure CN120358288B_ABST
Patent Text Reader

Abstract

The present invention proposes a rapid data acquisition method based on the integration of intelligent data processing and industrial connection technology. The method obtains the industrial protocols of several devices and the corresponding configuration parameters of the devices, and generates an acquisition model for each device; generates a structural fingerprint for each acquisition model; based on the structural fingerprint, performs a matching analysis on the several acquisition models through multimodal similarity calculation, and screens the several acquisition models according to the matching analysis results to perform uniqueness management, and associates and maps the devices with the corresponding acquisition models to obtain deduplicated acquisition models; the deduplicated acquisition models are accurately distributed to the edge gateway configuration through a bidirectional transmission channel, and the consistency is guaranteed by combining real-time structural fingerprint feedback and matching analysis. The present invention significantly improves the modeling efficiency, information consistency, and distribution capabilities of the data acquisition system through AI-driven process orchestration and semantic understanding technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial data collection, and in particular to a fast data collection method based on the integration of intelligent data collection and industrial connection technology. Background Art

[0002] The current digital transformation of the manufacturing industry faces significant challenges. Due to the diverse brands and models of production equipment and the varying standards of industrial communication protocols, there are significant differences in the configuration, entry, and organization of industrial control protocols. Traditional data collection methods rely heavily on manual operations. The entire process, from protocol adaptation management to parameter distribution to the gateway, requires manual completion by technicians. This is not only inefficient but also prone to configuration errors. This backward manual management model severely restricts data collection efficiency, often causing delays in production data acquisition, which in turn affects the timeliness of enterprises' real-time monitoring, quality analysis, and decision-making optimization. There is an urgent need to use intelligent technologies to achieve automatic protocol adaptation and rapid configuration distribution to improve the accuracy and timeliness of industrial data collection.

[0003] In the current industrial data collection field, existing technical solutions primarily rely on manual configuration and traditional protocol parsing. This approach requires technicians to manually identify and determine the communication protocols of various industrial devices (such as PLCs, CNCs, and sensors), and to develop corresponding parsing and conversion modules for each protocol. Data collection management requires manual creation of device data models, definition of data point mapping relationships, and maintenance of complex data point tables. This process is not only time-consuming but also prone to configuration errors due to human error. Edge gateway adaptation still requires manual configuration of communication parameters, protocol conversion rules, and data distribution strategies, making the entire data collection process inefficient and scalable. Furthermore, due to the diverse range of manufacturing equipment and the high heterogeneity of protocols, traditional solutions struggle with rapid deployment and flexible adjustment, severely hindering the large-scale application of the Industrial Internet of Things.

[0004] The existing data collection solutions have the following main shortcomings and technical roots:

[0005] 1. Low protocol adaptation efficiency. There are numerous industrial protocols, each employing different communication mechanisms, such as register addressing and message structure. Existing technologies lack a unified protocol abstraction layer, requiring the development of separate parsing code for each protocol, resulting in long development cycles and high maintenance costs.

[0006] 2. Manual configuration has a high error rate. Device data models (such as point tables and variable mappings) rely on manual input after interpreting device documentation. Due to the lack of automated modeling tools, human misreading of specifications and input errors are difficult to avoid. This error rate increases significantly when working with complex devices (such as PLC DB block data). Summary of the Invention

[0007] In light of the above, the main purpose of this invention is to propose a rapid data acquisition method and system based on the integration of intelligent data processing and industrial connection technology to address the configuration complexity and inefficient manual modeling issues caused by differences in industrial control protocols in industrial data acquisition. This eliminates the need for specific development for different industrial control protocols and enables automated conversion and unified management of protocol configurations. Through automated parsing of device documentation, the intelligent generation of device digital models and point tables is achieved, reducing manual intervention. Through intelligent verification and conflict detection mechanisms, the configuration and modeling error rate is controlled below 2%, significantly improving the reliability of data acquisition.

[0008] The present invention proposes a fast data acquisition method based on the integration of intelligent data and industrial connection technology, the method comprising the following steps:

[0009] Step 1: Obtain the industrial protocols of several devices and the corresponding configuration parameters of the devices, perform semantic parsing and structural conversion on the industrial protocols and configuration parameters, and generate a collection model for each device;

[0010] Step 2: Generate multi-dimensional fingerprint data based on the hash value, path structure, and semantic information of the acquisition model, and use the multi-dimensional fingerprint data as the structural fingerprint of the corresponding acquisition model;

[0011] Step 3: Based on the structural fingerprint, a matching analysis is performed on several acquisition models through multimodal similarity calculation to obtain a matching analysis result;

[0012] Based on the matching analysis results, several acquisition models are screened for uniqueness management, and the devices are associated with the corresponding acquisition models to obtain the deduplicated acquisition models.

[0013] Step 4: The deduplicated collection model is accurately distributed to the edge gateway through a bidirectional transmission channel, combined with real-time structural fingerprint feedback and matching analysis to ensure consistency.

[0014] This invention significantly improves the modeling efficiency, information consistency and distribution capabilities of the data acquisition system by focusing on the three core capabilities of "intelligent identification of acquisition protocols, rapid generation of acquisition models, and automatic distribution by edge gateways" through AI-driven process orchestration and semantic understanding technology.

[0015] Compared with the prior art, the present invention has the following beneficial effects:

[0016] 1. Application of intelligent data query for rapid implementation of edge data collection. This invention uses FastGPT's AI process orchestration engine to drive intelligent data query guidance, enabling the import of raw device collection point tables with high diversity across various industrial control protocols into the system. It automatically parses field and structural relationships and generates a standard collection model. This reduces the modeling and configuration time for a single device from one hour for traditional manual operations to five minutes, improving overall data collection efficiency by over 90%.

[0017] 2. Structural fingerprint and semantic embedding algorithms enable product model reuse. This invention uses dual algorithms, structural fingerprint and semantic embedding, to compare the similarities between device configuration models, automatically identifying and recommending reuse of similar device models, avoiding duplicate modeling and effectively increasing the reuse rate of similar device models to over 98%.

[0018] 3. Highly coordinated collection system and edge gateway. Heterogeneous device collection configuration information can be sent to the edge gateway with a single click, supporting both full and incremental synchronization. Combined with the gateway's backhaul capabilities, the system automatically compares fingerprints to determine model consistency or the need for redeployment, ensuring complete synchronization of edge and central configurations with 100% consistency accuracy.

[0019] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flow chart of the fast data acquisition method based on the integration of intelligent data processing and industrial connection technology proposed by the present invention;

[0021] Figure 2 This is the logical tree structure diagram of the present invention. DETAILED DESCRIPTION

[0022] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0023] These and other aspects of the embodiments of the present invention will become clear with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0024] See also Figure 1 and Figure 2This embodiment provides a fast data collection method based on the integration of intelligent data and industrial connection technology, the method comprising the following steps:

[0025] Step 1: Obtain the industrial protocols of several devices and the corresponding configuration parameters of the devices, perform semantic parsing and structural conversion on the industrial protocols and configuration parameters, and generate a collection model for each device;

[0026] As a preferred embodiment of the present invention, the industrial protocols of several devices and the configuration parameters corresponding to the devices are obtained, and the industrial protocols and configuration parameters are semantically parsed and structurally converted to generate a collection model for each device, specifically including the following steps:

[0027] Acquire the table file data of the configuration parameters, and verify the table file data to obtain the verified table file data;

[0028] An AI model is used to perform semantic recognition on the column header names of the inspected table file data. The protocol parameter logic corresponding to each column is determined in combination with the industrial protocol dictionary library to obtain a structured data table with semantic labels.

[0029] Perform data cleansing on structured data tables with semantic tags. Based on the semantics of column headers and the relationship between adjacent fields in the structured data tables with semantic tags, the dependency rules between fields are inferred to obtain a standardized configuration parameter table.

[0030] Obtain the industrial protocol, and decompose the standardized configuration parameter table into independent configuration units according to the type of the industrial protocol to obtain a structured collection element set;

[0031] Input the structured collection feature set into the pre-trained task orchestration model (FastGPT) to generate configuration steps in a logical order;

[0032] Insert user confirmation nodes into key steps in the configuration process to obtain an interactive task flow diagram;

[0033] Convert interactive task flow diagrams into natural language question-answering processes, guiding users to supplement missing parameters or correct logical conflicts, and obtain a complete set of configuration parameters confirmed by the user;

[0034] Select a preset template based on the type of industrial protocol, fill in the complete configuration parameter set confirmed by the user according to the template structure, and obtain the original configuration content;

[0035] The original configuration content is uniformly converted into a format supported by the target system to generate a standard file with a version identifier to obtain the acquisition model.

[0036] In this solution, the present invention utilizes the AI ​​process orchestration tool FastGPT to transform the traditional manually configured "collection and implementation of basic information" into an interactive, intelligent query process. This approach integrates step-by-step decomposition and task flow generation: The acquisition elements (such as device addresses, register mappings, and sampling periods) in industrial protocols like Modbus and OPC-UA are structured and decomposed. AI-powered task flow graphs are then constructed to generate continuous, interactive dialog tasks.

[0037] It also has the function of automatically identifying and formatting original Excel files; it supports importing original protocol information (such as Excel tables provided by manufacturers), and AI automatically parses based on column header semantics and data structure to identify the configuration meaning and parameter logic between fields;

[0038] After the dialogue process is completed, the system automatically generates standard collection configuration content, outputs it in a unified Excel or JSON format, and can be directly connected to the data collection system for one-click import.

[0039] Step 2: Generate multi-dimensional fingerprint data based on the hash value, path structure, and semantic information of the acquisition model, and use the multi-dimensional fingerprint data as the structural fingerprint of the corresponding acquisition model;

[0040] As a preferred embodiment of the present invention, generating multi-dimensional fingerprint data according to the hash value, path structure and semantic information of the acquisition model, and using the multi-dimensional fingerprint data as the structural fingerprint of the corresponding acquisition model specifically includes the following steps:

[0041] The basic information, attribute information, data type, and unit of each device in the acquisition model are sorted and then the hash value is calculated;

[0042] Combine the fields in the collection model into a logical tree structure and perform nested hierarchical coding to obtain the structural path coding;

[0043] Use AI semantic models to convert device names and device attribute information in the acquisition model into semantic vectors;

[0044] The hash value, structural path encoding and semantic vector are combined into a unified structure to form multi-dimensional fingerprint data, and the multi-dimensional fingerprint data is used as the structural fingerprint of the corresponding acquisition model.

[0045] Furthermore, the basic information, attribute information, data type, and unit of each device in the acquisition model are sorted and then the hash value is calculated, specifically including the following steps:

[0046] Normalize the basic information, attribute information, data type, and unit of each device in the acquisition model to obtain a standardized field string;

[0047] Split the standardized field string into independent fields according to the delimiter and sort them to obtain an ordered field sequence;

[0048] Use a unified delimiter to connect the ordered field sequences in sequence to obtain a complete concatenated string;

[0049] Apply a cryptographic hash algorithm to the concatenated complete string to obtain the hash value of the device field.

[0050] Furthermore, the fields in the collection model are combined into a logical tree structure and nested hierarchical encoding is performed to obtain the structural path encoding, which specifically includes the following steps:

[0051] Combine the fields in the collection model into a logical tree structure, and use the fields of the nodes in the logical tree structure as the encoding results;

[0052] Get the encoding results of all child nodes of the current node in the logical tree structure, and then sort them in lexicographic order to obtain the encoding set of the child nodes;

[0053] Concatenate all sub-node codes in the sub-node code set into a single string in order, and use a cryptographic hash algorithm to obtain the hash value of the code;

[0054] Take the first 4 digits of the encoded hash value as the unique identifier to obtain the aggregate hash value of the sub-node structure;

[0055] The node fields are concatenated with the aggregate hash value of the child node structure. During the concatenation process, if the current node is not the root node, the parent node path separator is concatenated to form a complete hierarchical path and obtain the structural path encoding.

[0056] Furthermore, using the AI ​​semantic model, the device name and device attribute information in the acquisition model are converted into semantic vectors, which specifically includes the following steps:

[0057] Obtain the device name in the acquisition model, and perform standardization on the device name to obtain a standardized device name;

[0058] Obtain the device attribute information in the acquisition model, reorganize the device attribute information in the form of "attribute name: value + unit" to obtain a structured attribute description string;

[0059] Concatenate the standardized device name and the structured attribute description string as a complete device description text sequence;

[0060] The complete device description text sequence is segmented and input into the pre-trained BERT model. The hidden vector corresponding to the CLS flag output by the BERT model is used as the semantic representation of the entire text to obtain the semantic encoding vector of the text.

[0061] Through a learnable industrial domain adaptation matrix, the semantic encoding vector of the text is mapped to the industrial equipment parameter feature space, and bias adjustment is performed to obtain the domain-adapted semantic vector;

[0062] The domain-adapted semantic vector is subjected to nonlinear transformation and normalization operations in sequence to obtain the final semantic vector.

[0063] In the above solution, in order to ensure the consistency and standardization of the acquisition configuration model, the system design uses an acquisition model structural feature fingerprint comparison algorithm to perform similarity matching and automatic association between the newly imported acquisition model and the existing platform models. The core process includes the following steps:

[0064] Structural feature fingerprint generation. The system constructs structural features, or fingerprints, for each acquisition model. Each acquisition model ultimately forms a set of multi-dimensional fingerprint data consisting of field-level hash signatures + structural paths + semantic vectors, constituting its "structural identity." Specific content includes:

[0065] 1. Field-level hash signature: The basic information, attribute information, data type, unit, etc. of each device are sorted and processed, and then the hash value is calculated;

[0066] 2. Field structure path: Combine the fields in the model into a logical tree structure or path sequence, and perform nested hierarchical encoding on its structure path, such as Figure 2 As shown;

[0067] 3. Semantic embedding generation: Using AI semantic models, device names and device attribute information are converted into semantic vectors to reflect the business semantic associations between different device models and determine whether the device models are "similar and reusable." For example:

[0068] Vector("Ball Mill #1")≈Vector("Ball Mill #2"), Vector("Ball Mill #1.Speed")≈Vector("Ball Mill #2.Spindle Speed");

[0069] Similarity comparison engine. The system determines whether a model is duplicate or reusable by combining multi-dimensional feature similarity metrics. The main strategies are as follows:

[0070] 1. Hash overlap comparison: If the overlap between the newly constructed field hash value and the existing model is higher than the set threshold (e.g., 85%), it is considered a strong suspected duplication.

[0071] 2. Structural path edit distance calculation: Calculate the edit distance (such as Levenshtein distance) of the structural path sequence to determine the similarity of the model's organizational structure;

[0072] 3. Semantic vector similarity comparison: Cosine similarity is used to determine the semantic overlap of device model names to ensure consistency in meaning rather than just superficial consistency.

[0073] 4. The comparison algorithm aggregates multiple indicators in a weighted fusion manner to form a final model similarity score and output recommended actions (such as: reuse / overwrite / merge / create).

[0074] Step 3: Based on the structural fingerprint, a matching analysis is performed on several acquisition models through multimodal similarity calculation to obtain a matching analysis result;

[0075] Based on the matching analysis results, several acquisition models are screened for uniqueness management, and the devices are associated with the corresponding acquisition models to obtain the deduplicated acquisition models.

[0076] As a preferred embodiment of the present invention, based on the structural fingerprint, performing matching analysis on several acquisition models through multimodal similarity calculation specifically includes the following steps:

[0077] Align the multi-dimensional fingerprint data by establishing field mapping relationships based on field names or business logic to obtain a structurally aligned multimodal feature dataset;

[0078] Traverse the hash value sets in the structurally aligned multimodal feature datasets of all acquisition models, count the proportion of identical hash values, and use the proportion of identical hash values ​​as the hash similarity score;

[0079] Convert the structural path encoding in the structure-aligned multimodal feature dataset into a string sequence;

[0080] The Levenshtein algorithm is used to calculate the minimum edit distance between the string sequences of all collected models and normalize them to obtain the structural path similarity score;

[0081] The semantic vectors in the structure-aligned multimodal feature dataset are projected into a unified semantic space using a pre-trained semantic model;

[0082] In the unified semantic space, the cosine similarity of the field vectors with the same name in all collected models is calculated to obtain the cosine similarity of all fields;

[0083] Assign field weights of different sizes according to their importance, and use the field weights to perform a weighted average operation on the cosine similarity of all fields to obtain the semantic similarity score;

[0084] Obtain the required business scenario type and the user's historical operation preferences, dynamically generate multimodal weights based on the required business scenario type and the user's historical operation preferences, and obtain the multimodal weight coefficient;

[0085] The hash similarity score, structural path similarity score and semantic similarity score are weighted and fused using the multimodal weight coefficient and normalized to obtain a multimodal comprehensive similarity score.

[0086] As a preferred embodiment of the present invention, screening a plurality of acquisition models according to the matching analysis results to perform uniqueness management, and obtaining a duplicated acquisition model specifically includes the following steps:

[0087] The multimodal comprehensive similarity score is divided into score intervals according to preset rules and given corresponding labels to obtain a list of candidate models with classification labels; among them, the high match interval is greater than or equal to 95% and is marked as "strong reuse candidate"; the gray area is 70%-95% and is marked as "needs manual verification"; the low match interval is less than 70% and is marked as "independent new creation";

[0088] Obtain the collection model set marked as "strong reuse candidate" in the candidate model list with classification labels and perform uniqueness determination. If the same target model only matches one candidate model, directly associate and reuse it;

[0089] If multiple candidate models match the same target acquisition model, they are dynamically sorted by acquisition model version time, user historical selection frequency, and field completeness. The acquisition model with the highest priority is selected as the primary version from the sorted results, and the remaining models are marked as "historical copies." Fields missing from the primary version in the "historical copies" are extracted to generate patch files for manual confirmation, resulting in a deduplicated primary acquisition model set.

[0090] Obtain a set of acquisition models marked as "requiring manual verification" from the candidate acquisition model list with classification labels, visualize the differences between similar or conflicting parts, and guide the user to correct and update the acquisition models marked as "requiring manual verification" to obtain a corrected acquisition model set;

[0091] The collection model set marked as "independent new creation", the main collection model set after deduplication, and the revised collection model set are associated and mapped with the corresponding devices to obtain the deduplication collection model.

[0092] In the above solution, the present invention displays the standard collection configuration content through the dialogue process in the form of an online table, and recommends an operation based on the similarity mark in each row. The user makes the final confirmation. After confirmation, the model reuse association or the creation and association of the new model are automatically completed to establish a mapping relationship:

[0093] If the similarity score is higher than the set threshold (95%), the system will determine that the model already exists and recommend reuse;

[0094] If the score is in the gray area (70%~95%), the system will highlight the difference field and suggest the user to confirm;

[0095] If the score is in other ranges (less than 70%), the system will determine that there is no existing model and suggest the user to create a new one for confirmation;

[0096] After the user's final confirmation is completed, a mapping relationship is automatically established between the collection point and the original product model or the newly created product model. For the newly created product model, the model fingerprint will be automatically saved to support fast matching and mapping in the subsequent collection configuration process.

[0097] Step 4: The deduplicated collection model is accurately distributed to the edge gateway through a bidirectional transmission channel, combined with real-time structural fingerprint feedback and matching analysis to ensure consistency.

[0098] As a preferred embodiment of the present invention, the deduplicated collection model is used to accurately deliver the edge gateway configuration through a bidirectional transmission channel, and the real-time structure fingerprint return and matching analysis are combined to ensure consistency. Specifically, the following steps are included:

[0099] Check whether the model fields cover all protocol parameters of the target device for integrity verification; compare the deployed model of the target gateway based on the structural identity fingerprint to perform conflict pre-detection;

[0100] After the integrity check and conflict pre-detection pass, the verified model is encapsulated as attached structural identity metadata according to the gateway protocol requirements to obtain a configuration package;

[0101] Based on the size and field complexity of the configuration package, different compression algorithms are selected for compression processing. Based on the current network environment status, corresponding transmission rules are selected to send the configuration package to the edge gateway.

[0102] When sending the configuration package to the edge gateway, the deployment status is transmitted in real time through a bidirectional channel. Based on the deployment status, it is confirmed whether to enable incremental retransmission. The retransmission conflict field is used to ensure the integrity of the collection model deployment and complete the collection model deployment.

[0103] After the edge gateway completes the deployment of the collection model, it obtains the actual configuration information of the edge gateway, extracts the field-level hash signature and structure path of the gateway configuration according to the preset period, and generates a new fingerprint structure;

[0104] Compare the differences between the original fingerprint structure and the new fingerprint structure when they are issued, mark the modified fields, and obtain real-time structural identity data and difference marks;

[0105] The difference fields in the real-time structural identity data are obtained based on the difference tags, and multimodal similarity calculations are performed on the difference fields for incremental comparison. A real-time consistency analysis report is obtained, and abnormal configuration repair and policy optimization are performed based on the real-time consistency analysis report.

[0106] In the above solution, the present invention realizes efficient synchronization and consistency verification of data acquisition configuration information with the edge gateway, including:

[0107] Centralized organization and confirmation of configuration information

[0108] The system displays the standard collection configuration content through the dialogue process in the form of an online table, and recommends actions based on the similarity tags in each row, and the user makes the final confirmation.

[0109] Gateway-delivered policy configuration and execution

[0110] After confirming the accuracy of the configuration, the system supports synchronizing the collection model to the edge gateway on demand. The distribution mechanism includes the following key steps:

[0111] A. Gateway information management. The system supports manual entry of gateway connection information through the gateway management function and tests the accuracy of the connection;

[0112] B. Target gateway selection and configuration delivery. Users can select the target gateway in the interface, deliver the JSON configuration in a unified structure, and complete edge push via MQTT topics or API interfaces. The push process supports log echo and failure retry mechanism. The system provides the following two delivery strategies:

[0113] Full delivery: All collection configurations of the current device are delivered at one time;

[0114] Selected delivery: Users filter some devices or collection points for delivery.

[0115] Gateway information feedback and consistency verification

[0116] To prevent configuration drift between the edge gateway and the central system, the system supports the gateway to actively transmit collection configuration information and perform the following operations:

[0117] A. Return field structure fingerprint generation

[0118] Extract field structure and device parameter information from the returned content to generate a new structural feature fingerprint.

[0119] B. Compare with the platform model fingerprint

[0120] The system compares the returned model fingerprint with the existing model in the system to see if it is consistent:

[0121] Completely consistent: Confirm that the synchronization is successful and mark it as "Deployed";

[0122] If the fingerprint is close but not identical: the message is marked as "Gateway-side configuration may have been manually modified" and the user is prompted for confirmation.

[0123] Completely inconsistent: The system records the exception and prompts the user to redeploy.

[0124] C. Avoid duplicate mapping

[0125] If the gateway is manually configured with device information of the same structure as in an existing model, the system can automatically determine and prevent repeated mapping and repeated distribution through fingerprint comparison, thereby keeping the model clear and the data unique.

[0126] In order to verify the effectiveness of the present invention, take copper processing equipment data collection as an example:

[0127] The system automatically identified the S7 protocol and Modbus RTU pressure sensor of the Siemens S7-1200 PLC in the production line, and uniformly mapped the PLC's DB block address and sensor register address into standardized variables.

[0128] The system automatically generates a digital model containing parameters such as temperature and pressure by parsing equipment technical documentation. It also adds metadata such as units and alarm thresholds based on a metallurgical industry template. The system also automatically detects and corrects configuration conflicts. When the system detects conflicts, such as duplicate register definitions, it automatically corrects the configuration.

[0129] It should be understood that, although the various steps in the flow chart of each embodiment of the present invention are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0130] It should be understood that various components of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0131] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0132] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A fast data acquisition method based on the integration of intelligent data and industrial connection technology, characterized in that: The method comprises the following steps: Step 1: Obtain the industrial protocols of several devices and the corresponding configuration parameters of the devices, perform semantic parsing and structural conversion on the industrial protocols and configuration parameters, and generate a collection model for each device; Step 2: Generate multi-dimensional fingerprint data based on the hash value, path structure, and semantic information of the acquisition model, and use the multi-dimensional fingerprint data as the structural fingerprint of the corresponding acquisition model; Step 3: Based on the structural fingerprint, a matching analysis is performed on several acquisition models through multimodal similarity calculation to obtain a matching analysis result; Based on the matching analysis results, several acquisition models are screened for uniqueness management, and the devices are associated with the corresponding acquisition models to obtain the deduplicated acquisition models. Step 4: The deduplicated collection model is transmitted through a bidirectional transmission channel to accurately deliver the edge gateway configuration. This is combined with real-time structural fingerprint feedback and matching analysis to ensure consistency. In step 2, generating multi-dimensional fingerprint data according to the hash value, path structure and semantic information of the acquisition model, and using the multi-dimensional fingerprint data as the structural fingerprint of the corresponding acquisition model specifically includes the following steps: The basic information, attribute information, data type, and unit of each device in the acquisition model are sorted and then the hash value is calculated; Combine the fields in the collection model into a logical tree structure and perform nested hierarchical coding to obtain the structural path coding; Use AI semantic models to convert device names and device attribute information in the acquisition model into semantic vectors; Combine hash values, structural path codes, and semantic vectors into a unified structure to form multi-dimensional fingerprint data, and use the multi-dimensional fingerprint data as the structural fingerprint of the corresponding acquisition model; The fields in the acquisition model are combined into a logical tree structure and nested hierarchical coding is performed to obtain the structural path coding, which specifically includes the following steps: Combine the fields in the collection model into a logical tree structure, and use the fields of the nodes in the logical tree structure as the encoding results; Get the encoding results of all child nodes of the current node in the logical tree structure, and then sort them in lexicographic order to obtain the encoding set of the child nodes; Concatenate all sub-node codes in the sub-node code set into a single string in order, and use a cryptographic hash algorithm to obtain the hash value of the code; Take the first 4 digits of the encoded hash value as the unique identifier to obtain the aggregate hash value of the sub-node structure; Concatenate the node's fields with the aggregate hash value of the child node structure. During the concatenation process, if the current node is not the root node, concatenate the parent node path separator to form a complete hierarchical path and obtain the structural path encoding. In step 3, based on the structural fingerprint, the matching analysis of several acquisition models is performed by multimodal similarity calculation, which specifically includes the following steps: Align the multi-dimensional fingerprint data by establishing field mapping relationships based on field names or business logic to obtain a structurally aligned multimodal feature dataset; Traverse the hash value sets in the structurally aligned multimodal feature datasets of all acquisition models, count the proportion of identical hash values, and use the proportion of identical hash values ​​as the hash similarity score; Convert the structural path encoding in the structure-aligned multimodal feature dataset into a string sequence; The Levenshtein algorithm is used to calculate the minimum edit distance between the string sequences of all collected models and normalize them to obtain the structural path similarity score; The semantic vectors in the structure-aligned multimodal feature dataset are projected into a unified semantic space using a pre-trained semantic model; In the unified semantic space, the cosine similarity of the field vectors with the same name in all collected models is calculated to obtain the cosine similarity of all fields; Assign field weights of different sizes according to their importance, and use the field weights to perform a weighted average operation on the cosine similarity of all fields to obtain the semantic similarity score; Obtain the required business scenario type and the user's historical operation preferences, dynamically generate multimodal weights based on the required business scenario type and the user's historical operation preferences, and obtain the multimodal weight coefficient; The hash similarity score, structural path similarity score and semantic similarity score are weighted and fused using the multimodal weight coefficient and normalized to obtain a multimodal comprehensive similarity score. Based on the matching analysis results, several acquisition models are screened for uniqueness management, and the devices are mapped to the corresponding acquisition models. The specific steps include the following: The multimodal comprehensive similarity score is divided into score intervals according to preset rules and given corresponding labels to obtain a list of candidate models with classification labels. Among them, the high match interval is greater than or equal to 95%, and is marked as "strong reuse candidate"; the gray area is between 70% and 95%, and is marked as "needs manual verification"; and the low match interval is less than 70%, and is marked as "independent new creation"; Obtain the collection model set marked as "strong reuse candidate" in the candidate model list with classification labels and perform uniqueness determination. If the same target model only matches one candidate model, directly associate and reuse it. If multiple candidate models match the same target collection model, they are dynamically sorted by collection model version time, user historical selection frequency, and field completeness. The collection model with the highest priority is selected as the primary version from the sorted results, and the remaining models are marked as "historical copies." Fields missing from the primary version in the "historical copies" are extracted to generate patch files for manual verification, resulting in a deduplicated set of primary collection models. Obtain a set of acquisition models marked as "requires manual verification" from the candidate acquisition model list with classification labels, visualize the differences between similar or conflicting parts, and guide users to correct and update the acquisition models marked as "requires manual verification" to obtain a set of corrected acquisition models; The collection model set marked as "independent new creation", the deduplicated main collection model set, and the revised collection model set are associated and mapped with the corresponding devices to obtain the deduplicated collection model.

2. The rapid data acquisition method based on the integration of intelligent data and industrial connection technology according to claim 1 is characterized in that: In step 1, the industrial protocols of several devices and the configuration parameters corresponding to the devices are obtained, and the industrial protocols and configuration parameters are semantically parsed and structurally converted to generate an acquisition model for each device. Specifically, the steps include: Acquire the table file data of the configuration parameters, and verify the table file data to obtain the verified table file data; An AI model is used to perform semantic recognition on the column header names of the inspected table file data. The protocol parameter logic corresponding to each column is determined in combination with the industrial protocol dictionary library to obtain a structured data table with semantic labels. Perform data cleansing on structured data tables with semantic tags. Based on the semantics of column headers and the relationship between adjacent fields in the structured data tables with semantic tags, the dependency rules between fields are inferred to obtain a standardized configuration parameter table. Obtain the industrial protocol, and decompose the standardized configuration parameter table into independent configuration units according to the type of the industrial protocol to obtain a structured collection element set; Input the structured collection element set into the pre-trained task orchestration model to generate configuration steps in a logical order; Insert user confirmation nodes into key steps in the configuration process to obtain an interactive task flow diagram; Convert interactive task flow diagrams into natural language question-answering processes, guiding users to supplement missing parameters or correct logical conflicts, and obtain a complete set of configuration parameters confirmed by the user; Select a preset template based on the type of industrial protocol, fill in the complete configuration parameter set confirmed by the user according to the template structure, and obtain the original configuration content; The original configuration content is uniformly converted into a format supported by the target system to generate a standard file with a version identifier to obtain the acquisition model.

3. The rapid data acquisition method based on the integration of intelligent data and industrial connection technology according to claim 1 is characterized in that: The basic information, attribute information, data type, and unit of each device in the acquisition model are sorted and then the hash value is calculated, which specifically includes the following steps: Normalize the basic information, attribute information, data type, and unit of each device in the acquisition model to obtain a standardized field string; Split the standardized field string into independent fields according to the delimiter and sort them to obtain an ordered field sequence; Use a unified delimiter to connect the ordered field sequences in sequence to obtain a complete concatenated string; Apply a cryptographic hash algorithm to the concatenated complete string to obtain the hash value of the device field.

4. The rapid data acquisition method based on the integration of intelligent data and industrial connection technology according to claim 1 is characterized in that: Using the AI ​​semantic model, converting the device name and device attribute information in the acquisition model into a semantic vector involves the following steps: Obtain the device name in the acquisition model, and perform standardization on the device name to obtain a standardized device name; Obtain device attribute information from the acquisition model, reorganize the device attribute information in the form of "attribute name: value + unit" to obtain a structured attribute description string; Concatenate the standardized device name and the structured attribute description string as a complete device description text sequence; The complete device description text sequence is segmented and input into the pre-trained BERT model. The hidden vector corresponding to the CLS flag output by the BERT model is used as the semantic representation of the entire text to obtain the semantic encoding vector of the text. Through a learnable industrial domain adaptation matrix, the semantic encoding vector of the text is mapped to the industrial equipment parameter feature space, and bias adjustment is performed to obtain the domain-adapted semantic vector; The domain-adapted semantic vector is subjected to nonlinear transformation and normalization operations in sequence to obtain the final semantic vector.

5. The rapid data acquisition method based on the integration of intelligent data and industrial connection technology according to claim 1 is characterized in that: In step 4, the deduplicated collection model is used to accurately transmit the edge gateway configuration through a bidirectional transmission channel. The real-time structure fingerprint return and matching analysis are combined to ensure consistency. The specific steps include the following: Check whether the model fields cover all protocol parameters of the target device for integrity verification; compare the deployed model of the target gateway based on the structural identity fingerprint to perform conflict pre-detection; After the integrity check and conflict pre-detection pass, the verified model is encapsulated as attached structural identity metadata according to the gateway protocol requirements to obtain a configuration package; Based on the size and field complexity of the configuration package, different compression algorithms are selected for compression processing. Based on the current network environment status, corresponding transmission rules are selected to send the configuration package to the edge gateway. When sending the configuration package to the edge gateway, the deployment status is transmitted in real time through a bidirectional channel. Based on the deployment status, it is confirmed whether to enable incremental retransmission. The retransmission conflict field is used to ensure the integrity of the collection model deployment and complete the collection model deployment. After the edge gateway completes the deployment of the collection model, it obtains the actual configuration information of the edge gateway, extracts the field-level hash signature and structure path of the gateway configuration according to the preset period, and generates a new fingerprint structure; Compare the differences between the original fingerprint structure and the new fingerprint structure when they are issued, mark the modified fields, and obtain real-time structural identity data and difference marks; The difference fields in the real-time structural identity data are obtained based on the difference tags, and multimodal similarity calculations are performed on the difference fields for incremental comparison. A real-time consistency analysis report is obtained, and abnormal configuration repair and policy optimization are performed based on the real-time consistency analysis report.

Citation Information

Patent Citations

  • Relay protection device communication model automatic matching system and method

    CN115712839A

  • Information acquisition and analysis method, device, equipment, medium and product

    CN120106232A