A symbol fast retrieval method for super-large multi-station PLC project
Patent Information
- Application Number
- CN202610805141.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-05
AI Technical Summary
(1)检索延迟高:传统全量遍历搜索耗时往往超过10秒,严重打断工程师的编程思路,无法实现输入即响应
[0014] This embodiment provides a fast symbol retrieval method for ultra-large multi-station PLC projects. It performs data parsing, symbol metadata standardization, and fragment preprocessing on the PLC project file to obtain standardized symbol metadata objects (SymbolMeta) fragmented data. Based on the standardized SymbolMeta fragmented data, a three-level hierarchical index architecture is constructed, consisting of a station-level hash routing index, a task-level compressed prefix Trie tree index, and a symbol-level differential encoded inverted list. According to the user's query conditions, station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation are performed sequentially within the three-level hierarchical index architecture to generate structured search results, which are then rendered on the front end. Incremental updates are performed on the three-level hierarchical index architecture based on the PLC project's modification instructions. This solves the problems of high latency, weak fuzzy matching capability, high resource consumption, and delayed updates in retrieval of symbols exceeding 100,000, achieving millisecond-level accurate positioning.
Smart Images

Figure CN122332617B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of PLC programming technology, and in particular to a method for rapid symbol retrieval for ultra-large multi-station PLC projects. Background Technology
[0002] In modern large-scale industrial projects, PLC projects often involve dozens of physical sites, hundreds of program tasks, and more than 100,000 I / O symbols, intermediate variables, and function blocks. When dealing with ultra-large-scale projects, existing PLC programming software typically uses linear traversal or simple hash lookup for symbol search. When the project scale reaches 100,000+ symbols, the following problems exist: (1) High retrieval latency: Traditional full traversal search often takes more than 10 seconds, severely interrupting the engineer's programming train of thought and making it impossible to achieve input-to-response. (2) Weak fuzzy matching capability: It is difficult to handle spelling errors or ambiguous symbol names, and lacks an efficient fault tolerance mechanism. (3) High resource consumption: The index structure that resides in memory to accelerate the search often occupies a large amount of RAM, causing the IDE to lag. (4) Lagging update: After the project is modified, the index reconstruction usually takes several seconds or even longer, resulting in a lag in the display of search results. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, this application provides a fast symbol retrieval method for ultra-large multi-station PLC projects, which can significantly improve the symbol retrieval speed in massive projects while maintaining retrieval accuracy.
[0004] Firstly, this application provides a method for rapid symbol retrieval in ultra-large multi-station PLC projects, the method comprising the following steps: Data parsing, symbol metadata standardization, and fragmentation preprocessing are performed on the PLC project file to obtain standardized symbol metadata object SymbolMeta fragmented data; Based on standardized SymbolMeta sharded data, a three-level hierarchical index architecture is constructed, consisting of a station-level hash routing index, a task-level compressed prefix Trie tree index, and a symbol-level differentially encoded inverted list. Based on the user's query conditions, the three-level hierarchical index architecture sequentially performs station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation to generate structured search results, which are then rendered on the front end. The PLC project's modification instructions are used to perform incremental updates on the three-level hierarchical index architecture.
[0005] In one possible implementation, the process of parsing data, standardizing symbolic metadata, and preprocessing fragments in the PLC project file to obtain standardized symbolic metadata object (SymbolMeta) fragment data includes the following steps: Scan the PLC project root directory, build the project dependency tree, and traverse the project dependency tree, reading the project files in a streaming, block-based manner. The program parses the read project files, extracts the raw data of all sites, tasks, program organization units, global variable tables, and local variable tables; and simultaneously collects the usage locations of variables in the code during the parsing process to form a preliminary cross-reference table. The extracted raw data is uniformly encapsulated into a symbolic metadata object SymbolMeta, and a globally unique identifier UniqueID is generated using the SHA-256 algorithm; The SymbolMeta object is hashed and sharded in two levels according to the site and task, with each shard corresponding to an independent index building unit.
[0006] In one possible implementation, a station-level hash routing index is constructed by the following steps: Build a thread-safe concurrent hash table in memory, using the site's unique identifier as the key and the site's root node object as the value, and perform unified registration and management of the index entry for each site; Iterate through the SymbolMeta shards and count the total number of symbols, data type distribution information, and latest modification timestamp for each site. Set up a dynamic task pointer array inside the site root node, and determine the array index position based on the hash value of the unique task identifier; Configure independent filters for each site and enter the hash values corresponding to all symbol names within the site into the filters; when a search keyword fails the filter verification during the retrieval process, the subsequent retrieval process for that site is interrupted, and the first level of retrieval scope filtering is performed.
[0007] In one possible implementation, a task-level compressed prefix Trie tree index is constructed by the following steps: A character mapping table is established for commonly used characters in PLC variables, and a prefix tree node structure containing a child node array, data offset, and node identifier is defined. A radix prefix tree construction method is used to perform merging and compression on single non-terminal branch nodes; Configure a failure pointer on each node and point it to the node corresponding to the longest true suffix; when a character match fails, jump to the search based on this pointer.
[0008] In one possible implementation, the symbol-level differential coding inverted list is constructed by the following steps: Collect the identifier information within the task block corresponding to the associated symbols of the leaf nodes in the prefix tree, and construct an identifier sequence list; The data in the identifier sequence list is sorted in a monotonically increasing order to form an ordered absolute identifier sequence. Convert the ordered absolute identifier sequence into an adjacent difference sequence; Separate the detailed attributes of symbols from the index structure and store them in a columnar manner.
[0009] In one possible implementation, the step of generating structured search results by sequentially performing site-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation in a three-level hierarchical index architecture based on user query conditions includes the following steps: The system performs intent recognition and standardized preprocessing on the user's input query conditions to generate a retrieval execution plan. According to the retrieval execution plan, station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation are performed to obtain the retrieval result item SearchResultItem object; The SearchResultItem object is extracted and highlighted based on its context, and paginated according to the front-end request, and pushed to the front-end in a streaming manner.
[0010] In one possible implementation, the search results are rendered on the front end, including syntax highlighting, keyword positioning, and cross-reference display; the modification instructions for the PLC project include add, modify, and delete operations.
[0011] Secondly, this application provides a symbol rapid retrieval device for ultra-large multi-station PLC projects, the device comprising: The preprocessing module is used to perform data parsing, symbol metadata standardization, and fragment preprocessing on PLC project files to obtain standardized symbol metadata object SymbolMeta fragment data. The building module is used to construct a three-level hierarchical index architecture based on standardized SymbolMeta sharded data: a station-level hash routing index, a task-level compressed prefix Trie tree index, and a symbol-level differentially encoded inverted list. The query module is used to perform site-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation in a three-level hierarchical index architecture according to the user's query conditions, generate structured search results, and perform front-end rendering. The update module is used to perform incremental updates to the three-level hierarchical index architecture based on the modification instructions of the PLC project.
[0012] Thirdly, this application provides an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the symbol fast retrieval method for ultra-large multi-station PLC projects as described in any of the first aspects are performed.
[0013] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the symbol fast retrieval method for ultra-large multi-station PLC projects as described in any of the first aspects.
[0014] This embodiment provides a fast symbol retrieval method for ultra-large multi-station PLC projects. It performs data parsing, symbol metadata standardization, and fragment preprocessing on the PLC project file to obtain standardized symbol metadata objects (SymbolMeta) fragmented data. Based on the standardized SymbolMeta fragmented data, a three-level hierarchical index architecture is constructed, consisting of a station-level hash routing index, a task-level compressed prefix Trie tree index, and a symbol-level differential encoded inverted list. According to the user's query conditions, station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation are performed sequentially within the three-level hierarchical index architecture to generate structured search results, which are then rendered on the front end. Incremental updates are performed on the three-level hierarchical index architecture based on the PLC project's modification instructions. This solves the problems of high latency, weak fuzzy matching capability, high resource consumption, and delayed updates in retrieval of symbols exceeding 100,000, achieving millisecond-level accurate positioning. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A flowchart of a symbol fast retrieval method for ultra-large multi-station PLC projects according to an embodiment of this application is shown; Figure 2 This invention illustrates a flowchart of obtaining standardized symbol metadata object (SymbolMeta) fragment data according to an embodiment of this application; Figure 3 A flowchart illustrating the construction of a three-level hierarchical index architecture according to an embodiment of this application is shown; Figure 4This document illustrates a flowchart of an embodiment of the present application that generates structured search results based on user query criteria. Figure 5 This paper shows a schematic diagram of the symbol fast retrieval device for ultra-large multi-station PLC projects according to an embodiment of this application; Figure 6 A structural block diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0018] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0019] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0020] In view of the technical problems raised in the background, this application provides a fast symbol retrieval method for ultra-large multi-station PLC projects, which can significantly improve the symbol retrieval speed in massive projects while taking into account the retrieval accuracy.
[0021] See the instruction manual appendix Figure 1 This application provides a method for rapid symbol retrieval in ultra-large multi-station PLC projects, the method comprising the following steps: S1. Perform data parsing, symbol metadata standardization, and fragmentation preprocessing on the PLC project file to obtain standardized symbol metadata object SymbolMeta fragmented data. S2. Based on standardized SymbolMeta sharded data, construct a three-level hierarchical index architecture consisting of a station-level hash routing index, a task-level compressed prefix Trie tree index, and a symbol-level differential encoded inverted list. S3. Based on the user's query conditions, the three-level hierarchical index architecture sequentially performs station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation to generate structured search results, which are then rendered on the front end. S4. Based on the modification instructions of the PLC project, perform incremental updates on the three-level hierarchical index architecture.
[0022] Step S1 mainly involves converting the scattered, heterogeneous, and multi-site PLC project raw files into a unified, globally unique, and shardable set of symbolic metadata, providing a stable, clean, and parallelizable input for the efficient construction of the subsequent three-level hierarchical index.
[0023] See the instruction manual appendix Figure 2 The process of parsing data, standardizing symbolic metadata, and preprocessing fragments in the PLC project file to obtain standardized symbolic metadata object (SymbolMeta) fragment data includes the following steps: S101. Scan the PLC project root directory, construct the project dependency tree, and traverse the project dependency tree, reading the project files in a streaming block manner. S102. Parse the read project file to extract the original data of all sites, tasks, program organization units, global variable tables, and local variable tables; and simultaneously collect the usage locations of variables in the code during the parsing process to form a preliminary cross-reference table. S103. The extracted raw data is uniformly encapsulated into a symbolic metadata object SymbolMeta, and a globally unique identifier UniqueID is generated using the SHA-256 algorithm; S104. Perform two-level hash sharding on the SymbolMeta object according to the site and task, with each shard corresponding to an independent index building unit.
[0024] Specifically, in step S101, a monitoring process is constructed to monitor the PLC project storage directory, recursively scan the specified project root directory, and identify the main project file and all sub-project files. The project configuration file is parsed to construct a project dependency tree, clarifying the reference relationships between sites and ensuring that the subsequent parsing order conforms to the dependency logic. The project dependency tree is traversed, and project files are read one by one. For very large project files, streaming reading is used instead of loading them into memory all at once. The file is divided into fixed-size data blocks, and using a producer-consumer model, the reading tasks are placed in a blocking queue, where they are asynchronously consumed by the background parsing thread pool to prevent the main thread from blocking.
[0025] In step S102, lexical analysis is performed on the read data stream to identify keywords, identifiers, data types, and comments, generating a token stream. Invalid whitespace characters and compiler instructions are filtered out, retaining tokens with semantic value. Based on predefined PLC language grammar rules, the token stream is constructed into an abstract syntax tree, with node types including: ProgramNode, BlockNode, VariableListNode, VariableNode, and ProtocolVariableNode. A scope stack is maintained; when a new program block or data structure is entered, a new scope is pushed onto it. When a VariableNode is identified, it is mounted under the current top scope of the stack, and its complete hierarchical path context is recorded. The original data of all stations, tasks, program organization units, global variable tables, and local variable tables are extracted. Furthermore, during the traversal, the usage location of variables in the code is collected synchronously, including Read and Write operations, forming a preliminary cross-reference table for subsequent calculation of symbol heat weights and variable logical usage location.
[0026] In step S103, the extracted raw data is mapped to a unified symbol metadata object, SymbolMeta. Symbol names containing special characters or reserved words are escaped to ensure the validity of index key values. Based on the symbol's absolute path, data type fingerprint, and the hash value of the project it belongs to, a globally unique UniqueID is generated using the SHA-256 algorithm for accurate symbol location in subsequent incremental updates. Specific data types are mapped to a system-wide unified standard type enumeration. The address string is parsed to extract the base address, bit offset, and data length, which are stored as numeric fields for subsequent address range retrieval. Comment content is extracted, HTML tags or rich text formatting are removed, leaving only plain text, and word segmentation preprocessing is performed. A SymbolMeta object is instantiated, standardized fields are populated, and metadata such as creation time, last modifier, and whether it is hidden by an optimization block is appended.
[0027] Standardized fields include: UniqueID: A globally unique identifier, UUID; FullName: The standardized full name; Address: Memory address or I / O mapped address; DataType: Data type; Comment: Comment text; ScopeLevel: Scope level. 0: Global, 1: Site, 2: Task, 3: Local.
[0028] In step S104, a two-level hash sharding system is first constructed, consisting of stations and tasks. The first-level sharding uses the station's StationID as the hash key to allocate all symbols to different station buckets, achieving data isolation between stations. The second-level sharding further subdivides each station bucket into task shards using TaskID as the secondary key. Next, within each shard, the SymbolMeta list is pre-sorted according to the lexicographical order of symbol names or address continuity to improve compression and cache hit rates during subsequent Trie tree construction and differential encoding. Then, null values are filtered out, and redundant symbols are merged. Unassigned symbols, symbols of type Void, or compilation temporary symbols are removed. Completely duplicate symbol definitions are detected and merged, retaining only the primary definition with the highest reference count, and marking the rest as alias references. Finally, the generated task shards are encapsulated into independent index building tasks. A global work queue is established, and each independent index building task is placed in the work queue, awaiting scheduling execution in step S2, ensuring that multi-core CPUs can process index building for different stations in parallel without interference.
[0029] Step S2 primarily involves constructing a three-tiered index architecture based on the standardized SymbolMeta sharded data output from Step S1, using the site and task dimensions. This architecture comprises a site-level hash route, a task-level compressed prefix Trie tree, and a symbol-level differentially encoded inverted index. The site-level index enables rapid pruning and isolation of cross-site data; the task-level Trie tree utilizes compressed storage with common character prefixes to accelerate name matching; and the symbol-level inverted index significantly reduces memory usage through differential encoding and variable-length byte encoding techniques. Finally, the resulting compact index structure is persisted via memory mapping.
[0030] The following is in conjunction with the instruction manual appendix. Figure 3 This section explains how to construct a three-level hierarchical index architecture.
[0031] Specifically, building a Level-1 Routing Index (a site-level hash routing index) includes the following steps: (1) Global hash routing table initialization. A thread-safe concurrent hash table, ConcurrentHashMap, is initialized in memory.<StationID, StationNode> StationID is the unique identifier for the site, and StationNode is the root node object of the site.
[0032] (2) Site metadata aggregation. Traverse all input fragments and count the total number of symbols (TotalCount), data type distribution (TypeDistribution), and latest timestamp (LastModified) for each site.
[0033] (3) Task pointer array allocation. Inside each StationNode, a dynamic array TaskPtrArray is created. The array index is generated by mapping the hash value of TaskID, and the array elements store the memory offset Offset or pointer to the root node of the next level task index.
[0034] (4) Fast Pruning Bitmap Generation. A filter is generated for each site, and the hash values of all symbol names under that site are filled in. During retrieval, if the query term fails the filter test of a certain site, the subsequent traversal of that site is skipped directly, realizing the first level of millisecond-level pruning.
[0035] Constructing a Level-2 Compressed Trie index includes the following steps: (1) Character set mapping and node definition. For the commonly used character sets of PLC variable names (AZ, az, 0-9, _, .), a compact character mapping table CharMap is established to map ASCII codes to consecutive integers from 0 to N in order to reduce the sparsity of the child node array. The Trie node structure TrieNode is defined, which includes: children[N], representing the array of child node pointers; payloadOffset, if it is the end of a word, the offset pointing to the Level-3 data; flags, indicating whether it is a leaf node.
[0036] (2) Path compression optimization. A radix Trie tree construction strategy is adopted. When inserting a symbol name, if it is found that there is only one child node on the path from the current node to the child node and the child node is not the end of the word, the current node and the child node are merged and the common prefix string is stored, thereby reducing the height of the tree and reducing the number of memory accesses.
[0037] (3) Construction of fuzzy matching auxiliary chain. An additional FailureLink pointer is maintained in each node, pointing to the longest true suffix node in the current path. This is used to quickly jump to a possible suffix matching position when a character mismatch occurs in the fuzzy search, avoiding backtracking to the root node.
[0038] Constructing a symbol-level inverted list and differential encoding (Level-3 Inverted List & Compression) includes the following steps: (1) Inverted list generation. For each symbol attached to a leaf node of the Trie, collect all its occurrence instances to form a list L=[DocID_1, DocID_2, ..., DocID_n]. Where DocID is the relative address or line number of the symbol within a specific task block.
[0039] (2) Monotonic sorting. Ensure that DocID in list L is strictly monotonically increasing. If the original data is unordered, perform quicksort first.
[0040] (3) Differential encoding. Convert the absolute DocID sequence into a difference sequence D:
[0041] ( ) (4) Variable-length byte encoding. The difference sequence D is compressed using VByte: each difference is split into 7-bit groups. The most significant bit (MSB) of each byte is used as a continuation flag: 1 indicates that there are more bytes to follow, and 0 indicates that the current byte is the last byte.
[0042] (5) Attribute columnar storage. The detailed attributes of symbols are separated from the index structure and packaged separately using a columnar storage format. Only the pointer to the starting position of the attribute block is retained in the index, and the compression ratio is further improved by utilizing the local similarity of attributes.
[0043] Furthermore, memory mapping is used to achieve index persistence and efficient access. A contiguous memory block layout is designed: [Header | Level-1 Table | Level-2 Trie Blobs | Level-3 Compressed Lists | Property Blocks]. Precise byte offsets for each part are calculated to eliminate pointer dependencies, making the index file position-independent. This memory structure is written to a temporary index file as a byte stream. A CRC32 checksum is calculated for the entire index file and written to the file header for integrity verification at startup. Operating system APIs are called to directly map the index file to the process's virtual address space. The mapping attributes are set to PROT_READ (read-only) and MAP_SHARED, utilizing the operating system's page caching mechanism to manage memory and achieve zero-copy reading. A page fault is triggered only when accessing a specific data page to load from disk, achieving on-demand loading.
[0044] Step S3 mainly involves receiving mixed query conditions input by the user, supporting wildcards, regular expressions, and address ranges. First, a site-level filter is used to perform millisecond-level global pruning to exclude irrelevant sites. Then, multi-path parallel matching is performed in the task-level compressed prefix tree Trie to quickly locate the candidate symbol set. Attribute information is aggregated in real time using differential decoding technology of the inverted list. The matching results are deduplicated and context-highlighted. Finally, the structured search results are streamed to the front end using zero-copy memory mapping technology.
[0045] See the instruction manual appendix Figure 4 The first step is to perform intent recognition and standardization preprocessing on the user-input query conditions to generate a retrieval execution plan. Specifically, the first step is to automatically identify the patterns in the user input: Exact mode: Variable names that are exactly matched.
[0046] Wildcard pattern: .
[0047] Regular expression pattern: a regular expression enclosed in / ... / .
[0048] Address range mode: Recognizes hardware address format and parses it into a numerical range of start and end addresses.
[0049] Natural Language Modeling: Identify annotation keyword searches and extract core nouns and type constraints.
[0050] The input is then uniformly converted to the system's internal standard format, ensuring consistent capitalization and delimiters. If no exact match is found in the index, the edit distance between the input string and high-frequency symbols in the dictionary is immediately calculated. If the distance is less than a threshold, it is ignored. An optimal execution plan is then generated based on the query type. For address range queries, the Trie tree traversal is skipped, and routing is directly to the address inverted index; for prefix matching, a depth-first search path is planned for the Trie tree.
[0051] The second step involves executing station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation according to the retrieval execution plan to obtain the SearchResultItem object. Specifically, L1 station-level Bloom filter fast pruning is performed first. The query keywords are hashed, and the Bloom filters for all stations are tested in parallel. If a station returns "not found," it is directly removed from the candidate station list without accessing the disk data blocks of that station.
[0052] Then, an L2 task-level Trie tree multi-path traversal is performed. For exact matching and prefix matching, the traversal starts from the root node and proceeds downwards along the character edges. If a wildcard is encountered... This triggers branch explosion processing, traversing all subtrees of the current node in parallel. If a character mismatch occurs during traversal and fuzzy search is enabled, jump along FailureLink to the longest suffix node to continue matching, avoiding backtracking to the root node for recalculation.
[0053] Next, L3 inverted index differential decoding and attribute reassembly are performed. The offset of the compressed inverted index pointed to by the Trie leaf nodes is obtained. The VByte-encoded difference sequence is decoded in parallel using the SIMD instruction set to restore the actual DocID list. Based on the DocID, the physical offset in the columnar storage attribute block is calculated, and the type, address, comments, and program block information of symbols are read in batches, assembling them into a complete SearchResultItem object.
[0054] It also identifies symbols with the same memory address but different names, displaying only the main defined name by default, and collapsing the rest to avoid redundancy. Temporary compiler symbols located in "optimized blocks" or "hidden namespaces" are removed, resulting in a processed SearchResultItem object.
[0055] The third step involves extracting the context and highlighting the SearchResultItem object, then paginating and truncating the data according to the frontend request, and pushing it to the frontend in a streaming manner. Specifically, for the SearchResultItem results, code snippets (5 lines before and after) of variable definitions are extracted from the cache of the original project file. Syntax highlighting is applied to these code snippets, and matching keywords are marked with HTML tags for easy frontend rendering. A logical cursor pointing to a memory-mapped file is maintained, and the corresponding number of SearchResultItems are truncated from the cursor position based on the page number and page size requested by the frontend. The paginated data is pushed to the frontend using a binary stream. While pushing the first page of data, the backend asynchronously calculates and returns the total hit count, without waiting for the full search to complete.
[0056] The system includes a dynamic list of results rendered on the front end, featuring syntax highlighting, fuzzy keyword location, and cross-reference navigation. It also provides functionality to navigate to the source code and view dependencies. Specific steps include: (1) Virtual list rendering and context highlighting. First, virtual scrolling adaptation is performed. For the possible return of tens of thousands of result sets, the front end maintains a visual window with a fixed height. The SearchResultItem component in the current viewport (e.g., 50 items before and after) is dynamically calculated and rendered according to the scroll bar position, keeping the number of nodes constant at the O(1) level. The HighlightRange metadata in the results is parsed, and the highlighting style is applied to the variable names and matching words in the comments. For long comments or complex expressions, the intelligent truncation algorithm is automatically executed to retain the matching words and N characters before and after them, and replace them with ellipses in the middle to avoid the single line content being too long and breaking the layout.
[0057] (2) Listen to user operations. When the user clicks the “Locate” button, jump precisely to the site, program block and specific line number defined by the symbol, and select the line of code.
[0058] (3) Based on the usage location information in the inverted index, dynamically generate the call relationship table for the symbol. The list displays in which program blocks the variable is read, written, or passed as a parameter.
[0059] Step S4 primarily involves listening for user modification commands, asynchronously writing the changed data back to the underlying project files and triggering incremental index updates, extracting the changed symbol set, performing local reconstruction on the affected parts, performing insertion operations on newly added symbols and updating the Bloom filter for newly added symbols; marking deleted symbols as "logical deletion" in the inverted list, and triggering background cleanup when the compression ratio reaches a threshold. It also involves updating the columnar storage attribute blocks.
[0060] To clearly illustrate the actual implementation process and application advantages of the symbol fast retrieval method for ultra-large multi-station PLC projects proposed in this application, a detailed explanation is provided below in conjunction with a specific application scenario.
[0061] The application case selected is the main control program of the main control system of a 10MW offshore wind farm unit. The specific PLC control program includes the pitch system, yaw system, converter, hydraulic system and SCADA communication module.
[0062] Code size: A single unit contains 400+ function blocks (FBs), with a total of over 600,000 lines of code. There are over 850,000 variables.
[0063] After adopting the method of this application, a software test program was built, including 849,540 variables such as station variables, task variables, and POU local variables. The specific time results for functional tests such as lookup and cross-reference are shown in Table 1 below.
[0064] Table 1 This application provides a fast symbol retrieval method for ultra-large multi-station PLC projects. Through a three-level hierarchical index, station-level hash routing, task-level compressed Trie tree, and symbol-level differential inverted index, it achieves millisecond-level accurate retrieval, reducing symbol query complexity. In actual testing with 850,000 variables, the response time is ≤300ms, completely solving the latency problem. By triggering local index reconstruction through changes, logical deletion, and background cleanup mechanisms, update latency is significantly reduced. Front-end dynamic highlighting, cross-reference navigation, and streaming result rendering all have response times <500ms, providing a virtually imperceptible experience for the user.
[0065] Based on the same inventive concept, this application also provides a symbol fast retrieval device for ultra-large multi-station PLC projects. Since the principle of the device in this application is similar to the symbol fast retrieval method for ultra-large multi-station PLC projects described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0066] As per the instruction manual Figure 5 As shown in the figure, this application provides a symbol rapid retrieval device for ultra-large multi-station PLC projects, the device comprising: Preprocessing module 501 is used to perform data parsing, symbol metadata standardization, and fragment preprocessing on PLC project files to obtain standardized symbol metadata object SymbolMeta fragment data; Module 502 is used to build a three-level hierarchical index architecture based on standardized SymbolMeta sharded data, consisting of a station-level hash routing index, a task-level compressed prefix Trie tree index, and a symbol-level differentially encoded inverted list. The query module 503 is used to perform site-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation in a three-level hierarchical index architecture according to the user's query conditions, generate structured search results, and perform front-end rendering. Update module 504 is used to perform incremental updates on the three-level hierarchical index architecture based on the modification instructions of the PLC project.
[0067] This application provides a rapid symbol retrieval device for ultra-large multi-station PLC projects. Addressing the issues of large data volume, scattered distribution, and low retrieval efficiency of symbols in ultra-large multi-station PLC projects, it establishes a three-level hierarchical index system at the station, task, and symbol levels. Hash routing enables rapid station filtering, compressed prefix trees optimize name matching efficiency, and differential encoding completes efficient compressed storage of the inverted list, significantly reducing the memory space occupied by the index. The retrieval process filters data level by level, reducing invalid data traversal and significantly shortening query response time. It also includes result highlighting and cross-referencing display functions, not only meeting the needs for rapid and accurate retrieval of large batches of symbols but also intuitively presenting the associations between symbols. Adaptable to large-scale industrial engineering operation and maintenance scenarios, it offers enhanced practicality and operational stability.
[0068] Based on the same concept of the present invention, the specification is attached. Figure 6 As shown in the figure, an embodiment of this application provides the structure of an electronic device 600, which includes: at least one processor 601, at least one network interface 604 or other user interface 603, memory 605, and at least one communication bus 602. The communication bus 602 is used to realize the connection and communication between these components. The electronic device 600 may optionally include a user interface 603, including a display (e.g., touch screen, LCD, CRT, holographic imaging, or projector, etc.), a keyboard, or a clicking device (e.g., mouse, trackball, touchpad, or touch screen, etc.).
[0069] Memory 605 may include read-only memory and random access memory, and provides instructions and data to processor 601. A portion of memory 605 may also include non-volatile random access memory (NVRAM).
[0070] In some implementations, memory 605 stores executable modules or data structures, or subsets thereof, or extended sets thereof: The 6051 operating system contains various system programs used to implement various basic business functions and handle hardware-based tasks. Application module 6052 contains various applications, such as desktop (launcher), media player (MediaPlayer), browser (Browser), etc., to implement various application services.
[0071] In this embodiment, by calling the program or instructions stored in the memory 605, the processor 601 executes steps such as in a fast symbol retrieval method for ultra-large multi-station PLC projects, which can significantly improve the symbol retrieval speed in massive projects while maintaining retrieval accuracy.
[0072] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs steps such as those in a fast symbol retrieval method for ultra-large multi-station PLC projects.
[0073] Specifically, the storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the storage medium is run, it can execute the above-mentioned symbol fast retrieval method for ultra-large multi-station PLC projects.
[0074] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0075] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0076] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0077] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0078] Finally, it should be noted that the above embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for rapid symbol retrieval in ultra-large multi-station PLC projects, characterized in that, The method includes the following steps: Data parsing, symbol metadata standardization, and fragmentation preprocessing are performed on the PLC project file to obtain standardized symbol metadata object SymbolMeta fragmented data; Based on standardized SymbolMeta sharded data, a three-level hierarchical index architecture is constructed, consisting of a station-level hash routing index, a task-level compressed prefix Trie tree index, and a symbol-level differentially encoded inverted list. The station-level hash routing index is constructed through the following steps: A thread-safe concurrent hash table is built in memory, using the station's unique identifier as the key and the station's root node object as the value, for unified registration and management of index entries for each station; the SymbolMeta shards are traversed to count the total number of symbols, data type distribution information, and the latest modification timestamp for each station; a dynamic task pointer array is set within the station's root node, with the array index position determined based on the task's unique identifier hash value; independent filters are configured for each station, and the hash values corresponding to all symbol names within the station are entered into the filters; when a search keyword fails the filter validation during the retrieval process, the subsequent retrieval process for that station is terminated, and the first-level retrieval scope is filtered. The task-level compressed prefix Trie tree index is constructed as follows, including the following steps: A character mapping table is established for commonly used characters in PLC variables, and a prefix tree node structure containing a child node array, data offset, and node identifier is defined; A radix prefix tree construction method is used to perform merging and compression on single non-terminal branch nodes; A failure pointer is configured on each node and points to the node corresponding to the longest true suffix; When a character match fails, a jump search is performed based on this pointer. The symbol-level differential coding inverted list is constructed as follows, including the following steps: collecting the identification information within the task block corresponding to the symbols associated with the leaf nodes of the prefix tree, and constructing an identification sequence list; performing monotonically increasing sorting on the data in the identification sequence list to form an ordered absolute identification sequence; converting the ordered absolute identification sequence into an adjacent difference sequence; separating the symbol detailed attributes from the index structure and storing them in columnar format. Based on the user's query conditions, the three-level hierarchical index architecture sequentially performs station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation to generate structured search results, which are then rendered on the front end. The PLC project's modification instructions are used to perform incremental updates on the three-level hierarchical index architecture.
2. The symbol fast retrieval method for ultra-large multi-station PLC projects according to claim 1, characterized in that, The process of parsing data, standardizing symbolic metadata, and preprocessing fragments in the PLC project file to obtain standardized symbolic metadata object (SymbolMeta) fragment data includes the following steps: Scan the PLC project root directory, build the project dependency tree, and traverse the project dependency tree, reading the project files in a streaming, block-based manner. The program parses the read project files, extracts the raw data of all sites, tasks, program organization units, global variable tables, and local variable tables; and simultaneously collects the usage locations of variables in the code during the parsing process to form a preliminary cross-reference table. The extracted raw data is uniformly encapsulated into a symbolic metadata object SymbolMeta, and a globally unique identifier UniqueID is generated using the SHA-256 algorithm; The SymbolMeta object is hashed and sharded in two levels according to the site and task, with each shard corresponding to an independent index building unit.
3. The symbol fast retrieval method for ultra-large multi-station PLC projects according to claim 2, characterized in that, The process of generating structured search results by performing site-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation sequentially within a three-level hierarchical index architecture based on user query conditions includes the following steps: The system performs intent recognition and standardized preprocessing on the user's input query conditions to generate a retrieval execution plan. According to the retrieval execution plan, station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation are performed to obtain the retrieval result item SearchResultItem object; The SearchResultItem object is extracted and highlighted based on its context, and paginated according to the front-end request, and pushed to the front-end in a streaming manner.
4. The symbol fast retrieval method for ultra-large multi-station PLC projects according to claim 3, characterized in that, in, The search results are rendered on the front end, including syntax highlighting, keyword positioning, and cross-reference display; the modification instructions for PLC projects include add, modify, and delete operations.
5. A symbol rapid retrieval device for ultra-large multi-station PLC projects, characterized in that, The device includes: The preprocessing module is used to perform data parsing, symbol metadata standardization, and fragment preprocessing on PLC project files to obtain standardized symbol metadata object SymbolMeta fragment data. The module is used to construct a three-level hierarchical index architecture based on standardized SymbolMeta sharded data, consisting of a station-level hash routing index, a task-level compressed prefix Trie tree index, and a symbol-level differentially encoded inverted list. The station-level hash routing index is constructed as follows: a thread-safe concurrent hash table is built in memory, using the station's unique identifier as the key and the station's root node object as the value, for unified registration and management of index entries for each station; the SymbolMeta shards are traversed to count the total number of symbols, data type distribution information, and the latest modification timestamp for each station; a dynamic task pointer array is set up inside the station's root node, with the array index position determined by the hash value of the task's unique identifier; independent filters are configured for each station, and the hash values corresponding to all symbol names within the station are entered into the filters; when a search keyword fails the filter validation during the retrieval process, the subsequent retrieval process for that station is terminated, and the first-level retrieval scope is filtered. The task-level compressed prefix Trie tree index is constructed as follows: a character mapping table is established for commonly used characters in PLC variables, and a prefix tree node structure containing a child node array, data offset, and node identifier is defined; a radix prefix tree construction method is used to perform merging and compression on single non-terminal branch nodes; a failure pointer is configured on each node and points to the node corresponding to the longest true suffix; when a character match fails, a jump search is performed based on this pointer. The symbol-level differential coding inverted list is constructed as follows: collecting the identification information within the task block corresponding to the symbols associated with the leaf nodes of the prefix tree, and constructing an identification sequence list; performing monotonically increasing sorting on the data in the identification sequence list to form an ordered absolute identification sequence; converting the ordered absolute identification sequence into an adjacent difference sequence; and separating the detailed symbol attributes from the index structure and storing them in columnar format. The query module is used to perform site-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation in a three-level hierarchical index architecture according to the user's query conditions, generate structured search results, and perform front-end rendering. The update module is used to perform incremental updates to the three-level hierarchical index architecture based on the modification instructions of the PLC project.
6. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the symbol fast retrieval method for ultra-large multi-station PLC projects as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the symbol fast retrieval method for ultra-large multi-station PLC projects as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Quick matching method for fuzzy keywords based on Trie tree
CN121614652A
KR1026573730000B1