Method and system for tracking production information of a cigarette product
By generating unique codes during the cigarette production process and combining them with data collected by IoT devices, a full-process association graph is constructed using a stream processing engine and graph database to generate a cigarette product production flow trajectory chain. This solves the problems of data collection lag and accuracy in the cigarette production process, and realizes refined management and visual tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONGTA TOBACCO (GROUP) CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-26
AI Technical Summary
In the cigarette production process, the existing technology has a lag in the collection of data throughout the entire process from cigarette production to industrial and commercial handover. Data verification consumes a lot of manpower and resources, and the accuracy of the data is difficult to guarantee, making it difficult for cigarette factories to achieve refined management.
By generating unique codes based on preset coding rules, combining real-time collection of timestamps, status identifiers, and location coordinates from industrial IoT devices, and using the Apache Flink stream processing engine for real-time association and cleaning, a full-process association graph of the Neo4j graph database is constructed. The production flow trajectory chain of cigarette products is generated through dynamic programming algorithms, and finally spatial arrangement is performed in a visualization engine.
It enables real-time data tracking of the entire process from cigarette production to industrial and commercial handover, ensuring data continuity and accuracy, reducing manual intervention, improving information reading efficiency, and forming an intuitive and visual production flow trajectory map to support the refined management of cigarette factories.
Smart Images

Figure CN122089334A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of tobacco technology, and in particular to a method and system for tracking cigarette product production information. Background Technology
[0002] In the cigarette production process, the rolling and packaging stage completes the processes of cigarette stick forming, cigarette pack forming, carton forming, and package forming; packages of cigarettes are transported to the logistics workshop for palletized stacking, palletized cigarette storage, and inventory management; then, according to orders, palletized cigarettes are released from the warehouse, packages of cigarettes are unpacked, loaded onto trucks, and shipped, finally completing the handover between the manufacturer and distributor. All of the above stages are managed internally within the cigarette factory and belong to the cigarette production process.
[0003] Currently, with the implementation of tobacco industry-specific conditions and regulatory projects, the cigarette production process has achieved one code per pack, one code per carton, and one code per case, and has completed the associated management of cigarette packs, cartons, and cases. However, throughout the entire process from cigarette production to handover to the industrial and commercial authorities, production information is incomplete and scattered across different systems, failing to form a complete and effective information association. Cigarette factories find it difficult to grasp the status information of each pack of cigarettes in real time. In addition, during the production statistics, inventory checks, and logistics delivery processes, a large amount of data collection, entry, and verification work is still required. Data collection has a certain lag, and data verification consumes a lot of manpower and resources. The accuracy of the data is difficult to guarantee, making it difficult to achieve refined management of finished cigarette products in cigarette factories. Summary of the Invention
[0004] The main purpose of this application is to provide a method and system for tracking cigarette product production information, in order to solve the problems in the existing technology where there is a certain lag in data collection throughout the entire process from cigarette production to industrial and commercial handover, data verification consumes a lot of manpower and resources, and the accuracy of data is difficult to guarantee.
[0005] To achieve the above objectives, this application provides the following technical solution: A method for tracking cigarette product production information, the tracking method comprising: Step S1: Generate a unique code for each pack, carton, case, and tray of cigarette products based on preset coding rules. Step S2: Collect the timestamps, status identifiers, and location coordinates of the unique code in real time throughout the entire process from production to industrial and commercial handover using industrial IoT devices to form the raw data stream; Step S3: Input the original data stream into the Apache Flink-based stream processing engine. The window aggregation operator of the stream processing engine performs real-time association and cleaning of timestamps, status identifiers, and location coordinates according to the encoding dimension of the original data stream to obtain a structured event sequence. Step S4: Import the structured event sequence into a graph database, and construct multidimensional relationship edges between the encoded nodes and process nodes of the structured events using Neo4j's Cypher query language to form a full-process association graph; Step S5: Using a dynamic programming algorithm, traverse the shortest paths between the coded nodes and process nodes of the entire process association graph to generate a cigarette product production flow trajectory chain from the perspective of any code. Step S6: Map the cigarette product production flow trajectory chain to the visualization engine, and use the force-directed graph layout algorithm to spatially arrange the nodes and edges in the cigarette product production flow trajectory chain to obtain an interactive production flow trajectory graph.
[0006] Beneficial effects of steps S1 to S6: By using predefined coding rules to uniquely identify each level of cigarette products, a structured identity binding is achieved from individual packs to trays, laying the foundation for subsequent data association. Industrial IoT devices continuously collect the timestamps, status identifiers, and location coordinates corresponding to the codes, forming a continuous and traceable raw data stream to ensure that the physical behavior of the entire production process is digitally recorded. The stream processing engine is based on Apache. Flink performs real-time grouping, time-series alignment, and state merging on the raw data stream to eliminate redundancy and jumps, outputting a highly consistent structured event sequence. The graph database uses the Cypher language to transform the event sequence into a multi-dimensional relationship graph of coded nodes and process nodes, constructing a semantically complete full-process topology. The dynamic programming algorithm traverses this graph, extracting the complete flow path of any coded element according to the time sequence, generating a production trajectory chain without breaks or redundancy. Finally, the force-directed graph layout algorithm spatially rearranges the nodes and edges in the trajectory chain, achieving a natural balance in node distribution through physical simulation, outputting an interactive visualization, presenting complex production relationships in an intuitive form, and improving information reading efficiency. The entire process achieves end-to-end mapping from physical entities to digital trajectories, without relying on human intervention, and is completely driven by data.
[0007] As a further improvement to this application, step S1 involves generating a unique code for each pack, each carton, each case, and each tray of cigarette products based on preset coding rules, including: Step S1.1: Define a coding system architecture consisting of four fixed fields concatenated in sequence: product level identifier, cigarette brand origin code, unique serial number, and check code; Step S1.2: Generate a product level identifier B for a single pack of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single pack of cigarettes to form a unique code for the single pack; Step S1.3: Generate a product level identifier T for a single carton of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single carton of cigarettes to form a unique code for each carton; Step S1.4: In the cigarette production process, the unique code of each carton is collected by an external scanning device, and the set of unique codes of all individual boxes in the cigarette packaging is collected simultaneously to establish a parent-child relationship between the unique code of each carton and the set of unique codes of each individual box. Step S1.5: Generate a product level identifier C for a single cigarette through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single cigarette to form a unique code for each cigarette. Step S1.6: In the cigarette production and packaging process, the unique code of each piece is collected by an external scanning device, and all unique codes in the cigarette packaging box are collected in sequence according to the preset piece position number order, so as to establish the parent-child relationship between the unique code of each piece and the set of unique codes. Step S1.7: Generate a product level identifier P for a single tray of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single tray of cigarettes to form a unique code for each tray. Step S1.8: In the pallet stacking process, the unique code of each pallet is collected by an external scanning device, and all the unique codes of each item on the pallet of cigarettes are collected in sequence according to the preset stack position number order, so as to establish the parent-child relationship between the unique code of each pallet and the set of unique codes of each item.
[0008] Beneficial effects of steps S1.1 to S1.8: By defining a fixed coding architecture that includes product level identifiers, cigarette brand and origin codes, unique serial numbers, and check codes, standardized identification of each level of cigarette products is achieved, ensuring a unified and resolvable coding structure. Individual boxes, cartons, cases, and trays are assigned corresponding level identifiers B, T, C, and P, respectively, forming a hierarchical coding system to support accurate classification and traceability of subsequent data. During the production stages of cartons, cases, and trays, scanning devices collect the lower-level code sets, establishing clear parent-child relationships between upper and lower-level codes, forming a structured hierarchical relationship chain. This process does not rely on manual judgment and is complete. Driven entirely by device acquisition and encoding rules, the objectivity and consistency of the relationships are ensured. The encoding system expands progressively from the smallest to the largest unit, with each level strictly following the same architecture, guaranteeing the connectivity and scalability of cross-level data. The establishment of parent-child relationships is based on physical acquisition behavior rather than logical inference, giving the relationships a mapping basis to the real production process. The entire encoding generation and association process forms a closed and self-consistent data chain, providing reliable identity anchors for subsequent stream processing, graph construction, and trajectory tracking, without introducing additional semantics, and completing the digital representation of product entities solely through structured encoding.
[0009] As a further improvement to this application, step S2 involves using industrial IoT devices to collect in real-time timestamps, status identifiers, and location coordinates of the unique code throughout the entire process from production to industrial and commercial handover, forming a raw data stream, including: Step S2.1: Identify the unique code and record the collection timestamp by an external RFID reader installed at the exit of the cigarette production line, and bind it with the unique code to form an original data tuple; Step S2.2: Align the original data tuples with the acquisition events from different devices using a time series alignment algorithm to obtain time-consistent code-timestamp pairs. Step S2.3: Collect the uniquely encoded physical location through an external industrial vision sensor, combine it with the three-dimensional coordinate system model of the factory area, and convert the pixel coordinates of the visual recognition into latitude and longitude values in the global geographic coordinate system through a coordinate encoding mapping algorithm to obtain the location coordinates; Step S2.4: Merge the code-timestamp pair with the location coordinates and match the status identifier corresponding to the current process node to obtain the original data stream including the unique code, the timestamp, the location coordinates, and the status identifier.
[0010] Beneficial effects of steps S2.1 to S2.4: By using RFID readers to identify the unique codes of cigarette products at the production line exit and simultaneously recording the collection time, the initial binding of physical entities and digital identifiers is achieved, providing a raw data entry point for subsequent tracking. A time-series alignment algorithm unifies the collected events from different devices to the same time base, eliminating timing discrepancies caused by clock deviations and ensuring accurate and reliable correspondence between codes and timestamps. Industrial vision sensors, combined with a 3D coordinate system model of the factory area, convert the pixel positions obtained from image recognition into latitude and longitude values in a global geographic coordinate system, completing the digital mapping of physical spatial locations and giving product locations calculable and comparable geographic semantics. Finally, the code-timestamp pairs are merged with the location coordinates and associated with the status identifier corresponding to the current process node, forming a complete raw data stream containing four elements: code, time, location, and status. This enables structured recording of dynamic product behavior in the production process. The entire process does not rely on manual intervention; data collection and association are entirely automated by equipment and algorithms, providing a consistent, complete, and traceable input foundation for subsequent stream processing.
[0011] As a further improvement to this application, in step S3, the original data stream is input into a stream processing engine based on Apache Flink. The window aggregation operator of the stream processing engine performs real-time correlation and cleaning of the timestamps, status identifiers, and location coordinates according to the encoding dimensions of the original data stream, resulting in a structured event sequence, including: Step S3.1: Partition the original data stream according to the encoded primary key, and perform stateful grouping according to the unique encoding using Apache Flink's KeyedProcessFunction; Step S3.2: Synchronize the event time sequence of each data record after grouping with the same unique code across devices using the sliding window time sequence alignment algorithm to obtain a time sequence consistent event sequence; Step S3.3: Input the time-consistent event sequence into the state machine-driven event merging engine, eliminate repeated state transitions according to preset state transition rules, and obtain a simplified event stream; Step S3.4: Input the simplified event stream into the hash deduplication engine, calculate the unique fingerprint of each event using the MurmurHash3 algorithm, remove duplicate event records and retain unique event instances to obtain the processed event sequence; Step S3.5: The same unique encoding of the processed event sequence is used to aggregate all timestamps, status identifiers and location coordinates within the window into a structured event sequence through the window aggregation algorithm.
[0012] Beneficial effects of steps S3.1 to S3.5: The KeyedProcessFunction partitions the raw data stream by unique encoding, enabling independent state management for each encoding's corresponding event stream and ensuring scalable parallel processing capabilities. A sliding window mechanism synchronizes the timing of events collected across devices, eliminating sequence errors caused by device clock differences and maintaining logical consistency of events with the same encoding over time. A state machine-driven event merging engine identifies and removes invalid state transitions based on preset state transition rules, reducing data noise and improving the semantic accuracy of the event stream. A hash deduplication engine uses the MurmurHash3 algorithm to generate a fixed-length fingerprint for each event, quickly identifying and removing duplicate records to ensure the uniqueness of event instances. A window aggregation operator aggregates the deduplicated event stream into a structured event sequence containing complete timestamps, state identifiers, and location coordinates, completing the transformation from raw data collection to highly consistent data units. The entire process runs continuously in a streaming environment, without relying on batch processing or introducing manual intervention; it achieves precise data cleaning and structured output solely through algorithms and state mechanisms.
[0013] As a further improvement to this application, step S4 involves importing the structured event sequence into a graph database, and using Neo4j's Cypher query language to construct multidimensional relationship edges between the encoded nodes and process nodes of the structured events, forming a full-process association graph, including: Step S4.1: Group the structured event sequence according to the unique code, and map the code, status identifier, and location coordinates in each record of the structured event sequence to node attributes in the graph database using a graph entity recognition algorithm to obtain a standardized node dataset; Step S4.2: Classify the standardized node dataset by node type, and assign four entity types to the nodes: “package”, “item”, “piece”, and “pallet”, through encoding hierarchy identifiers, to form a type-labeled node set; Step S4.3: Identify the three types of parent-child relationships in the set of labeled nodes: package-strip, strip-item, and item-pallet, and generate candidate relationship triples; Step S4.4: The time consistency verification algorithm is used to filter out the candidate relation triplets and the relation pairs whose timestamps satisfy a strict chronological order, so as to obtain the set of time-series legal association relations; Step S4.5: Merge the set of temporal legal associations with the set of type-labeled nodes, and identify the continuous flow path across levels using a graph pattern matching algorithm to obtain the complete production sequence sub-image segment; Step S4.6: Merge the continuous "package-condition-component" three-node links of the complete production sequence sub-image segment into a single composite edge to form a compressed sub-graph; Step S4.7: Write the compressed subgraph into the Neo4j graph database through the graph transaction commit protocol to obtain the immutable graph state; Step S4.8: The immutable graph state is used to generate a global inverted index based on unique encoding through a graph index building engine to obtain the full-process association graph.
[0014] Beneficial effects of steps S4.1 to S4.8: Events are grouped by unique encoding, and a standardized node dataset is established by mapping encoding, status identifiers, and location coordinates as node attributes using a graph entity recognition algorithm, thus unifying data representation. Nodes are assigned four entity types—package, bar, item, and tray—based on their encoding hierarchy, forming a type-labeled node set and clarifying node identity boundaries. Package-bar, bar-item, and item-tray parent-child relationships are identified, generating candidate relationship triples and extracting potential flow relationships. A time consistency verification algorithm is used to filter candidate relationships for pairs whose timestamps conform to the synchronization order, resulting in a set of time-series valid association relationships, ensuring the logical correctness of the relationships. These relationships are then merged. The relationship set and type-labeled node set are used to identify continuous flow paths across levels using a graph pattern matching algorithm, forming complete production sequence sub-image segments and outlining the basic flow context. Continuous package-condition-component three-node links in the sub-graph are merged into a single composite edge, compressing the sub-graph structure and simplifying relationship expression. The compressed sub-graph is written to Neo4j through a graph transaction commit protocol, forming an immutable graph state and ensuring the atomicity and consistency of data storage. A graph index building engine generates a global inverted index based on unique encoding, ultimately obtaining a full-process association graph, providing a structured foundation for subsequent path traversal and relationship queries.
[0015] As a further improvement to this application, step S5 involves using a dynamic programming algorithm to traverse the shortest paths between the encoded nodes and process nodes of the entire process association graph, generating a cigarette product production flow trajectory chain from the perspective of any encoding, including: Step S5.1: Select any coded node in the full-process association graph as the starting node through the graph traversal engine, and extract all outgoing edges of the starting node to form an initial path candidate set; Step S5.2: Sort each outgoing edge in the initial path candidate set in ascending order according to the timestamp field, following the causal order of time. Step S5.3: Take the node pointed to by the first outgoing edge after sorting as the current traversal node, recursively expand all outgoing edges of the current traversal node using the depth-first search algorithm, and retain the edges whose timestamps are greater than the timestamp of the current node. Step S5.4: During the recursive traversal, perform a path splicing operation on each selected edge, concatenating the current node code and the next node code in timestamp order to form an intermediate trajectory segment; Step S5.5: When the depth-first search reaches a leaf node with no outgoing edges or whose outgoing edge timestamps do not meet the increasing condition, terminate the current branch traversal and output the complete trajectory fragment of the current branch. Step S5.6: Aggregate the trajectory segments of all terminating branches according to the uniqueness of the starting encoding, eliminate duplicate node sequences through the path merging algorithm, and retain the unique time series continuous trajectory chain; Step S5.7: Complete the missing node timestamps in the trajectory chain by linear interpolation to obtain a trajectory chain without time breakpoints; Step S5.8: The time-breakpoint-free trajectory chain is serialized into a linear string format according to the encoded primary key to obtain a standardized cigarette product production flow trajectory chain.
[0016] Beneficial effects of steps S5.1 to S5.8: By selecting any coded node in the graph as the starting node, all its outgoing edges are extracted to form an initial path candidate set, establishing the traversal starting point. The outgoing edges of the candidate set are sorted in ascending order by timestamp, following the causal order of time to ensure the reasonable temporal sequence of the paths. The node pointed to by the first outgoing edge after sorting is taken as the current traversal node, and the outgoing edges are recursively expanded using depth-first search, retaining only edges with timestamps greater than the current node, focusing on effective flow relationships. Path splicing is performed in the recursion, and the node codes are concatenated according to timestamps to form intermediate trajectory fragments, accumulating flow information. When DFS reaches a leaf node with no outgoing edges or whose outgoing edge timestamps are not increasing, the branch terminates, and the complete trajectory fragment of that branch is output. All terminated branch fragments are aggregated, and the path merging algorithm is used to eliminate duplicate node sequences according to the starting code, retaining a unique, continuous trajectory chain with a unique time sequence. The missing node timestamps of the trajectory chain are completed by linear interpolation to eliminate time breakpoints. The trajectory chain without breakpoints is serialized into a linear string according to the encoded primary key to form a standardized flow trajectory chain.
[0017] As a further improvement to this application, step S6 maps the cigarette product production flow trajectory chain to a visualization engine, and uses a force-directed graph layout algorithm to spatially arrange the nodes and edges in the cigarette product production flow trajectory chain to obtain an interactive production flow trajectory diagram, including: Step S6.1: Extract all node identifiers and edge connections in the cigarette product production flow trajectory chain using a graph structure parsing engine to obtain a standardized list of nodes and an edge list; Step S6.2: Assign non-overlapping initial two-dimensional plane coordinates to each node in the node list using a uniform distribution algorithm while maintaining topological scalability; Step S6.3: Input the edge list with the initial two-dimensional plane coordinates into the force-guided graph layout algorithm, drive the interaction between each node with Coulomb force and Hooke force, and iteratively calculate the displacement vector of each node; Step S6.4: In each iteration, the displacement vectors of all nodes are normalized using the Euclidean distance normalization algorithm to constrain the displacement amplitude within a preset step size range. Step S6.5: Merge the updated node coordinates with the edge list to obtain a node coordinate sequence; Step S6.6: Spatial clustering compression is performed on the node coordinate sequence, and the clustered node coordinates and edge relationships are input into the graphics serialization engine. A structured layout description file including node ID, coordinates, and edge connection relationships is output in a fixed format. Step S6.7: Input the layout description file into the visualization rendering engine through the API interface, and dynamically bind it with the interactive production flow trajectory diagram to obtain the interactive production flow trajectory diagram.
[0018] Beneficial effects of steps S6.1 to S6.7: By extracting node identifiers and edge connections in the trajectory chain using a graph structure analysis engine, a standardized node list and edge list are obtained, clarifying the basic structure of the visualized object. A uniform distribution algorithm is used to assign non-overlapping initial two-dimensional plane coordinates to the node list while maintaining topological scalability, laying the foundation for the layout. The edge list and initial coordinates are input into a force-guided graph layout algorithm, using Coulomb forces and Hooke forces to drive node interactions, iteratively calculating the displacement vector of each node to optimize the spatial distribution. In each iteration, the displacement vector is normalized using an Euclidean distance normalization algorithm, constraining the displacement amplitude within a preset step size range to avoid layout oscillations. The updated node coordinates and edge lists are merged to obtain a node coordinate sequence. This sequence is spatially clustered and compressed, and then input into a graph serialization engine to output a structured layout description file containing node IDs, coordinates, and edge connections in a fixed format. This file is then input into a visualization rendering engine via an API interface and dynamically bound to an interactive production flow trajectory map, ultimately achieving an intuitive spatial arrangement of trajectory chain nodes and edges.
[0019] To achieve the above objectives, this application also provides the following technical solutions: A tracking system for cigarette product production information, the tracking system being applied to the tracking method described above, the tracking system comprising: The cigarette product coding generation module is used to generate unique codes for each pack, each carton, each case, and each tray of cigarette products based on preset coding rules. The cigarette product flow detection module is used to collect the timestamps, status identifiers, and location coordinates of the unique code in real time from production to industrial and commercial handover through industrial IoT devices, forming a raw data stream. The cigarette product flow data preprocessing module is used to input the raw data stream into an Apache Flink-based stream processing engine. The window aggregation operator of the stream processing engine performs real-time association and cleaning of timestamps, status identifiers, and location coordinates according to the encoding dimension of the raw data stream to obtain a structured event sequence. The full-process association graph generation module is used to import the structured event sequence into a graph database, and construct multi-dimensional relationship edges between the encoded nodes and process nodes of the structured events using Neo4j's Cypher query language to form a full-process association graph. The production flow trajectory chain generation module is used to traverse the shortest path between the coded nodes and process nodes of the full-process association graph using a dynamic programming algorithm to generate a cigarette product production flow trajectory chain from the perspective of any code. The production flow trajectory visualization module is used to map the cigarette product production flow trajectory to the visualization engine. The force-directed graph layout algorithm is used to spatially arrange the nodes and edges in the cigarette product production flow trajectory to obtain an interactive production flow trajectory graph.
[0020] To achieve the above objectives, this application also provides the following technical solutions: An electronic device includes a processor and a memory coupled to the processor, the memory storing program instructions executable by the processor; when the processor executes the program instructions stored in the memory, it implements the tracing method described above.
[0021] To achieve the above objectives, this application also provides the following technical solutions: A computer-readable storage medium storing program instructions that, when executed by a processor, enable the tracing method described above. Attached Figure Description
[0022] Figure 1 This is a schematic flowchart illustrating the steps of one embodiment of a method for tracking cigarette product production information according to this application; Figure 2 This is a schematic diagram of the functional modules of an embodiment of a cigarette product production information tracking system according to this application; Figure 3 This is a schematic diagram of the structure of an embodiment of the electronic device of this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the storage medium of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0024] The terms "first," "second," and "third" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationships and movements between components in a specific orientation (as shown in the figures). If the specific orientation changes, the directional indications also change accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0025] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0026] like Figure 1 As shown, this embodiment provides an example of a method for tracking cigarette product production information. In this embodiment, the tracking method includes the following steps: Step S1: Generate a unique code for each pack, carton, piece, and tray of cigarette products based on preset coding rules.
[0027] Further, step S1 involves generating a unique code for each pack, carton, case, and tray of cigarettes based on preset coding rules, specifically including the following steps: Step S1.1 defines a coding system architecture consisting of four fixed fields: product level identifier, cigarette brand origin code, unique serial number, and check code, which are sequentially concatenated.
[0028] Preferably, the product level identifier is a single-character code using the ASCII printable character set, with the package level being B, the bar level being T, the piece level being C, and the tray level being P, and the length being 1 byte; the cigarette brand origin code can refer to the brand code specified in the national standard GB / T-5606.1, and the brand origin code is 6 bytes long, with spaces added to the right if it is less than 6 bytes.
[0029] Preferably, the unique serial number is generated as a 64-bit integer (including a sign bit) using the Snowflake Algorithm. The structure of this 64-bit integer is: a 41-bit timestamp (millisecond level, the starting epoch can be set to 2020-01-01 00:00:00UTC) + a 10-bit machine ID (assigned according to production line number, ranging from 0 to 1023) + a 12-bit serial number (incrementing every millisecond for a single node, ranging from 0 to 4095), converted to a decimal string and the last 16 bits are extracted. The checksum is calculated using the CRC32 algorithm (polynomial 0xEDB88320) to obtain a 32-bit check value from the concatenated string of the first three elements, converted to a hexadecimal string and the lower 8 bits are extracted.
[0030] Step S1.2: Generate a product level identifier B for a single pack of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single pack of cigarettes to form a unique code for the single pack.
[0031] Preferably, the product level identifier for a single pack of cigarettes is set to B, the cigarette brand origin code is taken from the current production brand configuration file, and the unique serial number can be obtained by calling the snowflake algorithm example mentioned above (the machine ID is bound to the current packaging machine number, and the serial number is initially 0). The verification code can be obtained by using the CRC32 algorithm mentioned above.
[0032] Preferably, one example of single-box uniqueness encoding is shown in the following pseudocode: def generate_box_code(brand_code): layer ='B' seq = snowflake.generate() # Generate a 64-bit serial number using the snowflake algorithm, then convert it to a 16-bit string. data = f"{layer}{brand_code}{seq}" checksum = crc32(data.encode())&0xFFFFFFFF # CRC32 calculation return f"{data}{checksum:08x}" # Concatenate the checksum (8-digit hexadecimal) Step S1.3: Generate a product level identifier T for a single carton of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single carton of cigarettes to form a unique code for each carton.
[0033] Preferably, the product level identifier for a single pack of cigarettes is set to T, and the cigarette brand and place of origin code are the same as those for the single pack to ensure consistency within the same batch. The calculation method for the unique serial number and check code is the same as above and will not be repeated here.
[0034] Step S1.4: In the cigarette production process, the unique code of each carton is collected by an external scanning device, and the set of unique codes of all individual boxes inside the cigarette packaging is collected simultaneously to establish a parent-child relationship between the unique code of each carton and the set of unique codes of each box.
[0035] Preferably, in the cigarette production process, an industrial barcode scanner with a resolution of 1280×800 and a frame rate of 30fps can be deployed on the carton packaging machine. Data is transmitted via Modbus TCP protocol (port 502). The barcode scanner identifies the individual barcode printed on the surface of the carton and simultaneously triggers the built-in camera to collect the individual box codes of 10 packs inside the carton. The codes are visually identified through a transparent window or by identifying the 10 individual boxes separately before packing the carton. At the same time, the individual barcode is used as the key and the set of individual box codes (JSON array) is used as the value for association and storage.
[0036] Step S1.5: Generate a product level identifier (C) for each cigarette through the coding system architecture, and combine it with the cigarette brand origin code, unique serial number, and check code of the cigarette to form a unique code for each cigarette.
[0037] Preferably, the product level identifier for a single pack of cigarettes is set to C, and the cigarette brand and place of origin code are the same as those for the single pack to ensure consistency within the same batch. The calculation method for the unique serial number and check code is the same as above and will not be repeated here.
[0038] Step S1.6: In the cigarette production and packaging process, the unique code of each piece is collected by an external barcode scanning device, and all unique codes in the cigarette packaging box are collected in sequence according to the preset piece location number, so as to establish the parent-child relationship between the unique code of each piece and the set of unique codes.
[0039] Preferably, a vision system can be deployed at the entrance of the packing machine, with a field of view of 500×500mm and an accuracy of 0.1mm / pixel. It communicates via EtherNet / IP protocol and collects 50 individual codes sequentially according to the preset item position number, such as position 1-50, corresponding to the arrangement order of the cigarette packs. This forms an ordered set, and the individual item code is used as the key, with the ordered array of individual codes as the value, for association and storage in a MySQL database.
[0040] Step S1.7: Generate a product level identifier P for a single tray of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single tray of cigarettes to form a unique code for each tray.
[0041] Preferably, the product level identifier for a single tray of cigarettes is set to P, and the cigarette brand and place of origin code are the same as those of the single box to ensure consistency within the same batch. The calculation method for the unique serial number and check code is the same as above and will not be repeated here.
[0042] Step S1.8: In the pallet stacking process, the unique code of each pallet is collected by an external scanning device, and the unique codes of all individual cigarettes on each pallet are collected in sequence according to the preset stack position number, so as to establish the parent-child relationship between the unique code of each pallet and the set of unique codes of each individual cigarette.
[0043] Preferably, a barcode scanner can be deployed at the palletizing station, with a reading distance of 0.1-2m. The scanner transmits data via the PROFINET protocol and collects the codes of 20 individual cigarettes in sequence according to the preset pallet position number, such as positions 1-20, forming an ordered set. The barcode is then associated with the pallet code as the key and the ordered array of individual cigarette codes as the value and stored in a MySQL database.
[0044] Beneficial effects of steps S1.1 to S1.8: By defining a fixed coding architecture that includes product level identifiers, cigarette brand and origin codes, unique serial numbers, and check codes, standardized identification of each level of cigarette products is achieved, ensuring a unified and resolvable coding structure. Individual boxes, cartons, cases, and trays are assigned corresponding level identifiers B, T, C, and P, respectively, forming a hierarchical coding system to support accurate classification and traceability of subsequent data. During the production stages of cartons, cases, and trays, scanning devices collect the lower-level code sets, establishing clear parent-child relationships between upper and lower-level codes, forming a structured hierarchical relationship chain. This process does not rely on manual judgment and is complete. Driven entirely by device acquisition and encoding rules, the objectivity and consistency of the relationships are ensured. The encoding system expands progressively from the smallest to the largest unit, with each level strictly following the same architecture, guaranteeing the connectivity and scalability of cross-level data. The establishment of parent-child relationships is based on physical acquisition behavior rather than logical inference, giving the relationships a mapping basis to the real production process. The entire encoding generation and association process forms a closed and self-consistent data chain, providing reliable identity anchors for subsequent stream processing, graph construction, and trajectory tracking, without introducing additional semantics, and completing the digital representation of product entities solely through structured encoding.
[0045] Step S2 involves using industrial IoT devices to collect timestamps, status identifiers, and location coordinates of unique codes throughout the entire process from production to industrial and commercial handover, forming a raw data stream.
[0046] Further, step S2 involves using industrial IoT devices to collect timestamps, status identifiers, and location coordinates of unique codes throughout the entire process from production to industrial and commercial handover, forming a raw data stream. This specifically includes the following steps: Step S2.1: Identify the unique code and record the collection timestamp by using an external RFID reader installed at the exit of the cigarette production line, and bind it with the unique code to form the original data tuple.
[0047] Preferably, ultra-high frequency RFID readers can be installed at the exits of various nodes in the cigarette production line, such as tobacco processing, cigarette making, packaging, warehousing, and transportation. The operating frequency is 920MHz to 925MHz, and the reading distance is 0m to 8m. The readers can have an NTP client built in.
[0048] Preferably, when a carrier carrying a unique code, such as a cigarette pack, carton, box, or tray, passes through the reader's sensing area, the reader reads the code and synchronously records the UTC timestamp of the collection time. Then, it is bound to the code and the device ID to form the original data tuple.
[0049] Preferably, the tuple structure may include unique_code, timestamp_utc, and device_id, where timestamp_utc is a 13-bit integer that is sent to the message queue in real time via the Kafka Producer.
[0050] Step S2.2: Align the original data tuples with the collected events from different devices using a time series alignment algorithm to obtain time-consistent code-timestamp pairs.
[0051] Preferably, the Time Series Alignment Algorithm (NTP) can be set to a single-device time drift threshold of ≤50ms. If this threshold is exceeded, an alarm will be triggered and resynchronization will be initiated. The maximum allowable timing deviation for cross-device event acquisition is ≤100ms.
[0052] Preferably, the NTP alignment logic involves grouping the raw data tuples reported by different devices (such as RFID tags at production line exits and barcode scanners at warehouse entrances) by unique_code, and using the reference timestamp as the horizontal axis, through a linear interpolation formula t. aligned =t local +Δt (Δt is the NTP synchronization offset) corrects the local timestamp, removes data with deviations exceeding the limit, and considers data with a deviation of less than 0.1% as noise. Outputs time-consistent code-timestamp pairs (unique_code, t_aligned).
[0053] Step S2.3: Collect uniquely coded physical locations through external industrial vision sensors, combine them with the three-dimensional coordinate system model of the factory area, and convert the pixel coordinates of visual recognition into latitude and longitude values in the global geographic coordinate system through a coordinate encoding mapping algorithm to obtain the location coordinates.
[0054] Preferably, industrial vision sensors with a resolution of 2048×1536, a frame rate of 60fps, OCR recognition encoding, and an 8mm fixed-focus lens can be deployed at key locations such as packaging machines, palletizing areas, warehouse shelves, and loading and unloading points of transport vehicles. The installation height is 2.5m and the shooting angle is 45°.
[0055] Preferably, the three-dimensional coordinate system model of the factory area can be constructed based on the factory area CAD drawings (the origin O is set at the center of the main gate of the factory area, the X-axis is positive east, the Y-axis is positive north, and the Z-axis is positive upward, in meters). The coordinates of key nodes are measured by total station, for example, the coordinates of the packaging machine outlet are (X=120.5, Y=85.3, Z=0.2).
[0056] Preferably, the coordinate encoding and mapping algorithm uses a perspective transformation model to convert the pixel coordinates (u,v) of visual recognition into physical coordinates (X,Y,Z) in a three-dimensional coordinate system. The model parameters are obtained by using the Zhang Zhengyou calibration method to obtain the camera intrinsic parameter matrix. Then, the physical coordinates (X,Y) are converted into latitude and longitude (λ,φ) in a global geographic coordinate system through UTM projection, with an accuracy of ±5cm, and the output position coordinates (longitude, latitude) are given.
[0057] Step S2.4: Merge the code-timestamp pair with the location coordinates and match the status identifier corresponding to the current process node to obtain the original data stream including unique code, timestamp, location coordinates and status identifier.
[0058] Preferably, the encoding-timestamp pair (unique_code, t_aligned) from step S2.2 can be associated with the location coordinates (longitude, latitude) from step S2.3 by unique_code, and near real-time merging can be achieved through the JOIN operation of Flink SQL (time window of 5 seconds).
[0059] Preferably, the matching of status identifiers can be achieved through a predefined process node-status identifier mapping table, such as "Fiber production completed"=1, "Rolling completed"=2, "Packaging completed"=3, "Warehousing and storage"=4, "Transportation handover"=5. The current status identifier is automatically matched by the device ID (such as "LINE1_PACKING_EXIT_RFID" corresponding to "Packaging completed"). If no match is found, it is marked as 0, representing an abnormal status.
[0060] Preferably, the merged data includes unique_code, t_aligned, longitude, latitude, and status_id, and is output to the topic "raw_data_stream" in real time via Kafka Streams to form the raw data stream.
[0061] Preferably, in practical applications, barcode readers are distributed throughout the entire production process, not just at the production line exit. A large number of barcode readers are industrial cameras that only read and do not write (the codes for small boxes / bars are directly printed, and the codes for cartons are affixed online). RFID + industrial cameras are only used during palletizing. In this embodiment, it is only necessary to obtain the code at each stage, and the acquisition method is not limited.
[0062] Beneficial effects of steps S2.1 to S2.4: By using RFID readers to identify the unique codes of cigarette products at the production line exit and simultaneously recording the collection time, the initial binding of physical entities and digital identifiers is achieved, providing a raw data entry point for subsequent tracking. A time-series alignment algorithm unifies the collected events from different devices to the same time base, eliminating timing discrepancies caused by clock deviations and ensuring accurate and reliable correspondence between codes and timestamps. Industrial vision sensors, combined with a 3D coordinate system model of the factory area, convert the pixel positions obtained from image recognition into latitude and longitude values in a global geographic coordinate system, completing the digital mapping of physical spatial locations and giving product locations calculable and comparable geographic semantics. Finally, the code-timestamp pairs are merged with the location coordinates and associated with the status identifier corresponding to the current process node, forming a complete raw data stream containing four elements: code, time, location, and status. This enables structured recording of dynamic product behavior in the production process. The entire process does not rely on manual intervention; data collection and association are entirely automated by equipment and algorithms, providing a consistent, complete, and traceable input foundation for subsequent stream processing.
[0063] Step S3: Input the raw data stream into the Apache Flink-based stream processing engine. The window aggregation operator of the stream processing engine performs real-time correlation and cleaning of timestamps, status identifiers, and location coordinates according to the encoding dimensions of the raw data stream to obtain a structured event sequence.
[0064] Further, in step S3, the raw data stream is input into an Apache Flink-based stream processing engine. The window aggregation operator of the stream processing engine performs real-time correlation and cleaning of timestamps, status identifiers, and location coordinates according to the encoding dimensions of the raw data stream to obtain a structured event sequence. Specifically, this includes the following steps: Step S3.1: Partition the original data stream according to the encoded primary key, and perform stateful grouping according to the unique encoding using Apache Flink's KeyedProcessFunction.
[0065] Preferably, the partitioning logic uses the unique_code in the original data stream as the primary key, and divides the data stream into multiple logical partitions (each partition corresponds to a subset of events with a unique code) using Flink's keyBy() operator. The number of partitions is set to twice the number of CPU cores, for example, 32 partitions for a 16-core server. Stateful grouping can be implemented using Apache Flink's KeyedProcessFunction for stateful management, maintaining a ValueState state object for each partition (storing the latest event snapshot of the current code), and setting the state timeout to 24 hours (matching the production cycle).
[0066] Preferably, the pseudocode for the core logic is as follows: public class CodeGroupingFunction extends KeyedProcessFunction<String, RawData, GroupedData> { private ValueState <rawdata>latestEventState; / / Stores the latest event currently being encoded @Override public void open(Configuration parameters) { ValueStateDescriptor <rawdata>descriptor = newValueStateDescriptor<>("latestEvent", RawData.class); descriptor.enableTimeToLive(StateTtlConfig.newBuilder(Time.hours(24)).build()); / / State timeout is 24 hours latestEventState = getRuntimeContext().getState(descriptor); } @Override public void processElement(RawData event, Context ctx, Collector <groupeddata>out) throws Exception { latestEventState.update(event); / / Update the event state to the latest event. out.collect(new GroupedData(event.getUniqueCode(),latestEventState.value())); / / Output the grouped data } } Step S3.2: Synchronize the event timing of each data record after grouping with the same unique code under the same unique code through the sliding window timing alignment algorithm to obtain a timing-consistent event sequence.
[0067] Preferably, a sliding window timing alignment algorithm can be implemented based on Flink's SlidingEventTimeWindows. The window size can be set to 5 seconds to cover the single device acquisition interval, and the sliding step size is 1 second to ensure timing continuity. The watermark generation strategy is BoundedOutOfOrdernessTimestampExtractor, and out-of-order time is allowed for 5 seconds.
[0068] Preferably, for each partition (uniquely coded) within a window, the events are sorted in ascending order by the timestamp field, and then interpolated using the linear interpolation formula t. sync =t min +(i−1)×Δt (where Δt is the average event interval within the window, and i is the event number, which is not interchangeable with the symbols mentioned above) corrects the timing deviation of cross-device acquisition and outputs a timing-consistent event sequence (the number of events is ≤100 / window, and the window is split if the limit is exceeded).
[0069] Step S3.3: Input the time-consistent event sequence into the state machine-driven event merging engine, eliminate duplicate state transitions according to preset state transition rules, and obtain a simplified event stream.
[0070] Preferably, the state machine model can be set as a finite state machine (FSM), with the state set being a predefined state identifier (1-5, corresponding to the completion of yarn production to transportation handover), and the transition rule being "only the state number is allowed to increase or remain unchanged" (e.g., state 2 cannot be rolled back to 1), with the initial state being 0 (no production).
[0071] Preferably, the merging logic of the event merging engine is to input a sequence of events with consistent timing, and apply a state machine to consecutive events: if the current event state equals the previous event state, and the time interval is less than 10 seconds (a threshold to avoid misjudgment due to short pauses), then it is considered a repeated transition and discarded; otherwise, the state is updated and recorded. The core pseudocode is shown below: def merge_events(events): merged = [] prev_state, prev_time = 0, 0 for e in events: if e.status_id <prev_state or (e.status_id == prev_state ande.timestamp - prev_time<10 * 1000): continue # Repeated jumps or backs are discarded. merged.append(e) prev_state, prev_time = e.status_id, e.timestamp return merged # Output a simplified event stream Step S3.4: Input the simplified event stream into the hash deduplication engine, calculate the unique fingerprint of each event using the MurmurHash3 algorithm, remove duplicate event records and retain unique event instances to obtain the processed event sequence.
[0072] Preferably, the deduplication logic is as follows: for each event in the simplified event stream, concatenate the unique_code, timestamp, status_id, and location fields, calculate a 64-bit fingerprint using MurmurHash3, store the fingerprints that have appeared in Flink's HashSet state, remove the events corresponding to duplicate fingerprints, and retain the event instance that appears for the first time to obtain the processed event sequence.
[0073] Step S3.5: The same unique encoding of the processed event sequence is aggregated into a structured event sequence by all timestamps, status identifiers and location coordinates within the window using the window aggregation algorithm.
[0074] Preferably, the window aggregation algorithm can use a tumbling window with a window size of 10 minutes (covering the complete process of a single production stage), and aggregate after secondary grouping by unique_code.
[0075] Preferably, the aggregation rule is to take the earliest MIN (timestamp) and latest MAX (timestamp) timestamps for all events with the same code within the window, take the latest LAST (status_id) status identifier, and take the weighted average position coordinates, with the weight being the reciprocal of the event timestamp from the center of the window.
[0076] Preferably, the aggregated data includes unique_code, start_time, end_time, final_status, and agg_location, which are encapsulated into a structured event sequence in JSON format.
[0077] Beneficial effects of steps S3.1 to S3.5: The KeyedProcessFunction partitions the raw data stream by unique encoding, enabling independent state management for each encoding's corresponding event stream and ensuring scalable parallel processing capabilities. A sliding window mechanism synchronizes the timing of events collected across devices, eliminating sequence errors caused by device clock differences and maintaining logical consistency of events with the same encoding over time. A state machine-driven event merging engine identifies and removes invalid state transitions based on preset state transition rules, reducing data noise and improving the semantic accuracy of the event stream. A hash deduplication engine uses the MurmurHash3 algorithm to generate a fixed-length fingerprint for each event, quickly identifying and removing duplicate records to ensure the uniqueness of event instances. A window aggregation operator aggregates the deduplicated event stream into a structured event sequence containing complete timestamps, state identifiers, and location coordinates, completing the transformation from raw data collection to highly consistent data units. The entire process runs continuously in a streaming environment, without relying on batch processing or introducing manual intervention; it achieves precise data cleaning and structured output solely through algorithms and state mechanisms.
[0078] Step S4: Import the structured event sequence into the graph database, and use Neo4j's Cypher query language to construct multidimensional relationship edges between the encoded nodes and process nodes of the structured events, forming a full-process association graph.
[0079] Further, in step S4, the structured event sequence is imported into a graph database, and multidimensional relationship edges between the encoded nodes and process nodes of the structured events are constructed using Neo4j's Cypher query language to form a full-process association graph. This specifically includes the following steps: Step S4.1: Group the structured event sequences according to their unique codes, and use a graph entity recognition algorithm to map the codes, status identifiers, and location coordinates in each record of the structured event sequence to node attributes in the graph database to obtain a standardized node dataset.
[0080] Preferably, the grouping logic is to group events according to the unique_code field in the structured event sequence. Each group contains full-lifecycle events with the same code (such as timestamps, status identifiers, and location coordinates), and the grouping key is set to the 16-byte MD5 hash value of the unique_code.
[0081] Preferably, the graph entity recognition algorithm can adopt the Attribute-Matching Graph Entity Recognition algorithm to map the unique_code in the event to the node ID, and map the status_id (status identifier), agg_location (aggregation location coordinates), start_time / end_time (timestamp) to the node attributes; the node attribute types are node_id (string, 32-64 bytes in length), status (integer, value range 1-5), location (floating-point array, longitude [-180, 180], latitude [-90, 90]), time_range (long integer array, millisecond-level timestamp).
[0082] In step S4.2, classify the node types of the standardized node dataset, and assign four entity types of "package", "strip", "case", and "pallet" to the nodes through the encoding level identifier to form a type-annotated node set.
[0083] Preferably, the node type is assigned based on the first character of the product level identifier of unique_code, and the mapping dictionary is "B" to "package", "T" to "strip", "C" to "case", "P" to "pallet", and other characters are regarded as invalid nodes and filtered, that is, the level identifier must be a single character and within {B, T, C, P}.
[0084] Preferably, an entity_type attribute can be added to each node, with the string value of "package", "strip", "case", "pallet", to form a type-annotated node set, which is used for the direction constraint of the subsequent relationship edges. For example, "package" can only be a child node of "strip".
[0085] In step S4.3, identify the three types of parent-child association relationships of package-strip, strip-case, and case-pallet in the type-annotated node set, and generate candidate relationship triples.
[0086] Preferably, the association types are three types of parent-child relationships: package-strip (B to T), strip-case (T to C), case-pallet (C to P), and the relationship types are respectively denoted as CONTAINS_STRIP, CONTAINS_CASE, CONTAINS_PALLET.
[0087] Preferably, the generation of candidate triples can traverse the type-annotated node set. For nodes in the same production batch, if the entity_type of the parent node is "strip" and the child node is "package", or the parent is "case" and the child is "strip", or the parent is "pallet" and the child is "case", then generate candidate triples (parent_node_id, child_node_id, relation_type).
[0088] Step S4.4: The time consistency check algorithm is used to filter out the candidate relation triplet pairs whose timestamps satisfy a strict chronological order, thus obtaining the set of time-series legal association relations.
[0089] Preferably, the time consistency check algorithm can adopt the dual-timestamp consistency check algorithm, which requires that the parent node time_range[1] (end time) < the child node time_range[0] (start time) and the time difference > 1 second threshold to avoid device synchronization error.
[0090] The filtering logic involves verifying each candidate triplet, retaining triplets that meet the conditions, and removing relationships with reversed time (e.g., child event earlier than parent event) or overlapping time (e.g., child event started before parent event ended) to obtain a set of legal temporal association relationships.
[0091] Step S4.5: Merge the set of legal temporal associations with the set of type-labeled nodes, and use a graph pattern matching algorithm to identify continuous flow paths across levels to obtain complete production sequence sub-image segments.
[0092] Preferably, the set of temporally valid associations can be associated with the set of type-labeled nodes by node_id to construct a temporary graph structure of node set + edge set.
[0093] Preferably, the graph pattern matching algorithm can use depth-first search (DFS) pattern matching, defining the target path pattern as package, strip, piece, pallet (four consecutive nodes with relation types CONTAINS_STRIP, CONTAINS_CASE, and CONTAINS_PALLET respectively), traversing the temporary graph to identify all consecutive paths that match the pattern, and outputting a complete production sequence sub-image segment including node sequence and edge sequence.
[0094] Step S4.6: Merge the continuous "package-condition-component" three-node links of the complete production sequence sub-image segment into a single composite edge to form a compressed subgraph.
[0095] Preferably, in the complete production sequence sub-image segment, identify three consecutive node links (packets, strips, and components) (with relational edges CONTAINS_STRIP and CONTAINS_CASE), merge them into a single composite edge, and record the relation type as CONTAINS_CASE_VIA_STRIP (indicating that the component contains the packet through the strip). At the same time, delete the original intermediate node "strip" and two single edges to form a compressed subgraph, reducing the number of nodes by 33% and the number of edges by 50%.
[0096] Step S4.7: Write the compressed subgraph into the Neo4j graph database through the graph transaction commit protocol to obtain the immutable graph state.
[0097] Preferably, Neo4j's Bolt protocol and Graph Transaction Commit Protocol can be used to encapsulate the nodes (CREATE (:Entity {node_id, entity_type, status,location, time_range})) and edges (CREATE (p)-[:REL_TYPE]->(c)) of the compressed subgraph into transaction scripts, and write them in batches using Cypher statements.
[0098] Preferably, the immutable graph state can be written using Neo4j's read-only copy mode, which prohibits modification of historical nodes / edges after writing, and only allows the addition of new versions. At the same time, the version attribute is used to mark the initial version v1 to ensure data atomicity and traceability.
[0099] Step S4.8: Generate a global inverted index based on unique encoding from the immutable graph state using the graph index building engine to obtain the full-process association graph.
[0100] Preferably, the indexing engine can be an Elasticsearch graph indexing engine, using `node_id` as the keyword and node attributes (`entity_type`, `time_range`) as auxiliary fields to build a global inverted index. The index structure is `{keyword: node_id, attributes: {type, start_time, end_time}}`. The compressed subgraph is then bound to the inverted index, and the index is registered using the Neo4j `apoc.index` plugin. This supports fast retrieval of nodes and relationships by code, ultimately resulting in a complete relational graph, including a three-layer structure of nodes, edges, and indexes.
[0101] Beneficial effects of steps S4.1 to S4.8: Events are grouped by unique encoding, and a standardized node dataset is established by mapping encoding, status identifiers, and location coordinates as node attributes using a graph entity recognition algorithm, thus unifying data representation. Nodes are assigned four entity types—package, bar, item, and tray—based on their encoding hierarchy, forming a type-labeled node set and clarifying node identity boundaries. Package-bar, bar-item, and item-tray parent-child relationships are identified, generating candidate relationship triples and extracting potential flow relationships. A time consistency verification algorithm is used to filter candidate relationships for pairs whose timestamps conform to the synchronization order, resulting in a set of time-series valid association relationships, ensuring the logical correctness of the relationships. These relationships are then merged. The relationship set and type-labeled node set are used to identify continuous flow paths across levels using a graph pattern matching algorithm, forming complete production sequence sub-image segments and outlining the basic flow context. Continuous package-condition-component three-node links in the sub-graph are merged into a single composite edge, compressing the sub-graph structure and simplifying relationship expression. The compressed sub-graph is written to Neo4j through a graph transaction commit protocol, forming an immutable graph state and ensuring the atomicity and consistency of data storage. A graph index building engine generates a global inverted index based on unique encoding, ultimately obtaining a full-process association graph, providing a structured foundation for subsequent path traversal and relationship queries.
[0102] Step S5: Use dynamic programming algorithm to traverse the shortest paths between the coding nodes and process nodes of the entire process association graph to generate a cigarette product production flow trajectory chain from the perspective of any coding.
[0103] Further, in step S5, the shortest paths between the encoded nodes and process nodes of the entire process association graph are traversed using a dynamic programming algorithm to generate a cigarette product production flow trajectory chain from the perspective of any encoding. This specifically includes the following steps: Step S5.1: Select any coded node in the entire process association graph as the starting node through the graph traversal engine, and extract all outgoing edges of the starting node to form an initial path candidate set.
[0104] Preferably, the graph traversal engine can use Neo4j's Traversal Framework, which implements node access logic based on the Java API.
[0105] Preferably, the starting node can be selected by receiving a unique code from an external input and querying the corresponding coded node in the location map using Cypher (MATCH (n:Entity {node_id: $code}) RETURN n) to use it as the starting node.
[0106] Preferably, the initial path candidate set is generated by extracting all outgoing edges of the starting node (relationship type is CONTAINS_STRIP / CONTAINS_CASE / CONTAINS_PALLET / CONTAINS_CASE_VIA_STRIP). Each edge contains the target node ID and the relationship timestamp (take the time_range[0] when the edge is created) to form the initial path candidate set. The data structure is [(edge_id, target_node_id, rel_timestamp), ...].
[0107] Step S5.2: Sort each outgoing edge in the initial path candidate set in ascending order according to the timestamp field, following the causal order of time.
[0108] Preferably, the ascending order that follows the temporal causal order can be achieved by using a merge sort to sort the initial path candidate set in ascending order by the rel_timestamp field, thus ensuring the temporal causal order.
[0109] Step S5.3: Take the node pointed to by the first outgoing edge after sorting as the current traversal node, recursively expand all outgoing edges of the current traversal node using the depth-first search algorithm, and retain the edges whose timestamps are greater than the timestamp of the current node.
[0110] Preferably, the Depth-First Search (DFS) algorithm recursively traverses the outgoing edges of a node, with the constraint that "the next node's timestamp > the current node's timestamp". The recursive logic is as follows: Input: current node current_node, current path path (including node ID sequence), current timestamp current_ts; traverse all outgoing edges of current_node, filter edges where rel_timestamp > current_ts (time difference ≥ 1ms); sort the filtered edges according to the rules in step S5.2, set the target node as the new current_node, update current_ts to the edge timestamp, and recursively expand; the core pseudocode is shown below: def dfs(current_node, path, current_ts, graph): edges = [e for e in graph.out_edges(current_node) if e.rel_timestamp>current_ts] edges_sorted = merge_sort(edges, key=lambda x: x.rel_timestamp) # Merge sort for edge in edges_sorted: new_path = path + [edge.target_node_id] dfs(edge.target_node, new_path, edge.rel_timestamp, graph) # Recursive expansion Step S5.4: During the recursive traversal, perform a path concatenation operation on each selected edge, concatenating the current node code and the next node code in timestamp order to form an intermediate trajectory segment.
[0111] Preferably, during the recursive traversal, a concatenation operation is performed on the selected edges, concatenating the current node code and the next node code in timestamp order, which can be connected using the separator "→" to form intermediate trajectory segments. The segment structure is [start_code,(next_code, ts1), (next_code2, ts2), ...], where ts is the edge timestamp in milliseconds.
[0112] Step S5.5: When the depth-first search reaches a leaf node with no outgoing edges or whose outgoing edge timestamps do not meet the increasing condition, terminate the current branch traversal and output the complete trajectory segment of the current branch.
[0113] Preferably, the current branch terminates when DFS reaches the following nodes: ① No outgoing edges, leaf nodes, such as pallet nodes, which have no subsequent packaging process.
[0114] ② If the timestamps of all outgoing edges are less than or equal to the timestamp of the current node, i.e. the times are reversed, they are considered abnormal branches.
[0115] Preferably, the path at termination (node ID sequence + corresponding timestamp) is encapsulated into a complete trajectory fragment and stored as a triple (start_code, fragment_id, trajectory_fragment). The fragment ID is generated by the starting code + recursion depth, for example, Bxxx_depth3, where B is a single pack of cigarettes.
[0116] Step S5.6: Aggregate the trajectory segments of all terminating branches according to the uniqueness of the starting encoding, eliminate duplicate node sequences through the path merging algorithm, and retain the unique time series continuous trajectory chain.
[0117] Preferably, the path merging algorithm can be a timestampoverlap-based path merging algorithm. First, the longest common subsequence (LCS) of the node sequences between segments is calculated, and the similarity threshold can be set to 90%. Then, similar segments are merged, and unique sequences with consecutive timestamps are retained to remove duplicate nodes. If the same code appears multiple times in a segment, the earliest timestamp is taken.
[0118] Preferably, the merged trajectory chain is a unique time series continuous chain with the structure (start_code, merged_trajectory).
[0119] Step S5.7: Complete the missing node timestamps in the trajectory chain by linear interpolation to obtain a trajectory chain without time breakpoints.
[0120] Preferably, linear interpolation is used to fill in the intermediate timestamps when there are gaps (interval > 5 minutes, threshold) between adjacent node timestamps in the trajectory chain. It is worth noting that linear interpolation is a mature existing technology, and the formula and calculation process of linear interpolation will not be described in detail in this embodiment.
[0121] Step S5.8: Serialize the time-breakpoint-free trajectory chain into a linear string format according to the encoded primary key to obtain a standardized cigarette product production flow trajectory chain.
[0122] Preferably, the node IDs of the trajectory chain without time breakpoints can be arranged in ascending order of timestamps, encoded with commas as separators, and enclosed in square brackets to form a linear string, such as [Bxxx,Tyyy,Czzz,Pwww].
[0123] Beneficial effects of steps S5.1 to S5.8: By selecting any coded node in the graph as the starting node, all its outgoing edges are extracted to form an initial path candidate set, establishing the traversal starting point. The outgoing edges of the candidate set are sorted in ascending order by timestamp, following the causal order of time to ensure the reasonable temporal sequence of the paths. The node pointed to by the first outgoing edge after sorting is taken as the current traversal node, and the outgoing edges are recursively expanded using depth-first search, retaining only edges with timestamps greater than the current node, focusing on effective flow relationships. Path splicing is performed in the recursion, and the node codes are concatenated according to timestamps to form intermediate trajectory fragments, accumulating flow information. When DFS reaches a leaf node with no outgoing edges or whose outgoing edge timestamps are not increasing, the branch terminates, and the complete trajectory fragment of that branch is output. All terminated branch fragments are aggregated, and the path merging algorithm is used to eliminate duplicate node sequences according to the starting code, retaining a unique, continuous trajectory chain with a unique time sequence. The missing node timestamps of the trajectory chain are completed by linear interpolation to eliminate time breakpoints. The trajectory chain without breakpoints is serialized into a linear string according to the encoded primary key to form a standardized flow trajectory chain.
[0124] Step S6: Map the cigarette product production flow trajectory chain to the visualization engine, and use the force-directed graph layout algorithm to spatially arrange the nodes and edges in the cigarette product production flow trajectory chain to obtain an interactive production flow trajectory diagram.
[0125] Further, in step S6, the cigarette product production flow trajectory chain is mapped to the visualization engine, and the nodes and edges in the cigarette product production flow trajectory chain are spatially arranged using a force-directed graph layout algorithm to obtain an interactive production flow trajectory diagram. This specifically includes the following steps: Step S6.1: Extract all node identifiers and edge connections in the cigarette product production flow trajectory chain using a graph structure parsing engine to obtain a standardized list of nodes and edges.
[0126] Preferably, the graph structure parsing engine can use D3.js Graph Parser as the graph structure parsing engine, and it supports parsing the linear strings [Bxxx,Tyyy,Czzz,Pwww] output by S5.8.
[0127] Preferably, the extraction logic involves capturing the node ID sequence using the regular expression \[(.*?)\], splitting it by commas to obtain a node list nodes = [id1, id2, ..., idn]; the edge list edges consists of adjacent node pairs, i.e., (id1, id2), (id2, id3), ..., (id(n-1), idn), with the relationship type being "flow order". The output node list includes node_id (a string) and type (inherited from entity_type in S4.2, such as "package" or "item"); the edge list includes source_id, target_id, and relation (a fixed value "flow"), forming a standardized structure {nodes: [...], edges: [...]}.
[0128] Step S6.2: Assign initial two-dimensional plane coordinates to each node in the node list using a uniform distribution algorithm that does not overlap and maintains topological scalability.
[0129] Preferably, the uniform distribution algorithm can employ the Spiral Uniform Distribution Algorithm to avoid initial node overlap and maintain topology scalability.
[0130] The canvas size is set to 1920×1080 pixels, the minimum node spacing is 50 pixels, the starting radius of the spiral parameters is r=100px, the angular velocity is ω=0.5rad / node, and the radial increment is Δr=30px / turn.
[0131] Preferably, the coordinates of the i-th node (x) i ,y i Formula: x i =r i cos(θ i ), y i =r i sin(θ i ).
[0132] Where r i =100+30×⌊i / 20⌋, meaning an additional cycle is added every 20 nodes, θ i =i×0.5 (radians) to ensure that the nodes are evenly distributed along the Archimedean spiral.
[0133] Step S6.3: Input the edge list and the initial two-dimensional plane coordinates into the force-guided graph layout algorithm, drive the interaction between each node with Coulomb force and Hooke force, and iteratively calculate the displacement vector of each node.
[0134] Preferably, the force-directed graph layout algorithm can adopt the Fruchterman-Reingold force-directed layout algorithm, which uses Coulomb repulsion and Hooke attraction between edge-connecting nodes to drive the layout. The repulsion coefficient can be set to 200, the attraction coefficient to 0.1, the number of iterations to 500, and the convergence threshold to 0.01 (stopping when the total node displacement is < 0.01px). The temperature parameter is initially 1000px and linearly decays to 0 with each iteration.
[0135] Preferably, the iterative logic is represented by the following core pseudocode: function fruchtermanReingold(nodes, edges, iterations=500) { let temp = 1000; for (let iter=0; iter <iterations; iter++) { nodes.forEach(node => calculateForce(node, nodes, edges, k_r, k_a)); / / Calculate the resultant force nodes.forEach(node => updatePosition(node, temp)); / / Update position (temperature-constrained) temp *= 0.95; / / Cool down } } Step S6.4: In each iteration, the displacement vectors of all nodes are normalized using the Euclidean distance normalization algorithm to constrain the displacement amplitude within a preset step size range.
[0136] Preferably, the preset step size can be set to 5px, which is the maximum displacement in a single iteration.
[0137] Preferably, for the nodal displacement vector (dx, dy), the Euclidean distance d is calculated to be equal to the square root of (dx, dy). 2 +dy 2 If d is greater than step, then the normalized displacement (dx', dy') = (dx / d*step, dy / d*step); otherwise, the original displacement is retained.
[0138] Step S6.5: Merge the updated node coordinates with the edge list to obtain the node coordinate sequence.
[0139] Preferably, the merging logic is to associate the node coordinates (including node_id, x, y) after iteration in step S6.4 with the edge list in step S6.1 according to node_id to form a node coordinate sequence.
[0140] Preferably, the structure of the node coordinate sequence is: {nodes: [{id, x, y, type}, ...], edges:[{source, target}, ...]}, where x and y are the final layout coordinates, floating-point type, with a precision of 0.1px.
[0141] Step S6.6: Perform spatial clustering compression on the node coordinate sequence, and input the clustered node coordinates and edge relationships into the graph serialization engine, outputting a structured layout description file including node ID, coordinates, and edge connection relationships in a fixed format.
[0142] Preferably, spatial clustering compression can be achieved using the DBSCAN density clustering algorithm, where the parameters eps=30px (neighborhood radius), minPts=2 (minimum number of points), and dense nodes with a distance <30px are merged, such as multiple package nodes in the same process.
[0143] Preferably, the graphics serialization engine can define the layout description file format through Protocol Buffers (proto3 syntax), with fields including node_id (string), x / y (float), type (enum), and edge_source / edge_target (string).
[0144] Step S6.7: Input the layout description file into the visualization rendering engine through the API interface, and dynamically bind it with the interactive production workflow map to obtain the interactive production workflow map.
[0145] Preferably, the API interface can be a RESTful API (POST / api / v1 / trajectory / render), with the request body being the layout description file from step S6.6, and the response returning the rendering session ID.
[0146] Preferably, the visualization rendering engine can be ECharts GL, which supports large-scale node rendering accelerated by WebGL.
[0147] Preferably, dynamic binding can be achieved by parsing the layout description file through the engine, drawing nodes (circles, which can be colored according to type: package-blue, bar-green, item-orange, tray-gray) and edges (gray solid lines) in the form of a force-directed graph, binding interactive events (clicking a node displays node_id and time_range, dragging to pan the canvas, and scrolling with the wheel), and finally obtaining an interactive production flow trajectory diagram.
[0148] Beneficial effects of steps S6.1 to S6.7: By extracting node identifiers and edge connections in the trajectory chain using a graph structure analysis engine, a standardized node list and edge list are obtained, clarifying the basic structure of the visualized object. A uniform distribution algorithm is used to assign non-overlapping initial two-dimensional plane coordinates to the node list while maintaining topological scalability, laying the foundation for the layout. The edge list and initial coordinates are input into a force-guided graph layout algorithm, using Coulomb forces and Hooke forces to drive node interactions, iteratively calculating the displacement vector of each node to optimize the spatial distribution. In each iteration, the displacement vector is normalized using an Euclidean distance normalization algorithm, constraining the displacement amplitude within a preset step size range to avoid layout oscillations. The updated node coordinates and edge lists are merged to obtain a node coordinate sequence. This sequence is spatially clustered and compressed, and then input into a graph serialization engine to output a structured layout description file containing node IDs, coordinates, and edge connections in a fixed format. This file is then input into a visualization rendering engine via an API interface and dynamically bound to an interactive production flow trajectory map, ultimately achieving an intuitive spatial arrangement of trajectory chain nodes and edges.
[0149] Beneficial effects of steps S1 to S6: By using predefined coding rules to uniquely identify each level of cigarette products, a structured identity binding is achieved from individual packs to trays, laying the foundation for subsequent data association. Industrial IoT devices continuously collect the timestamps, status identifiers, and location coordinates corresponding to the codes, forming a continuous and traceable raw data stream to ensure that the physical behavior of the entire production process is digitally recorded. The stream processing engine is based on Apache. Flink performs real-time grouping, time-series alignment, and state merging on the raw data stream to eliminate redundancy and jumps, outputting a highly consistent structured event sequence. The graph database uses the Cypher language to transform the event sequence into a multi-dimensional relationship graph of coded nodes and process nodes, constructing a semantically complete full-process topology. The dynamic programming algorithm traverses this graph, extracting the complete flow path of any coded element according to the time sequence, generating a production trajectory chain without breaks or redundancy. Finally, the force-directed graph layout algorithm spatially rearranges the nodes and edges in the trajectory chain, achieving a natural balance in node distribution through physical simulation, outputting an interactive visualization, presenting complex production relationships in an intuitive form, and improving information reading efficiency. The entire process achieves end-to-end mapping from physical entities to digital trajectories, without relying on human intervention, and is completely driven by data.
[0150] like Figure 2 As shown, this embodiment provides an example of a tracking system for cigarette product production information. In this embodiment, the tracking system is applied to the tracking method described in the above embodiment.
[0151] Specifically, the tracking system includes a cigarette product coding generation module 1, a cigarette product flow detection module 2, a cigarette product flow data preprocessing module 3, a full-process correlation map generation module 4, a production flow trajectory chain generation module 5, and a production flow trajectory chain visualization module 6, which are connected electrically or by signal in sequence.
[0152] The cigarette product coding generation module 1 generates unique codes for each pack, carton, case, and pallet of cigarettes based on preset coding rules. The cigarette product flow detection module 2 uses industrial IoT devices to collect timestamps, status indicators, and location coordinates of these unique codes throughout the entire process from production to industrial and commercial handover, forming a raw data stream. The cigarette product flow data preprocessing module 3 inputs the raw data stream into an Apache Flink-based stream processing engine, using the engine's window aggregation operator to perform real-time association and cleaning of timestamps, status indicators, and location coordinates according to the coding dimensions of the raw data stream, resulting in a structured event sequence. This process is then used for full-process association. The graph generation module 4 imports structured event sequences into a graph database and uses Neo4j's Cypher query language to construct multidimensional relationship edges between the encoded nodes and process nodes of the structured events, forming a full-process association graph. The production flow trajectory chain generation module 5 uses a dynamic programming algorithm to traverse the shortest paths between the encoded nodes and process nodes of the full-process association graph, generating a cigarette product production flow trajectory chain from any encoding perspective. The production flow trajectory chain visualization module 6 maps the cigarette product production flow trajectory chain to a visualization engine and uses a force-directed graph layout algorithm to spatially arrange the nodes and edges in the cigarette product production flow trajectory chain, obtaining an interactive production flow trajectory diagram.
[0153] It should be noted that this embodiment is a functional module embodiment based on the above method embodiment. For additional content such as extensions, optimizations, limitations, examples, principle explanations, and beneficial effects of this embodiment, please refer to the above embodiments. This embodiment will not repeat them here.
[0154] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Figure 3 As shown, the electronic device 7 includes a processor 71 and a memory 72 coupled to the processor 71.
[0155] The memory 72 stores program instructions for implementing the federated learning-based collaborative energy-saving method for government data clusters in any of the above embodiments.
[0156] The processor 71 is used to execute program instructions stored in the memory 72 for collaborative energy saving of government data clusters based on federated learning.
[0157] The processor 71 can also be referred to as a CPU (Central Processing Unit). The processor 71 may be an integrated circuit chip with signal processing capabilities. The processor 71 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor.
[0158] Furthermore, Figure 4 This is a schematic diagram of the structure of a storage medium according to an embodiment of this application. See also: Figure 4 In this embodiment of the application, the storage medium 8 stores program instructions 81 capable of implementing all the above methods. These program instructions 81 can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets.
[0159] In the several embodiments provided in this application, it should be understood that the disclosed systems, methods, and approaches can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, signal, or other forms.
[0160] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.< / groupeddata> < / rawdata> < / rawdata>
Claims
1. A method for tracking cigarette product production information, characterized in that, The tracking method includes: Step S1: Generate a unique code for each pack, carton, case, and tray of cigarette products based on preset coding rules. Step S2: Collect the timestamps, status identifiers, and location coordinates of the unique code in real time throughout the entire process from production to industrial and commercial handover using industrial IoT devices to form the raw data stream; Step S3: Input the original data stream into the Apache Flink-based stream processing engine. The window aggregation operator of the stream processing engine performs real-time association and cleaning of timestamps, status identifiers, and location coordinates according to the encoding dimension of the original data stream to obtain a structured event sequence. Step S4: Import the structured event sequence into a graph database, and construct multidimensional relationship edges between the encoded nodes and process nodes of the structured events using Neo4j's Cypher query language to form a full-process association graph; Step S5: Using a dynamic programming algorithm, traverse the shortest paths between the coded nodes and process nodes of the entire process association graph to generate a cigarette product production flow trajectory chain from the perspective of any code. Step S6: Map the cigarette product production flow trajectory chain to the visualization engine, and use the force-directed graph layout algorithm to spatially arrange the nodes and edges in the cigarette product production flow trajectory chain to obtain an interactive production flow trajectory graph.
2. The tracking method according to claim 1, characterized in that, Step S1: Generate a unique code for each pack, carton, case, and tray of cigarettes based on preset coding rules, including: Step S1.1: Define a coding system architecture consisting of four fixed fields concatenated in sequence: product level identifier, cigarette brand origin code, unique serial number, and check code; Step S1.2: Generate a product level identifier B for a single pack of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single pack of cigarettes to form a unique code for the single pack; Step S1.3: Generate a product level identifier T for a single carton of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single carton of cigarettes to form a unique code for each carton; Step S1.4: In the cigarette production process, the unique code of each carton is collected by an external scanning device, and the set of unique codes of all individual boxes in the cigarette packaging is collected simultaneously to establish a parent-child relationship between the unique code of each carton and the set of unique codes of each individual box. Step S1.5: Generate a product level identifier C for a single cigarette through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single cigarette to form a unique code for each cigarette. Step S1.6: In the cigarette production and packaging process, the unique code of each piece is collected by an external scanning device, and all unique codes in the cigarette packaging box are collected in sequence according to the preset piece position number order, so as to establish the parent-child relationship between the unique code of each piece and the set of unique codes. Step S1.7: Generate a product level identifier P for a single tray of cigarettes through the coding system architecture, and combine the product level identifier with the cigarette brand origin code, unique serial number, and check code of the single tray of cigarettes to form a unique code for each tray. Step S1.8: In the pallet stacking process, the unique code of each pallet is collected by an external scanning device, and all the unique codes of each item on the pallet of cigarettes are collected in sequence according to the preset stack position number order, so as to establish the parent-child relationship between the unique code of each pallet and the set of unique codes of each item.
3. The tracking method according to claim 1, characterized in that, Step S2: Real-time collection of timestamps, status identifiers, and location coordinates of the unique code throughout the entire process from production to industrial and commercial handover using industrial IoT devices, forming a raw data stream, including: Step S2.1: Identify the unique code and record the collection timestamp by an external RFID reader installed at the exit of the cigarette production line, and bind it with the unique code to form an original data tuple; Step S2.2: Align the original data tuples with the acquisition events from different devices using a time series alignment algorithm to obtain time-consistent code-timestamp pairs. Step S2.3: Collect the uniquely encoded physical location through an external industrial vision sensor, combine it with the three-dimensional coordinate system model of the factory area, and convert the pixel coordinates of the visual recognition into latitude and longitude values in the global geographic coordinate system through a coordinate encoding mapping algorithm to obtain the location coordinates; Step S2.4: Merge the code-timestamp pair with the location coordinates and match the status identifier corresponding to the current process node to obtain the original data stream including the unique code, the timestamp, the location coordinates, and the status identifier.
4. The tracking method according to claim 1, characterized in that, Step S3: Input the raw data stream into an Apache Flink-based stream processing engine. The engine's window aggregation operator performs real-time correlation and cleaning of timestamps, status identifiers, and location coordinates according to the encoding dimensions of the raw data stream, resulting in a structured event sequence, including: Step S3.1: Partition the original data stream according to the encoded primary key, and perform stateful grouping according to the unique encoding using Apache Flink's KeyedProcessFunction; Step S3.2: Synchronize the event time sequence of each data record after grouping with the same unique code across devices using the sliding window time sequence alignment algorithm to obtain a time sequence consistent event sequence; Step S3.3: Input the time-consistent event sequence into the state machine-driven event merging engine, eliminate repeated state transitions according to preset state transition rules, and obtain a simplified event stream; Step S3.4: Input the simplified event stream into the hash deduplication engine, calculate the unique fingerprint of each event using the MurmurHash3 algorithm, remove duplicate event records and retain unique event instances to obtain the processed event sequence; Step S3.5: The same unique encoding of the processed event sequence is used to aggregate all timestamps, status identifiers and location coordinates within the window into a structured event sequence through the window aggregation algorithm.
5. The tracking method according to claim 1, characterized in that, Step S4: Import the structured event sequence into a graph database, and construct multidimensional relationship edges between the encoded nodes and process nodes of the structured events using Neo4j's Cypher query language to form a full-process association graph, including: Step S4.1: Group the structured event sequence according to the unique code, and map the code, status identifier, and location coordinates in each record of the structured event sequence to node attributes in the graph database using a graph entity recognition algorithm to obtain a standardized node dataset; Step S4.2: Classify the standardized node dataset by node type, and assign four entity types to the nodes: "package", "item", "piece" and "tray" through encoding hierarchy identifiers, forming a type-labeled node set; Step S4.3: Identify the three types of parent-child relationships in the set of labeled nodes: package-strip, strip-item, and item-pallet, and generate candidate relationship triples; Step S4.4: The time consistency verification algorithm is used to filter out the candidate relation triplets and the relation pairs whose timestamps satisfy a strict chronological order, so as to obtain the set of time-series legal association relations; Step S4.5: Merge the set of temporal legal associations with the set of type-labeled nodes, and identify the continuous flow path across levels using a graph pattern matching algorithm to obtain the complete production sequence sub-image segment; Step S4.6: Merge the continuous "package-condition-component" three-node links of the complete production sequence sub-image segment into a single composite edge to form a compressed sub-graph; Step S4.7: Write the compressed subgraph into the Neo4j graph database through the graph transaction commit protocol to obtain the immutable graph state; Step S4.8: The immutable graph state is used to generate a global inverted index based on unique encoding through a graph index building engine to obtain the full-process association graph.
6. The tracking method according to claim 1, characterized in that, Step S5: Using a dynamic programming algorithm, traverse the shortest paths between the encoded nodes and process nodes of the entire process association graph to generate a cigarette product production flow trajectory chain from the perspective of any encoding, including: Step S5.1: Select any coded node in the full-process association graph as the starting node through the graph traversal engine, and extract all outgoing edges of the starting node to form an initial path candidate set; Step S5.2: Sort each outgoing edge in the initial path candidate set in ascending order according to the timestamp field, following the causal order of time. Step S5.3: Take the node pointed to by the first outgoing edge after sorting as the current traversal node, recursively expand all outgoing edges of the current traversal node using the depth-first search algorithm, and retain the edges whose timestamps are greater than the timestamp of the current node. Step S5.4: During the recursive traversal, perform a path splicing operation on each selected edge, concatenating the current node code and the next node code in timestamp order to form an intermediate trajectory segment; Step S5.5: When the depth-first search reaches a leaf node with no outgoing edges or whose outgoing edge timestamps do not meet the increasing condition, terminate the current branch traversal and output the complete trajectory fragment of the current branch. Step S5.6: Aggregate the trajectory segments of all terminating branches according to the uniqueness of the starting encoding, eliminate duplicate node sequences through the path merging algorithm, and retain the unique time series continuous trajectory chain; Step S5.7: Complete the missing node timestamps in the trajectory chain by linear interpolation to obtain a trajectory chain without time breakpoints; Step S5.8: The time-breakpoint-free trajectory chain is serialized into a linear string format according to the encoded primary key to obtain a standardized cigarette product production flow trajectory chain.
7. The tracking method according to claim 1, characterized in that, Step S6: Map the cigarette product production flow trajectory chain to the visualization engine, and use a force-directed graph layout algorithm to spatially arrange the nodes and edges in the cigarette product production flow trajectory chain to obtain an interactive production flow trajectory diagram, including: Step S6.1: Extract all node identifiers and edge connections in the cigarette product production flow trajectory chain using a graph structure parsing engine to obtain a standardized list of nodes and an edge list; Step S6.2: Assign non-overlapping initial two-dimensional plane coordinates to each node in the node list using a uniform distribution algorithm while maintaining topological scalability; Step S6.3: Input the edge list with the initial two-dimensional plane coordinates into the force-guided graph layout algorithm, drive the interaction between each node with Coulomb force and Hooke force, and iteratively calculate the displacement vector of each node; Step S6.4: In each iteration, the displacement vectors of all nodes are normalized using the Euclidean distance normalization algorithm to constrain the displacement amplitude within a preset step size range. Step S6.5: Merge the updated node coordinates with the edge list to obtain a node coordinate sequence; Step S6.6: Spatial clustering compression is performed on the node coordinate sequence, and the clustered node coordinates and edge relationships are input into the graphics serialization engine. A structured layout description file including node ID, coordinates, and edge connection relationships is output in a fixed format. Step S6.7: Input the layout description file into the visualization rendering engine through the API interface, and dynamically bind it with the interactive production flow trajectory diagram to obtain the interactive production flow trajectory diagram.
8. A tracking system for cigarette product production information, wherein the tracking system is applied to the tracking method as described in any one of claims 1 to 7, characterized in that, The tracking system includes: The cigarette product coding generation module is used to generate unique codes for each pack, each carton, each case, and each tray of cigarette products based on preset coding rules. The cigarette product flow detection module is used to collect the timestamps, status identifiers, and location coordinates of the unique code in real time from production to industrial and commercial handover through industrial IoT devices, forming a raw data stream. The cigarette product flow data preprocessing module is used to input the raw data stream into an Apache Flink-based stream processing engine. The window aggregation operator of the stream processing engine performs real-time association and cleaning of timestamps, status identifiers, and location coordinates according to the encoding dimension of the raw data stream to obtain a structured event sequence. The full-process association graph generation module is used to import the structured event sequence into a graph database, and construct multi-dimensional relationship edges between the encoded nodes and process nodes of the structured events using Neo4j's Cypher query language to form a full-process association graph. The production flow trajectory chain generation module is used to traverse the shortest path between the coded nodes and process nodes of the full-process association graph using a dynamic programming algorithm to generate a cigarette product production flow trajectory chain from the perspective of any code. The production flow trajectory visualization module is used to map the cigarette product production flow trajectory to the visualization engine. The force-directed graph layout algorithm is used to spatially arrange the nodes and edges in the cigarette product production flow trajectory to obtain an interactive production flow trajectory graph.
9. An electronic device, characterized in that, The method includes a processor and a memory coupled to the processor, the memory storing program instructions executable by the processor; when the processor executes the program instructions stored in the memory, it implements the tracing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when executed by a processor, enable the tracing method as described in any one of claims 1 to 7.