Language model multi-bit embedding-based imperceptible layered watermark embedding method
By employing a language model-based multi-bit embedding method and utilizing instruction logic trees and dynamic hash values to calculate dynamic sampling offset operators, the problem of logical integrity and traceability reliability in watermark embedding in industrial instruction sequences is solved, achieving a combination of high-reliability traceability and parameter accuracy in short-sequence text.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI PUDONG CRYPTOZOOLOGY RESEARCH INSTITUTE
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies struggle to embed highly reliable, imperceptible, and layered traceability-supporting multi-bit watermarks into industrial instruction sequences while maintaining logical integrity. This leads to failures in retrieval of traceability information or logical collapse, impacting the engineering usability of industrial production.
By employing a language model-based multi-bit embedding method, the dynamic sampling offset operator is calculated using the hierarchical dependency relationship of the instruction logic tree and dynamic hash value. This constructs a probabilistic mapping relationship between watermark bits and key instruction fields, performs dimension remapping correction and sampling space partitioning modulation, and ensures that the watermark information is embedded within a high-entropy semantic range, thus avoiding interference with numerical parameters.
It enables highly reliable traceability in short-sequence text, ensuring the logical integrity and parameter accuracy of industrial instructions, and improving the attack resistance and traceability of industrial digital assets.
Smart Images

Figure CN121997303A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of industrial data processing technology, and in particular relates to an imperceptible layered watermark embedding method based on language model multi-bit embedding. Background Technology
[0002] With the increasing penetration of large language models into industrial data processing, the use of generative artificial intelligence to generate industrial design scripts, processing instructions, and process logic descriptions has become a mainstream trend. To ensure the security of industrial design assets, embedding invisible watermarks in the generated text to achieve data ownership confirmation and traceability is the core means of current industrial data security processing. Industrial instruction data has high logical atomicity and parameter sensitivity. Existing multi-bit watermark embedding schemes usually adopt a probability offset mechanism based on statistical distribution, modulating the log probability during the vocabulary sampling stage. However, such schemes have a conflict between traceability information capacity and data logic precision in actual industrial scenarios. Industrial instruction sequences are often short and have a high repetition rate. Traditional statistical schemes cannot accumulate sufficient statistical significance within a very short character window, resulting in the failure of traceability information extraction. At the same time, in order to embed multi-bit identifiers within a limited space, existing probability modulation strategies do not consider the logical constraints within industrial parameters, causing offsets in numerical bits or structural definitions in the instruction script.
[0003] To address the difficulty of tracing the source of short-sequence texts, conventional improvement paths typically enhance signal strength by increasing the probability offset amplitude. However, logical deduction shows that simple parameter enhancement can induce distortions in text expression, leading to logical collapse in subsequent simulation execution of industrial modeling scripts or processing instructions, thus losing their engineering usability as input sources for industrial production. Furthermore, if the control logic of watermark embedding algorithms lacks coupling with industrial domain knowledge, it can cause engineering usability issues. For example, Chinese invention patent CN119939544B discloses a method for generating content detection based on sentence semantic injection watermarks using a large language model. While this method utilizes sentence semantic features to establish a marker mapping to improve anti-tampering capabilities, it suffers from adaptation defects in industrial data processing: the semantic encoding mechanism identifies macro-level semantic tendencies but ignores the hierarchical dependencies of industrial instruction logic tree nodes, resulting in a disconnect between the watermark signal distribution and the design skeleton, leading to low source tracing confidence when processing short-sequence instructions; when selecting sentences for watermark injection, the lack of geometric constraint numerical nonlinearity avoidance mechanisms and the pursuit of semantic marker uniqueness cause disturbances in key dimension parameters or rotation angles. Conflicts between the semantic space and parameter space restrict the injection capacity of multi-bit source tracing information, threatening the accuracy of CAD modeling and CNC machining instructions.
[0004] Therefore, the technical problem to be solved by this invention is how to embed highly reliable, imperceptible, and hierarchical traceability identifiers into short-sequence text while maintaining the logical integrity of industrial instructions, so that the generated industrial digital assets have source traceability capabilities under different length conditions. Summary of the Invention
[0005] This invention provides an imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model, comprising the following steps: Step S1: Obtain the industrial design script sequence to be processed, the candidate tag list corresponding to the industrial design script sequence to be processed, and the instruction logic tree corresponding to the industrial design script sequence to be processed. The instruction logic tree includes a global reference node, a geometric entity definition node, a size association constraint node, and a manufacturing parameter instruction node. Step S2: Based on the hierarchical dependency relationship of each node in the instruction logic tree, determine the logical topology depth value of the candidate label to be generated relative to the global baseline node in the instruction logic tree. The logical topology depth value is used to characterize the logical hierarchy position of the current instruction in the parameterized modeling process. Step S3: Using the preset private key and the dynamic hash value generated by the industrial design script sequence to be processed, combined with the logical topology depth value, calculate the dynamic sampling offset operator to establish the probabilistic mapping relationship between the watermark bits and the key instruction fields. Step S4: Before performing sampling space modulation, search the high-entropy semantic interval of the candidate tag list to see if there are candidate tags containing geometric constraint values; if there are candidate tags containing geometric constraint values, use the dynamic sampling offset operator to perform dimension remapping correction on the probability distribution vector of the candidate tag list, and map the watermark information to the instruction field corresponding to the geometric constraint values, so as to limit the probability distribution offset of the candidate tag list to within the preset numerical fluctuation range. Step S5: Based on the value of the dynamic sampling offset operator, perform sampling space partitioning modulation on the high-entropy semantic interval of the candidate tag list, so that the generated candidate tags carry a multi-bit superimposed watermark containing model source information and user traceability information.
[0006] Preferably, step S2 specifically includes: traversing the instruction logic tree to identify the geometric variable node or process parameter node to which the candidate marker to be generated belongs; calculating the logical path length of the geometric variable node or process parameter node relative to the initially defined node; calculating the logical topology depth value based on the logical path length to quantify the semantic contribution weight of the candidate marker to be generated in the parametric modeling process; and reducing the logical topology depth value when the logical path length exceeds a preset length threshold to reduce the sampling offset intensity at the corresponding candidate marker.
[0007] Preferably, in step S3, the dynamic sampling offset operator The calculation rules are as follows: Where H is based on a preset private key With the industrial design script sequence to be processed The hash operation is performed, where d is the logical topology depth value, α is the preset topology response coefficient, and the value of α ranges from 0.05 to 0.15.
[0008] Preferably, step S5 specifically includes: dividing the model code sampling interval and the identity identifier sampling interval within the high-entropy semantic interval according to the value of the dynamic sampling offset operator; when the logical topology depth value of the candidate tag to be generated is detected to be greater than the preset depth threshold, the identity identifier sampling interval is nested and mapped into the model code sampling interval, and the nonlinear embedding of multi-bit superimposed watermark is achieved by changing the probability distribution of the candidate tag.
[0009] Preferably, the method further includes a hierarchical detection step: step S51, counting the total number of instructions in the industrial script to be detected; step S52, if the total number of instructions is lower than a preset first quantity threshold, reconstructing the hash projection space according to the preset private key and performing model identifier matching; step S53, if the total number of instructions is not lower than the first quantity threshold, performing user source identification by inversely deducing the bit offset features in the high-entropy semantic interval based on the dynamic sampling offset operator.
[0010] Preferably, the high-entropy semantic interval is determined by: calculating the unnormalized predicted probability distribution of each candidate tag in the candidate tag list. ; Select a set of candidate tags whose information entropy exceeds a preset entropy weight threshold to construct a high-entropy vocabulary; Use dynamic hash values to rearrange the high-entropy vocabulary in random order and define a sampling interval for carrying multi-bit superimposed watermarks.
[0011] Preferably, step S5 is followed by a consistency verification step: retrieving the probability distribution offset corresponding to the candidate marker after sampling space partitioning modulation; if the probability distribution offset causes the change in geometric constraint value to exceed the preset tolerance range, then negative compensation is performed on the dynamic sampling offset operator until the output candidate marker meets the monotonicity constraint of the industrial design specification.
[0012] Preferably, before step S3, a security certificate loading step is included: retrieving the security encryption certificate corresponding to the current industrial design task from a third-party server; extracting the feature vector from the security encryption certificate, and injecting the feature vector as an auxiliary factor into the initialization vector of the hash operation.
[0013] Preferably, step S53 specifically includes: extracting key instruction segments from the industrial script to be detected and restoring the local topological features corresponding to the key instruction segments; calculating the logical correlation between the local topological features and the dynamic sampling offset operator; and reconstructing the bitstream hidden in the word probability distribution based on the logical correlation to obtain embedded user traceability information.
[0014] Preferably, the industrial design script sequence to be processed is a script generated based on a parametric modeling language, and the dimension-related constraint nodes include units of mm for limiting the geometric spacing and degrees for limiting the rotation angle.
[0015] Compared with existing technologies, the imperceptible hierarchical watermarking embedding method based on multi-bit embedding of language models in this invention has the following advantages: 1. In the embedding of imperceptible layered watermarks, this invention resolves the compatibility contradiction between multi-bit traceability information and short instruction sequence windows. It adopts a logical field architecture with unequal embedding probabilities, distributing model identification codes, user identity identifiers, and time information across sampling windows with progressive probability gradients. When processing short-text-length industrial instruction sequences, the high-weighted first field, through a higher sampling trigger frequency, quickly accumulates statistical significance within a limited character window, achieving instantaneous determination of the source of generation. This avoids the limitation of traditional multi-bit watermarks that rely on long text sequences to extract effective information, ensuring traceability reliability in short-sequence conditions such as industrial production scripts.
[0016] 2. To eliminate the numerical offset interference caused by the watermark embedding process to industrial modeling parameters, this invention constructs a semantic protection mechanism based on low-entropy vocabulary sampling avoidance. Before performing probability modulation, it pre-identifies and locks numerical variables, coordinate parameters, and grammatical reserved words in industrial instructions. By maintaining the probabilistic stability of such highly sensitive logical nodes during vocabulary sampling, this invention ensures that watermark information is only implanted in high-entropy semantic regions, thus physically blocking the interference of watermark signals to industrial design logic, preventing CAD modeling collapse or manufacturing instruction failure due to parameter drift, and ensuring the absolute accuracy of industrial assisted design data.
[0017] 3. Constructing anti-rewrite stability based on logical coupling and hash mapping: This invention drives the probabilistic offset operator by generating a dynamic hash value from the private key and the preceding token sequence, making the watermark bit implicitly coupled with the logical skeleton of the industrial text. Under this mechanism, the watermark information is no longer arranged in a linear manner, but is deeply embedded in the selection logic of high-entropy words. Even if attackers try to remove the identifier through synonym replacement or local instruction reorganization, as long as the core semantic topology of the industrial design script remains unchanged, the probabilistic offset features hidden in the statistical distribution can still be accurately reconstructed, thereby improving the anti-attack strength of industrial digital assets during the circulation process. Attached Figure Description
[0018] Figure 1 This is a flowchart of the layered watermark embedding process based on the logical topology depth and dynamic sampling of this invention; Figure 2 This is a curve comparing the watermark extraction accuracy of three schemes under different sequence lengths according to the present invention; Figure 3 This is an architecture diagram of the industrial data watermarking embedding system integrating an AI inference engine, as described in this invention. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0020] It should be noted that all directional and positional terms used in this invention, such as: up, down, left, right, front, back, vertical, horizontal, inner, outer, top, bottom, transverse, longitudinal, center, etc., are only used to explain the relative positional relationship and connection between components in a specific state (as shown in the accompanying drawings). They are only for the convenience of describing this invention and do not require that this invention be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention. In addition, the descriptions of "first," "second," etc., in this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated.
[0021] In the description of this invention, unless otherwise explicitly specified and limited, the terms installation, connection, and linking should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections; they can refer to direct connections or indirect connections through an intermediate medium; they can refer to the internal connection of two components. For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances.
[0022] In the description of this specification, references to the terms "an embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example, and the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0023] This invention provides an imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model. By performing watermarking based on logical topology depth on the industrial design script sequence to be processed, it addresses the low traceability confidence and parameter offset risks of industrial instruction data in short-sequence scenarios. The method mainly includes stages such as acquiring the industrial design script sequence to be processed and its corresponding instruction logic tree, calculating the logical topology depth value, constructing a dynamic sampling offset operator, performing dimension remapping correction, and sampling space partitioning modulation. Addressing the challenges of highly logical atomicity and structural sensitivity in industrial instruction scripts, the system executes a parsing procedure. The processor calls a lexical parser to process the industrial design script sequence to be processed, constructing an instruction logic tree. The instruction logic tree includes a global reference node, geometric entity definition nodes, size-related constraint nodes, and manufacturing parameter instruction nodes. The processor traverses the instruction logic tree, identifying the geometric variable node or process parameter node to which the current candidate mark belongs. By calculating the logical path length of the geometric variable node or process parameter node relative to the global reference node, the logical topology depth value d is determined. The logical topology depth value d is used to characterize the logical hierarchical order of the current instruction in the parametric modeling process, determining the probabilistic mapping relationship between the watermark bits and key instruction fields.
[0024] To establish coupling between the watermark bits and the logical skeleton of the industrial text, the system executes an operator construction procedure. The processor uses a preset private key and a dynamically generated hash value from the industrial design script sequence to be processed, combined with the logical topology depth value. Calculate the dynamic sampling offset operator Dynamic sampling offset operator The calculation formula is as follows: Where H is based on a preset private key With the industrial design script sequence to be processed The hash operation result, d, represents the logical topology depth value, and α is the preset topology response coefficient, with α ranging from 0.05 to 0.15. By introducing topology depth as an offset perturbation factor, the distribution state of the watermark information is bound to the logical center of gravity of the industrial instructions, improving the traceability reliability under short sequence conditions with fewer than 50 tags. Based on the impact of numerical variables and coordinate parameters in the industrial design script on logical integrity, the system performs a correction procedure before sampling modulation. The processor searches the high-entropy semantic interval of the candidate tag list to see if there are candidate tags containing geometric constraint values. If there are candidate tags containing geometric constraint values, the dynamic sampling offset operator is used. The probability distribution vector of the candidate tag list is subjected to dimensional remapping correction. The specific logic for performing dimensional remapping correction is as follows: The processor first traverses the 32000-word vocabulary index, identifies and extracts all 1024 sensitive word index positions corresponding to digits from 0 to 9, decimal points, and length units mm and angle units deg using a preset regular expression library. Then, the processor forcibly clears the probability offset increment corresponding to these 1024 sensitive word index positions in the probability distribution vector to zero, and adjusts the probability bias weight that should have been superimposed on that position according to a 1.0 to 5... A scaling factor of 0.0 is linearly allocated to the 4096 non-numerical descriptive semantic tags indexed in the high-entropy vocabulary. Through this index reorientation mechanism, the system ensures that the energy field of the watermark signal avoids the instruction field containing key dimensions, mapping the watermark information to the instruction field corresponding to the geometric constraint value. This keeps the probability distribution offset of the candidate tag list within a preset 5% numerical fluctuation range. This operation uses a non-linear correction mechanism to avoid geometric parameters, ensuring the accuracy of parameters such as unit mm and rotation angle, and preventing the computer-aided design modeling logic from failing in subsequent simulations.
[0025] The processor extracts the top N high-entropy markers from the unnormalized predicted probability distribution in the current generation step, sorts them in descending order of their original log-probability values, and divides them into M equal-length probability intervals corresponding to the number of bits to be embedded in the watermark. Then, it uses a dynamic sampling offset operator... The scalar value determines the target offset direction, when When a value falls within the range defined by the k-th interval, a fixed bias increment δ is added to the logarithmic probability of all markers within that interval. For the dynamic sampling offset operator, variable N is the number of high-entropy markers, variable M is the number of probability intervals, variable δ is the log-probability bias increment, and variable k is the index of the interval. The bias increment δ is 5% of the absolute value of the average log-probability of the high-entropy markers, generating a probability space distribution orientation offset. Before performing sampling space partitioning modulation, dimension remapping correction is performed. The processor identifies geometric parameter-related items in the candidate marker list through a preset regular expression library. The library contains rules for matching numerical strings in floating-point format as well as rules for length and angle units. When a candidate marker in a high-entropy semantic interval is detected to match the rules, the marker is removed from the set to be modulated and the corresponding probability offset increment is forcibly cleared to zero. The affected probability weights are proportionally redistributed to non-numerical descriptive semantic markers. The execution deviation of the corresponding instruction field in the parametric modeling process is limited to the 5% fluctuation range allowed by industrial design specifications. The high-entropy semantic interval is generated by calculating the unnormalized predicted probability distribution of each candidate marker and filtering the set whose information entropy exceeds the preset entropy weight threshold.
[0026] Based on the dynamic sampling offset operator The processor performs sampling space partitioning modulation on the high-entropy semantic interval of the candidate tag list, generating candidate tags that carry multi-bit superimposed watermarks. These watermarks include model identification codes, user identity identifiers, and time information. In this embodiment, the watermark information is divided into four fields with progressive probability gradients, with embedding probabilities of 50%, 25%, 18.75%, and 6.25%, respectively. When the logical topology depth d of the current tag is detected to be greater than a preset depth threshold, the identity identifier sampling interval is nested and mapped into the model code sampling interval. By changing the probability distribution of the candidate tags, nonlinear embedding of multi-bit superimposed cryptography is achieved, enabling a single tag to carry bit features of multiple fields. This allows for determining the model source in short texts and completely tracing the user identity identifier in long texts. During the detection phase, the system switches detection strategies based on the total number of instructions in the industrial script to be detected. If the total number of instructions is less than a preset first threshold, the processor reconstructs the hash projection space using a preset private key and performs model identifier matching. If the total number of instructions is not less than the first threshold, the processor uses a dynamic sampling offset operator... By reverse-engineering the bit offset features in the high-entropy semantic interval, user identity and generation time are identified. This hierarchical detection method supports a smooth transition from basic model recognition to in-depth user tracing, providing a technical evidence chain for the full lifecycle management of industrial design assets.
[0027] In industrial design environments involving multi-level supply chain flows, the processor executes a hierarchical extraction procedure based on bit energy accumulation to address the risk of truncation of the industrial design script during forwarding. This is achieved by traversing each marker in the industrial script to be detected and calculating the corresponding logical topology depth value d, using a preset private key. The hash projection space is reconstructed to map the watermark energy of each bit, and a priority competition operator is established between the model identification code field and the user identity field. If the number of instructions in the detection window is less than 80 tags, the system freezes the decoding register of the identity field and converges all hash mapping energy to the sampling interval where the 11-bit model identification code is located. The discrimination threshold of these 80 tags is obtained by recursively testing 500 sets of benchmark industrial scripts with lengths ranging from 20 to 300 tags during the calibration phase of system power-on initialization. When the instruction sequence length is less than 80 tags, the extraction signal-to-noise ratio of the identity field will fall below the 3.5dB confidence threshold, resulting in a decoding error rate exceeding 1%. Therefore, the system sets an 8-bit counter register to count the current instruction stream in real time. Once the count value is less than 80, the logic gate circuit is triggered to switch the hash mapping energy to the model code recognition interval with higher weight. To determine the stability of the detection judgment logic, the processor performs energy normalization processing based on the distribution variance of each bit and outputs the corresponding confidence index. When the average watermark energy of the model identification code field is less than 80 tags, the system performs the watermark projection space to map the watermark energy to the sampling interval where the watermark energy of the model identification code is less than 80 tags. When the logical judgment threshold of 0.75 is exceeded, the system outputs the source identification result and initiates the recursive reconstruction logic of user identity information. By searching the statistical consistency of the 34-bit user identity identifier under different topological depth nodes, and using parity bits to correct and compensate for damaged bits, the system maintains a user identity tracing accuracy of 99.3% in a 15dB signal-to-noise ratio environment in a complete script containing 260 tags. Among them, the average watermark energy is... The calculation formula is as follows: ,in, Let N be the average watermark energy, and N be the total number of tags within the sampling window. Let be the projected energy value of the i-th mark on the corresponding bit.
[0028] Example 1: In an industrial design scenario involving additive manufacturing of high-precision aerospace components, the system faces the task of generating a parameterized path instruction script with 40 markers. Because this type of script contains numerous geometric constraint variables affecting the structural strength of the parts, and the text entropy space is limited, if a conventional text watermarking scheme is used to embed traceability information exceeding 30 bits, it will induce random offsets of coordinate values greater than 0.05mm, leading to dimensional deviations in the manufactured entity. To address this application-layer challenge, this example first utilizes the processor to perform lexical analysis on the OpenSCAD format industrial design script sequence to be processed, constructing an instruction logic tree containing global reference nodes and geometric entity definition nodes. By traversing the instruction logic tree, the system identifies that the candidate marker to be generated belongs to a rotation angle constraint node and calculates its logical path length relative to the initial definition node, determining the logical topology depth value d to be 3. Subsequently, the system calls a preset private key. The dynamic sampling offset operator is calculated by combining the dynamic hash value H generated from the first four tags in the industrial design script sequence to be processed with the logical topology depth value d. Specifically, the dynamic sampling offset operator The calculation formula is as follows: Where H is based on a preset private key With the industrial design script sequence to be processed The hash operation result, where d is the logical topology depth and α is the preset topology response coefficient, is used in this case with the topology response coefficient α set to 0.12. The calculated dynamic sampling offset operator... With a value of 0.68, this operator nonlinearly couples the distribution weight of the watermark bits with the logical hierarchy depth of the instruction, so that even in the case of short text, the watermark signal can be anchored to a topological node with high logical stability, thereby solving the contradiction between the confidence of watermark extraction and the reliability of detection in short sequence scenarios.
[0029] Before performing sampling spatial modulation, the system performs a dimension remapping correction procedure. The processor searches the current candidate marker list and determines if it contains a length variable in mm. The system then utilizes a dynamic sampling offset operator. A dimension remapping operation is performed on the probability distribution vector, forcibly mapping the offset containing watermark information to a non-numerical semantic description field, while limiting the probability distribution offset to within 5%. This operation, through quantization compensation for execution deviations, ensures that the generated coordinate values remain consistent with the initial design logic. This changes the problem nature of traditional solutions involving perturbations in the numerical domain, ensuring that high-precision geometric constraints are no longer affected by watermark embedding behavior under the new logical framework. Finally, the processor uses a dynamic sampling offset operator... The high-entropy semantic interval of the candidate tag list is subjected to sampling space partitioning modulation, and a superimposed watermark consisting of an 11-bit model identification code and a 34-bit user identity low-order bit is embedded. By nesting and mapping the low-order sampling interval of the user identity to the sampling interval of the model identification code, a single sampling action can carry multiple logical field bit features. This hierarchical probability weight allocation mechanism has achieved the implantation of 45-bit watermark information in a short sequence of 40 tags. After verification by reconstructing the hash projection space at the detection end, the accuracy of the extracted traceability information is no less than 99%. The final industrial design script output by the system, while maintaining the accuracy of physical size definition, achieves robust user identity traceability in an extremely short instruction stream through a positive feedback loop of logical topology depth and sampling offset.
[0030] Example 2: In a test environment equipped with a distributed computing cluster, 1200 segments of OpenSCAD format industrial design instruction streams were used as raw input data. The performance of watermark embedding under semantic perturbation conditions was recorded. The instruction streams covered multi-level part declarations and coordinate constraint logic. In the experiment, the signal-to-noise ratio gradient was set to three levels: 35dB, 25dB, and 15dB. The parameter α was set between 0.05 and 0.15 to establish a constraint relationship between watermark robustness and instruction semantic fidelity. When the system processed instruction sequences with up to 40 tags, the value of parameter α was selected as 0.12 to achieve statistical significance of the bit signal. Under strong noise conditions with a signal-to-noise ratio of 15dB, the control group used an unordered hash modulation method, and its probability consistency index of key intermediate features dropped to 42.6%, resulting in a corresponding decrease in the model recognition code extraction accuracy to 58.4%. However, using the sample group scheme of this invention, due to the introduction of logical topology depth value d probability anchoring, the candidate tag list forms a probability enhancement at the geometric entity definition node. At this time, the dynamic sampling offset operator The measured values stabilized within the range of 0.65 to 0.71, indicating that the dynamic sampling offset operator... The calculation formula is as follows: ,in, For dynamic sampling offset operators, H is based on a preset private key. With the industrial design script sequence to be processed The result of the hash function is given, where d is the logical topology depth and α is the preset topology response coefficient.
[0031] When performing boundary tests on parameter α, if α is selected as an upper limit value of 0.25, the thread pitch coordinate value, originally 2.50mm, in the generated industrial design script will drift dimensionally and output as 2.62mm due to the Logits offset disturbance. This causes the interference check in the computer-aided design modeling process to fail. However, when α is set to 0.12, the probability distribution fluctuation is limited to within 4.2%, and the generated physical quantity coordinates are consistent with the initial design logic. This data reflects that the value range of 0.05 to 0.15 avoids geometric parameters through a nonlinear correction mechanism, suppressing the execution deviation caused by bit embedding. In the instruction stream of 40 marked characters, by executing the multi-field probability interval collapse instruction, a single sampling behavior carries bit features including an 11-bit model identification code and a 34-bit user identity identifier. The detection end extracts the embedded 45-bit watermark information by reconstructing the hash projection space. The measured traceability accuracy is no less than 99.1%. This result corresponds to the positive feedback coupling between the watermark information and the logical topology depth of the industrial instructions, enabling the extracted traceability data and the preset private key to work together. By matching the corresponding identity information, user identity can be traced while maintaining the accuracy of the manufacturing script values.
[0032] Example 3: This example combines Figures 1 to 3 This section describes an imperceptible hierarchical watermarking embedding method based on multi-bit embedding of a language model, as follows: Figure 1 As shown, step S1 is executed first to obtain the industrial design script sequence to be processed, the corresponding candidate tag list, and the instruction logic tree containing multiple nodes. Then, step S2 is executed to determine the logical topology depth value of the current tag relative to the global reference node based on the hierarchical dependency relationship, which is used to characterize the instruction logic hierarchy position. Next, step S3 is executed to calculate the dynamic sampling offset operator by combining the preset private key, dynamic hash value, and logical topology depth value to establish a probability mapping relationship. Based on this, step S4 is executed to retrieve the high-entropy interval and use the operator to perform dimension remapping correction on the tag with geometric constraint values to limit the probability distribution offset. Finally, step S5 is executed to perform sampling space partitioning modulation on the high-entropy semantic interval according to the dynamic sampling offset operator to generate a multi-bit superimposed watermark.
[0033] like Figure 2As shown in the chart, the horizontal axis represents the text length (number of tags), with a scale range of 20 to 260, and the vertical axis represents the watermark extraction accuracy, with a percentage scale range of 30 to 100. The chart contains three curves: the solid line represents the method of this invention, the dashed line represents the traditional statistical scheme, and the dotted line represents unordered hash modulation. The data shows that when the text length reaches 40 tags, the watermark extraction accuracy of the method of this invention is close to 100% and remains stable in subsequent lengths. In contrast, the accuracy of the traditional statistical scheme and unordered hash modulation is lower at the same length and only gradually improves with the increase of text length, indicating that the present invention has a performance advantage in short sequences.
[0034] like Figure 3 As shown, the system consists of an industrial design workstation on the left as the input end, which includes a design software interaction front-end and a raw script sending component. The raw instruction script is transmitted to an AI server equipped with a GPU via an encrypted network connection (SSL / TLS). Its core is an imperceptible watermark embedding inference engine, which integrates a language model base (responsible for generating candidates), a logical topology deep analyzer, a dynamic sampling offset operator calculator, a geometric constraint detection and correction module, and a sampling space partition modulator. The AI server connects to the key and basic data center through an internal high-speed secure channel, reading the private key and hash value secure storage, the instruction logic tree knowledge base, and historical and raw data in the industrial script database. The finally generated watermarked script is sent to the CNC machine tool and production equipment on the right through an industrial-grade reliable transmission protocol, where the watermarked script execution unit completes the processing operation.
[0035] Example 4: In an industrial setting involving the generation of path instructions for multi-axis CNC machine tools, the processor processes 150 lines of G-Code instruction scripts contained in the industrial design script sequence to be processed. The system uses a lexical unit scanner to identify the main function words, coordinate values, and auxiliary parameters in the instructions. It maps G-class motion instructions to motion geometry branches in the instruction logic tree, and F-class feed rate parameters to dimension-related constraint nodes. Based on the line index and parameter dependencies of each instruction line, the system calculates the logical path length, transforming the linear character sequence into a discrete set of nodes with a topological structure. Based on this, the system determines... The logical topology depth value d of the candidate marker quantifies the temporal position of the current processing action in the overall process route, and is used to establish the probabilistic mapping relationship between watermark bits and key instruction fields. To determine the parameter value, the system performs matching of the topology response coefficient α. The processor determines the instruction density ρ by statistically counting the number of nodes in the motion geometry branches within a preset sliding window. When the instruction density ρ is in the range of 0.8 to 1.2, the topology response coefficient α is set to 0.12. If the instruction density ρ is lower than 0.5, the system adjusts the topology response coefficient α to 0.15 to enhance the sampling offset intensity. The dynamic sampling offset operator... The calculation formula is as follows: ,in, For dynamic sampling offset operators, H is based on a preset private key. With the industrial design script sequence to be processed The result of the running hash function, where d is the logical topology depth value, α is the topology response coefficient, the functional specifications of the implementation environment meet real-time requirements, and the instruction stream throughput is not less than 10k / s.
[0036] After determining the operator values, the processor obtains the probability distribution vector P output by the large language model, locates the high-entropy region in the candidate label list that belongs to non-numerical semantic fields, and uses the dynamic sampling offset operator. Generate a Logits bias mask vector with the same dimension as vector P. The processor will vector Superimposed onto the original Logits space, this causes a directional shift in the probability of the marker corresponding to the target watermark bit. A dimensionality remapping correction procedure locks in a maximum shift of 5%. During the physical mapping of G-Code instructions, the processor compares the difference between the sampled marker and the original instruction in real time. If the deviation of the sampled F-type feed parameter or S-type spindle speed parameter exceeds 0.005 mm / min or 1.0 rpm, the system automatically triggers a negative compensation interrupt. Under this interrupt logic, the processor gradually reduces the dynamic sampling offset operator with a compensation coefficient of 0.1. The bias strength of the probability space is adjusted until the output mark completely returns to the tolerance range allowed by the industrial design specifications, thereby realizing closed-loop precision control from the logical probability space to the physical execution level of the CNC machine tool. This procedure avoids the F and S parameters in the G-Code instructions, ensuring that the final output machining script maintains a feed accuracy of 25,000 mm and a speed definition of 3000 rpm. The confidence level of its hidden bit sequence is 99.5%. This process realizes the implantation of watermark signals without changing the control logic of the CNC machine tool, completing the logical closed loop from topology perception to physical instruction protection.
[0037] Example 5: In a scenario where a parametric simulation script generation system is deployed, the processor performs offline calibration on the semantic features of the industrial design script sequence to be processed. Since the scripts generated for different industrial applications have different feature distributions in the probability space, the system performs a pre-deployment calibration process to determine the initial value of the topological response coefficient α. By collecting 5000 segments of original simulation instructions without watermarks and constructing the corresponding instruction logic tree, the average statistical entropy of candidate labels at each level is calculated. And establish the topological response coefficient α and the mean statistical entropy. The mapping constraint, when the statistical entropy mean of the motion geometry branches in the simulation script When the value is in the range of 3.5 to 4.2, the processor determines the initial calibration value of the topological response coefficient α to be 0.08 based on the distribution of the candidate label list.
[0038] When the system processes a manufacturing instruction stream containing multiple layers of circular dependencies, the processor executes a dynamic sampling offset operator based on the parameters obtained from the calibration process described above. The operation is performed, and when the logical topology depth value d of each node changes, the dimension remapping mask in the register is automatically retrieved to correct the probability distribution vector P. If the extracted watermark feature energy value is lower than the preset threshold of 0.75, the processor performs feedback compensation to make the topology response coefficient α approach 0.15 along a monotonically increasing law. Under the condition of maintaining a simulation step size of 0.001ms physical accuracy, the statistical consistency probability of the watermark sequence of the generated candidate mark stream is maintained at 99.8% within a sliding window of 100 marks. Dynamic sampling offset operator The calculation formula is as follows: ,in, For dynamic sampling offset operators, H is based on a preset private key. With the industrial design script sequence to be processed The result of the hash function is given, where d is the logical topology depth and α is the topology response coefficient.
[0039] Example 6: In a watermarking system deployment scenario targeting heterogeneous language models, the processor performs pre-calibration to determine the initial baseline for the topological response coefficient α. A baseline dataset is constructed by collecting 10,000 segments of industrial scripts covering motion instructions and loop logic. The original probability distribution vector is extracted from the model's generation head, and the average information entropy of the candidate label list is calculated. Subsequently, the correlation function was used to establish the relationship between the topological response coefficient α and the average information entropy. Given the constraint relationships, the formula for calculating the topological response coefficient α is as follows: Where α is the topological response coefficient. Let be the average information entropy with one dimension, and the average information entropy The value ranges from 2.5 to 5.0. The calculated coefficient value is written into the processor's initialization register as the physical reference for subsequent sampling offset.
[0040] During model inference, the system executes an adaptive energy compensation process to address fluctuations in source traceability confidence caused by changes in the length of the generated sequence. The processor monitors the average bit energy value within the sliding window in real time. If the average bit energy value is detected If the threshold value is below 0.70, the processor reduces the sampling width of low-probability terms in the candidate tag list and simultaneously adjusts the offset in the dimension remapping correction process upward by 1%. This operation, without changing the internal weight parameters of the language model, uses the bias adjustment of the probability space at the generation end to offset the noise interference generated by the long sequence, so that the final output industrial design script maintains a simulation step size of 0.001ms physical accuracy while the watermark bit extraction energy value rises back to above 0.85.
[0041] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit of this application and the scope of protection of this invention, and all of these forms are within the protection scope of this application.
Claims
1. An imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model, characterized in that, Includes the following steps: Step S1: Obtain the industrial design script sequence to be processed, the candidate tag list corresponding to the industrial design script sequence to be processed, and the instruction logic tree corresponding to the industrial design script sequence to be processed. The instruction logic tree includes a global reference node, a geometric entity definition node, a size association constraint node, and a manufacturing parameter instruction node. Step S2: Based on the hierarchical dependency relationship of each node in the instruction logic tree, determine the logical topology depth value of the candidate label to be generated relative to the global baseline node in the instruction logic tree. The logical topology depth value is used to characterize the logical hierarchy position of the current instruction in the parameterized modeling process. Step S3: Using the preset private key and the dynamic hash value generated by the industrial design script sequence to be processed, combined with the logical topology depth value, calculate the dynamic sampling offset operator to establish the probabilistic mapping relationship between the watermark bits and the key instruction fields. Step S4: Before performing sampling space modulation, search the high-entropy semantic interval of the candidate tag list to see if there are candidate tags containing geometric constraint values; if there are candidate tags containing geometric constraint values, use the dynamic sampling offset operator to perform dimension remapping correction on the probability distribution vector of the candidate tag list, and map the watermark information to the instruction field corresponding to the geometric constraint values, so as to limit the probability distribution offset of the candidate tag list to within the preset numerical fluctuation range. Step S5: Based on the value of the dynamic sampling offset operator, perform sampling space partitioning modulation on the high-entropy semantic interval of the candidate tag list, so that the generated candidate tags carry a multi-bit superimposed watermark containing model source information and user traceability information.
2. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 1, characterized in that, Step S2 specifically includes: traversing the instruction logic tree to identify the geometric variable node or process parameter node to which the candidate mark to be generated belongs; calculating the logical path length of the geometric variable node or process parameter node relative to the initially defined node; calculating the logical topology depth value based on the logical path length; and reducing the logical topology depth value when the logical path length exceeds a preset length threshold to reduce the sampling offset intensity at the corresponding candidate mark.
3. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 1, characterized in that, In step S3, the dynamic sampling offset operator The calculation rules are as follows: Where H is based on a preset private key With the industrial design script sequence to be processed The hash operation is performed, where d is the logical topology depth value, α is the preset topology response coefficient, and the value of α ranges from 0.05 to 0.
15.
4. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 1, characterized in that, Step S5 specifically includes: dividing the model code sampling interval and the identity identifier sampling interval within the high-entropy semantic interval according to the value of the dynamic sampling offset operator; when the logical topology depth value of the candidate tag to be generated is detected to be greater than the preset depth threshold, the identity identifier sampling interval is nested and mapped to the model code sampling interval, and the nonlinear embedding of multi-bit superimposed watermark is achieved by changing the probability distribution of the candidate tag.
5. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 1, characterized in that, It also includes a layered detection step: Step S51, count the total number of instructions in the industrial script to be detected; Step S52, if the total number of instructions is lower than the preset first quantity threshold, reconstruct the hash projection space according to the preset private key and perform model identifier matching; Step S53, if the total number of instructions is not lower than the first quantity threshold, reverse the bit offset features in the high-entropy semantic interval based on the dynamic sampling offset operator and perform user source identification.
6. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 1, characterized in that, High-entropy semantic intervals are determined by calculating the unnormalized predicted probability distribution of each candidate label in the candidate label list. ; Select a set of candidate tags whose information entropy exceeds a preset entropy weight threshold to construct a high-entropy vocabulary; Use dynamic hash values to rearrange the high-entropy vocabulary in random order and define a sampling interval for carrying multi-bit superimposed watermarks.
7. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 1, characterized in that, Step S5 is followed by a consistency verification step: retrieve the probability distribution offset corresponding to the candidate marker after sampling space partitioning modulation; if the probability distribution offset causes the change in geometric constraint value to exceed the preset tolerance range, then perform negative compensation on the dynamic sampling offset operator until the output candidate marker meets the monotonicity constraint of the industrial design specification.
8. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 1, characterized in that, Before step S3, there is also a security certificate loading step: retrieve the security encryption certificate corresponding to the current industrial design task from the third-party server; extract the feature vector from the security encryption certificate, and inject the feature vector as an auxiliary factor into the initialization vector of the hash operation.
9. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 5, characterized in that, Step S53 specifically includes: extracting key instruction segments from the industrial script to be detected and restoring the local topological features corresponding to the key instruction segments; calculating the logical correlation between the local topological features and the dynamic sampling offset operator; and reconstructing the bitstream hidden in the word probability distribution based on the logical correlation to obtain embedded user source information.
10. The imperceptible hierarchical watermark embedding method based on multi-bit embedding of a language model according to claim 1, characterized in that, The industrial design script sequence to be processed is a script generated based on a parametric modeling language, and the dimension-related constraint nodes include units of mm for limiting the spacing between geometric objects and degrees for limiting the rotation angle.
Citation Information
Patent Citations
Content detection method based on large language model generation based on sentence semantic watermarking
CN119939544B
Security guarantee method for large language model generated text
CN120524471A
Large language model watermark embedding method and system based on entropy adaptive adjustment
CN120705842A
Controllable text watermark embedding method based on reinforcement learning strategy model
CN120974466A
Large model generation content traceability technology based on model copyright ID watermark embedding
CN121302334A