Method and system for secure sampling of large language models based on quantum entropy sources
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-08-11
AI Technical Summary
一方面,软件生成的伪随机序列在特定环境下面临内部状态或初始种子被推演的风险,其抗预测能力有待提升,在较高安全要求的场景中适用性受限
1.本发明通过获取底层硬件的原始量子信号生成目标随机数主数据流,为大语言模型采样过程提供物理熵源。相比于基于确定性算法的伪随机数生成机制,该方案引入了光电二极管散粒噪声等物理过程的内在不确定性,为模型提供了不具备可重复性的随机数序列,从而提高了文本生成环节的抗预测能力,降低了模型输出内容被反向推断或操控的风险,增加了系统在高安全性要求场景下的应用可靠性。
Smart Images

Figure CN122549598A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of large language model data processing, and relates to a secure sampling method and system for large language models based on quantum entropy sources. Background Technology
[0002] When large language models perform text generation tasks, the diversity and creativity of the output results largely depend on the randomness introduced in the sampling process. Therefore, when selecting from the model's predicted probability distribution, a random number source needs to be introduced for assistance.
[0003] Existing large language model inference frameworks generally employ software pseudo-random number generators built into the computing system, combined with predetermined strategies such as greedy search or core sampling, to generate text. These pseudo-random number generators are based on deterministic algorithms, using an initial seed to generate corresponding numerical sequences. Meanwhile, hyperparameters in existing sampling strategies, such as probability thresholds or temperature parameters controlling distribution smoothness, are typically set to fixed static values before model inference begins. These values are manually adjusted and remain constant throughout the entire text generation cycle.
[0004] The aforementioned pseudo-random number generation mechanism based on deterministic algorithms and its static parameter settings have certain limitations. On the one hand, software-generated pseudo-random sequences face the risk of internal states or initial seeds being deduced under specific environments, and their anti-prediction ability needs improvement, limiting their applicability in scenarios with high security requirements. On the other hand, the fixed sampling hyperparameter mechanism lacks the ability to respond to dynamic changes in the operating environment. When the underlying hardware experiences performance fluctuations due to interference, the upper-level sampling algorithm struggles to perceive changes in the physical state and make corresponding adjustments. This easily leads to the introduction of low-quality random signals into the generation process and makes it difficult to adaptively schedule resources based on real-time supply and demand, exposing the technical shortcoming of information isolation between the model sampling logic and the underlying physical entropy source state. Summary of the Invention
[0005] In a first aspect, the present invention provides a secure sampling method for large language models based on quantum entropy sources, comprising the following steps: S1. Obtain the original quantum signal at the hardware level, perform randomization processing on the original quantum signal, and generate the target random number main data stream and physical state signal; S2. Obtain the difference between the read and write pointers of the circular buffer, combine it with the physical state signal to perform normalized array encapsulation, and generate a quantum entropy state vector; S3. Extract the entity segment data of the target random number main data stream according to the preset data structure specifications, and perform a header binding mechanism on the quantum entropy state vector and the entity segment data to generate a joint data packet; S4. Intercept the preset underlying random number call logic of the large language model inference program, extract the quantum entropy state vector from the joint data packet for verification, and activate the adaptive sampling kernel function when the verification is successful. S5. Analyze the original entropy generation rate in the quantum entropy state vector according to the adaptive sampling kernel function. If the original entropy generation rate is lower than the preset hardware failure threshold, generate and output a deterministic decision signal to switch to the non-random algorithm and skip the probabilistic processing step for the preset log probability matrix. S6. If the original entropy generation rate is not lower than the preset hardware failure threshold, extract the mass damage degree and buffer filling speed from the quantum entropy state vector, adjust the sampling temperature and probability mass threshold of the large language model sampling strategy accordingly, and use the adjusted sampling temperature and probability mass threshold to perform restriction processing on the preset log probability matrix to generate the conditional probability distribution matrix. S7. Call the entity segmentation data to match the conditional probability distribution matrix to perform targeted mapping calculation, generate and output target words.
[0006] A further aspect of the present invention, step S1, includes the following steps: Based on a preset sampling frequency, the physical components of an independent quantum random number generator are monitored to capture physical noise variation sequences to form the original quantum signal; The original quantum signal is input into a preset cryptographic confusion function to extract a high-density data stream. Parallel unbiased verification items are performed on the high-density data stream, and the data that passes the verification items are spliced together to form the target random number main data stream. The failure rate parameter of the statistical unbiased verification project and the real-time bit generation parameter of the quantum random number generator are encapsulated to generate physical state signals.
[0007] A further aspect of the present invention, step S2, includes the following steps: The current buffer data storage volume is calculated based on the difference between the read and write pointers in the current time window. Time difference operation is performed and combined with smoothing filtering to generate a stable instantaneous filling speed. Extract the mass damage degree and original entropy generation rate contained in the physical state signal; Based on preset hardware theoretical parameters, the instantaneous filling speed and the original entropy generation rate are normalized to obtain scaling indexes. The scaling indexes and the mass damage degree are then reordered according to preset levels to form a quantum entropy state vector.
[0008] A further aspect of the present invention, step S3, includes the following steps: The memory space is divided into a metadata header area and a dynamic length payload area according to the preset data structure specifications. Extract the entity segment data for the current sampling request period; The quantum entropy state vector is written to the starting position of the allocated metadata header area, and the entity segment data is loaded into the dynamic length payload area to form a joint data packet with continuously distributed memory addresses.
[0009] A further aspect of the present invention, step S4, includes the following steps: Obtain the complete combined data packet through data dequeueing operations; Parse the metadata header of the combined data packet to extract the preset integrity check code comparison field and the quantum entropy state vector; The current integrity check code is calculated for the combined sequence of the extracted quantum entropy state vector and the corresponding entity segment data. The data consistency of the comparison field is compared with the current integrity check code to complete the verification.
[0010] A further aspect of the present invention, step S5, includes the following steps: Extract the original entropy generation rate from the quantum entropy state vector; The failure status of the physical entropy source is determined by comparing the original entropy generation rate with the preset hardware failure threshold. If the original entropy generation rate is lower than the preset hardware failure threshold, a physical source alarm command is generated and sent to the large language model inference program to bypass the probability calculation logic and switch to the preset greedy search algorithm to generate a deterministic decision signal.
[0011] A further aspect of the present invention, step S6, includes the following steps: Extract the mass damage degree from the quantum entropy state vector; The attenuation ratio is obtained by multiplying the quality damage degree by a preset adjustment coefficient; based on the preset reference temperature and the preset global absolute temperature lower limit, the amplitude limiting peak finding calculation is performed in combination with the attenuation ratio to generate the reconstructed sampling temperature value; The pre-defined logarithmic probability matrix is scaled and smoothed using the reconstructed sampled temperature values, and the output conditional probability distribution matrix is transformed.
[0012] In a further embodiment of the present invention, step S6 further includes the following steps: Extract the buffer filling velocity from the quantum entropy state vector and introduce it into the preset core threshold calculation equation to generate a dynamic probability mass clipping point; The candidate word probability components are accumulated in descending order according to the temporary probability distribution obtained by transforming the log probability matrix. Remove the tail distribution components of the cumulative result set to zero that exceed the dynamic probability quality clipping point, and perform a second normalization.
[0013] A further aspect of the present invention, step S7, includes the following steps: Construct an ascending probability ladder sequence by accumulating the conditional probability distribution matrix; The preset bits of the entity segment data are extracted into integer values and a division mapping calculation is performed to convert them into equivalent uniformly distributed floating-point values, which are then used as the target falling pointer. Traverse the ascending probability ladder sequence to find the smallest ladder element index that exceeds the target falling pointer; The minimum step element index is converted into an actual string by calling the preset dictionary mapping table, and the actual string is output as the target word.
[0014] Secondly, this invention provides a secure sampling system for large language models based on quantum entropy sources, comprising the following modules: The quantum signal processing module is used to acquire the original quantum signal at the hardware level, perform randomization processing on the original quantum signal, and generate the target random number main data stream and physical state signal; The entropy state vectorization module is used to obtain the difference between the read and write pointers of the circular buffer, combine it with the physical state signal to perform normalized array encapsulation, and generate a quantum entropy state vector. The data encapsulation module is used to extract entity segment data from the target random number main data stream according to the preset data structure specifications, and to perform a header binding mechanism on the quantum entropy state vector and entity segment data to generate a joint data packet; The sampling call interception module is used to intercept the preset underlying random number call logic of the large language model inference program, extract the quantum entropy state vector from the joint data packet for verification, and activate the adaptive sampling kernel function when the verification passes. The hardware failure monitoring module is used to analyze the original entropy generation rate in the quantum entropy state vector according to the adaptive sampling kernel function. If the original entropy generation rate is lower than the preset hardware failure threshold, a deterministic decision signal is generated and output to switch to a non-random algorithm and skip the probabilistic processing step for the preset log probability matrix. If the original entropy generation rate is not lower than the preset hardware failure threshold, the sampling hyperparameter modulation module extracts the mass damage degree and buffer filling speed from the quantum entropy state vector. Based on this, it adjusts the sampling temperature and probability mass threshold of the large language model sampling strategy, and uses the adjusted sampling temperature and probability mass threshold to perform restriction processing on the preset log probability matrix to generate a conditional probability distribution matrix. The quantum random mapping module is used to call the entity segment data to match the conditional probability distribution matrix to perform targeted mapping calculations, generate and output target words.
[0015] In summary, the present invention has the following beneficial technical effects: 1. This invention generates a target random number master data stream by acquiring the original quantum signal from the underlying hardware, providing a physical entropy source for the large language model sampling process. Compared to pseudo-random number generation mechanisms based on deterministic algorithms, this scheme introduces inherent uncertainties in physical processes such as photodiode shot noise, providing the model with a non-repeatable random number sequence. This improves the anti-prediction capability of the text generation stage, reduces the risk of reverse inference or manipulation of the model output, and increases the reliability of the system in scenarios with high security requirements.
[0016] 2. This invention constructs a dynamic feedback adjustment mechanism based on hardware physical state. By acquiring the statistical quality impairment degree of the entropy source in real time, it adjusts the sampling temperature parameter of the large language model inversely according to a preset negative correlation mapping relationship. When fluctuations in the quality of the physical random number stream are detected, this mechanism can dynamically lower the sampling temperature to shrink the output probability distribution, reduce the model's dependence on the random signal during that period, and tend to output high-confidence terms. This design achieves adaptive adaptation of the model sampling strategy to the underlying hardware health status, which helps maintain the logical rationality and output stability of the text generation results when the physical entropy source state changes.
[0017] 3. This invention quantifies the real-time supply and demand relationship of physical entropy resources by extracting the filling speed index of the circular buffer, and dynamically adjusts the probability quality threshold of the core sampling of the large language model accordingly. When the supply of underlying entropy resources is sufficient, the system increases the threshold to increase the range of candidate lexical units and the diversity of generation; when the supply of entropy resources is tight, the threshold is appropriately reduced to shrink the candidate set and thus save physical entropy consumption. This mechanism realizes the linkage scheduling of underlying hardware computing resources and upper-layer model generation strategies, enabling the system to achieve a dynamic balance between text generation diversity and physical resource utilization efficiency, and enhancing the system's robustness in the face of complex operating environments. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. The drawings are used to provide a further understanding of the present invention.
[0019] Figure 1 A flowchart illustrating the secure sampling method for large language models based on quantum entropy sources in the embodiments of this application is disclosed.
[0020] Figure 2 The present application discloses a curve showing the trend of the reconstructed sampling temperature value as a function of the degree of quality damage in the embodiments.
[0021] Figure 3 The present application discloses a graph showing the trend of the dynamic probabilistic quality cutoff point as a function of the buffer filling speed in the embodiments of this application.
[0022] Figure 4 A schematic diagram of the structure of a secure sampling system for a large language model based on a quantum entropy source, as described in an embodiment of this application, is disclosed. Detailed Implementation
[0023] The following is in conjunction with the appendix Figures 1-4 A preferred description of the present invention is provided below.
[0024] See attached document Figure 1 This invention proposes a secure sampling method for large language models based on quantum entropy sources, comprising the following steps: S1. Obtain the original quantum signal at the hardware level, perform randomization processing on the original quantum signal, and generate the target random number main data stream and physical state signal; S2. Obtain the difference between the read and write pointers of the circular buffer, combine it with the physical state signal to perform normalized array encapsulation, and generate a quantum entropy state vector; S3. Extract the entity segment data of the target random number main data stream according to the preset data structure specifications, and perform a header binding mechanism on the quantum entropy state vector and the entity segment data to generate a joint data packet; S4. Intercept the preset underlying random number call logic of the large language model inference program, extract the quantum entropy state vector from the joint data packet for verification, and activate the adaptive sampling kernel function when the verification is successful. S5. Analyze the original entropy generation rate in the quantum entropy state vector according to the adaptive sampling kernel function. If the original entropy generation rate is lower than the preset hardware failure threshold, generate and output a deterministic decision signal to switch to the non-random algorithm and skip the probabilistic processing step for the preset log probability matrix. S6. If the original entropy generation rate is not lower than the preset hardware failure threshold, extract the mass damage degree and buffer filling speed from the quantum entropy state vector, adjust the sampling temperature and probability mass threshold of the large language model sampling strategy accordingly, and use the adjusted sampling temperature and probability mass threshold to perform restriction processing on the preset log probability matrix to generate the conditional probability distribution matrix. S7. Call the entity segmentation data to match the conditional probability distribution matrix to perform targeted mapping calculation, generate and output target words.
[0025] In one embodiment of the present invention, step S1 includes the following steps: Based on a preset sampling frequency, the physical components of an independent quantum random number generator are monitored to capture physical noise variation sequences to form the original quantum signal. The original quantum signal is input into a preset cryptographic confusion function to extract a high-density data stream. Parallel unbiased verification items are performed on the high-density data stream, and the data that passes the verification items are spliced together to form the target random number main data stream. The failure rate parameter of the unbiased verification items and the real-time bit generation parameter of the quantum random number generator are statistically analyzed and encapsulated to generate a physical state signal.
[0026] Specifically, this step is executed in the hardware driver or operating system kernel-level daemon used for quantum entropy source management, and is designed to continuously extract the raw physical signal from the independent quantum random number generator (QRNG) hardware chip and process it into two closely related data products.
[0027] First, the driver accesses the two core physical entropy sources on the QRNG hardware chip in parallel via memory-mapped I / O or a proprietary device interface: The first entropy source is a photodiode. The driver performs analog-to-digital conversion on the current flowing through the diode at a sampling rate in the gigahertz range, capturing the shot noise amplitude generated by the Poisson distribution characteristics of the photons, and forming the first original signal sequence.
[0028] The second entropy source is a quantum tunneling junction. The driver synchronously monitors the voltage across its terminals and captures the voltage fluctuation frequency caused by the probabilistic tunneling of electrons through the insulating barrier, forming a second original signal sequence.
[0029] The two digitized raw signal sequences are initially mixed using a bitwise XOR operation to enhance resistance to common-mode noise interference, thereby generating a unified raw quantum signal. The raw quantum signal is then fed into a multi-stage processing pipeline: Entropy extraction and whitening: The driver calls an entropy extractor based on the secure hash algorithm SHA-256 to input the original quantum signal in blocks. Through the avalanche effect of the cryptographic confusion function, the statistical bias and predictable patterns that may exist in the input signal are smoothed, and the output binary data stream with higher entropy density is produced.
[0030] In the statistical health test, the whitened data stream is fed in real-time into a standard-compliant test suite. This suite performs a set of unbiased checks in parallel, including frequency testing, intra-block frequency testing, run-length testing, and approximate entropy testing. A data block is only certified as qualified if it passes all or most of the preset checks.
[0031] All authenticated data blocks are concatenated to form the final target random number master data stream. The driver then writes this master data stream to a circular buffer located in fixed-memory memory on the host machine. This circular buffer uses write pointers to track the latest write position and read pointers to track positions that have been consumed. The fixed-memory allocation strategy ensures that this memory region is not paged or swapped by the operating system, laying the foundation for efficient, asynchronous data transfer via direct memory access technology by devices such as the graphics processor in subsequent steps.
[0032] During the execution of statistical health tests, the driver synchronously generates physical state signals. These signals contain two key dimensions. The first dimension is the statistical quality impairment, calculated using the following formula: In the formula, statistical quality damage degree It is a normalized value ranging from 0 to 1, used to quantify the degree of degradation in the quality of the random number stream. In the preset time window The number of failed validation items in the test suite. This represents the total number of verification items executed within the same time window.
[0033] The second dimension is the raw entropy generation rate, which is the total number of bits of the raw quantum signal generated by the QRNG hardware chip per unit time without post-processing. The driver program encapsulates the calculated statistical quality impairment degree and the measured raw entropy generation rate into a structured data pair, forming a physical state signal that is precisely aligned in time with the target random number master data stream. This signal objectively records the physical health status and randomness quality level of the entropy source when generating the corresponding data stream.
[0034] It should be noted that the physical state signal is a two-dimensional data structure, with the following format: It records the underlying state of the entropy source hardware within a specific time window. A circular buffer is a first-in, first-out (FIFO) data structure that achieves efficient data access by moving read / write pointers within a fixed-size memory block, avoiding the overhead of frequent memory allocation and deallocation. In scenarios like financial transaction encryption, where security requirements are high, any failure of a core test can be considered unacceptable, even for extremely small errors. Any value may trigger a system alarm.
[0035] Original entropy generation rate It is a core metric for measuring the performance of the QRNG hardware physical layer, measured in bps. Its value reflects the raw rate at which random events are generated by the physical process. Based on the technological level of commercially available QRNG devices, Typical values range from 500 Mbps to 4 Gbps. This parameter is based on the publicly available technical specifications of quantum random number generators currently on the market that are based on semiconductor or optical technologies.
[0036] In a specific example of this invention, the original entropy generation rate of the QRNG hardware chip The speed is constant at 1Gbps. Within a 1ms observation window, the system captures... The original quantum signal of the bit. This signal is processed by the SHA-256 entropy extractor and, according to a compression ratio of 2:1, generates... The data stream to be tested consisted of bits. This data stream was fed into a NIST test suite containing 15 independent verification items. In this test, 14 items passed successfully, but the run-length test item failed. A value below the significance level of 0.01 is considered a failure. Therefore, the number of failed projects... The total number of projects is 1. The value is 15. Based on the formula, the statistical quality damage degree is calculated. Calculated as Ultimately, the system performs two output actions: first, it outputs that... The random data bits are determined to be valid target random number main data streams and sequentially written to the current write pointer position of the circular buffer in the host's fixed memory; secondly, a physical state signal associated with this batch of data is generated, the content of which is a structure containing two floating-point numbers, namely [1.0×10 9 [0.067] is used to characterize the original rate and quality damage of the entropy source when generating this batch of random numbers.
[0037] In one embodiment of the present invention, step S2 includes the following steps: The current buffer data storage volume is calculated based on the difference between the read and write pointers in the current time window. Time difference operation is performed and combined with smoothing filtering to generate a stable instantaneous filling speed. The mass impairment degree and the original entropy generation rate contained in the physical state signal are extracted. Based on the preset hardware theoretical parameters, the instantaneous filling speed and the original entropy generation rate are normalized to obtain the scaling index. The scaling index and the mass impairment degree are reordered according to the preset level to form the quantum entropy state vector.
[0038] This step is performed by a state vector synthesizer deployed within the host operating system. It periodically evaluates the dynamic state of the circular buffer and fuses the physical state signals generated in the preceding steps to ultimately generate a structured vector that can comprehensively characterize the current quantum entropy source supply capability.
[0039] Specifically, the state vector synthesizer uses a preset high-frequency time interval. The metadata of the circular buffer is polled. At each time point... It obtains the write pointer through system calls or direct memory reads. and read pointer The current address of each pointer. The address values of both pointers represent their positions within the buffer in bits or bytes.
[0040] The state vector synthesizer calculates the amount of valid random data stored in the buffer at the current time step. To handle boundary conditions of the circular buffer, when the underlying read / write pointers are limited to... When dealing with relative addresses within a range, assume the total capacity of the circular buffer is... The effective data volume is calculated using modulo logic, and its value is: This calculation ensures that even if the write pointer becomes less than the read pointer due to relative address loops, the correct, non-negative absolute value of the valid data amount can still be output. It should be noted that if the underlying system implementation uses an unsigned absolute counter that does not overflow for the read and write pointers, the valid data amount can also be directly obtained through the absolute difference. The calculation yielded the result.
[0041] To quantify the instantaneous supply and demand of random numbers, the state vector synthesizer then processes the effective random data volume. Execution time differential calculation.
[0042] Specifically, it first uses the backward difference method to calculate the first-order time derivative to assess the original rate of change of the buffer data volume. To prevent drastic numerical fluctuations caused by a sudden batch of random number requests from downstream large language models, the system applies an exponential moving average or sliding time window averaging algorithm to smooth the original rate of change, thereby obtaining a stable instantaneous filling speed. .
[0043] The second-order time derivative, i.e., the filling acceleration, is also calculated. This is used to predict trends in data volume changes. Although the second derivative is calculated for use in more advanced predictive scheduling algorithms, this step primarily uses the smoothed first derivative as the core metric. This instantaneous fill rate... It is further processed by dividing by a preset maximum theoretical fill rate. Normalization yields a dimensionless parameter, namely the buffer filling rate. This normalization operation decouples the metric from the specific hardware implementation and allows for comparison and combination with other normalized parameters.
[0044] After completing the dynamic quantization of the buffer, the state vector synthesizer reads the physical state signal closest to the current time window. The statistical quality impairment is then extracted from this signal. With the original entropy generation rate Two parameters. The rate of generation of original entropy. It is also normalized, that is, its value is divided by the theoretical maximum entropy generation rate of the QRNG hardware design. The normalized original entropy generation rate is obtained. Finally, the state vector synthesizer will use the newly calculated buffer fill rate. Extracted statistical quality impairment and the normalized original entropy generation rate These three key parameters are rearranged into an array according to a predefined data structure specification. The three independent scalar values are organized into a three-dimensional row vector, thus structurally encapsulating them into a multi-dimensional quantum entropy state vector. This vector, as an indivisible data unit, provides a comprehensive snapshot of the current entropy source system, from physical layer performance and data flow quality to application layer supply margin, for subsequent steps.
[0045] The quantum entropy state vector is a three-dimensional vector, and its standard form is: It integrates information from three dimensions: the supply rate, quality, and initial production capacity of the entropy source. (Buffer fill rate) This is a normalized value between -1 and 1, reflecting the net rate of change of the random number stock in the circular buffer. A positive value indicates that supply exceeds consumption, and inventory is increasing; a negative value indicates that consumption exceeds supply, and inventory is decreasing; its absolute value indicates the degree of drastic change. Time interval This is the time step for performing differential calculations. Its value needs to strike a balance between capturing system dynamics and smoothing noise, based on typical high-performance computing scenarios. The preferred range is 1ms to 10ms.
[0046] The original rate of change is unsmoothed. This is a smoothing factor, typically ranging from 0.1 to 0.3, used to control sensitivity to changes in the latest data; maximum theoretical fill rate. This is the baseline for normalized calculations, and its value is set based on hardware limitations such as system bus bandwidth and memory copy speed. For example, for systems connected via a PCIe 4.0 bus, its preferred value is 10Gbps. The normalized raw entropy generation rate. This is also a dimensionless parameter used for cross-platform comparisons. Theoretical maximum entropy generation rate. Provided by the QRNG hardware manufacturer as the upper limit of its device performance.
[0047] In a specific example of the present invention, following the example of the preceding steps, the system has obtained a physical state signal, the content of which is [1.0 × 10 9 [0.067] represents the original entropy generation rate. For 1Gbps, the statistical quality impairment is... The value is 0.067. Now, the state vector synthesizer begins executing this step. Assume the system parameters are set as follows: time interval... Maximum theoretical fill rate Theoretical maximum entropy generation rate At the current point in time The synthesizer measures the effective data volume within the circular buffer. for Bits. It retrieves the previous point in time from historical records. At that time, the amount of effective data for Bits. First, calculate the original rate of change. For simplicity, assume the system is running smoothly. The instantaneous fill rate after smoothing filtering is... Equal to the original rate, that is Next, this speed is normalized to obtain the buffer filling speed. Then, the original entropy generation rate is normalized to obtain... Finally, these three parameters are arranged according to... The order of the states is rearranged into an array to generate the final quantum entropy state vector, with specific values of [0.01, 0.067, 0.25]. This vector is then passed to subsequent processing steps.
[0048] In one embodiment of the present invention, step S3 includes the following steps: According to the preset data structure specifications, the memory space is divided into a metadata header area and a dynamic length payload area; the entity segment data of the current sampling request period is extracted; the quantum entropy state vector is written to the starting position of the allocated metadata header area, and the entity segment data is loaded into the dynamic length payload area to form a joint data packet with continuously distributed memory addresses.
[0049] Specifically, this step is executed by the data encapsulation module located within the host operating system. This module aims to atomically bind the discrete data streams generated in the preceding steps, generating a combined data packet that can be used as a standard transport primitive in low-level network or inter-process communication. First, during the system initialization phase, this module defines the memory layout of the combined data packet according to a set of preset data structure specifications. This specification divides a contiguous memory space into two logical regions: a fixed-length metadata header specifically used to store structured state information; and a subsequent variable-length entity data payload region used to carry the actual random number bit stream.
[0050] During runtime, the data encapsulation module operates synchronously with the read operations of the circular buffer. When a downstream application requests a batch of random numbers, the data encapsulation module extracts a predetermined length of data starting from the read pointer position of the circular buffer. The target random number is segmented from the main data stream. This segment constitutes the entity data within the current time window.
[0051] The data encapsulation module acquires the quantum entropy state vectors within the same time window. After acquisition, it performs a binding operation: allocating a new memory block with a size equal to the metadata header length. Segment length of the data just extracted The sum of these is then used. Subsequently, the contents of the quantum entropy state vector are forcibly injected into the starting position of the newly allocated memory block, i.e., the metadata header area, through a direct memory copy operation.
[0052] After the header injection is completed, the extracted target random number main data stream is copied in segments to the entity data payload area immediately following the header.
[0053] This series of operations ensures that the metadata describing the physical state of the entropy source and the random number content itself are continuous and adjacent in physical memory, forming a joint data packet that is both logically and physically inseparable.
[0054] Before assembly is complete, the data encapsulation module calculates a global integrity check code for the newly generated joint data packet and writes it into a reserved field in the metadata header to ensure that the physical state information and the corresponding random number data will not be damaged or tampered with in subsequent transmission. The joint data packet contains a quantum entropy state vector and an entity data payload area to solidify the data with its generated context, i.e., the physical state, to prevent the two from being separated or tampered with during transmission or processing.
[0055] The data structure specification is a blueprint defining the binary format of the data packet, detailing the byte offset, data type, and length of each field. For example, a typical specification might design the metadata header to be 36 bytes, containing three 8-byte double-precision floating-point numbers to store the three components of the quantum entropy state vector, an 8-byte 64-bit integer to store a high-precision timestamp, and a 4-byte integrity checksum. The entity data payload area is the portion of the data packet used to store the main data stream segment of the target random numbers. The metadata header stores the quantum entropy state vector and other control information, such as the timestamp, data packet length, and checksum. Forced injection refers to the low-level operation of directly writing the binary representation of the quantum entropy state vector to a specified memory location without logical transformation, ensuring the originality and integrity of the metadata. The time window refers to a discrete time period, the length of which is the same as the frequency of calculating the quantum entropy state vector in step S2, ensuring a strict causal correspondence between the random number segment in each joint data packet and the state vector in its header.
[0056] In a specific example of this invention, the system has generated a quantum entropy state vector with values of [0.01, 0.067, 0.25]. Simultaneously, a circular buffer stores the corresponding target random number main data stream. A downstream application requests a data packet, with the following data structure specifications: metadata header length... The requested data segment length includes three 8-byte double-precision floating-point numbers forming a QESV, an 8-byte timestamp, and a 4-byte integrity checksum. The data encapsulation module begins its operation. It first reads 64 bytes of random data from the circular buffer, assuming the first 8 bytes of their hexadecimal representation are 0xD3, 0x07, 0x54, 0x9C, 0x1A, 0x8B, 0xE2, 0x4F, ... . Simultaneously, it acquires the corresponding quantum entropy state vector [0.01, 0.067, 0.25] and the current high-precision timestamp, such as 0x01D8A4B6C2D1E8F0. Next, the module allocates a contiguous block of memory of size 100 bytes. It converts the values 0.01, 0.067, and 0.25 into 8-byte double-precision floating-point binary representations and writes them sequentially to bytes 0 through 23 of the memory block. Subsequently, it writes the timestamp 0x01D8A4B6C2D1E8F0 to bytes 24 through 31. Finally, the 64 bytes of random data 0xD3, 0x07, ... are copied completely to the position starting from byte 36. After the data copy is complete, the module calculates the integrity check code for the entire data from bytes 0 to 31 and bytes 36 to 99, and writes the result to the positions of bytes 32 to 35. At this point, a 100-byte combined data packet is generated, with its memory layout consisting of a 36-byte metadata header and a 64-byte entity data payload area, which are closely connected. This data packet is then sent out as a whole.
[0057] In one embodiment of the present invention, step S4 includes the following steps: The complete joint data packet is obtained through data dequeue operation; the metadata header of the joint data packet is parsed to extract the preset integrity check code comparison field and the quantum entropy state vector; the current integrity check code is calculated on the combination sequence of the extracted quantum entropy state vector and the corresponding entity segment data, and the data consistency between the current integrity check code and the preset integrity check code comparison field is compared to complete the verification.
[0058] This step, executed within a runtime environment tailored for the large language model inference program, redirects the data stream from the standard pseudo-random number generator to a secure physical entropy source generated in the preceding steps by redirecting function calls to the underlying computation framework. Before the large language model inference engine loads at startup, the operating system forces the loading of a specially crafted shared library by setting the LD_PRELOAD environment variable or using an equivalent dynamic link library injection technique. This shared library implements alternative functions with the same function signatures as the standard pseudo-random number call interfaces used internally by the target computation framework, such as PyTorch or TensorFlow. Therefore, when the large language model inference engine requests random numbers at runtime, its original call to the standard pseudo-random number generation function is transparently intercepted by the dynamic linking mechanism, and the alternative function implemented in the shared library is executed instead.
[0059] Once the substitution function is invoked, its execution logic is redirected. It first accesses the output of a shared memory region in the host's fixed memory specifically designated for this invention—the circular buffer. Through a data dequeue operation, such as an atomic dequeue operation, the complete combined data packet is retrieved from this buffer.
[0060] After acquiring the data packet, the substitution function parses it according to a predefined data structure specification. It reads a fixed-length metadata header from the beginning of the data packet and extracts the quantum entropy state vector from it. To meet the characteristics of batch requests for random numbers in the underlying tensor operations of large language models, the substitution function supports extracting multiple joint data packets within the same time window in a loop or in batches, based on the target request size.
[0061] Before the data is used, the system immediately performs a global integrity check on the combined data packets to ensure that the data has not been corrupted.
[0062] Specifically, this verification is achieved by calculating the integrity check code of the binary representation of the concatenated quantum entropy state vector and the entity data payload area, such as the cyclic redundancy check (CRC32). A substitution function calculates the current CRC32 value of the actually extracted combined byte sequence and compares the result with the original CRC32 value stored in another reserved field in the metadata header. If the two integrity check code values match, the quantum entropy state vector is considered intact, and the verification passes.
[0063] Passing the verification is a mandatory prerequisite for activating subsequent advanced processing logic. Once the verification is successful, the main logic of the substitution function is completed, and the underlying adaptive sampling kernel function will be formally called. During the call, the verified quantum entropy state vector and a memory pointer pointing to the starting position of the entity data payload area in the same joint data packet are passed as parameters to the adaptive sampling kernel function, thereby activating the kernel function and providing state input.
[0064] It should be noted that dynamic linking is an operating system-level function that allows functions from external shared libraries to be linked to the main program at runtime. If a function in the shared library has the same name as a standard library function that the main program depends on, the latter can be overridden or replaced, thereby intercepting the call. The standard pseudo-random number call interface refers to the specific function ultimately called in the underlying C++ or CUDA code when the deep learning framework performs operations requiring randomness, such as top-p sampling. Here, the CRC32 algorithm is used, which can detect common bit errors in data transmission with a high probability. The adaptive sampling kernel function is the specific executor of the subsequent steps of this invention. It is a complex function that receives the quantum entropy state vector and the target random number main data stream, and dynamically adjusts the sampling strategy accordingly to generate the target tokens.
[0065] In a specific example of this invention, following the example of the preceding steps, the system has already generated a 100-byte joint data packet, the header of which contains a quantum entropy state vector [0.01, 0.067, 0.25]. Assume that bytes 32 to 35 of the metadata header of this data packet also store a global integrity checksum for the entire data packet, pre-calculated and stored in step S3, with a value of 0x89ABCDEF. At this time, when the large language model based on PyTorch performs text generation, it needs random numbers to perform top-p sampling, triggering a call to the underlying random number generation function. Due to the existence of the LD_PRELOAD mechanism, this call is intercepted and processed by the substitution function. The substitution function retrieves the aforementioned 100-byte joint data packet from the circular buffer. It first reads bytes 0 to 23, extracting the three double-precision floating-point values of the quantum entropy state vector. The substitution function runs a preset checksum algorithm, such as the CRC32 algorithm, on the remaining 96 bytes of data excluding bytes 32-35, and the calculated checksum is also exactly 0x89ABCDEF. The calculated value matches the stored value 0x89ABCDEF read from bytes 32 to 35, thus the integrity check passes. After the check passes, the substitution function immediately calls the adaptive sampling kernel function, passing it two parameters: the first is a quantum entropy state vector containing the values [0.01, 0.067, 0.25]; the second is a memory pointer to the starting position of byte 36 in the 100-byte data packet, which is precisely the beginning of the target random number main data stream segment. At this point, the adaptive sampling kernel function is successfully activated and has obtained all the inputs needed to execute subsequent steps.
[0066] In one embodiment of the present invention, step S5 includes the following steps: Extract the original entropy generation rate from the quantum entropy state vector; compare the original entropy generation rate with the preset hardware failure threshold to determine the failure state of the physical entropy source; if the original entropy generation rate is lower than the preset hardware failure threshold, generate a physical source alarm command and send it to the large language model inference program to bypass the probability calculation logic and switch to the preset greedy search algorithm to generate a deterministic decision signal.
[0067] This step is executed within the adaptive sampling kernel function activated by the preceding step, and is used to implement a fast failure safety mechanism. Before proceeding to any complex probabilistic calculations, the adaptive sampling kernel function first performs a high-priority check: a failure assessment. This assessment directly utilizes the parameters contained in the joint data packet to evaluate the physical health of the quantum entropy source.
[0068] Specifically, the normalized original entropy generation rate is directly extracted from the quantum entropy state vector, which serves as an input parameter. The adaptive sampling kernel function extracts the normalized raw entropy generation rate and compares it with the preset hardware failure threshold in the system configuration. Perform real-time comparison.
[0069] If the value of the normalized raw entropy generation rate is detected to be lower than or equal to the preset hardware failure threshold, If the physical entropy source fails, it is determined that the hardware failure is due to a serious malfunction or a sharp decline in performance of the QRNG hardware, rendering it unable to provide sufficiently reliable physical randomness.
[0070] In response to this hardware failure determination, the system immediately generates a clear alarm signal. The control flow logic inside the adaptive sampling kernel function encapsulates this alarm signal into a low-level physical source alarm instruction and sends this instruction to the large language model inference program. Once this alarm instruction is received, the behavior of the adaptive sampling kernel function will undergo a fundamental change to bypass the probability calculation logic and switch to a preset greedy search algorithm.
[0071] Specifically, it will bypass all subsequent layers involving probabilistic processing of the logarithmic feature array of the large language model's prediction output, including softmax transformation, candidate word set selection, and random number-based sampling, thus directly generating deterministic decision signals.
[0072] From detecting that the physical layer parameters are below the threshold to finally switching to a deterministic algorithm, a complete failure protection closed loop is formed. This ensures that in the extreme case of physical entropy source failure, the system can automatically degrade to a predictable but still safe operating mode, thereby reducing the risk of uncontrollable output due to the use of unreliable random sources.
[0073] It should be noted that failure assessment is a preliminary check performed before any sampling calculations to determine whether the entropy source is in a hardware failure state. Hardware failure threshold. This is a dimensionless normalization threshold, preferably between 0 and 1. The setting of this value requires a trade-off between security and availability; a typical value is 0.05, meaning that when the actual entropy generation rate is less than 5% of the theoretical maximum, the hardware is considered to be in an unreliable shutdown state. This setting is based on research into the failure modes of QRNG hardware, specifically that its output rate typically decreases non-linearly when a severe failure occurs.
[0074] The underlying physical source alarm command is a specific state flag within the adaptive sampling kernel function, used to transmit critical hardware failure information between modules within the function. The large language model's logistic feature distribution refers to the raw, unnormalized log probability vector output by the model for predicting the next word, representing the model's perceived likelihood of each candidate word. The deterministic decision signal is the instruction guiding the sampling process to use a non-random algorithm; in this embodiment, this signal is directly mapped to the command to execute a greedy search.
[0075] In a specific example of the present invention, following the example of the preceding steps, the adaptive sampling kernel function has received an input containing a quantum entropy state vector [0.01, 0.067, 0.25]. The system's preset hardware failure threshold... The value is 0.1. First, the kernel function performs a failure assessment. It extracts the normalized original entropy generation rate from the quantum entropy state vector. Its value is 0.25. Since the normalized original entropy generation rate is not lower than the hardware failure threshold, it is not judged as a hardware failure. The kernel function determines that the physical entropy source state is normal and continues to execute the subsequent normal probability sampling logic. In another hypothetical scenario, due to the QRNG hardware overheating causing a sudden drop in performance, the quantum entropy state vector in the newly generated joint data packet becomes [..., ..., 0.08], where the normalized original entropy generation rate is... The value is 0.08. In this case, when the kernel function performs the comparison, it finds that 0.08 < 0.1, and determines that it is a hardware failure. The kernel function immediately generates an internal low-level physical source alarm instruction. This instruction causes it to skip all planned probability calculation steps and directly execute the preset failure safety procedure. This procedure generates a deterministic decision signal, which calls the greedy search algorithm. Therefore, for the current large language model inference step, the model will find the index with the largest value from its output logits vector, output the corresponding word, and then end the current sampling loop.
[0076] In one embodiment of the present invention, step S6 includes the following steps: Extract the mass damage degree from the quantum entropy state vector; multiply the mass damage degree by a preset adjustment coefficient to obtain the attenuation ratio; based on a preset reference temperature and a preset global absolute temperature lower limit, perform amplitude limiting and peak finding calculation in combination with the attenuation ratio to generate a reconstructed sampling temperature value; use the reconstructed sampling temperature value to perform scaling and smoothing operations on a preset logarithmic probability matrix, and transform the output conditional probability distribution matrix; extract the buffer filling velocity from the quantum entropy state vector, introduce it into a preset core threshold calculation equation, and generate a dynamic probability mass clipping point; accumulate the candidate word probability components of the temporary probability distribution obtained from the logarithmic probability matrix in descending order; strip and zero out the tail distribution components that exceed the position of the dynamic probability mass clipping point, and perform secondary normalization.
[0077] After determining in the preceding steps that no hardware failure has occurred in the physical entropy source, the adaptive sampling kernel function extracts the statistical mass impairment from the quantum entropy state vector, which serves as the input. This value directly reflects the quality level of the current batch of random numbers. For example... Figure 2 As shown, the kernel function uses this value to inversely adjust the sampling temperature parameter, which serves as the smoothness controller for the model's probability distribution, through a preset negative correlation mapping relationship. The specific calculations are as follows: In the formula, It is the reference temperature, which is the default value of the system under ideal conditions; It is the coefficient for adjusting the degree of quality damage to temperature, and its value is strictly limited to... Within the interval, to ensure the non-negativity of the internal decay factor; This is a global absolute temperature lower limit set to prevent computational overflow or model generation from getting stuck in an infinite loop due to excessively low temperature. When the decayed computational temperature is lower than this value, the system will forcibly output this lower limit value.
[0078] This formula ensures that when the statistical quality impairment is... When the sampling temperature parameter increases, i.e., when the quality of random numbers decreases, the sampling temperature parameter... It will be dynamically reduced according to the proportion of quality damage; at the same time, through the outer nested peak-finding function A hard fallback mechanism is implemented to prevent extreme quality damage from causing the temperature to drop to an illegal range. This makes the probability distribution calculated by the softmax function more acute, and the model's output will be more inclined to high-probability deterministic terms, thereby automatically suppressing noise and risks that low-quality randomness may introduce at the algorithmic level.
[0079] See Figure 3 Meanwhile, the adaptive sampling kernel function extracts the buffer filling rate from the quantum entropy state vector. This value precisely quantifies the changing trend of entropy resource inventory. The kernel function incorporates it into the predefined core threshold calculation equation, continuously adjusting the probability quality threshold used for core sampling based on this value. This generates dynamic probabilistic quality cutoff points. The specific calculation is as follows: In the formula, the core probability quality threshold These are parameters of the core sampling algorithm, defining the total probability quality that the candidate lexical set should cover. It is the baseline threshold; It is the adjustment coefficient of the filling speed on the threshold; and It is a reasonable range set for the threshold, for example ; The function will change the variable The value is limited to Within the closed interval.
[0080] This formula makes the buffer fill speed... When the value is positive and large, it indicates that the entropy inventory is abundant, which is close to the probability quality threshold. This will increase accordingly, thereby expanding the size of the candidate word set and encouraging the model to generate more exploratory terms. Conversely, when the buffer filling rate is negative, it indicates that the entropy inventory is being depleted, and the threshold... This will dynamically shrink, reducing the candidate word set and making the model more conservative in order to conserve physical entropy resources.
[0081] The adaptive sampling kernel function dynamically modulates the sampled temperature parameters generated through the above process. With core probability quality threshold The log-probability matrix, or logits, is applied to the current inference step of the large language model. It first uses the modulated sampling temperature parameter. The logits are scaled and then converted into a temporary probability distribution using the softmax function.
[0082] Next, the modulated core probability quality threshold is used. Perform a core sampling operation on this temporary probability distribution. Specifically, the probability components of each candidate word in the temporary probability distribution are accumulated in descending order to select the smallest set of words whose cumulative probability sum just exceeds the dynamic probability quality pruning point. Then, the tail-end distribution components whose cumulative results exceed this position are stripped and zeroed, and the probabilities of the words in the retained set are normalized twice. Finally, the conditional probability distribution matrix is calculated and generated. This matrix serves as the final decision basis for the current sampling step, and its internal probability distribution shape has been jointly shaped by the physical and logical states of the entropy source.
[0083] Global control hyperparameters refer to configuration parameters that have a global impact on the sampling behavior of large language models, mainly including sampling temperature and core probability quality threshold. Sampling temperature parameter It is a scalar used to adjust the smoothness of the softmax function output; lower temperatures make the probability distribution sharper, while higher temperatures make it smoother.
[0084] The conditional probability distribution matrix is a probability vector of the same length as the vocabulary, and the sum of its components is... , representing the model's final prediction probability for the next word under the current entropy source state. The baseline values and coefficients of these parameters were obtained through offline testing and empirical optimization on a large number of samples under various simulated entropy source states.
[0085] In a specific example of the present invention, following the example of the preceding steps, the adaptive sampling kernel function has received the quantum entropy state vector [0.01, 0.067, 0.25] and has passed the failure assessment. The system's preset hyperparameter adjustment parameters are: reference temperature. Temperature regulation coefficient minimum temperature Baseline probability threshold Threshold adjustment coefficient The threshold range is First, the kernel function extracts the statistical quality impairment. To calculate the sampling temperature parameters Substituting into the formula, we get: Next, extract the buffer fill speed. To calculate the core probability quality threshold Substituting into the formula, we get: Assume the current logits output by the large language model are [2.5, 5.1, 1.3, 4.8]. The kernel function first uses... Scaling the logits yields [3.233, 6.596, 1.681, 6.208]. Applying the softmax function to this scaled logits gives a temporary probability distribution, for example, [0.020, 0.581, 0.004, 0.395]. Then, it applies... Core sampling is performed. Since the sum of the probabilities of the two words with the highest probabilities accumulated in descending order is 0.976, which exceeds the dynamic probability quality pruning point, the candidate set only includes these two words, and the tail distribution components are stripped and zeroed out. Finally, the remaining word probabilities are normalized twice to obtain the final conditional probability distribution matrix, for example, [0, 0.595, 0, 0.405].
[0086] In one embodiment of the present invention, step S7 includes the following steps: Accumulate the conditional probability distribution matrix to construct an ascending probability ladder sequence; extract the preset bits of the entity segment data into integer values and perform division mapping calculation to convert them into equivalent uniformly distributed floating-point values, which are used as the target falling pointer; traverse the ascending probability ladder sequence to find the smallest ladder element index that exceeds the target falling pointer; call the preset dictionary mapping lookup table to convert the smallest ladder element index into an actual string, and output the actual string as the target word; After outputting the actual string as the target term, the output also includes: Write the generated target lexical units into the pre-defined large language model context dialogue pool; clear and release the temporary shared memory record block occupied by the joint data packet; reset the local computation variable accumulator of the adaptive sampling kernel function to prepare for the next round of model inference propagation loop.
[0087] This step is performed in the final stage of the adaptive sampling kernel function to complete the final lexical selection within the probability space carefully constructed in the preceding steps.
[0088] First, the adaptive sampling kernel function analyzes the input conditional probability distribution matrix. It treats the matrix as a one-dimensional vector and calculates the cumulative probability ladder values for each component. This process starts with the first component and accumulates the probability values one by one, generating a monotonically increasing sequence of length equal to the vocabulary, with the last value of the sequence being 1. For example, if the probability distribution is... The cumulative probability ladder value is This ladder sequence logically constructs... The interval is divided into multiple sub-intervals, and the length of each sub-interval is exactly equal to the probability of the corresponding word.
[0089] Simultaneously, the adaptive sampling kernel function extracts a memory pointer from the parameters passed to it, pointing to the starting position of the entity data payload region in the current joint data packet. Since the quantum entropy state vector in this joint data packet is strictly bound to the target random number main data stream segment during generation within the same time window, the extracted random number segment is synchronized with the state used to generate the conditional probability distribution matrix. The kernel function reads a sufficient number of bits from the starting position of this data segment and treats it as an unsigned integer value. And by performing a division mapping calculation, it is precisely transformed into a distribution that strictly follows a uniform distribution. Double-precision floating-point numbers within the range: in, This represents the maximum number of states that can be represented by 64 bits. This step completes the operation of segmenting the target random number main data stream and projecting it onto the specified dimension of the candidate character word probability interval.
[0090] The kernel function compares the random floating-point value generated in the previous step, derived from the physical entropy source, with the previously constructed sequence of cumulative probability ladder values. It checks each element in the sequence sequentially, starting from the first, to find the first ladder value greater than or equal to that random floating-point number. The index of this ladder value in the sequence corresponds to the index of a specific character in the dictionary mapping table.
[0091] In this way, the random number's landing point is uniquely mapped to the candidate lexical. This selected lexical is the final product of this sampling step-by-step loop: the target lexical. The adaptive sampling kernel function returns this target lexical as its return value to the calling program, namely the inference engine of the large language model. After receiving the lexical, the inference engine writes it into the preset large language model context dialogue pool as subsequent input.
[0092] Subsequently, the temporary shared memory record block occupied by the joint data packet is cleared and released, while the local computation variable accumulator of the adaptive sampling kernel function is reset to prepare for the next round of model inference propagation loop, thereby driving the entire text inference process forward. This complete process, from parsing the probability matrix to returning the target word, constitutes a safe sampling stepping loop for large language models.
[0093] The cumulative probability ladder is an auxiliary data structure used in the roulette wheel selection algorithm. It transforms the probability distribution into segmented continuous intervals, facilitating sampling using a random number. The probability interval for candidate character words is this... The logical representation of a continuous interval is such that the probability of each word determines the length of the line segment it occupies in that interval.
[0094] The target random number main data stream segmentation refers to the portion of the truly random bit sequence stored in the data payload area of the combined data packet entity. Projection refers to the process of mapping a high-dimensional bit sequence to a low-dimensional space through numerical interpretation.
[0095] The dictionary mapping table is a data structure that connects lexical indices to their specific character representations, and it is an implementation of the vocabulary of a large language model. The target lexical is an output lexical generated under the protection of the method of this invention, and its selection process depends on physical randomness.
[0096] In a specific example of the present invention, following the example of the preceding steps, the adaptive sampling kernel function has generated a conditional probability distribution matrix with values [0, 0.595, 0, 0.405]. First, the kernel function calculates its cumulative probability ladder value, obtaining the sequence [0, 0.595, 0.595, 1.0]. This constructs the probability interval of the candidate character vocabulary, where index 1 corresponds to the interval... The interval corresponding to index 3 Simultaneously, the kernel function extracts an 8-byte target random number main data stream segment from the entity data payload area of the current combined data packet, its hexadecimal representation being 0xD307549C1A8BE24F. It extracts this 64-bit data as an unsigned integer and divides it by... Perform mapping calculations to transform it into The calculated value for the interval is a double-precision floating-point number, approximately 0.8243. Next, mapping is performed. This random number, 0.8243, is projected into the probability interval of the candidate character vocabulary. The kernel function compares the following: step value 0 is less than 0.8243; step value 0.595 is less than 0.8243; step value 0.595 is less than 0.8243; but step value 1.0 is greater than 0.8243. Therefore, the random number falls into the fourth interval. The index of this interval is 3. The kernel function queries the dictionary mapping table to find the term at index 3, assuming this term is "quantum". Therefore, "quantum" is determined as the target term for this sampling. The kernel function ultimately returns "quantum", completing one model sampling step loop. After receiving "quantum", the large language model writes it into the preset large language model context dialogue pool as the new context. Simultaneously, the system clears and releases the temporary shared memory record block occupied by this federated data packet and resets the local computation variable accumulator, preparing to begin the next round of model inference propagation loop.
[0097] See appendix Figure 4 This invention also proposes a secure sampling system for large language models based on quantum entropy sources, comprising the following modules: The quantum signal processing module is used to acquire the original quantum signal at the hardware level, perform randomization processing on the original quantum signal, and generate the target random number main data stream and physical state signal; The entropy state vectorization module is used to obtain the difference between the read and write pointers of the circular buffer, combine it with the physical state signal to perform normalized array encapsulation, and generate a quantum entropy state vector. The data encapsulation module is used to extract entity segment data from the target random number main data stream according to the preset data structure specifications, and to perform a header binding mechanism on the quantum entropy state vector and entity segment data to generate a joint data packet; The sampling call interception module is used to intercept the preset underlying random number call logic of the large language model inference program, extract the quantum entropy state vector from the joint data packet for verification, and activate the adaptive sampling kernel function when the verification passes. The hardware failure monitoring module is used to analyze the original entropy generation rate in the quantum entropy state vector according to the adaptive sampling kernel function. If the original entropy generation rate is lower than the preset hardware failure threshold, a deterministic decision signal is generated and output to switch to a non-random algorithm and skip the probabilistic processing step for the preset log probability matrix. If the original entropy generation rate is not lower than the preset hardware failure threshold, the sampling hyperparameter modulation module extracts the mass damage degree and buffer filling speed from the quantum entropy state vector. Based on this, it adjusts the sampling temperature and probability mass threshold of the large language model sampling strategy, and uses the adjusted sampling temperature and probability mass threshold to perform restriction processing on the preset log probability matrix to generate a conditional probability distribution matrix. The quantum random mapping module is used to call the entity segment data to match the conditional probability distribution matrix to perform targeted mapping calculations, generate and output target words.
[0098] Each of the modules can be implemented in whole or in part through software, hardware, or a combination thereof. It supports hardware embedded in or independent of the processor in the computer device, and also supports software stored in the memory of the computer device, so that the processor can call and execute the operations corresponding to each of the above modules.
[0099] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for secure sampling of large language models based on quantum entropy sources, characterized in that, Includes the following steps: S1. Obtain the original quantum signal at the hardware level, perform randomization processing on the original quantum signal, and generate the target random number main data stream and physical state signal; S2. Obtain the difference between the read and write pointers of the circular buffer, combine it with the physical state signal to perform normalized array encapsulation, and generate a quantum entropy state vector; S3. Extract the entity segment data of the target random number main data stream according to the preset data structure specifications, and perform a header binding mechanism on the quantum entropy state vector and the entity segment data to generate a joint data packet; S4. Intercept the preset underlying random number call logic of the large language model inference program, extract the quantum entropy state vector from the joint data packet for verification, and activate the adaptive sampling kernel function when the verification is successful. S5. Analyze the original entropy generation rate in the quantum entropy state vector according to the adaptive sampling kernel function. If the original entropy generation rate is lower than the preset hardware failure threshold, generate and output a deterministic decision signal to switch to the non-random algorithm and skip the probabilistic processing step for the preset log probability matrix. S6. If the original entropy generation rate is not lower than the preset hardware failure threshold, extract the mass damage degree and buffer filling speed from the quantum entropy state vector, adjust the sampling temperature and probability mass threshold of the large language model sampling strategy accordingly, and use the adjusted sampling temperature and probability mass threshold to perform restriction processing on the preset log probability matrix to generate the conditional probability distribution matrix. S7. Call the entity segmentation data to match the conditional probability distribution matrix to perform targeted mapping calculation, generate and output target words.
2. The quantum-entropy-source-based large language model secure sampling method according to claim 1, characterized in that, Step S1 includes the following steps: Based on a preset sampling frequency, the physical components of an independent quantum random number generator are monitored to capture physical noise variation sequences to form the original quantum signal; The original quantum signal is input into a preset cryptographic confusion function to extract a high-density data stream. Parallel unbiased verification items are performed on the high-density data stream, and the data that passes the verification items are spliced together to form the target random number main data stream. The failure rate parameter of the statistical unbiased verification project and the real-time bit generation parameter of the quantum random number generator are encapsulated to generate physical state signals.
3. The quantum-entropy-source-based large language model secure sampling method according to claim 1, wherein, Step S2 includes the following steps: The current buffer data storage volume is calculated based on the difference between the read and write pointers in the current time window. Time difference operation is performed and combined with smoothing filtering to generate a stable instantaneous filling speed. Extract the mass damage degree and original entropy generation rate contained in the physical state signal; Based on preset hardware theoretical parameters, the instantaneous filling speed and the original entropy generation rate are normalized to obtain scaling indexes. The scaling indexes and the mass damage degree are then reordered according to preset levels to form a quantum entropy state vector.
4. The quantum-entropy-source-based large language model secure sampling method according to claim 1, wherein, Step S3 includes the following steps: The memory space is divided into a metadata header area and a dynamic length payload area according to the preset data structure specifications. Extract the entity segment data for the current sampling request period; The quantum entropy state vector is written to the starting position of the allocated metadata header area, and the entity segment data is loaded into the dynamic length payload area to form a joint data packet with continuously distributed memory addresses.
5. The quantum-entropy-source-based large language model secure sampling method according to claim 1, wherein, Step S4 includes the following steps: Obtain the complete combined data packet through data dequeueing operations; Parse the metadata header of the combined data packet to extract the preset integrity check code comparison field and the quantum entropy state vector; The current integrity check code is calculated for the combined sequence of the extracted quantum entropy state vector and the corresponding entity segment data. The data consistency of the comparison field is compared with the current integrity check code to complete the verification.
6. The quantum-entropy-source-based large language model secure sampling method according to claim 1, wherein, Step S5 includes the following steps: Extract the original entropy generation rate from the quantum entropy state vector; The failure status of the physical entropy source is determined by comparing the original entropy generation rate with the preset hardware failure threshold. If the original entropy generation rate is lower than the preset hardware failure threshold, a physical source alarm command is generated and sent to the large language model inference program to bypass the probability calculation logic and switch to the preset greedy search algorithm to generate a deterministic decision signal.
7. The quantum-entropy-source-based large language model secure sampling method according to claim 1, wherein, Step S6 includes the following steps: Extract the mass damage degree from the quantum entropy state vector; The attenuation ratio is obtained by multiplying the quality damage degree by a preset adjustment coefficient; based on the preset reference temperature and the preset global absolute temperature lower limit, the amplitude limiting peak finding calculation is performed in combination with the attenuation ratio to generate the reconstructed sampling temperature value; The pre-defined logarithmic probability matrix is scaled and smoothed using the reconstructed sampled temperature values, and the output conditional probability distribution matrix is transformed.
8. The quantum-entropy-source-based large language model secure sampling method according to claim 1, wherein, Step S6 also includes the following steps: Extract the buffer filling velocity from the quantum entropy state vector and introduce it into the preset core threshold calculation equation to generate a dynamic probability mass clipping point; The candidate word probability components are accumulated in descending order according to the temporary probability distribution obtained by transforming the log probability matrix. Remove the tail distribution components of the cumulative result set to zero that exceed the dynamic probability quality clipping point, and perform a second normalization.
9. The quantum-entropy-source-based large language model secure sampling method according to claim 1, wherein, Step S7 includes the following steps: Construct an ascending probability ladder sequence by accumulating the conditional probability distribution matrix; The preset bits of the entity segment data are extracted into integer values and a division mapping calculation is performed to convert them into equivalent uniformly distributed floating-point values, which are then used as the target falling pointer. Traverse the ascending probability ladder sequence to find the smallest ladder element index that exceeds the target falling pointer; The minimum step element index is converted into an actual string by calling the preset dictionary mapping table, and the actual string is output as the target word.
10. A large language model secure sampling system based on quantum entropy sources, characterized in that, Includes the following modules: The quantum signal processing module is used to acquire the original quantum signal at the hardware level, perform randomization processing on the original quantum signal, and generate the target random number main data stream and physical state signal; The entropy state vectorization module is used to obtain the difference between the read and write pointers of the circular buffer, combine it with the physical state signal to perform normalized array encapsulation, and generate a quantum entropy state vector. The data encapsulation module is used to extract entity segment data from the target random number main data stream according to the preset data structure specifications, and to perform a header binding mechanism on the quantum entropy state vector and entity segment data to generate a joint data packet; The sampling call interception module is used to intercept the preset underlying random number call logic of the large language model inference program, extract the quantum entropy state vector from the joint data packet for verification, and activate the adaptive sampling kernel function when the verification passes. The hardware failure monitoring module is used to analyze the original entropy generation rate in the quantum entropy state vector according to the adaptive sampling kernel function. If the original entropy generation rate is lower than the preset hardware failure threshold, a deterministic decision signal is generated and output to switch to a non-random algorithm and skip the probabilistic processing step for the preset log probability matrix. If the original entropy generation rate is not lower than the preset hardware failure threshold, the sampling hyperparameter modulation module extracts the mass damage degree and buffer filling speed from the quantum entropy state vector. Based on this, it adjusts the sampling temperature and probability mass threshold of the large language model sampling strategy, and uses the adjusted sampling temperature and probability mass threshold to perform restriction processing on the preset log probability matrix to generate a conditional probability distribution matrix. The quantum random mapping module is used to call the entity segment data to match the conditional probability distribution matrix to perform targeted mapping calculations, generate and output target words.